How RTP is verified
The certification standards say what they check but not how much data it takes — the word “confidence” appears once in GLI-11. Here is the sample-size mathematics they leave out: how many spins 96.0% actually needs.
“Certified RTP” sounds like a measured quantity, checked to some precision. Read the actual standards and a gap opens up: they are exact about what a laboratory must test and almost silent about how much data it takes. This chapter fills the gap with the sample-size mathematics — and lets you turn the crank yourself.
What GLI-11 actually requires
Theoretical RTP is verified against the game’s model, not by empirical play. GLI-11 defines the theoretical return as based on “mathematical calculations or simulations”, and sets one substantive floor on the number: each game “shall theoretically payout a minimum of seventy-five percent (75%) during the expected lifetime of the game,” a requirement that “shall be met for all wagering configurations.” For the random source, the standard requires that the laboratory’s chosen tests “shall be evaluated, collectively, at a 99% confidence level,” and that a data-collection run gather “at least 10,000 game outcomes” (adjustable up or down case by case).
What it deliberately does not require
Notice what is absent. There is no spin count, no tolerance and no confidence interval attached to the return figure — no rule of the form “confirm 96.0% to within ±0.1%.” The word “confidence” appears exactly once in the whole of GLI-11, in that single 99% line about RNG tests. The choice of which statistical tests to run is delegated to the laboratory case-by-case; ongoing-monitoring language in the online-systems standards is procedural (investigate large variances) rather than a statistical definition of “large.” None of this is a scandal — it is a reasonable division of labour, because the theoretical number comes from the model. But it does mean the phrase “certified to 96.0%” is a statement about a model that met a floor, not about an empirical measurement to a stated precision.
The mathematics the standards leave out
Suppose you did want to confirm a return empirically — from spins alone, to a stated tolerance. That is a confidence-interval problem for the mean of the per-spin return. If each spin’s return has standard deviation σ (in bet units) and you want to pin the mean to within a tolerance tol at confidence level z, the number of spins needed is:
For a medium-volatility slot (σ ≈ 5 bet units), a tolerance of ±0.1% (tol = 0.001) and 95% confidence (z = 1.96): N = (1.96 × 5 / 0.001)² = 9,800² = 96,040,000 spins. Roughly 96 million — to confirm one game to a tenth of a percent.
The structure is unforgiving. Because σ is squared, doubling volatility quadruples the spin count; because tol is squared, asking for half the tolerance quadruples it too. Jackpot-heavy games are the extreme case: their return lives in very rare, very large hits, so their true standard deviation is far larger than any everyday slot, and the honest sample size explodes accordingly. Sequential testing methods can stop early when a result is clearly inside or outside the band, but they do not repeal the arithmetic — a tight confirmation still needs an enormous amount of data.
Turn the crank
Change the volatility, tolerance and confidence and watch the required spin count move. This is the number the standards never print.
Normal approximation — a confidence interval on the mean per-spin return, N = ⌈(z · σ / tolerance)²⌉. Jackpot-heavy games hide their return in rare giant hits, so their true σ is far larger than any preset here and they need dramatically more spins than this figure suggests. See the chapter for why.
How to read “certified RTP,” then
A certificate tells you a game’s model was evaluated and met the standard’s requirements, including its payout floor and its RNG-test confidence level. It does not tell you the return was measured, by spinning, to the decimal places printed — because empirically pinning those decimals would take the tens or hundreds of millions of spins the calculator shows. Read the figure as a verified property of the design, which is exactly what it is, and no more. The companion question — what a certificate does and does not assert about the random number generator — is taken up in the RNG pillar.
Next, the same games often ship in several return configurations at once. What the certificate covers when a title exists at 96%, 94% and 92% is the subject of configurable RTP.
Common questions
How many spins does it take to verify a 96.0% RTP?
Far more than most people expect. Pinning the mean return of a medium-volatility slot (standard deviation about 5 bet units) to within ±0.1% at 95% confidence takes roughly 96 million spins, from N = (z·σ / tolerance)². Higher volatility or a tighter tolerance pushes the figure up sharply, and jackpot games — whose return hides in very rare, very large hits — need dramatically more still.
Do laboratories actually spin the game millions of times to check its RTP?
Not to establish the theoretical RTP — that is verified against the game’s mathematical model (calculations or simulations), not by empirical spinning to a tight tolerance. Standards such as GLI-11 require a data-collection sample (at least 10,000 game outcomes, adjustable case by case) and evaluate the RNG’s tests “collectively, at a 99% confidence level,” but they specify no spin count, tolerance or confidence interval for confirming the return figure itself.
Sources (1)
Education, not advice. This chapter explains how return-to-player is defined, computed and checked so you can read the number honestly. It is not a system, and nothing here treats gambling as a way to make money — over enough play the mathematics favours the house. 18+.
Next in the pathRTP rules by jurisdiction