How the artefact was caught. The rider's mean residual
correlated +0.887 with the rider's mean predicted value
— the control had plainly failed. The underlying model
under-predicted, and not by a constant: the shortfall ran from +£84
per ride in the cheapest decile to +£1,562 in the dearest. This
even survives a naive within-horse test, because a stable's best rider
takes the horse to its valuable engagements while others
partner it in handicaps.
What survives, on the rebuilt control sweep. Recentring
within predicted-value band, season, country and code, then comparing
the same horse in races of similar value: the permutation p runs
0.000 recalibrated, 0.458 once you add the
same horse, and 0.500 once you add the race value too
— that is, by the time both controls are on, a permutation test
can no longer distinguish the spread of rider effects from what
shuffling the labels produces. Read that as a weak test failing to
detect, not as proof there is nothing to detect. The permutation
measures the overall spread and is low-powered against heavy tails,
and heavy tails at low ride counts are exactly what breaks the variance
decomposition described below — one cause plausibly explains
both. It is also why the two tightest specifications swap places
between data vintages: that is instability in a low-powered test, not
a finding about which specification is right. Split-half agreement on
that tightest money-scale
specification is +0.129, and the between-rider spread falls
to £522.6 a run. Shrinkage is severe throughout: at the
current threshold a rider needs about 1,616 rides before
even half of their raw estimate is retained. There
is still more here than noise — individual bootstrap
intervals separate about a hundred riders — but it is not a
ranking and it will not be presented as one.
One row of that sweep is deliberately blank, and the reason is
instructive. The uncontrolled population figure has no rebuilt
value, because the code that produces the sweep only reports on the
recalibrated residual. An earlier draft filled the gap with +0.346
— which is a different model's recalibrated split-half, a
different quantity entirely. It was caught before publication. We
record it because a blank cell is honest and a plausible number in a
blank cell is not, and because this page's own worst error was the
same shape.
The estimator behind the earlier version of this finding turned
out to be unstable, and that is the strongest thing on this page
against naming riders. Until 2026-08-15 we published "452 of 576
riders cannot be told apart", measured with a 100-ride floor. On the
rebuilt spine that threshold returns a negative between-rider
variance (-56,855) — meaning no signal at all, and no
estimates that can legitimately exist. The cause is not the prize
correction, which moves the figure about 10% and does not move its
sign; it is heavy tails at low ride counts. 77 riders
with 100–150 rides carry a within-rider variance around 119
million — one large prize on a rare good mount — and
those terms alone swamp the noise estimate for the whole population.
Seven more days of data flipped it.
What that would have produced if nobody had guarded it. Not
an error and not an empty table: a full table of large, plausible,
sign-flipped figures against named professionals. The estimator
now refuses and prints why. This is the clearest illustration we have
of why the ranking is not published — the failure mode is not
"the numbers look wrong", it is "the numbers look right".
The current basis. A 200-ride floor, which is where the
estimator is stable across thresholds and across both prize columns:
459 riders, 34 clear of zero above and
68 below. The older "38 above, 86 below, of 576"
cannot be reproduced at the threshold that produced it, and the two
sets of counts are not comparable — so they are not compared
here, and no trend should be read between them.
The names, however, do compare — and they say the
instability is all at the tail. Counts measured at two different
thresholds cannot be set against each other, but asking which
riders appear on both lists is well defined either way.
27 of the 38 survive; 11 drop and 7 are new. Of
the 11 that drop, 4 fall below the new ride floor
mechanically — all of them worth £20–30 a ride, which is
to say nothing — and 7 are still in the data with
an interval that no longer clears zero. Every one of the drops is
from the bottom of the list. The head does not move: the seven
largest estimates all remain clear of zero at similar magnitude, the
top one going from £1,062 to £912 a ride.
So the honest reading is that the head is real and the tail is
noise-dominated, rather than that the whole thing is mush. That is
a narrower claim than "these estimates do not survive", and it is the
true one. It is also why the conclusion does not soften: the marginal
entries were never findings, and they are precisely the ones that
evaporate when the estimator is put on firmer ground. A list whose
bottom half rearranges itself between data vintages is not something to
publish as a ranking, even when its top half holds. No rider is
named on this page in either direction.
Two things we will not paper over. The permutation test and
the bootstrap disagree, and on the rebuilt spine they disagree
more: on the money scale the permutation p moved from 0.050
to 0.500, which is squarely at chance, while the
individual bootstrap intervals still separate about a hundred riders.
The permutation tests the overall spread and is low-powered against
heavy tails; the bootstrap tests individual riders. We are not going to
pretend that is resolved. And the surviving ordering matching expert
consensus is reassuring rather than evidence — exactly
the kind of face validity that makes a confounded result feel true.
A separate question with a cleaner answer. Asked whether a
rider beats their own market price, the answer is that no rider
persistently does. That is not a weak effect, it is no effect: the
split-half agreement between an early and a late period is
-0.002 across 449 riders — indistinguishable
from zero. It is the expected result, because the Betfair price
already prices the jockey.
This page said "the split-half is negative" until 2026-08-15,
and that was wrong. The published figure was −0.243 and was
read as riders being anti-predictive of their own market, with
a regression-to-the-mean story built on top. Re-measured with the same
estimator, it is -0.002. The cause was inside the script that
produced it: a blanket "multiply Irish rows by 0.600" — the same
era-blind correction this project's own errata condemns — still
live there, injecting a distortion correlated with both jurisdiction
and period, which is exactly the shape that manufactures
early-versus-late structure. The conclusion is unchanged and
cleaner; the anti-predictive reading is withdrawn. One residual is
real and unexplained: British riders split-half at
-0.159 against Irish riders at +0.273.
Why this reached the page at all is worth recording. The
figure was never quoted here as a number, only as the phrase "the
split-half is negative" — so a search of this page for the
published value found nothing, and the claim survived a check designed
to catch exactly this. A story can carry a withdrawn finding after its
number has been dropped.
The standing rule this produced. Any comparison between two
groups of riders is a comparison of their mounts until you
condition on the horse. It cost 85% of the raw
prize-money effect measured here, and then it reversed the
sign of the claimer result. The size of that correction does not carry
across models: on the win-probability work behind the booking grid
above, adding a horse control moves the spread by about 11 to 14%, not
by 85%. What generalises is the direction, not the
magnitude — assume the rule applies to every new rider comparison
until shown otherwise, and measure the cost separately each time.