← PlusEVData

Sales research library

Everything we have tested for the horses-in-training and yearling sales work, in one place — including, and especially, the things that turned out not to be there. Each finding carries the number that supports it, the population it was measured on, and the date it was run.

Last rebuilt 15 August 2026 · Racing data to 2026-08-13

How to read this page. Each finding opens with one plain sentence and the number that carries it, with the denominator attached. Everything underneath — the method, the confidence intervals, the things we could not control for — is one click away rather than deleted. If a caveat is inside a drawer it is because the page is long, never because it is inconvenient.

Section 3 is not an appendix. It is what we tested and did not find, and it is here because it is the more useful half. Two of those nulls are the reason a plausible-looking product was not built, and one of them — a model that scored 0.206 on a headline metric — was withdrawn after we worked out what it was really measuring. A page that only showed the positives would be a different and much less honest page.

Nothing here is a causal claim. Where a yard or a rider appears with a number beside it, the number describes what has happened to horses in that situation, not what that person does to a horse. Where the difference matters we say so on the finding itself.

Section 1

What pays

Findings that change what you would spend money on. Each is a comparison between things you can actually choose between — a yard, a horse type, a lot to bid on.

The cheapest competent yard almost always wins, and the arithmetic is not close

−£7,594
what the dearer of two real yards costs over a 20-month career, after crediting it with every pound of uplift we can measure

A pound of rating is worth real money, and we can say how much: about £14.13 a month for a 90–110 rated horse, rising to £30.78 at the top of the handicap. But a stronger yard buys you a few pounds of rating, not thirty. Take two real Irish yards: the uplift on a 95-rated horse with 20 months of career left is worth about £85 a month, and the dearer yard costs £456 a month more. To break even you would need roughly 32 pounds of uplift. The best yard we measure is under six.

Basis 5,111 careers, prize actually banked (prize_money_won), fitted in log space and evaluated at the median so it describes a typical horse. Rebuilt 2026-08-15.

Method, caveats and confidence intervals

The worked example, in full. A +6 lb yard, a 95-rated horse, 20 months of career left, 90% chance of racing again:

  • Extra earnings: 6 lb × £14.13 × 20 months × 0.9 = £1,526
  • Fee premium: (£1,565 − £1,109) × 20 months = £9,120
  • Net: −£7,594

The break-even, with the arithmetic shown. At the 90–110 band rate the example uses:

  • What 6 lb buys: 6 × £14.13 = £85 a month.
  • What the premium demands: £456 ÷ £14.13 per lb per month = 32 lb of uplift.
  • What is actually available: the corrected estimator's best yard is +5.93 lb pooled across codes and +5.95 lb on the Flat. So the premium needs roughly five and a half times the largest effect that exists anywhere in the measured set, not merely more than this yard offers.

This conclusion got stronger on 2026-08-15, not weaker. When the arithmetic was first done, the destination-yard estimator was contaminated by a units bug and reported a best yard of +9.6 lb, and the "+6 lb yard" was described as near the top of 140 measured yards. Both descriptions are stale. On the corrected estimator a +6 lb yard is above the ceiling of the whole market. The uplift side of the trade is about 40% smaller than it looked; the fee side has not moved.

Why it survives imprecision. The gap is a factor of three to five, so the conclusion holds through an order of magnitude of error in the £-per-pound estimate in either direction. That is the only reason it is publishable at all: the estimate itself is not precise. The two fee quotes are also the only two real ones we hold — this is one real pair of yards, not a fee survey.

The sub-70 band is published so that its exclusion is checkable. It fails its own currency control — £1.58 against £4.21 on British runs only, a 166% disagreement where every band above 70 agrees within ±10%. The source's own instruction is "do not quote a trainer-uplift value for a sub-70 horse", and this page follows it.

What this is fitted on. Deliberately not the expected-value columns, which are built on prize_fund_winner and run about 67% high for Irish lots. This uses money actually banked.

Why the live pages show slightly different numbers, and why that is correct. The figures above come from a single fit over the whole career set. The sales pages refit the same function per sale, using only moves that happened before that sale's date — most recently £2.21 / £22.34 / £12.01 / £29.00 on 5,150 careers. That is as-of discipline working, not drift: a page about a sale must not be fitted on data from after it. The 40–70 band is unusable in both, and the worked example above uses the whole-set figures throughout.

Rating band£ per lb per month CareersBritish-only control DisagreementCurrency check
40–70£1.581,433£4.21+166%Unusable — noise
70–90£18.831,772£20.61+9.5%Passes
90–110£14.131,212£13.07−7.5%Passes
110–200£30.78694£32.73+6.3%Passes

The control column refits on British runs only. Prize money is a mix of sterling and euro, so a band whose British-only figure disagrees is telling you the conversion is carrying the result rather than the horses. Source: NOW.md, rating-to-money section, 2026-08-15.

Store horses have returned about three times as much prize money per pound spent as yearlings — and the usual objection to that makes the gap bigger, not smaller

£0.99 vs £0.35
prize money per £1 of purchase price, 48 months from sale — 1,680 stores against 6,404 yearlings

The obvious objection is that these are different slices of a career: a yearling is 1.75 when it sells and a store is 3.49, so a fixed 48-month window is not watching the same thing. That objection is right about the mechanism and wrong about the direction. The window captures 66% of a yearling's lifetime earnings and only 24% of a store's, so the published figure is a floor for the store, not a flattering one. Every age-honest treatment we tried widened the gap — up to 7.0× when both populations are compared at the same ages.

Basis Sales corpus joined to the GB and Ireland racing spine, 740,853 runs, 2020-01-01 to 2026-08-13. Non-runners counted as zeros. Audited for age censoring 2026-08-15.

Method, caveats and confidence intervals

Treat the headline as a conservative floor for stores. A yearling peaks at three and is finished; a store peaks at seven and is still earning at twelve. Given its whole observable life from age 2 to 8 — the most generous treatment possible, crediting it even after it has been sold on — the yearling returns 0.35, the published figure exactly. There is nothing outside the window to find on that side. The store figure roughly triples.

Sensitivity to where a lifetime is cut off. Through age 5 the ratio is 1.7×; through age 8, 5.0×; through age 10, 6.9×; through age 12, 8.7× (on 19 stores, thin and flagged). The narrowest defensible framing still leaves stores ahead.

Do not compare the raw columns without standardising price. Yearlings average £70,024 and stores £35,201, earnings per pound falls steeply with price in both populations, and the published statistic is a pooled sum-over-sums. So a raw comparison partly compares price levels rather than horse types. Standardised to a common price mix, the gap is +2.49 and the ratio is 3.1× against an unstandardised 3.5×. Stores lead in every one of the five price bands, from 4.2× under £5k to 1.9× at £60–150k.

Controls that could have overturned it and did not. Birth-year matched (same generation, same calendar): 1.16–1.42 for stores against 0.16–0.19 for yearlings across 2016, 2017 and 2018. Real deflated prices: gap widens. Resale proceeds included, which could have rescued the yearling: both populations resell at almost identical rates (32.9% against 31.6%) and the total per pound is 0.53 for yearlings against 1.89 for stores. Median lot rather than pooled ratio: yearling 0.19, store 1.01. Shuffling the population label within birth year gives a null of +0.013 ± 0.027 against an observed +1.054, p = 0.0000.

Two censorings that both favour the yearling, and neither can be quantified from this corpus. These are the honest limit on how far the finding can be pushed:

  • Breeding and stud value is invisible to both the sales feed and the racing spine, and it falls almost entirely on the flat (yearling) side. A colt going to stud can be worth a multiple of anything measured here. No number is guessed for it. This is the largest single unknown.
  • Export is invisible. The racing spine is GB and Ireland only. A horse exported to race abroad reads as a zero here while earning in fact, and Flat yearlings export far more than jumps stores do. The corpus cannot separate "exported" from "never made the track".

Also unpriced: timing and keep. Every figure is undiscounted, and the yearling buyer pays about 1.8 years earlier and carries 1.8 more years of keep — which makes the undiscounted comparison, if anything, generous to the yearling. Flat against National Hunt is not controlled and should not be; it is most of what distinguishes the two propositions.

One filter that produces a wrong answer, recorded so nobody repeats it. Restricting to "career-complete" horses deletes 98.3% of stores — only 10 of 605 have a finished career by age 8.6 — and the deleted ones earn 9.1× the kept ones, because the filter removes precisely the horses still racing. The store figure it produces (0.14) is an artefact and must not be quoted.

Among stores, what you pay barely predicts what the horse earns — and that is specific to stores

+0.104
rank correlation between price and earnings, 1,680 stores — against +0.205 for yearlings

Across a 26-fold range of store prices, the top price decile earns a median 1.5× the bottom — where for yearlings the same comparison is 15.1×. Price ranks yearling earnings about twice as well as store earnings. Lengthening the window does not turn price into a predictor of store earnings: holding the same 256 horses and moving only the window end, the correlation goes 0.175 → 0.104 → 0.221 and stays weak throughout, with overlapping confidence intervals.

Basis Same corpus. Published 48-month window; re-tested age-matched and at every window length from [4,6) to [4,10). Audited 2026-08-15.

Method, caveats and confidence intervals

Confidence intervals. Published window +0.104 [+0.054, +0.152]; age-matched [4,8) +0.116 [+0.057, +0.172]; longest available window [4,10) on birth-2016 stores +0.221 [+0.094, +0.339]. Weak throughout, and rising only slightly with window length.

When you buy the store changes the answer. Stores sold at age 3.0–3.5, the spring sales, give +0.057 [−0.027, +0.140] — not distinguishable from zero at all. Stores sold at 3.5–4.0, the autumn sales, give +0.233 [+0.160, +0.310]. "Price tells you nothing" is much the stronger claim for the spring half of the market.

Adding resale proceeds to the outcome barely moves it (+0.130 at ages [4,8)). Over each population's own full life slice the contrast holds: yearling +0.205 [+0.170, +0.241] against store +0.094 [+0.015, +0.173].

We can tell you which lots are most likely to go unsold, but not what the odds are

7.8×
spread in actual non-sale rate between the bottom and top fifth of the catalogue — 1.96% against 15.33%

The ordering is real and monotone: sort a catalogue by this model and the bottom fifth goes unsold 1.96% of the time against 15.33% for the top fifth. The probabilities attached to that ordering are not usable. The model under-predicts by about five times at the bottom of the range, and its skill against simply quoting the sale's own clearance rate is about zero. So this will be published as a relative band — "this lot is in the fifth of the catalogue least likely to sell" — and never as a percentage.

Basis Held-out calibration quintiles pooled over all folds. Lot count not stated in the source — a denominator has not been invented for it; reproduction queued. Ranking approved, not yet shipped.

Method, caveats and confidence intervals

The two numbers that decide how it can be presented. Mean-fold AUC 0.621, which is a real but modest ability to rank; and Brier skill against the sale's base rate of +0.002 — recorded elsewhere as "about zero", and +0.002 is what the measured figure actually is. Either way the probability adds essentially nothing over knowing the sale's historic clearance rate. A model can rank well and be worthless as a probability, and this is that model.

Why the calibration is off. Partly the model, and partly that the base rate is itself drifting down over the folds, from 11.0% to 7.0%. A probability fitted across a drifting base rate is answering a question about the average year rather than this one.

The fix is known and is not a research problem — a held-out isotonic recalibration. It has not been started, so until it lands the band is what there is.

Quintile Model saidActually happened
1st — least likely to go unsold0.36%1.96%
2nd0.97%2.53%
3rd1.97%4.52%
4th4.49%7.71%
5th — most likely to go unsold18.05%15.33%

Held-out quintiles, pooled over folds. The right-hand column is monotone, which is the publishable part. The left-hand column is wrong by up to five times, which is why no percentage is shown on the sales pages. Source: SALES_CORPUS_INVENTORY.md.

Section 2

Who to use

Findings that name a yard, or price a booking decision. Read the attribution note on each one: these describe what has happened to horses in a situation, not what a person does to a horse.

On the Flat, horses arriving at some yards have historically outperformed what their form predicted — by up to six pounds

+5.95 lb
best of 159 Flat yards with 15+ arrivals, from 7,951 same-code moves

When a horse changes yard, we predict where its rating should land from its own form history alone — the model never sees which trainer it left or which it joined. What is left over, averaged across every horse a yard received, is the number below. Horses arriving here have historically outperformed what their form predicted. That is all it says: it is not a measurement of training, because a yard that is simply good at buying produces exactly the same signal.

Basis Timeform spine, 740,853 runs, 2020-01-01 to 2026-08-13. Flat-to-Flat moves only. Rebuilt 2026-08-15.

Method, caveats and confidence intervals

What the number is. A baseline model predicts a horse's post-move rating from its pre-move level, last run, trend, age, gap since running, career stage, prize level and Betfair SP. It is blind to both the origin and the destination yard. The yard's effect is the mean residual across the horses it received, shrunk by n/(n+20) so a yard with sixteen arrivals is pulled harder toward zero than one with ninety. The ± column is the standard error of the shrunk figure shown, not of the raw mean.

Does it survive falsification? Shuffling the destination labels 30 times gives a noise floor; the real spread of effects is 2.08× that floor across all codes and 2.27× on the Flat alone. Splitting the data in time and re-estimating, a yard's effect in the earlier period predicts its effect in the later one at r = +0.662 overall and r = +0.698 on the Flat. Residuals do not correlate with the horse's own pre-move rating (r = −0.018), so mean reversion is absorbed rather than left in the answer.

What it is worth, and it is less than we used to say. Backing top-third yards over bottom-third yards is worth +1.90 lb against -1.22 lb on later moves — a spread of 3.12 lb. Until 2026-08-15 this page's research recorded 7.55 lb. The difference was a bug: pre-move and post-move ratings were being differenced across racing codes, and a Flat rating and a hurdle rating are not the same ruler. About 59% of the apparent value of this model was the ruler, not the yard.

The strong-yard benchmark, with its vintage attached. The figure the sales pages compare a lot against fell from +1.54 lb to +1.10 lb, a drop of 28.5%. That number is measured on moves before the ANH25 sale date (2025-10-27) against a spine ending 2026-08-13 (fingerprint 4cc4835ebcc04608), which is 9,373 moves rather than the full 10,561. It is not a constant and should not be quoted as one — it is refitted for every sale using only moves that happened before that sale, so it will move for the next catalogue simply because the cutoff moves. Quoted with its sale and spine attached, the next change reads as expected rather than as a regression. The served sales file carries this figure, verified against the file a browser is given rather than the code that writes it. Two other versions of the same statistic exist and are both correct about different questions: −22.0% describes the estimator over the whole population, and −29.0% is the validation suite's held-out figure. Neither belongs beside a lot.

Six of the previous top eight are no longer in the top eight, and one — a yard whose recorded figure rested on 7 same-code moves — drops below the reporting gate entirely and is not shown. The old list was substantially a ranking of which jumps yards take the most ex-Flat horses.

What is not controlled, and it is the big one. Selection on the origin side. Yards choose horses as much as horses choose yards; if a yard systematically receives horses this baseline misprices, its residual is non-zero for reasons that have nothing to do with training. A yard near the top may simply be good at buying. Two further limits: the same-code restriction removes 19.1% of moves and those are not random (Flat→Hurdle is ordinary jumps career progression, not an anomaly), and only a horse's first recorded move is used, so this measures what a yard does with a horse arriving from its first trainer.

Expect blanks. Comparing origin against destination needs both yards past the 15-move gate, which happens on 32.3% of Flat moves. Where it cannot be computed the sales page shows nothing, never zero.

#Destination yard Effect (lb)± se Unshrunk (lb)Moves
1John & Thady Gosden+5.95±0.80+7.7965
2William Haggas+4.04±0.70+8.0720
3Joseph Patrick O'Brien, Ireland+3.75±0.77+4.8270
4George Boughey+3.39±0.62+4.2282
5Andrew Balding+3.11±0.77+6.2120
6Harry Charlton+2.91±0.91+5.3324
7Ralph Beckett+2.85±0.68+6.4216
8Hugo Palmer+2.83±0.65+4.4934
9Harry & Roger Charlton+2.76±0.71+5.2722
10Charlie Johnston+2.72±0.55+3.4872
11P. Twomey, Ireland+2.71±0.82+5.1722
12Charlie & Mark Johnston+2.63±0.43+3.2191
13Marco Botti+2.58±0.85+4.5726
14Kevin Philippart de Foy+2.38±0.56+2.9584
15Stephen Thorne, Ireland+2.14±0.57+4.6617
16M. Halford & T. Collins, Ireland+2.06±0.77+3.8623
17Charlie Fellowes+2.04±0.65+3.4429
18David O'Meara+2.01±0.49+2.4788
19John Patrick Murtagh, Ireland+1.99±0.54+3.6424
20K. R. Burke+1.98±0.77+3.5026
21John & Sean Quinn+1.83±0.51+3.1927
22Richard & Peter Fahey+1.74±0.50+2.2271
23Edward Bethell+1.72±0.72+2.5044
24Ollie Sangster+1.66±0.57+3.4119
25Hamad Al Jehani+1.64±0.48+3.4518
26A. Slattery, Ireland+1.63±0.69+3.2520
27Robert Cowell+1.59±0.48+2.5234
28James Horton+1.50±0.67+3.4915
29George Scott+1.48±0.63+2.9021
30Freddie & Martyn Meade+1.47±0.82+3.2017
31Harry Eustace+1.39±0.55+2.7121
32Daniel & Claire Kubler+1.35±0.72+2.6321
33David & Nicola Barron+1.35±0.51+2.2829
34Jack Channon+1.32±0.58+2.1532
35Michael Bell+1.32±0.75+2.4224
36Archie Watson+1.30±0.52+1.7066
37Mrs J. Harrington, Ireland+1.27±0.56+2.6718
38James Ferguson+1.24±0.66+2.2824
39Gavin Patrick Cromwell, Ireland+1.20±0.48+2.4020
40Ed Walker+1.10±0.52+2.5815
41Michael Appleby+0.95±0.39+1.06173
42William Muir & Chris Grassick+0.95±0.50+1.7723
43Daniel James Murphy, Ireland+0.84±0.72+1.6421
44G. O'Leary, Ireland+0.82±0.63+1.3432
45James Owen+0.82±0.83+1.0863
46Jack W. Davison, Ireland+0.82±0.62+1.6819
47Grant Tuer+0.80±0.53+1.3032
48Adrian Murray, Ireland+0.78±0.63+1.5022
49Roger Fell & Sean Murray+0.76±0.55+1.3128
50Dylan Cunha+0.75±0.51+1.3226
51Jim Goldie+0.72±0.47+1.0741
52Richard Fahey+0.68±0.67+1.2026
53Harriet Bethell+0.62±0.74+1.3617
54Seb Spencer+0.62±0.63+1.3517
55Julie Camacho+0.60±0.51+0.7963
56Joseph Parr+0.57±0.45+0.9927
57Brian Ellison+0.56±1.04+1.0324
58Adrian McGuinness, Ireland+0.55±0.47+0.7266
59Jane Chapple-Hyam+0.55±0.65+0.8043
60Gemma Tutty+0.49±0.48+0.7538
61George Baker+0.44±0.50+0.7133
62Declan Carroll+0.44±0.59+0.8024
63Roger Teal+0.43±0.39+0.9616
64Ian Williams+0.36±0.56+0.4762
65Michael Dods+0.35±0.38+0.5535
66Dr Richard Newland & Jamie Insole+0.33±0.64+0.6124
67Stephen Hanlon+0.28±0.48+0.6615
68Nigel Tinkler+0.28±0.65+0.6316
69Eve Johnson Houghton+0.25±0.56+0.5318
70Ciaran Murphy, Ireland+0.20±0.62+0.3823
71Chris Dwyer+0.20±0.50+0.4615
72Dominic Ffrench Davis+0.12±0.52+0.2320
73Michael Keady+0.11±0.54+0.2122
74Julia & Shelley Birkett+0.11±0.41+0.2515
75Adrian Nicholls+0.10±0.53+0.2021
76Eddie & Patrick Harty, Ireland+0.09±0.56+0.2017
77Peter Fahey, Ireland+0.08±0.56+0.1816
78Darryll Holland+0.05±0.65+0.1024
79Stuart Williams+0.05±0.38+0.0675
80Rod Millman+0.04±0.32+0.1016
81Alan King+0.01±0.42+0.0118
82Katie Scott-0.01±0.60-0.0325
83Gay Kelleway-0.02±0.45-0.0519
84Tom Ward-0.02±0.72-0.0521
85David Loughnane-0.04±0.45-0.0555
86Richard John O'Brien, Ireland-0.04±0.50-0.0920
87Lemos De Souza-0.05±0.93-0.1117
88Iain Jardine-0.11±0.59-0.1649
89Mark Loughnane-0.12±0.43-0.1571
90Michael & David Easterby-0.12±0.35-0.1748
91Linda Perratt-0.14±0.72-0.2722
92Amy Murphy, France-0.16±0.54-0.3219
93Tristan Davidson-0.19±0.32-0.4117
94Craig Lidster-0.19±0.44-0.4217
95Ed Dunlop-0.20±0.43-0.4516
96Michael Herrington-0.20±0.44-0.3430
97Sam England-0.20±0.33-0.3922
98Patrick Chamings-0.22±0.59-0.3729
99Ross O'Sullivan, Ireland-0.23±0.27-0.4719
100Kevin Ryan-0.24±0.70-0.5118
101Alan Brown-0.28±0.45-0.6515
102Alice Haynes-0.28±0.47-0.3673
103Jack Jones-0.32±0.50-0.5133
104Adrian Keatley-0.32±0.51-0.7316
105Gary & Josh Moore-0.35±0.39-0.5337
106Derek Shaw-0.37±0.42-0.8715
107Simon Hodgson-0.40±0.48-0.9116
108Ruth Carr-0.44±0.37-0.5767
109Roger Fell-0.44±0.62-0.6935
110John Mackie-0.44±0.44-0.9916
111Simon Dow-0.44±0.53-0.8422
112Leanne Breen, Ireland-0.44±0.40-0.9318
113Brett Johnson-0.49±0.35-1.1415
114Jennie Candlish-0.49±0.47-0.7637
115David Thompson-0.54±0.37-1.2515
116Ivan Furtado-0.55±0.58-0.7457
117Robyn Brisland-0.56±0.56-1.0324
118Bernard Llewellyn-0.66±0.40-1.4018
119Mark Walford-0.68±0.59-1.2026
120Michael Wigham-0.69±0.38-1.2724
121Jim Boyle-0.71±0.36-1.2526
122Jessica Macey-0.77±0.47-1.4323
123Mike Murphy & Michael Keady-0.79±0.48-1.6119
124Denis Hogan, Ireland-0.81±0.37-1.0372
125James McAuley, Ireland-0.82±0.77-1.2241
126Jamie Osborne-0.84±0.50-1.1850
127Michael Attwater-0.84±0.49-1.2047
128Gary Moore-0.91±0.55-1.6226
129Gerard Keane, Ireland-0.92±0.49-2.1515
130Paul Midgley-0.93±0.43-1.2951
131Tim Easterby-0.94±0.45-1.3743
132John McConnell, Ireland-0.95±0.41-1.2957
133T. G. McCourt, Ireland-1.03±0.46-2.3216
134R. Mike Smith-1.05±0.46-2.1619
135Ewan Whillans-1.05±0.37-2.0122
136Antony Brittain-1.11±0.41-1.8430
137Charlie Wallis-1.11±0.31-1.9327
138Paul W. Flynn, Ireland-1.12±0.43-2.2420
139Micky Hammond-1.17±0.46-2.5417
140Robert Stephens-1.21±0.50-2.4120
141Philip Kirby-1.22±0.44-1.8837
142Patrick Morris-1.24±0.50-2.1527
143John Butler-1.25±0.41-1.6464
144Tony Carroll-1.34±0.34-1.59105
145Scott Dixon-1.34±0.38-1.7565
146Kevin Frost-1.34±0.46-2.1135
147Mark Usher-1.39±0.43-3.1316
148Adam West, France-1.42±0.64-2.8520
149David Evans-1.44±0.39-1.9359
150Grace Harris-1.46±0.44-2.4131
151Mike Murphy-1.57±0.49-3.0621
152Rebecca Menzies-1.62±0.55-2.3644
153Joe Ponting-1.85±0.36-3.9018
154Alexandra Dunn-1.89±0.46-3.8819
155Adrian Wintle-1.93±0.44-2.8741
156Deborah Faulkner-1.94±0.56-4.2217
157Stella Barclay-1.96±0.65-3.5924
158Liam Bailey-2.12±0.54-3.8225
159Phil McEntee-2.82±0.49-4.2639

All 159 Flat destination yards with 15 or more arrivals, strongest first. Nothing is filtered out of this list; yards below the 15-move gate are absent because the standard error on them is too wide to name a person, not because the number was unflattering. Source: research/trainer_rerun_2026_08/leaderboard_flat.csv, 2026-08-15.

Over jumps, we cannot certify the same thing — and the reason is the finding

2 of 32
jumps yards with enough moves to test whether their effect persists out of time

The jumps list below is a description of the past that has not been shown to hold in the future. Only 2 yards of the 32 that clear the reporting gate clear it in both halves of a time split, so the persistence test that certifies the Flat list simply cannot be run here. Its placebo ratio (1.82×) is the weakest of the four lists too. We publish the numbers because withholding them would be a different kind of dishonesty, but they carry a warning the Flat numbers do not.

Basis Same spine and estimator as the Flat list. 2,610 same-code National Hunt moves, 612 destination yards. Rebuilt 2026-08-15.

Method, caveats and confidence intervals

Why this is thin and the Flat list is not. Flat has three times the moves and five times the reportable yards. The received picture — that the top of this leaderboard is jumps-heavy — was backwards: it looked that way because cross-code moves are overwhelmingly Flat-to-hurdle horses arriving at jumps yards, and the ruler bug handed every one of them a spurious +14.20 lb. Once pre and post ratings are measured on the same ruler, a Flat lot is the well-supported case and a jumps lot is the weak one.

What would change this. More seasons, or a code-specific rating column so that Flat→hurdle progressions can be recovered properly instead of discarded. Lowering the 15-move gate would not count: the gate exists because these are named people.

Chase and bumper are not published at all. Five chase yards and zero bumper yards clear the gate. A five-name list of chase trainers would be an invitation to read rank differences the standard errors do not support, so it is not here. That absence is deliberate, not an oversight.

Coverage. For a jumps lot, the origin-to-destination comparison is blank about 93% of the time — both ends clear the gate on only 7.2% of jumps moves.

Do the two codes agree? Six yards clear the gate in both, and none is a major yard. Their effects correlate r = +0.944, which sounds reassuring and is worth almost nothing on six points spanning a two pound range. Treat the lists as separate estimates: there is no evidence here that a yard's Flat effect transfers to its jumps horses or the reverse.

#Destination yard Effect (lb)± se Unshrunk (lb)Moves
1W. P. Mullins, Ireland+4.79±0.77+9.5820
2Gordon Elliott, Ireland+2.94±0.71+4.0155
3Mrs Denise Foster, Ireland+2.77±0.83+3.5372
4Gavin Patrick Cromwell, Ireland+2.38±1.10+4.2825
5Olly Murphy+2.19±0.68+3.5632
6Fergal O'Brien+1.63±0.58+2.2652
7Jonjo & A.J. O'Neill+1.48±0.55+2.1842
8John Joseph Hanlon, Ireland+1.38±0.45+3.1116
9Philip Hobbs & Johnson White+1.33±0.64+2.4524
10Harry Derham+1.33±0.57+2.4823
11Dr Richard Newland+1.19±0.82+1.9631
12James Owen+1.14±0.70+2.4118
13Dan Skelton+1.13±0.58+1.7934
14Joel Parkinson & Sue Smith+0.91±0.95+1.8220
15Nigel & Willy Twiston-Davies+0.74±0.53+1.5219
16Kim Bailey & Mat Nicholls+0.72±0.77+1.6216
17Paul Nicholls+0.62±0.68+1.3916
18Neil Mulholland+0.57±0.62+1.0325
19Henry de Bromhead, Ireland+0.54±0.43+1.2216
20Gary & Josh Moore+0.36±0.58+0.6130
21Jennie Candlish+0.32±0.55+0.6818
22David Pipe+0.23±0.76+0.4818
23Lucinda Russell & Michael Scudamore+0.14±0.48+0.2236
24Joe Tizzard+0.14±0.82+0.2139
25Donald McCain+0.06±0.77+0.1321
26Micky Hammond-0.01±0.63-0.0221
27Oliver Greenall & Josh Guerriero-0.08±0.85-0.1331
28Robbie Llewellyn-0.20±0.87-0.4416
29John McConnell, Ireland-0.22±0.58-0.4519
30Ian Patrick Donoghue, Ireland-0.60±0.86-1.2419
31Andy Irvine-0.69±0.61-1.5017
32Ben Haslam-0.77±0.65-1.5819

All 32 National Hunt destination yards with 15 or more arrivals. Read the caveat above before using this table. Source: research/trainer_rerun_2026_08/leaderboard_nh.csv, 2026-08-15.

A better jockey is worth forty times more on a live chance in a valuable race than on an outsider in a bad one

2.9% of the prize
what upgrading rider tier is worth on a live 10% chance, horse-controlled — against 7.9% if you do not condition on the horse

A rider's edge is a constant shift in log-odds, which means its cash value is the change in win probability multiplied by the prize — and both terms move together across the card. Book the good rider where the chance is live and the prize is real; the cheap rider costs you almost nothing where it is neither. That shape is the finding and it is solid. The magnitude is a range, not a number. Comparing the same horse under different riders puts the upgrade at 2.9% of the winner's prize on a live chance; comparing rider tiers without conditioning on the horse puts it at 7.9%. The first is biased low and the second is biased high, and we publish both rather than choose.

Basis GB and Ireland, 2020-01-01 to 2026-08-06. Walk-forward by season. Model is blind to jockey identity and to market price. Built 2026-08-07.

Method, caveats and confidence intervals

Why this is publishable when the ranking was not. It is a rule about race types and it names nobody. Individual riders' estimates are not stable — one well-known rider moves from +0.52 to +0.28 log-odds depending on the control, and +0.52 would imply a 10% chance becoming 15.7%, which is not credible. Everything here works with tiers rather than named riders. The document's horse-controlled tier gap is 0.2016 log-odds between the 90th and 20th percentile points; the grid above instead compares the whole top decile against the whole bottom fifth, whose group means differ by 0.3385 log-odds, and it applies no horse control at all.

The prize column was also wrong, separately and less seriously. Irish prize money was never jurisdiction-corrected, and the Irish column changes meaning on 2025-01-01 — so the blanket correction proposed elsewhere in this project is itself wrong from 2025 onward. The correction was applied downstream first, and then at source: on 2026-08-15 the research spine was rebuilt with it and every parquet regenerated. It bites on 156,660 of 739,847 rides, 21.2%, and the conversion constant settled at 1.4479. The quantities it feeds moved very little, because most of them are ratios — the tier gap moved from 0.2028 to 0.2016 and the course-specialist split-half from +0.035 to +0.032. The grid above is read from the rebuilt figures. The shape of the finding survived all of this unchanged; it is the missing horse control, not the prize column, that determines how much to trust the magnitudes.

The functional form was decided, not assumed. Each rider gets one parameter fitted on 2022–23 and used to predict 2024–26. Multiplicative beats additive beats no-jockey-term on out-of-sample log-loss (0.328061 against 0.328303 against 0.329234). The shape confirms it independently: the gap between top-30 and bottom-30 riders grows six-fold from outsiders to live chances, where additive would be flat.

How much of it is real. The rider parameter is re-fitted with horse fixed effects and the horse penalty swept from "does nothing" to "bites hard". The distributional gap moves only about 14% between the extremes, and 95.1% of rides are on horses partnered by two or more riders, so identification is broad rather than a thin subset. Split-half of the horse-controlled parameter: pearson +0.549.

The "riders matter more over jumps" claim reverses, and this page used to repeat it. Uncontrolled, the tier gap rises with the code — Flat +0.0523, Hurdle +0.0733, Chase +0.0898 — which is where "the rider matters about 44% more over jumps" came from. Refit inside each code with horse fixed effects and the controlled gap falls: Flat +0.2206, Hurdle +0.1784, Chase +0.1315 log-odds. The ordering does not merely attenuate, it inverts. Over jumps the better riders sit on a narrower set of good horses, so the uncontrolled gap absorbs more of the horse there than it does on the Flat. We carried the 44% figure until 2026-08-15; it was the uncontrolled estimator, from the same table this section already warns about. Correcting the grid's magnitude and leaving a claim derived from the same estimator standing was our error, not the research's.

Bumper cannot be answered at all, and not for want of data: 7,662 out-of-sample rides exist, but only 3 riders in Bumper are ever compared against another rider on the same horse. Without that comparison there is nothing to identify. The booking tool refuses for Bumper rather than borrowing another code's number, and so does this page.

Britain and Ireland do not differ. Class adds nothing once the horse's chance is known.

On cost. Riding fees are a flat published scale, so the fee is not an upgrade cost — you pay it to whoever rides. The only incremental cost is the rider's percentage on the extra prize, about 9%. In pure expected-value terms the upgrade never loses; the real constraints are booking friction, availability and retainers, which are not in the published scale and are not modelled. That is why the grid is framed as "is this worth the effort", with a practical floor around £500 a ride.

What prize you need to clear that floor, and this is the number the correction moves most. On the horse-controlled estimator you need about £17,500 of winner's prize on a live 10% chance and about £40,400 on a 4% outsider. On the uncontrolled grid the same thresholds read £6,300 and £14,300 — which is what this document originally published. Read against the controlled estimator, the rule of thumb is roughly a third of its published value and the prize needed to justify chasing a booking is roughly three times as high. If you are deciding whether a booking is worth the effort, use the controlled column; it is the conservative one and the one the booking tool leads with.

Understated, not overstated. The blind model absorbs past jockey contribution through Timeform ratings, so a consistently good rider is already partly priced into the horse's rating. Irish and British fee scales differ and the fee table is still a placeholder.

Do not attach the "85% was selection" figure to this grid. That number comes from the prize-money model in the ranking work below, which is a different model on a different scale from the win-probability one used here. The two findings rhyme — both are about selection surviving an inadequate control — but the magnitude does not transfer.

Read this before the table: it does not control for which horse, and it overstates the gain by roughly 2.7 times. The same document's horse-controlled estimator puts the tier gap at +0.0205 in win probability where this grid observes +0.0555. About 44% of the difference is a like-for-like error — the controlled figure compares percentile points, the grid compares whole tier groups, whose means differ by 0.3385 log-odds — and 56% is unaccounted for. The likely remainder is this project's own standing rule applied to this table: inside a chance band, top-tier riders still get the better horses. Neither figure is being declared correct. The controlled one is biased downward, because a rider's past contribution is already inside the horse's rating; the uncontrolled one is biased upward. The truth is between them, and the table below is the upper end.
Horse's chance Under £4k£4–8k £8–20k£20k+
Under 5%£174£195£411£1,958
5–10%£261£374£787£2,487
10–20%£330£498£1,010£2,694
20% or better£551£852£1,075£6,393

Uncontrolled tier comparison — expected gain per ride, in pounds, from upgrading a bottom-tier rider to a top-tier one. Columns are the winner's prize. Do not read cell against cell. The £6,393 corner rests on 494 replacement-tier rides and carries ±£2,024 — and it demonstrates its own caveat: on 2026-08-15 the research spine was rebuilt with an era-aware prize correction and that same corner moved from £7,311 to £6,393, a fall of 12.6%, without leaving its own interval. The shape across the whole table is the finding; the value of any single cell is not. Read from the booking tool's own payload rather than transcribed, so the two pages cannot drift apart. Source: jockey_tool_data.json, built 2026-08-15, 2020-01-01 to 2026-08-13, GBR + IRE, 203,875 tier rides; method in JOCKEY_ALLOCATION.md §4 as corrected by its errata and addenda R1–R2 of 2026-08-15.

This correction is the attribution standard applied to our own research, and it is worth saying so. The rule that any comparison between two groups is a comparison of what they were given until you condition on the individual was written down here after it cost 85% of a jockey prize-money effect — a different model from this grid, on a different scale — and then reversed the claimer result. It then went unapplied to this very grid for eight days, and was caught by someone rebuilding the numbers from the parquets instead of transcribing them. The method works; it only works when somebody actually points it at the thing.
Open the jockey booking tool →

A claiming rider is at worst neutral on the same horse, and probably slightly positive over jumps

+0.00427
change in win probability, claimer against full-fee rider on the same horse, across 27,942 horses

This result reversed once we conditioned on the horse. The first version of this analysis said claimers are worse everywhere. It was wrong twice over: it assumed the weight column was net of the claim when it is the allotted weight, and it compared claimers against full-fee riders unconditionally — which measures claimers getting worse mounts, not their riding. Redone within horse, the effect is +0.00427 ± 0.00194, and it is a jumps effect (+0.00852 over jumps; +0.00141 on the Flat, not distinguishable from zero).

Basis 118,554 claimed and 254,549 unclaimed rides, GB and Ireland, 2020-2026. Within-horse comparison. Corrected 2026-08-08.

Method, caveats and confidence intervals

Read the direction, not the point estimate. The within-horse design controls for which horse but not for when, and yards do book a claimer deliberately when a horse is well treated, which would inflate this. By claim size, 3 lb (+0.00698) beats 7 lb (+0.00507) — the opposite of a pure weight effect, because claim size tracks inexperience rather than pounds.

The threshold table you would expect is deliberately absent, and its absence is the finding. A natural request is a grid of the prize and chance at which a 3, 5 or 7 lb claim outweighs the gap between an average and a top rider. Both the allowance and rider ability enter as shifts in log-odds, so their difference is scale-free: it does not depend on the horse's chance or on the prize. The better booking is the same choice everywhere on the card. A threshold table would imply a crossing point the functional form does not permit. The data agrees by being uninformative — of 15 populated cells, one had an interval clear of zero, which is what chance alone produces. Filling the other fourteen would have been fabrication.

A pound is not separable from the rider here. Isolating "the 5 lb" from "the rider" needs the causal value of a pound, and that is not identifiable from this data: two strategies both returned the wrong sign, because weight is assigned from the rating and so proxies for class.

This does not weaken the grid above. Claimers concentrate in the replacement tier (0.287 of rides against 0.204 in the top tier), so that tier is flattered by the unmodelled allowance and the 0.2016 log-odds gap is an under-estimate.

Whether a yard improves its horses is a question this data cannot answer, and the table that looks like an answer is measuring which horses a yard is sent

p = 0.127
the effect once the yard is included as a model feature — it largely disappears, against p = 0.047 when it is left out

The identification failed, and that is the finding. For jockeys there is a clean control — the same horse under different riders. For trainers there is not: a horse almost never changes yard inside this window, so "same horse, different trainer" barely exists in the data. The tell is decisive. Leave the yard out of the model and yards look strongly different (p = 0.047). Put the yard into the model and the effect largely disappears — p = 0.127, and the spread between yards falls from £254 a run to £138. A number that evaporates when you let the model know which yard it is looking at was describing the yard's horses, not the yard. The underlying table is below, inside the caveat rather than beside it.

Basis GB and Ireland 2020-2026, 721,072 runs. Trainer deliberately not a model feature. Irish prize converted at the ECB rate on the meeting date. Built 2026-08-07.

Method, caveats and confidence intervals

The researcher's own verdict was "do not publish the trainer table as it stands", and this page does not. What is published as the finding is the identification failure; the table is its illustration and sits inside this disclosure, under the objection, rather than beside it where it could be read, quoted or screenshot on its own. A number quietly withheld tends to be rediscovered later by someone with less context, which is why it is here at all.

This is not a ranking of trainers. It is a ranking of the horses trainers receive, with a trainer's name on it. There is no control here for which horse ran, and the yards at the top are the yards that get sent the best horses in Europe.
YardRuns Per run above expectation Value per year Net to owner per year
Aidan O'Brien2,583£3,541£1,858,608£1,672,747
W. P. Mullins4,312£1,814£1,613,578£1,452,221
William Haggas3,031£1,067£681,549£613,394
Joseph Patrick O'Brien4,337£500£483,686£435,318
Roger Varian2,445£870£453,205£407,884

All five rows the source records — there is no longer list, and nothing has been omitted for being unflattering. Source: JOCKEY_VALUE.md §6, model B1, 2026-08-07.

What is nevertheless real about it. The machinery does what it claims: permutation p = 0.047, split-half spearman +0.402, noticeably more persistent than the equivalent jockey signal. The figure beside Aidan O'Brien is not evidence that he adds £3,541 a run over 2,583 runs; it is mostly evidence that Coolmore sends him the best horses in Europe.

What would make it trustworthy. The switcher design — measure a horse before and after it changes yard against a mean-reversion baseline. That is exactly what the destination-yard work in this section does, and it is why that finding is presented as the usable one and this table is not. The two are not alternative views of the same thing: one has an identification strategy and one does not.

Currency. Irish prize money is euro-denominated and is converted at the European Central Bank rate on the meeting date, over a window where the rate ran 1.0756 to 1.2138. All figures are sterling.

Section 3

What we checked, and it isn't there

Every one of these could have gone the other way, and several of them looked like they had before we controlled properly. They are here at the same size as the positive findings because they are worth the same.

We cannot tell most jockeys apart, and a jockey ranking is not something we will publish

357 of 459
riders whose effect cannot be distinguished from an average rider, once you compare them on the same horse

The uncontrolled version of this analysis looked convincing: a clear £709 per ride separating riders, p = 0.000, and an ordering that matched expert consensus. About 85% of it was selection and model miscalibration rather than riding. Riders who sit on good horses in big-money races bank a large positive residual with no riding involved. Once you compare only the same horse in races of similar value under different riders, the effect shrinks by roughly an order of magnitude and 357 of 459 riders can no longer be told apart from average — 34 clear zero above and 68 below, against about 11 each expected by chance.

Basis GB and Ireland, 2020-01-01 to 2026-08-13. Cluster bootstrap over horses, riders with 200+ rides. Re-based 2026-08-15 on the rebuilt spine — see the threshold note below.

Method, caveats and confidence intervals

How the artefact was caught. The rider's mean residual correlated +0.887 with the rider's mean predicted value — the control had plainly failed. The underlying model under-predicted, and not by a constant: the shortfall ran from +£84 per ride in the cheapest decile to +£1,562 in the dearest. This even survives a naive within-horse test, because a stable's best rider takes the horse to its valuable engagements while others partner it in handicaps.

What survives, on the rebuilt control sweep. Recentring within predicted-value band, season, country and code, then comparing the same horse in races of similar value: the permutation p runs 0.000 recalibrated, 0.458 once you add the same horse, and 0.500 once you add the race value too — that is, by the time both controls are on, a permutation test can no longer distinguish the spread of rider effects from what shuffling the labels produces. Read that as a weak test failing to detect, not as proof there is nothing to detect. The permutation measures the overall spread and is low-powered against heavy tails, and heavy tails at low ride counts are exactly what breaks the variance decomposition described below — one cause plausibly explains both. It is also why the two tightest specifications swap places between data vintages: that is instability in a low-powered test, not a finding about which specification is right. Split-half agreement on that tightest money-scale specification is +0.129, and the between-rider spread falls to £522.6 a run. Shrinkage is severe throughout: at the current threshold a rider needs about 1,616 rides before even half of their raw estimate is retained. There is still more here than noise — individual bootstrap intervals separate about a hundred riders — but it is not a ranking and it will not be presented as one.

One row of that sweep is deliberately blank, and the reason is instructive. The uncontrolled population figure has no rebuilt value, because the code that produces the sweep only reports on the recalibrated residual. An earlier draft filled the gap with +0.346 — which is a different model's recalibrated split-half, a different quantity entirely. It was caught before publication. We record it because a blank cell is honest and a plausible number in a blank cell is not, and because this page's own worst error was the same shape.

The estimator behind the earlier version of this finding turned out to be unstable, and that is the strongest thing on this page against naming riders. Until 2026-08-15 we published "452 of 576 riders cannot be told apart", measured with a 100-ride floor. On the rebuilt spine that threshold returns a negative between-rider variance (-56,855) — meaning no signal at all, and no estimates that can legitimately exist. The cause is not the prize correction, which moves the figure about 10% and does not move its sign; it is heavy tails at low ride counts. 77 riders with 100–150 rides carry a within-rider variance around 119 million — one large prize on a rare good mount — and those terms alone swamp the noise estimate for the whole population. Seven more days of data flipped it.

What that would have produced if nobody had guarded it. Not an error and not an empty table: a full table of large, plausible, sign-flipped figures against named professionals. The estimator now refuses and prints why. This is the clearest illustration we have of why the ranking is not published — the failure mode is not "the numbers look wrong", it is "the numbers look right".

The current basis. A 200-ride floor, which is where the estimator is stable across thresholds and across both prize columns: 459 riders, 34 clear of zero above and 68 below. The older "38 above, 86 below, of 576" cannot be reproduced at the threshold that produced it, and the two sets of counts are not comparable — so they are not compared here, and no trend should be read between them.

The names, however, do compare — and they say the instability is all at the tail. Counts measured at two different thresholds cannot be set against each other, but asking which riders appear on both lists is well defined either way. 27 of the 38 survive; 11 drop and 7 are new. Of the 11 that drop, 4 fall below the new ride floor mechanically — all of them worth £20–30 a ride, which is to say nothing — and 7 are still in the data with an interval that no longer clears zero. Every one of the drops is from the bottom of the list. The head does not move: the seven largest estimates all remain clear of zero at similar magnitude, the top one going from £1,062 to £912 a ride.

So the honest reading is that the head is real and the tail is noise-dominated, rather than that the whole thing is mush. That is a narrower claim than "these estimates do not survive", and it is the true one. It is also why the conclusion does not soften: the marginal entries were never findings, and they are precisely the ones that evaporate when the estimator is put on firmer ground. A list whose bottom half rearranges itself between data vintages is not something to publish as a ranking, even when its top half holds. No rider is named on this page in either direction.

Two things we will not paper over. The permutation test and the bootstrap disagree, and on the rebuilt spine they disagree more: on the money scale the permutation p moved from 0.050 to 0.500, which is squarely at chance, while the individual bootstrap intervals still separate about a hundred riders. The permutation tests the overall spread and is low-powered against heavy tails; the bootstrap tests individual riders. We are not going to pretend that is resolved. And the surviving ordering matching expert consensus is reassuring rather than evidence — exactly the kind of face validity that makes a confounded result feel true.

A separate question with a cleaner answer. Asked whether a rider beats their own market price, the answer is that no rider persistently does. That is not a weak effect, it is no effect: the split-half agreement between an early and a late period is -0.002 across 449 riders — indistinguishable from zero. It is the expected result, because the Betfair price already prices the jockey.

This page said "the split-half is negative" until 2026-08-15, and that was wrong. The published figure was −0.243 and was read as riders being anti-predictive of their own market, with a regression-to-the-mean story built on top. Re-measured with the same estimator, it is -0.002. The cause was inside the script that produced it: a blanket "multiply Irish rows by 0.600" — the same era-blind correction this project's own errata condemns — still live there, injecting a distortion correlated with both jurisdiction and period, which is exactly the shape that manufactures early-versus-late structure. The conclusion is unchanged and cleaner; the anti-predictive reading is withdrawn. One residual is real and unexplained: British riders split-half at -0.159 against Irish riders at +0.273.

Why this reached the page at all is worth recording. The figure was never quoted here as a number, only as the phrase "the split-half is negative" — so a search of this page for the published value found nothing, and the claim survived a check designed to catch exactly this. A story can carry a withdrawn finding after its number has been dropped.

The standing rule this produced. Any comparison between two groups of riders is a comparison of their mounts until you condition on the horse. It cost 85% of the raw prize-money effect measured here, and then it reversed the sign of the claimer result. The size of that correction does not carry across models: on the win-probability work behind the booking grid above, adding a horse control moves the spread by about 11 to 14%, not by 85%. What generalises is the direction, not the magnitude — assume the rule applies to every new rider comparison until shown otherwise, and measure the cost separately each time.

Course specialists do not exist in this data

377 of 7,560
rider-and-course combinations that look real — against 378 expected from chance alone

Controlled twice over: the underlying model already prices the horse, and each rider's own overall effect is netted off, so what is left measures "better at this track than this rider is generally", not "wins a lot here". 377 of 7,560 cells clear zero, which is 4.99% against the 5.0% chance alone produces. And a course edge in 2022–23 tells you nothing about 2024–26: split-half agreement is +0.032.

Basis Every rider-course pair with 15 or more rides, GB and Ireland 2020-2026. Median cell 36 rides. Built 2026-08-07; figures recomputed from course_specialists.parquet each time this page is built.

Method, caveats and confidence intervals

The one thing that does look odd, and its answer. Of the 377 cells that clear zero, 319 are positive and only 58 negative. Pure noise would split them about evenly, so that asymmetry is real — but it is not a course specialism. It is the residual yard effect at riders' home tracks: a rider attached to a stable rides that stable's horses most often at the track nearest it, and whatever the model under-prices about those horses lands on that cell. The null being tested is whether the rate of significant cells exceeds chance, and at 4.99% against 5.0% it does not.

The tempting list is the noise. The strongest-looking cells are exactly the ones that would get published — a top rider at his home Irish track on 321 rides, another at Punchestown on 211, a third at Newmarket. The underlying data is retained so the null stays checkable, not so the table can be used.

Why we bothered. "Course specialist" is one of the oldest claims in the sport and it is cheap to appear to confirm: rank riders by strike rate at a track and the list looks authoritative. The test that matters is whether the edge repeats, and it does not. Split-half pearson +0.032, spearman +0.014.

These counts are not transcribed. They are recounted from course_specialists.parquet every time this page is built, and the build fails if they stop matching what the write-up claims. A number quoted from a sentence can be contradicted by a correction appended beneath that sentence; a number recomputed from the data cannot.

Which consignor sold the horse tells you nothing about how good the horse is

r = -0.05
split-half agreement of a consignor's apparent quality — a coin flip would score zero, and this scores slightly worse

Split a consignor's own lots in half at random, 200 times, and score each half on how its horses went on to perform against the market. If consignor quality were a real, stable property, the two halves would agree. They do not. The null was recorded twice: r = -0.05 and r = -0.024. Added to the price model as a feature — price level, clearance rate and lot count, all as of the sale date — consignor identity moves the ranking by +0.0013 at p = 0.84, which is nothing at all. One consignor property does survive, and it is the only one published: clearance rate, r = +0.43.

Basis Catalogue exports covering 36,348 catalogued rows across 32 sales. Two independent methods. 2026-08-10 and 2026-08-11.

Method, caveats and confidence intervals

This is the null that most nearly became a product. Consignor reputation is exactly the kind of feature a sales page is expected to have, it is cheap to compute, and the first version of it looked excellent. See the R-squared trap below, which is the same investigation.

Why clearance survives when quality does not. Clearance rate is a fact about how a consignor operates — what they enter, what reserve they set, whether they let a horse go. It is a property of the business. "Quality" is a property of the horses, and the horses are drawn from wherever the breeder happens to be. Only the first of those is stable enough to reproduce on a split half.

On the two recorded figures. -0.05 and -0.024 are two separate recordings of the same reliability test, written down by different people in different documents. They are not two different findings and they are not averaged here into a third number that neither document contains. Either way the answer is the same: a split-half correlation at or slightly below zero, where a real property would be clearly positive.

Buyers too. The same test on buyer identity gives +0.08 (also recorded as +0.079). No buyer quality ranking either.

No sample size is attached to the clearance figure in any document that records it. It is published as the bare correlation it was recorded as, and a denominator has not been invented for it.

A consignor feature scored 0.206 on our headline metric, and it was measuring which sale the horse was at

+0.206
R-squared, withdrawn — the number that came closest to reaching the page without being true

A consignor feature read +0.206 R², which would have been by some distance the largest single improvement anyone had found. It survived until the second-stage residual was demeaned within each prior sale — at which point it went to +0.0013, p = 0.84. A consignor nearly identifies which sale a lot is at, and lots at the same sale make similar money, so the feature was reading off the sale and being credited for the horse.

Basis Within-sale R-squared, one fold per sale, 11 catalogues. Investigated and withdrawn 2026-08-11.

Method, caveats and confidence intervals

This is here as a worked example, not as an anecdote. It is the reason nulls are published on this page at all. The failure mode is general: a categorical feature with many levels will silently encode any grouping that correlates with the levels, and the more levels it has the better it will score before you check. Nothing about the 0.206 looked wrong. It was found because someone demeaned within the obvious confounder and watched the number collapse, which is a check that has to be run deliberately — it does not announce itself.

The related finding it produced. The price model turned out not to know which sale it was pricing at all. Within-sale R-squared was negative on all eight runnings of the biggest yearling catalogue — worse than simply quoting that sale's own average — with the median lot 37 to 65% too cheap. The error ran both ways: one catalogue was priced at £22,559 against a sale whose median lot makes £2,625, another at £36,413 against £8,748. Adding an explicit sale-series level moved within-sale R-squared from −0.406 to +0.216 and mean absolute median-price error from 48% to 14%. Those are the figures recorded in the repository and they are the ones to cite. A second measurement of the same fix, taken during the working session that found it, records it as 0.114 → 0.596 — a larger gain on a different cut. That harness was never committed, so it is noted here for completeness and not used: where the two disagree, the repository figure stands.

And what that fix does not cover. A sale being run for the first time gets nothing from it — the Irish August store sale has no history in this population, so it is unchanged and carries a warning on the page. The fix corrects the level of a sale and barely moves the ordering within one.

One catalogue has a large, perfectly consistent lot-position effect, and we are not shipping it

+0.1185
gain in ranking accuracy from knowing a lot's position, in one sale series — 8 of 8 runnings, p = 0.000

Where a horse sits in the catalogue predicts nothing in three sale families and predicts a great deal in the fourth. +0.1185 in October Book 3, in every one of eight runnings, at p = 0.000 — against effects near zero and slightly negative in Book 1, Book 2 and Somerville. An effect that large in one series and absent in its immediate neighbours is as likely to be a structural property of how that catalogue is ordered — session grouping, vendor blocks, a supplementary section at the end — as anything about the horses.

Basis 27 settled catalogues decomposed by sale family, 2026-08-11. Compare against −0.008 and −0.009 in the neighbouring series.

Method, caveats and confidence intervals

Why this is in the nulls section when the number is real. Because "we found a large effect and cannot explain it" is a finding about our own understanding, not about the market. Shipping it would put a strong signal on the page whose mechanism nobody has established, and the most likely mechanism — that Book 3 is ordered by something that correlates with quality — would make it a restatement of the catalogue rather than a discovery about it.

What would unblock it. Establishing how Book 3 is actually ordered. That is a question about the catalogue and answerable by asking, not a modelling question.

The history is worth recording. Lot position was first rejected on 11 sales (rank correlation −0.019, p = 0.0010, 0 of 11 sales improved), then a 27-sale test reversed the rejection, and then decomposing by sale family reversed the reversal for every family except this one. Three passes, three different answers, and only the decomposition is trustworthy — the pooled results were averaging one strong effect against three null ones.

And it does not generalise. The original rejection was on Newmarket Flat yearling catalogues, where lots are ordered alphabetically by dam and ring position carries little about quality. An earlier test on a differently ordered sale found a gain. Lot position is not shown to be absent everywhere — it is shown to depend entirely on how the particular catalogue was built.

Sale series Change in rank accuracy Runnings improvedp
October Book 1−0.00832 of 80.073
October Book 2−0.00895 of 80.260
October Book 3+0.11858 of 80.000
Somerville−0.00382 of 30.517

Source: ASKS.md A5 addendum, 2026-08-11. None of these four is shipped.

Section 4

How honest the price model is

The model that produces predicted prices on the sales pages, described by where it fails rather than by where it works. If you are using those numbers, this is the section that tells you when not to.

The price model is least accurate on the cheapest quarter of the catalogue, which is where most buying happens

90.2%
median error on the cheapest quarter, against 31.5% at its best — 457 held-out lots

Overall the model is respectable: median absolute error 42.3%, rank correlation 0.738 [0.691, 0.782], and the typical lot makes 0.985 times what was predicted, so there is no level bias. That average hides the thing you need to know. Broken down by price band, the error on the cheapest quarter is more than double the error anywhere else. This is a tile on the sales page rather than a footnote, because the cheap end is where most people are actually bidding.

Basis Settled sale, sold lots only. Vendor buybacks, RNAs and withdrawn lots excluded and checked against the catalogue export rather than the feed's claim. 2026-08-11.

Method, caveats and confidence intervals

What is excluded and why. 93 vendor buybacks, 49 lots that failed to reach their reserve and 506 withdrawn lots are not scored, because a buyback price is not a market price. Which lots those were was taken from the catalogue export rather than from the feed's own status flag.

Rank against price. A typical lot moves 12.6 percentile points from its predicted position; 43.1% land within 10 points and 67.0% within 20. On the one online sale scored so far the error was 45.6% with a rank correlation of 0.646 [0.396, 0.809] on 42 genuine sales — a wide interval on a small sample, quoted as such.

The published sale median is a median. 49.2% of realised prices fall below the per-lot median we publish. The sale-level median we published was £20,205 against £18,900 realised, +6.9% — and those scored lots sit 20.0% above the median of all 997 sold lots, which verifies rather than undermines the page's standing caveat that selection is worth 15 to 25%.

Price band Median error 
Cheapest quarter90.2%Worst, and this is where most buying happens
Second quarter39.5%Roughly the page-wide average
Third quarter31.5%Best
Dearest quarter44.7%Compressed at the top — see below

Median absolute percentage error by quarter of the catalogue, 457 held-out sold lots. Source: ASKS.md A2, 2026-08-11.

Session record — not reproducible from the repo

The model ranks best at the two ends of the market and worst in the middle — and in one band it has been worse than guessing

0.26 / 0.11
rank correlation in the dearest quarter against the third quarter — the middle of the market is where it is weakest

Sorting a catalogue is what this model is actually for, and it does not do it evenly. Rank correlation runs 0.26 in the dearest quarter and 0.18 in the cheapest, against 0.11 and 0.07 in the two middle quarters — a U-shape, strongest at the extremes. In the third quarter it was worse than chance on 5 of the 12 sales. The ends of a catalogue are easy: a stakes-bred colt and a plain gelding differ on everything the model can see. The middle is where the horses look alike, and it is where an ordering would be worth most.

Basis 12 held-out sales, bands cut on predicted price, shuffled null. Session record 2026-08-12; reproduction queued.

Provenance. Measured 2026-08-12 across 12 held-out sales, on prediction-cut price bands, against a shuffled null. The harness was never committed, so unlike every other figure on this page these cannot be reproduced from the repository and no document on it corroborates them. They are published because they were really measured; a reproduction is queued, and until it lands they carry less weight than the figures around them.

Method, caveats and confidence intervals

The bands are cut on predicted price, and that is load-bearing. The same work established that bands cut on the realised price make the within-band correlation negative by construction — which is why an earlier by-band accuracy display was dropped rather than fixed. If someone re-runs this on realised-price bands and reports that the U-shape does not hold, they have measured a different thing. Prediction-cut bands are also the only ones available before the hammer falls, and so the only ones a buyer could act on.

What "worse than chance" means here. Negative rank correlation within that band on those sales: the model's ordering of those lots was, on average, slightly inverted against what they made. Not a large effect, and on a minority of sales — but it is the opposite of what the column claims to do, and it happens in the part of the catalogue with the most lots in it.

Why this one is flagged differently from everything else here. Every other figure on this page is traceable to a document in the repository and is checked against it character-for-character each time the page is built. These are not: the session that measured them never committed its harness. Dropping them would punish a real finding for a process failure; stating them unmarked would claim a provenance they do not have. So they are published, marked, and queued for reproduction.

The model has never predicted a million-pound horse, and it will not

£185,991
highest price the model predicted across 3,800 held-out lots, against an actual top price of £997,500

The range is compressed at the top. Actual prices exceed £150,000 for 2.05% of lots; the model predicts above £150,000 for 0.03% — one lot in 3,800. It retains 56% of the spread in log price (actual standard deviation 1.130, predicted 0.631). If you are looking at the very top of a catalogue, the predicted price is a floor on what the lot may make rather than an estimate of it. Across the whole range the slope of actual on predicted is 0.976 (se 0.013), close enough to one that no output rescaling has been applied — there is no single slope to force.

Basis 3,800 held-out lots across 12 sales for the range; 3,910 sold lots for the slope. Live model output confirms the ceiling separately. 2026-08-11.

Method, caveats and confidence intervals

Two independent measurements of the ceiling. The held-out analysis above puts the highest prediction at £185,991; checking the live model output separately puts it at £190,906. Different populations, same conclusion: the model has never priced a lot near the top of this market and will not.

A near-miss worth keeping, because it nearly went on the page. Binning the same lots by actual price produced something that looked damning: the model 2.9 times too high on the cheapest fifth and half too low on the dearest, which was almost written up as a warning on every value column. It is wrong, and wrong in a way this project's own guardrails warn about: binning on the outcome measures regression to the mean, not bias. Conditioned on predicted price — the only thing available before the hammer falls — the ratios are 0.83 / 1.00 / 0.95 / 1.08 / 0.92, with 80.0% interval coverage overall and 82.6% in the top predicted 5%. This distinction decides the next finding too: any band cut on the realised price is negative by construction, so a result stated on realised-price bands cannot be compared with one stated on predicted-price bands.

A low spread ratio is not by itself a defect. A conditional median must have less spread than the outcome it predicts — it is the same fact as the R-squared wearing a different hat. What makes the compression here worth stating is the specific consequence at the top: 0.03% of predictions above £150,000 against 2.05% of lots actually selling there.

The intervals are wide rather than overconfident. Across all lots together, the published 80% interval actually contains 86.0% of realised prices (7.2% fall below it, 6.8% above). That is miscalibrated, and in the safe direction. Narrowing it is a model change and was deliberately not smuggled in alongside a display change.

Two coverage numbers, and they are not in conflict. The 86.0% above is the unconditional figure — every lot pooled. The 80.0% quoted in the paragraph above, and on the sales page, is coverage conditioned on predicted price, computed within predicted-price bands. They answer different questions and both are correct: pooled, the interval is wider than advertised; band by band, it is close to nominal. If you have seen one of these numbers elsewhere on the site and the other here, that is why.

What the compression test does not cover, and it is the whole point. Every settled catalogue we hold is from the top or upper middle of the market. The page also prices the lower yearling books, the Irish September sales and the Irish August store sale, and we have no settled catalogue for any of them. So the bottom of the market is absent from the test set entirely. The suspicion that the model behaves differently down there is untested, not refuted, and the next step is fetching those historical runnings rather than changing the model.

Pooling hides real disagreement. The overall slope of 0.976 is an average of sale series pulling in opposite directions — one runs 0.66 to 1.07 with six of eight runnings below one, another runs 1.16 to 1.50 with all above. Retraining with the unsold lots as bounded observations moves the slope by +0.005 against a standard error of 0.013, so censoring does not explain it.