Accuracy
The real, current evaluation of the Open Lifts open-lift-count model — every split, the baseline it's measured against, and exactly how much evidence it's based on. This page is generated directly from the same evaluation artifacts the model itself was registered with — nothing here is written by hand.
Model reconstruction-hgb-v1-20260918 is currently labelled EXPERIMENTAL. The PRD requires ≥75.0% range hit rate, beating a plain baseline, on every held-out split below before a model may be called validated. This one has not cleared that bar yet — every prediction it produces on this site is marked experimental, not a booking-grade promise.
model held-out-date coverage 0.743 does not beat the resort-historical-mean baseline's 0.788 — PRD §2.9.1 requires both; held-out-date coverage 0.7428571428571429 < 0.75; held-out-resort evidence insufficient: only 3 eligible resorts (need >=10 per Phase 3E Sec 14's own criterion for a trustworthy estimate); held-out-season evidence insufficient: only 79 training rows (Phase 3E Sec 14 flags 300+ as the scale needed to trust this split); held-out-season coverage 0.5263157894736842 < 0.75
The three required splits
A range “hit” means the actual observed open-lift count fell inside the predicted range. Each split answers a different question about whether the model actually generalises, not just whether it fits data it has already seen.
A date the model never trained on, same resorts.
n = 35 n
A resort the model never trained on at all.
n = 3 resorts evaluated
A season the model never trained on at all.
n = 38 n
The baseline comparison
A baseline is the simplest reasonable guess with no model at all — here, a resort's own historical average for that time of year. A real model has to beat it, not just clear 75%, or it isn't adding anything.
This model currently does not beat the baseline on this split — the honest reading is that a simple historical average is at least as good a guess right now.
Baseline evaluated on 33 of 35 held-out test days (the rest had no historical data for that resort to average).
Why the gate isn't just a percentage
A coverage number on too little evidence isn't trustworthy on its own. Here's exactly how much real evidence this evaluation is based on:
Held-out-resort evaluation used: cauterets, galibier-thabor, valloire.
Live forecast verification
Separately from the training-time backtest above, every real forecast this model makes is checked against what actually happened once the date passes — the ongoing, real accuracy record.
Model reconstruction-hgb-v1-20260918, trained on data up to 16 June 2026, registered 18 September 2026. See how it works for what observed/reconstructed/forecast mean, or /data for source attribution.