A rider steps off the trainer after a 20-minute FTP test, fresh, fed, and rested, and that number becomes the one that runs the whole season — zones, targets, race-day pacing, all built off legs that had done nothing hard in the two hours before the test. Then the race happens, and by hour four, after three attacks, two feed-zone scrambles and a headwind section nobody warned them about, that same rider can't hold a number anywhere near what the test said they could. Their training partner, who tested a touch lower that same Tuesday morning, is still there. The test wasn't wrong. It just measured the wrong thing.
This is a genuinely uncomfortable gap for anyone who's built a season around a single fresh-legs number. FTP tests are convenient precisely because they're short, controlled, and repeatable — twenty minutes, a trainer, done. But convenient and decisive aren't the same thing. If the number that predicts who's still riding hard in the finale isn't the number a twenty-minute test produces, then a lot of training plans are optimising for the wrong Tuesday.
Section 01What the data actually shows
Maunder and colleagues gave this gap a name: durability — "the time of onset and magnitude of deterioration in physiological-profiling characteristics… during prolonged exercise" (Maunder et al. 2021). Their argument is direct: the thresholds, the efficiency numbers, and the power-duration relationship a lab test or an app captures on fresh legs all drift once a rider is deep into a long effort — and the review's central case is that conventional fresh-state profiling doesn't capture that drift at all, because it never asks a rider's legs to do it (Maunder et al. 2021). Two riders can post identical 20-minute power numbers in a Tuesday-morning test and still split apart by hour four of a Saturday race — the fresh test simply never measured the trait that decided the gap.
Spragg, Leo, and Swart tracked this directly in 30 under-23 professional cyclists across a full competitive season, comparing each rider's fresh power profile against their fatigued power profile — power output measured after they'd already burned through a meaningful amount of work that day (Spragg et al. 2023). Absolute 5-minute fatigued power, absolute 12-minute fatigued power, and relative 12-minute fatigued power were all significantly lower late in the season than early or mid-season (Spragg et al. 2023). More telling: the fatigued power profile varied more across the season than the fresh power profile did (Spragg et al. 2023). Fresh-state fitness held roughly steady across the group as the months passed. Durability moved.
Time below the first ventilatory threshold — genuinely easy riding — correlated with improvement in 2-minute fatigued power over the season (r = 0.43 for absolute power, r = 0.376 for relative power) (Spragg et al. 2023). A shift toward a more polarized training distribution — more truly easy, more truly hard, less in between — correlated with improved 12-minute fatigued power (r = 0.41) (Spragg et al. 2023). Neither of those training variables tracked meaningfully with fresh, rested power. They tracked with the fatigued number.
Section 02Why fresh legs don't tell the whole story
The mechanism is straightforward once you separate the two things a fitness test actually measures. A fresh-legs FTP or 20-minute test captures peak capacity under ideal conditions — full glycogen, no accumulated central fatigue, no thermoregulatory drift, no repeated neuromuscular recruitment stacked on top of itself. None of that is present on race day after three hours of racing. Durability sits on top of peak capacity as a separate trait: how well a rider's threshold, efficiency, and power-duration curve hold their shape once glycogen is down, core temperature is up, and the muscle has already done thousands of kilojoules of work (Maunder et al. 2021).
That mechanism is also why the finding splits by rider type rather than producing one universal "durability score." A climber's durability shows up in whether their 20-minute power is still there after the workload of a mountain stage; a sprinter's shows up in whether their sprint power is still there after 180 kilometres in the wheels. Durability isn't a single number any more than "power" is a single number — it's the shape of the decline curve for the duration that actually decides how that rider wins. Separate field analyses of professional race data have reached the same conclusion from the other direction: it's the ability to hold power deep into a race, not the peak number on a fresh test, that tracks with who actually finishes well — support for treating durability as its own trainable trait rather than a footnote to FTP.
None of this makes the fresh-legs test useless. It's still the fastest way to set a rough zone structure and to track raw ceiling fitness over a training block. The mistake is treating it as the whole picture — using a number measured in the first twenty minutes of a session to predict a performance that, for most road and gravel racing, actually needs to hold up for three, four, five hours past that point.
Section 03The protocol
- Test fatigued, not just freshDon't rely on a single rested FTP test to set your zones. After 2–3 hours of steady endurance riding, add a 5- or 12-minute effort at the tail end and compare it to your fresh number. That gap is your durability baseline — track it alongside FTP, not instead of it.
- Build the volume that builds durabilityMore time spent below first ventilatory threshold tracked with better fatigued-power retention across the season (Spragg et al. 2023). Durability is built on the easy volume underneath the block, not on stacking more threshold intervals onto an already-thin base.
- Shift the mix toward genuinely polarizedWithin riders, a shift toward a more polarized training distribution — more true-easy, more true-hard, less moderate "grey zone" — correlated with improvement in 12-minute fatigued power across the season (Spragg et al. 2023). It's a correlation from one cohort, not a guarantee, but it points the same direction as the easy-volume finding above.
- Know your durationIf your race gets decided by sustained climbing, durability-test your 12- and 20-minute power after a long day in the legs. If it gets decided by a sprint, test your short-power numbers under the same fatigue. The metric that matters is the one tied to how you actually finish a race.
- Retest across the season, not onceA single durability test in March tells you where you started. Durability moved more across the season than fresh power did in the tracked cohort (Spragg et al. 2023) — repeat the fatigued test every few weeks so your training plan is reacting to where you actually are, not where you were.
Section 04Where riders get this wrong
Assuming a bigger FTP means better durability
The two aren't the same trait. Maunder's review argues fresh-state profiling doesn't capture how a rider's numbers hold up under fatigue at all (Maunder et al. 2021) — a rider with a lower FTP but a flatter decline curve can plausibly out-last a rider with a higher FTP and a steep one.
Chasing one universal 'durability score'
The durability that matters is tied to the effort duration that decides your race, not a generic number lifted from someone else's result. A climber and a sprinter are tracking two different curves entirely.
Skipping the easy volume that builds it
Durability improvements correlated with time spent below first ventilatory threshold (Spragg et al. 2023). Time-crunched training that trims the easy hours to make room for more threshold work is trimming the exact volume the data ties to building the trait it's trying to protect.
Testing durability once and calling it done
A single fatigued-power test is a snapshot, not a trend. The whole point of the season-long data is that durability moves — track it the way you'd track FTP, on a repeating schedule, not as a one-off box to tick.
Section 05Applying it with ULTRA
Your season · Adapting daily
Durability, tracked over the season
Marco reconciles your plan against how you actually recovered and rode, every day — so the long, easy volume that builds durability doesn't quietly get replaced by more threshold work when a hard week runs long.
See your training planSection 06Glossary
Glossary · Terms in this article
The terms that matter.
Durability —
The time of onset and magnitude of deterioration in physiological characteristics — thresholds, efficiency, power-duration curve — during prolonged exercise.
Fatigued power profile MMPfatigue
Maximal mean power measured after a meaningful amount of prior work has already been done, as opposed to a fresh, rested test.
Polarized training —
A training distribution weighted toward genuinely easy and genuinely hard sessions, with less time spent in moderate-intensity "grey zone" effort.
First ventilatory threshold VT1
The intensity at which breathing pattern first shifts under rising effort — commonly used as the ceiling for "easy" endurance riding.
Central fatigue —
Fatigue originating in the nervous system's drive to the muscle, distinct from local muscular fatigue — one of several mechanisms thought to contribute to durability loss over a long ride.
Power-duration relationship —
The mathematical curve describing the trade-off between how much power a rider can produce and for how long — one of the fresh-state metrics durability research tracks for drift under fatigue.
Section 07Bottom line
FTP, VO2max, and every other fresh-legs number describe what a rider can do when nothing has happened to them yet. Durability is the separate trait that describes what's left once something has — the timing and size of deterioration in those fresh-state numbers during prolonged exercise (Maunder et al. 2021). Riders and training plans that only test and chase the fresh number are optimising for a scenario that doesn't exist in a real race. Easy volume and a genuinely polarized week are the two training variables this season's data linked to smaller fatigued-power losses (Spragg et al. 2023), and durability itself moves enough across a season that it's worth testing on a repeating schedule, not once and forgotten. Test yourself tired. That's the number that shows up when it counts.
Counterpoint · Read this before you rebuild your week
The other side of the evidence.
Durability is an emerging construct without a single standardised test protocol, so cross-study numbers aren't directly comparable. It's built largely by high-volume training that time-crunched riders can't always do. It complements, not replaces, fresh-state benchmarks.
Sources.
- 01Maunder et al. (2021). The importance of 'durability' in the physiological profiling of endurance athletes. Sports Medicine, 51(8), 1619–1628 DOI 10.1007/s40279-021-01459-0
- 02Spragg et al. (2023). The relationship between training characteristics and durability in professional cyclists across a competitive season. European Journal of Sport Science, 23(4), 489–498 DOI 10.1080/17461391.2022.2049886