⊕ Long-form · Cycling Science · 9 min

Critical Power isn't your FTP — and the gap quietly wrecks your intervals

Critical Power and 20-minute FTP get treated as the same threshold number.

9Read (min)
2Studies
0Protocols
2139Words
2026
Cover · art direction pending

Close on a bike computer mounted on the stem mid-interval, the rider's breath visible in cold air. A strip of tape on the top tube has one wattage number circled in pen and a second number crossed out beside it — the two competing numbers are the whole story.

Cover · Drivetrain loss · Bouillod 2022

Your training plan says 95% of CP for the next six weeks. Your head unit says FTP. You do the sensible-seeming thing — treat the two numbers as interchangeable, punch in the wattage, start the block. Three weeks in, sessions that were supposed to be sweet-spot are turning into all-out threshold efforts, your legs stop coming back between intervals, and the block gets quietly abandoned as "too aggressive for where I'm at right now."

Nothing about your fitness was wrong. The number you borrowed was.

This mix-up is common enough that it has its own research literature — and the literature says the gap between Critical Power and 20-minute FTP isn't a rounding error. It's real, it's individual, and it lands exactly in the range that turns a moderate interval into a threshold one.

Section 01What the data actually says

Karsten and colleagues put 17 trained cyclists and triathletes through both protocols in the same study (Karsten et al. 2021): a single-visit Critical Power test built from several maximal fixed-duration efforts feeding a hyperbolic power-duration model, and a standalone 20-minute time trial with FTP calculated as 95% of the 20-minute average. At the group level the two numbers looked close — CP averaged 256 ± 50 W, FTP averaged 249 ± 44 W, a 7 W difference. Group averages hide what actually happens rider by rider.

The 95% limits of agreement ran from −19 W to +33 W. That means somewhere in this study, a rider's CP undershot their own FTP by close to 20 watts, and elsewhere a rider's CP overshot FTP by more than 30 watts. Karsten's team put a number on the direction, too: a 91.7% probability that CP sat higher than FTP for any given individual — likely, not guaranteed. Despite a strong overall correlation between the two measures (r = 0.969), the authors' own conclusion was blunt: CP and FTP "generally should not be used interchangeably" (Karsten et al. 2021).

This isn't a one-lab artifact. A separate group of researchers ran the same comparison on a different cohort of highly-trained cyclists and triathletes (McGrath et al. 2021) and found an even wider gap — CP averaging 16 W above FTP (282 ± 53 W vs 266 ± 55 W), a difference that held at p < 0.001. Their conclusion matched Karsten's: in this group, too, the two values were not interchangeable. Two labs, two cohorts, two different gap sizes, one consistent verdict — CP and FTP are related, but they are not the same number, and exactly how far apart they sit is specific to you.

Translate the McGrath gap into an actual session. A rider prescribed 4 × 8 minutes at 90% of CP, who instead sets the target off their FTP number because that's what's on the head unit, is training at 90% of a number that's roughly 16 W light — about 14-15 extra watts sitting inside every single interval in this cohort. One rep, that's an uncomfortably-hard sweet-spot effort nudged toward genuine threshold work. By the third or fourth rep, the accumulated extra load looks nothing like the session the plan's author actually intended.

Section 02Why the two tests disagree

The disagreement isn't a flaw in either test. It's that CP and FTP are built from structurally different efforts. Critical Power comes from a handful of short, maximal, fixed-duration trials — commonly something like a 12-minute, a 7-minute, and a 3-minute all-out effort — fed into a model that estimates the power asymptote your body could theoretically sustain without depleting your finite work capacity above that line (your W′). Pacing barely enters into it: you're told the duration, you go as hard as you can hold for exactly that long, and the model extracts the asymptote from the curve.

FTP is a completely different kind of task. A 20-minute time trial is a pacing problem, not a maximal-effort problem. You have to hold back enough in the first five minutes to survive the last five, manage motivation and discomfort over a longer window, and the number you land on reflects pacing skill and tolerance for sustained discomfort almost as much as it reflects raw physiology. Scaling the 20-minute average by 95% is a practical adjustment — an attempt to estimate what you could hold for roughly 60 minutes from a shorter proxy — not a physiological correction.

Put the two side by side and the gap makes sense. CP is a modelled output from short, severe, minimally-paced efforts. FTP is a lightly-adjusted proxy from a single sustained, self-paced trial. They correlate strongly because they're both estimating the same rough neighbourhood of intensity — but they're not the identical physiological line, and how far apart they land for you depends on how you handle a 20-minute pacing task relative to how you handle short maximal ones. A rider who's a poor pacer but has huge top-end capacity will tend to under-produce FTP relative to CP. A rider who paces conservatively and starts fast tends to close the gap.

There's a second, practical reason the two models don't play nicely together in interval design. CP comes bundled with W′ — your finite store of work available above CP before you're forced to slow down. Ride above CP and W′ drains at a predictable rate; drop below CP and it recharges. That's the engine behind CP-based interval design: a coach who knows your CP and W′ can calculate exactly how many repeated surges above threshold you can absorb before you're cooked, which is genuinely useful for punchy, attack-heavy racing. FTP-based zone systems don't carry an equivalent finite-capacity term — they're just percentages of one number, with no accounting for how much "above-threshold" work you've already spent. Import a CP-based interval set and execute it off your FTP number, and you're not just off by a few watts — you're running a session designed around a work-capacity budget that was never calculated for the number you actually used.

Section 03How to actually run each test properly

Critical Power: the standard approach uses three to four maximal, fixed-duration efforts — commonly around 12, 7, and 3 minutes — each ridden completely flat-out for that exact duration, with 30-45 minutes of easy recovery between efforts, ideally split across more than one day if you can't manage full recovery within a session. The average power from each effort is fed into a two-parameter model (power vs. 1/time, or an equivalent hyperbolic fit) that extracts CP as the asymptote and W′ as the curvature term. Skip the recovery windows or shorten the durations and the model output drifts — the test only means what it's supposed to mean if the protocol is followed exactly.

FTP: the field-standard version is a single 20-minute maximal time trial, ridden as evenly as pacing allows after a proper warm-up, with FTP calculated as 95% of the 20-minute average power. Ramp tests (steadily increasing power to failure, FTP estimated as a percentage of the final one-minute power) are faster but rest on a different set of assumptions and produce numbers that aren't directly comparable to the 20-minute-TT-derived version — another reason "FTP" isn't automatically one consistent number even within the FTP model itself.

Whichever you run, both tests are also noisy on a single attempt — motivation, prior fatigue, heat, and pacing experience all shift the result session to session, independent of any CP-vs-FTP model question. That's a separate source of variance sitting on top of the model gap discussed above, and it's part of why a single test six months ago shouldn't be treated as gospel today regardless of which protocol produced it.

Section 04Fix isn't picking the "right" number. It's picking one and staying with it

Pick a model and commit to it. If the plan you're following — or writing — is expressed in %CP, test CP. If it's expressed in %FTP, test FTP. Don't convert on the fly between the two mid-block; convert the plan once, at the start, and re-derive every interval from your own number in that model.

Match the protocol, not just the label. CP is sensitive to which fixed durations you use for the trials. A CP estimated from 12/7/3-minute efforts is not the same number as one estimated from a 15/5-minute pair. If you're following someone else's CP-based plan, ask which protocol produced their zones before you assume yours will transfer.

Retest on a fixed cadence with the same protocol. Every 6-8 weeks, same test, same format, similar fatigue state going in. A drifting number is only meaningful if the measuring stick underneath it hasn't changed.

Don't borrow a training partner's zones across models. Their FTP isn't your CP even when the numbers look similar on the screen. If an indoor group ride's leader calls out "290 watts, hold it" and your CP happens to be 290, you may be sitting at their FTP-equivalent effort, not theirs — a materially different physiological ask.

When you change coaches or plans, ask the model question first. Before you touch a single interval, find out which threshold construct the plan is built on. It's a thirty-second question that prevents a six-week problem.

Wherethis breaks — the honest limits

This isn't licence to obsess over which number is "true." Both studies here tested trained cyclists and triathletes specifically — the size of the CP–FTP gap, and possibly even its direction, can shift with training status, so don't assume the same 7–16 W spread applies if you're new to structured training or considerably fitter or less fit than these cohorts. CP itself isn't one fixed number across labs, either. Change the fixed-duration efforts used to derive it and you get a different estimate, which means "my CP" really means "my CP under this specific protocol."

More fundamentally, neither number is a direct physiological measurement. FTP is a training construct invented to make interval prescription practical — nobody's body has a switch that flips at exactly 95% of a 20-minute average. Critical Power is a mathematical asymptote extracted from a curve fit across a handful of trials — an abstraction, not a sensor reading. Chasing "the correct number" between the two is the wrong question. What matters is testing the same way every time and building zones off a number that's internally consistent with the plan you're actually running.

Section 06Applying it with HELIOS

None of this is an argument against power-based testing — CP and FTP are both useful, as long as you know which one you're using and how it was derived. What HELIOS and the THRIVE training planner add is a way to catch the day the number stops matching the effort. The planner logs which protocol set your current zones — CP, ramp-test FTP, 20-minute TT — and Marco tracks how your heart rate, HRV, and perceived effort respond during a prescribed session relative to what that intensity should feel like for you.

When a session written as "moderate" sweet-spot starts producing a heart-rate and recovery signature that reads like threshold work, that's usually the first sign your number has drifted — or that the session you're running was actually written for the other model. The ring isn't measuring your watts. It's flagging the mismatch between what the plan expected and what your body is actually doing, so you retest before three more weeks go by built on the wrong number.

Section 07Bottom line

Critical Power and 20-minute FTP are close cousins, not the same threshold. Across two independent studies of trained cyclists, CP sat somewhere between 7 and 16 watts above FTP on average — and swung by as much as 20–30 watts rider to rider around that average. That's exactly the margin that turns a planned sweet-spot session into unplanned threshold work. Pick one model, test it the same way every time, and stop assuming any plan's wattage imports cleanly from a different threshold construct.


Caveat

Cohort was trained cyclists; the CP–FTP relationship shifts with fitness and with which CP durations you use. FTP is a construct, not a physiological line — neither number is "true," they're two estimates of the same ballpark. Internal consistency matters more than which model is "correct."

Counterpoint · Read this before you rebuild your week

The other side of the evidence.

Cohort was trained cyclists; the CP–FTP relationship shifts with fitness and with which CP durations you use. FTP is a construct, not a physiological line — neither number is 'true', they're two estimates of the same ballpark. Internal consistency matters more than which model is 'correct'.

Written by

THRIVE Cycling

Cycling Science Desk

THRIVE's Cycling Science desk translates peer-reviewed exercise-science literature into protocols riders can actually use. Every claim is checked against the primary source before publish, and every piece carries its counterpoint.

↗ 2 studies cited↗ Every claim source-checked↗ Updated 2026↗ Counterpoint included

About this article

Methodology & transparency.

Studies cited
2 peer-reviewed papers · Frontiers in Physiology, International Journal of Exercise Science
Cohort base
Trained cyclists tested on both a critical-power protocol (multiple fixed-duration maximal efforts) and a single 20-min FTP test.
Conflicts of interest
THRIVE Cycling publishes this article. Where HELIOS or ULTRA is mentioned, the underlying research claim stands independently of the product mention.
Last reviewed
2026 · verification: Cross-checked against primary sources via an independent research pass; corrections logged in the record history.
Reading time
9 min · 2139 words · 230 wpm average adult reading speed

Sources.

  1. 01Karsten et al. (2021). Relationship between the critical power test and a 20-min functional threshold power test in cycling. Frontiers in Physiology, 11, 613151. https://doi.org/10.3389/fphys.2020.613151 No DOI on record
  2. 02McGrath et al. (2021). Do critical and functional threshold powers equate in highly-trained athletes?. International Journal of Exercise Science, 14(4), 45–59. https://doi.org/10.70252/ISYH9512 No DOI on record

The Weekly Note

One study. One protocol.

Peer-reviewed cycling science, deconstructed and ready to apply — straight to your inbox. No filler, no promo.

One ring. One app.
Every answer.

The science only works if you measure what it asks you to. HELIOS reads your HRV, sleep, ride load, and recovery — and turns it into one daily call.

Get HELIOS Pro Pack →

Pre-order · ships Jul 31 – Aug 10