← Back to News
Why Your Training Plan Expired Before You Even Started It

Why Your Training Plan Expired Before You Even Started It

A twelve-week plan is a story told in advance, frozen before it met you. What changes when something can finally read every run you have ever done?

Chris MintzChris Mintz

Nobody starts a marathon plan exactly twelve weeks out.

You find the plan in October, and the race is in early January, and it wants sixteen weeks. So you skip the base and start at week five, which is where the volume already assumes a base you do not have. Or you find it in June for a November race and you invent two weeks of your own to bridge the gap, guessing at what belongs there. The plan cannot tell you which of those to do. It was written before it met you.

This is not a failure of the coach who wrote it. It is the training plan model working exactly as designed, and the design has expired.

What a training plan was actually for

Strip away the week grid and a training plan is a story told in advance.

Someone sat down, decided which weeks mattered, decided how much running answered them, and then froze that decision into a table. Week 1: 40km. Week 9: the first back-to-back. Week 14: taper. Each cell is a pre-computed narrative arc: here is where you build, here is where you sharpen, here is where you should feel tired and not worry about it.

That was a rational design. When a coach has forty athletes and can watch perhaps three of them run, aggregation is not a limitation, it is the entire point. The plan compressed everything the coach knew into something a stranger could follow without asking a single question.

The trouble with compression is that it is lossy, and you have to decide what to throw away before you know what you will need.

Every plan is a bet that your season will resemble the season it was written for. It is a bet that gets worse every week it stays up. Anyone who has followed one to week seven and then caught a cold, or moved house, or found that the Tuesday intervals were pitched for someone with a different threshold, has watched the bet lose in real time. The plan does not adapt, because it cannot. It has no idea you exist.

The constraint that disappeared

Here is what changed, and it is not subtle.

For as long as coaching has been a profession, the reason plans were written in advance and handed over as documents was that a coach could not read every run from every athlete. That was the binding constraint. Coaching time was expensive, sure, but the real bottleneck was attention. One person cannot hold thirty training histories in their head, notice that your last four easy runs have crept eleven seconds per kilometre faster, and connect that to the calf that started complaining on Sunday.

A language model can read every run. Not skim them. Read them, correlate them, notice that four sessions share a pattern and an unusual heart-rate drift, and explain why that matters in a paragraph. The bottleneck moved. And when the bottleneck moves, the architecture built around it stops making sense.

The old model was: compress everything a coach knows into a document, hand it over, and hope the athlete adapts it correctly. The new model is: leave the training history where it is, and talk to it.

The data was already there. Nothing could read it.

This would be a nice theory if the raw material did not exist. It does, and it has for a decade.

Your watch has been logging every run since you bought it. Pace, cadence, heart rate, elevation, and on the newer ones power, ground contact time, and an estimate of your vertical oscillation that nobody has ever acted on. The chest strap has HRV. The scale has weight and, if you let it, resting heart rate. Strava has the social layer and every route you have ever run. This is a genuinely rich longitudinal dataset about one human being, collected continuously, at a resolution no coach has ever had access to.

And overwhelmingly it is used in the dumbest possible way: as a scoreboard you look at after the run, and a weekly total you compare to a plan that was written by someone who has never seen any of it.

The conventional version of "the app knows something" is a readiness score. Your watch samples overnight HRV, resting heart rate, and sleep, and returns a number between one and a hundred with a colour attached. It is a dashboard tile. It is a pre-frozen question, "how recovered is this person, in general", computed by a model that has never been told what you are training for, when your race is, that the 34km long run on Sunday was deliberately hard, or that you always sleep badly the night after a hard session and it has never meant anything.

The categories nobody anticipated go unnoticed, because the score has no vocabulary for them. It cannot tell you that your easy runs have been drifting harder for a month, which is the single most common way a build quietly fails. It was not asked that question when it was designed.

Ask it directly and the questions are not tiles at all:

  • "Show me every run I logged as easy in the last eight weeks where my heart rate finished more than eight beats above where it started, and tell me whether that is getting worse."
  • "I have nine weeks until a fifty-mile race with 2,000 metres of climbing. Look at what I have actually done since March and tell me what is missing."
  • "My calf started hurting on the 14th. What changed in the two weeks before that?"

None of those are readiness scores. All of them are answerable today, against data you are already carrying on your wrist.

Why this matters more in ultra than anywhere else

Most sports have training plans. Ultra has events long enough and individual enough that a template cannot survive them.

A marathon is roughly the same problem for everybody: a known distance, a mostly known surface, three to five hours, one fuelling strategy that either holds or does not. A hundred miler is twenty-four to thirty-six hours of accumulating decisions, at an intensity low enough that fitness stops being the limiter and everything else starts being it: gut, feet, sleep, heat, the ability to eat at 3am. Two runners with identical VO2 max and identical weekly volume can have completely different races, and the reason is almost never the running.

That is exactly the situation a frozen document handles worst. The variables that decide your race are the ones the plan's author could not know: how your stomach behaves after hour eight, whether you overheat, how your sleep debt accumulates, what your feet do when wet. A template can give you the volume. It cannot give you the rehearsal, because the rehearsal has to be designed around your specific failure modes, and nobody has catalogued those but you and your watch.

The honest part

Two things would make this article a sales pitch if I left them out.

Pointing a language model at raw training data does not work. On Spider 1.0, the long-standing academic text-to-SQL benchmark, frontier models score above 90%. On Spider 2.0, which sets the same task against realistic schemas with thousands of columns, multiple dialects and multi-step workflows, execution accuracy collapses to roughly 21%. Enterprise-adapted variants of BIRD land near 39%. The failure mode is the dangerous kind: the query does not crash, it returns a plausible number that is wrong.

A raw Garmin export is exactly that kind of schema. Ask an unassisted model what your threshold pace is and it will tell you, fluently, confidently, and often wrong, because it has no definition of threshold, no idea which of the four heart-rate fields is trustworthy, no concept that a 2km warm-up should be excluded, and no way to know that the 6:12 kilometre in the middle of that run was a road crossing.

The fix is not a better model. It is a physiology layer. The gap between 21% and something you would let a runner act on is almost entirely made of modelled definitions: what "easy" means for this runner this month, which heart-rate zones are actually theirs rather than 220-minus-age, what counts as a training stimulus versus a commute, when a long run is a long run. Skip that step and you get confident hallucination with a training plan attached.

And conversation has its own failure mode. Researchers Ken Gu, Srishti Palani and Vidya Setlur presented work at CHI 2026 on what they call conversational debt: people conducting extended analytical conversations cannot find their way back to insights buried in the history. There is no search, no navigation, no structure. The more you explore, the harder it becomes to recover what you found.

Three months of chat about your calf, and no way back to the thing that actually worked. A plan, for all its rigidity, is at least persistent and addressable. You can point at week nine.

So what do you actually do

The thesis is not "throw away your training plan." It is that the plan has been demoted: from the thing you are handed at the start to the durable artefact of a conversation that already happened.

  1. Keep the raw runs, not the weekly totals. Weekly mileage is an aggregation, and it is the one that hides everything interesting. Every question worth asking lives at the level of the individual session.
  2. Build the physiology layer before the coaching. Define what easy, threshold and long mean for this runner, from this runner's own data, before anything is allowed to prescribe. This is the unglamorous 80% of the work and the entire difference between useful and dangerous.
  3. Let the plan be written last. When a question has been asked three times, whether that is how much vert, how long the back-to-backs, or what to eat, it has earned a place in a document. That is the correct order of operations, and the inverse of how every training plan has ever been produced.
  4. Keep the human on the decisions that hurt. Reading the data can be automatic. Whether to start a race on a sore Achilles, whether to drop at 100km, whether this is the season to try. Those are not analysis problems.

What we are building

We are working on this at Gateway, and it is not finished, so treat what follows as a progress note rather than a product. Most of the effort so far has gone precisely where this article says it should: not into the conversation, which is the easy part, but into the layer underneath it that decides what a runner's own numbers actually mean. That work is slower and less impressive to demo than a chat window, and it is the only part that determines whether the advice is worth following. We will write about what we get wrong.

The training plan's job was to tell you a story about your season. It expired before you even started it, because it was written under a constraint that no longer exists: that a coach could not read every run you had ever done.

Now they can. The story is already in the data. You just have to ask.


Sources & further reading

  • Lei et al., Spider 2.0: Towards Evaluating Text-to-SQL in Realistic Enterprise Settings (arXiv:2411.07763)
  • Li et al., BIRD: A Big Bench for Large-Scale Database Grounded Text-to-SQL
  • Gu, Palani & Setlur, "I Need to Find That One Chart", CHI 2026, on conversational debt
Chris Mintz

Chris Mintz

Head of Engineering

Chris is an ultrarunner with several dozen ultra finishes to his name, including Fat Dog 120, Bigfoot 200 and the QMT 135. When he isn't out on a course, he's helping put one on as a member of the race director team for the Pick Your Poison trail race. Off the trail, Chris brings over 15 years of experience in software architecture, engineering and data science to his projects. He holds a Bachelor of Science in Data Science from the University of Waterloo and a Master of Computer Science with distinction in Applied AI from the University of Hull, and is an AWS Certified Solutions Architect Associate and PCAP Certified Associate Python Programmer.