← Back to blog

The 20% Buffer Myth: Estimating by Uncertainty

·13 min read

The 20% Buffer Nobody Believes In

There's a ritual in almost every software planning meeting. Someone asks how long a feature will take. The engineer says three weeks. The project manager nods, then quietly writes down "four weeks." Three plus 20%. Done.

Then the feature takes five weeks anyway.

Sound familiar? The uniform buffer, that reflexive 20%, 25%, or "just add a sprint", is one of the most common planning habits in software. It's also one of the weakest, because it treats every kind of uncertainty as if it came in the same size. A payments integration into a system nobody has touched doesn't carry the same risk as a copy change on the settings page. Padding both by 20% doesn't make the estimate more accurate. It just makes one number slightly safer and the other one fictional.

The problem isn't that you're buffering. It's that you're buffering blindly.

The alternative is to stop thinking in percentages and start thinking in uncertainty classes. That's the argument I want to make here, that the uncertainty buffers habit quietly papers over the real problems (unclear scope, unknown dependencies, missing discovery) and produces teams that neither trust their own estimates nor learn anything from them.

Let's start with the myths that keep the ritual alive.

Myth #1: You Can Estimate High-Uncertainty Work Like Familiar Work

No, you can't. And pretending you can is where most estimate failures begin. A task built on known technology, stable requirements, and proven dependencies behaves nothing like one built on new tech, fuzzy acceptance criteria, or an external vendor's API.

Consider two tasks on the same sprint board. First: "Add a last_login field to the user table and show it on the admin panel." Second: "Figure out whether our billing provider supports usage-based metering for EU customers." The first one might take four hours. The second one might take four hours, or it might take two weeks, because the answer could be "no" and then you're redesigning the whole plan. On a board, they look similar. They aren't.

Research from DORA and others keeps pointing at the same thing: teams that classify work by uncertainty hit their windows more often than teams that don't. The framework most experienced leads use by instinct looks roughly like this:

  • Low uncertainty. Familiar stack, known interfaces, clear acceptance criteria. Estimate these normally.
  • Medium uncertainty. Some unknowns, moderate integration work, acceptance criteria still being refined. Estimate, then add a contingency in days, not a percentage.
  • High uncertainty. New technology, ambiguous requirements, external dependencies, or research-heavy work. Don't estimate delivery. Estimate a discovery task first.

A high-uncertainty task isn't really a task. It's a question that hasn't been answered yet. The moment you accept that, the 20% buffer stops looking like a solution and starts looking like a way to avoid the conversation.

Here's the practical version of the rule. When you genuinely can't tell whether something is medium or high uncertainty, you're really asking whether a discovery step is missing. The answer is almost always yes.

Myth #2: A Target Date Is a Commitment

No. A target date is a hope dressed as a plan. A forecast date is your current best guess. A commitment date is something you agreed to out loud with a stakeholder. Conflating all three is why "slipped deadlines" feel like betrayals instead of ordinary updates.

Most teams use one date for three different jobs, then act surprised when things go sideways. The fix isn't more discipline. It's clearer vocabulary.

  • Target date. What you'd like to hit. A north star, not a promise.
  • Forecast date. Your current best estimate given what you know today. It moves as facts change.
  • Commitment date. A date you've explicitly agreed with someone outside the team. This is the one you protect and escalate on.
  • Review date. The date you've set to re-check your assumptions. This is the one most teams skip.

That last one matters more than people think. Without a review date, a forecast is just a number that slowly becomes a lie. With one, the forecast can change without anyone losing trust, because you told people when you'd check, and then you checked.

A forecast isn't a guess you're stuck with. It's an update schedule.

I've watched this play out on small teams. A backend engineer wrote "Friday" on a migration task. That Friday, blockers surfaced, a permission model nobody had accounted for. Instead of quietly missing the date, the engineer updated the forecast to the following Wednesday and dropped a two-line note explaining why. Nobody panicked. The stakeholder replied "got it." That's what the framework buys you: room to be wrong on time.

The alternative, treating Friday like a blood oath, then missing it anyway, destroys trust faster than a moved forecast ever will. And this is where a task system with real keyboard-first capture earns its keep: the update takes ten seconds, so people actually make it. If revising a forecast means clicking through four screens, it won't happen, and the forecast rots in place. That's not a philosophy problem. It's a friction problem.

Myth #3: Discovery and Delivery Should Share One Deadline

They shouldn't. Discovery and delivery are different kinds of work, and folding them into one task under one deadline hides the real risk. Split the discovery out, and the delivery estimate suddenly becomes honest.

Here's a scenario I've seen more than once. A team gets a request: "Ship SSO for enterprise customers by end of quarter." Someone turns that into a single task, gives it a six-week estimate, and puts it on the roadmap. Three weeks in, the team discovers the identity provider's SAML implementation doesn't support the exact flow they designed for. Now they're rebuilding from scratch. Quarter blown.

Run it differently and the story changes:

  1. Discovery: determine whether the provider supports the required flow. Time-box: one week.
  2. Implementation: build the supported flow. Estimate only after discovery.
  3. Verification: test edge cases and failure handling.
  4. Release: deploy and monitor.

The discovery task might come back with "yes, and it's easy." Or it might come back with "no, and here's a different approach." Either answer is useful. What it won't do is sit invisibly inside a six-week estimate and blow it up in week three.

Discovery is the cheapest week you'll ever spend, because it's the week that stops you from spending six.

This is also why the definition of done matters so much. A task called "Improve onboarding" has no definition of done. A task called "Identify the top three onboarding drop-off points and propose one measurable experiment by Friday" does. The second one is a discovery task with a concrete outcome. The first one is a mood.

The "one task, one outcome" rule is the simplest version of this. If you can't write a sentence describing the outcome and a date when you'd know you were done, you've got a discovery task on your hands. So write that one instead. Harvard Business Review has published plenty on this theme over the years, the estimate you make before you've defined the outcome isn't an estimate, it's a placeholder.

Myth #4: Changing a Forecast Means the Project Is Failing

It doesn't. A changed forecast is information. A forecast that never changes is a sign nobody is looking.

There's a strange cultural tic in software. Slipping a date feels like admitting failure. So teams either (a) don't slip, they just quietly stop mentioning the date, or (b) slip so rarely that when it happens, it's a crisis.

Both patterns are bad. The healthier one is a forecast that moves a little, often, with a reason attached each time. It's boring. Boring is exactly what you want.

A useful protocol, assembled from what a lot of experienced planners already do:

  1. Set an initial target with the information you have.
  2. Write down the assumptions behind the target.
  3. Identify the single biggest uncertainty or dependency.
  4. Schedule a review date.
  5. When facts change, update the forecast, and say what changed.
  6. Escalate only when the commitment date is at risk.
  7. Keep the original date for retrospective analysis.

That last point is underrated. If you overwrite your old estimates, you never learn which kinds of work you systematically underestimate. The pattern is usually hiding in plain sight: migrations always run long, integrations always run long, UI polish usually runs short. You only see the pattern if you keep the receipts.

Preserve the original date. It's the only way to learn from your own planning.

Dynamic planning isn't the same as casual planning. Casual planning means throwing a date at the wall and hoping. Dynamic planning means checking in on schedule, revising the forecast when you have to, and explaining why, every single time. The change is the process, not the failure.

What Actually Happens When AI Swells the Backlog

AI coding tools make developers faster at writing code, but they also increase the volume of work flowing through the system, reviews, testing, coordination, follow-up tasks. Estimation gets harder, not easier, because the shape of the work has changed.

Bain & Company's 2026 Global Technology Report found that AI coding tools helped developers complete about 21% more tasks, while time spent reviewing AI-generated output jumped 91%. Developers also reported managing 47% more simultaneous workstreams than before. Meanwhile, engineering leaders expected release-cycle speed to improve by 148% over the next one to two years, while the realized gain sat closer to 20–27%.

That's a big gap. And the gap isn't a coding problem. It's a planning problem. Code is moving faster. Everything around the code, review queues, requirements clarification, handoffs, release coordination, is moving at roughly the same speed it always did. So the bottleneck just relocated downstream.

For estimation, this creates a specific hazard. People start pricing tasks as if the code-writing part is the whole task. It isn't. A feature isn't done when the code is written. It's done when it's verified and shipped, and AI-generated code often costs more to verify than code a person wrote, because the reviewer has to reconstruct the reasoning that produced it.

One team I know adopted a simple four-state rule for every AI-assisted task:

  • Requested, what outcome is needed?
  • Generated, what did the human or the agent produce?
  • Verified, has it passed review, tests, and acceptance criteria?
  • Shipped, is it deployed or otherwise complete?

The rule exists so nobody mistakes "code written" for "work finished." And once you see the four states, the buffer math changes completely. You're no longer buffering coding time. You're buffering verification and coordination time behind it.

If your estimates only cover the coding, you're only estimating part of the job.

How to Run Uncertainty-Based Buffers in Practice

Classify each task by uncertainty, convert the buffer from a percentage into a stated contingency in days, and review every forecast on a fixed cadence. It sounds bureaucratic. It takes about five minutes per planning session.

Here's how it looks in the wild.

Step 1: Tag every non-trivial task with an uncertainty class.

  • Low, familiar implementation, stable requirements, known dependencies.
  • Medium, some unknowns, moderate integration work, incomplete acceptance criteria.
  • High, new technology, ambiguous requirements, external dependencies, research-heavy work.

Most calls fit in a few seconds of thought, which is the point. If you can't decide between medium and high, treat it as high and add a discovery step.

Step 2: Estimate low-uncertainty work normally. Don't pad it. If it's a familiar operation, your team's historical data is probably fine.

Step 3: Convert buffers to days, not percentages. For medium-uncertainty work, name the specific risk and add a specific number of days for it: "two extra days in case the vendor's rate limit is lower than documented." A named contingency forces you to actually think about what could go wrong. A percentage doesn't.

Step 4: Split high-uncertainty work. Discovery first. Delivery estimate second. Discovery tasks themselves should be time-boxed, a week, two days, whatever fits, and should end with a written answer, not a vibe.

Step 5: Set a review date for anything with medium or high uncertainty. Put it on the calendar when you set the estimate. Review it whether or not anyone asks.

Step 6: Watch your flow metrics, not your effort. The useful numbers are lead time from request to production, cycle time from started to completed, review turnaround, and the count of blocked tasks. Activity metrics like commits or hours online will happily climb while nothing actually ships.

A note on measurement: developers seem to want this. One survey cited in the research found 89.2% of developers believe productivity can be measured with useful metrics, and 92.1% want to know how productive they've been. The catch is which metrics. Counting commits is easy. Counting how reliably work moves from intent to outcome is what people actually mean.

There's a tooling angle here, but it's a small one. A task system that's fast enough to update, keyboard-first, no mouse, no ceremony, is the one where forecasts actually get revised on time. If updating a task is annoying, forecasts rot. That's the whole mechanism.

The point of all six steps is the same: make uncertainty visible instead of averaging it away.

Frequently Asked Questions

Isn't this just adding a buffer with extra steps?

Sort of, but the important difference is where the buffer comes from. A flat 20% is a number you made up. A named contingency is a specific risk with a specific cost attached. When the risk doesn't materialize, you can see that, and next time you can estimate better. With a flat percentage, you never learn anything at all. You just keep paying the tax on work that didn't need it.

What if my team can't agree on what "high uncertainty" means?

Then you have a definition problem, not a buffer problem. Write down three examples from your own history for each class. Most teams find that once they look at past work, the classification is obvious in hindsight, it was only hard in advance because nobody had labeled anything before. Do the labeling once and the argument goes away.

Does this work for solo founders and freelancers?

Especially for them. A solo founder has no team to absorb an overrun and no project manager to renegotiate the date. Splitting discovery from delivery is the cheapest risk control available. A two-day scoping task before a three-week build is the best two days you'll spend all quarter. And when you're the only person on the project, a forecast you can defend is worth more than a date you hoped for.

Won't clients and stakeholders hate moving forecasts?

Only if the forecast moves without explanation. The protocol that works is boring and repetitive: change the date, say why, say what you know now that you didn't before. What stakeholders actually hate is a date that quietly stops being true. A forecast that moves with a reason attached reads as professional control. A forecast that silently goes stale reads as chaos.

How many days of contingency is "right"?

There isn't a universal number, and anyone selling you one is guessing. The right contingency is proportional to the size of the unknown. If you can't describe the unknown, you haven't done the analysis yet, go run a discovery task. This is also why vendor productivity claims should be handled with care. IBM reported average gains around 45% across modernization, security, and new development, and one team on its Instana product cut time on selected tasks by roughly 70%, about 10 hours a week per person. Those figures come from IBM's own case material and should be read as case studies, not benchmarks. Your results will differ. Run your own numbers.