Article
/ Agile Teams

The Predictability Trap: Why Reliable Agile Teams Aren't Built by Demanding Better Forecasts

Alexander Hilton |  7m 30s

Earn Scrum Education Units (SEUs)

Earn credit towards renewing your certifications.

0.25 SEUs Earned
Log in to earn SEUs Log in
My SEUs
0 0
Log in to earn SEUs Log in
SEUs
0.25 SEUs Earned
My SEUs
0 0

"We need the team to be more predictable."

It is a reasonable request. Leaders want confidence. Teams want to be trusted. Nobody is acting in bad faith.

But I have watched this request play out enough times that I can almost set a clock to it. The behavior changes show up on a reporting cadence. Within one or two cycles of predictability becoming a measured target, whether the label is say-do ratio, commitment accuracy, or sprint completion percentage, the same patterns appear. Testing activity that used to live inside a story gets split out as a separate work item, then split again. The board fills up with more cards that each look smaller and safer. Team members who were happy to pair or swarm start working alone, because solo work is easier to attribute to a forecast. Lead time creeps up. Time to market stretches out. The dashboards turn green. Less actually reaches the customer.

That last sentence is the one that matters, and it almost never makes it into the conversation with leadership. More money is being spent on less work being delivered. The team looks productive on paper. The P&L tells a different story.

Reliability is what we actually desire

The word leaders use is "predictability," but the thing they want is reliability. Those are not the same.

Predictability is easy to achieve. A train that arrives about an hour late every day is highly predictable. Nobody wants that train. What people want is a fast, reliable train that gets them where they are going, and when something does go wrong, gets back on schedule quickly.

The DORA research has been making this point empirically for over a decade. The metrics that distinguish elite-performing organizations are lead time for changes, deployment frequency, change failure rate, and time to restore service. Forecast accuracy is not on the list. None of the DORA measures ask how well a team predicted its own work. They ask whether the system delivers value quickly and recovers when something breaks. 

This surprises people. Reliability does not come from longer testing windows or more careful planning. It often comes from doing the opposite. Smaller, more frequent releases mean less can go wrong in any single change, and when something does break, the team can fix forward in minutes rather than hours. Variability drops, throughput improves, and predictability shows up as a side effect.

I tested this years ago with one of the most challenging teams I have worked with. Their quarterly release was the kind of event where nobody slept well the night before. Every conversation about it ended with someone asking for a longer testing window, more sign-offs, tighter change control. More predictability, in other words. We did the opposite. We shifted to daily releases. Change failure rates dropped sharply. Time to restore service collapsed from hours to minutes. The release that used to feel like walking a tightrope became something the team barely noticed. Smaller batches, less variability, more reliability. The predictability everyone had been chasing finally arrived, not because we demanded it, but because we stopped getting in its way.

Why the waste

The reason I can forecast the behavior change so reliably is that the underlying mechanics have been documented for decades. Eliyahu Goldratt described them clearly in his work on the Theory of Constraints. Every task carries uncertainty, and the higher the uncertainty, the longer the tail on its completion distribution. To finish on time with high confidence, you do not estimate at the 50% mark. You estimate somewhere past the 80% mark, well into the long tail.

Alex Hilton graph

Figure 1. The higher the uncertainty, the longer the tail. Each curve shows the range of times a task could take. The 50% line marks the honest midpoint estimate; the 80% line marks where someone would have to commit to be confident of finishing on time. The red arrow shows the extra padding required to move from a 50% estimate to an 80% commitment. When uncertainty is low, that padding is small. When uncertainty is high, it is several times larger, which is why teams quietly inflate estimates rather than admit how wide the range really is. Adapted from Goldratt, Beyond the Goal (2005).

Consider the position this puts a team member in. To be considered reliable, you must meet your commitments. To meet your commitments, you estimate above the 80% mark. But you also cannot be seen as exaggerating. The loop is unstable, and the way most people resolve it is by quietly padding while pretending they are not. Once predictability becomes the target, it stops being a useful measure. 

Where scrum masters, product owners, and delivery managers earn their keep

The real constraints behind reliability are rarely inside the team. They sit in the technical architecture, in the deployment pipeline, in handoffs between groups, in policies that require five approvals to ship a one-line change, in the way work is funded and prioritized. Solving for those constraints means being willing to ask uncomfortable questions, and to ask them upward.

Why does this deployment take six hours? Why does QA happen in a separate team three weeks after development? Why do we batch work into quarterly releases when the technology supports daily ones? Why are we measuring this team on commitment accuracy when the lead time data shows the constraint sits two teams over? A scrum master who only runs ceremonies is missing most of the job. A product owner who only grooms the backlog is leaving the most valuable work on the table. The leverage is in the questions that nobody else on the team has the standing or the framing to ask.

A useful starting point is variability. Where standard deviation in lead time is high, that is where your real constraint lives. That is the place worth investigating, not the team's commitment accuracy. And when leaders push for predictability anyway, the response is sequencing: improve lead time first, predictability follows. Never the other way around.

AI raises the cost of getting this wrong

AI compresses everything around delivery: development cycles, market feedback loops, competitor response times. A team optimizing for clean forecasts is buffering against the wrong variable while competitors using AI to shorten their build, measure, learn loop reach the market first. Pointed at the wrong target, AI produces better forecasting theater. Pointed at the right one, it shortens the reliability loop. The choice of target has always mattered. AI just makes choosing the wrong one more expensive and faster.

The inversion

Agile teams should be predictable because they have become reliable, not forced to look reliable because they have been measured into predictability. Nobody wants a predictably slow train. The goal of agile is not perfect prediction. It is reliable adaptation, and in an AI-accelerated market, reliable adaptation is the only kind of predictability worth having.

References

Forsgren, N., Humble, J., & Kim, G. (2018). Accelerate: The Science of Lean Software and DevOps: Building and Scaling High Performing Technology Organizations. Portland, OR: IT Revolution Press.

DORA. (2024). Accelerate State of DevOps Report 2024. Google Cloud / DevOps Research and Assessment Team.

Goldratt, E. M. (1997). Critical Chain. Great Barrington, MA: North River Press.

Goldratt, E. M. (2005). Beyond the Goal: Eliyahu Goldratt Speaks on the Theory of Constraints [Audio program]. Gildan Audio.

Strathern, M. (1997). "'Improving ratings': Audit in the British university system." European Review, 5(3), 305–321.

--

Lead your team toward reliable adaptation. Gain the skills to identify systemic constraints, shorten delivery loops, and foster true organizational agility. Learn more about the Certified ScrumMaster® course.

About the author

Alexander Hilton
Alexander D. Hilton, MBA, is a Certified Team Coach (CTC) through Scrum Alliance and a digital strategy and AI consultant working with Agile New England, an ACM Chapter. He has consulted with Global 2000 and FTSE 100 organizations across regulated industries, with a focus on the intersection of Theory of Constraints, delivery system design, throughput accounting, and AI value quantification, connecting how teams deliver to the financial outcomes that follow. He has presented at global and regional Scrum Gatherings, and his writing has appeared in California Management Review Insights (Berkeley Haas), CFO.com, and ACM SIGAI. Opinions are the author's own.