top of page

[2/2] Escaping the Firefighting Fallacy: A Practical Guide

1 day ago
9 min read

Previously I wrote about a reactive approach to quality that leads to organizations constantly fighting one fire after another. Companies often believe that they need to keep putting out fires and fix the cause later. However, tomorrow becomes today and more fires are popping up.


Photo by Jay Heike
Photo by Jay Heike

Over time, something subtle starts happening. More and more organizational energy gets consumed, teams become overloaded with coordination overhead and improvement gets postponed because there is never enough time.

The organization becomes trapped in permanent reaction mode. The Firefighting Fallacy.

This follow up guide is not a blueprint that will magically fix that situation. There is no universal framework that can fully account for the complexity, politics, incentives and tacit knowledge inside every organization. What this guide can do is help you set out on a path of fixing the system that led to firefighting fallacy.

Many organizations that are stuck firefighting are no longer operating based on reality. They are operating based on assumptions, abstractions, habits and local optimizations that slowly drift away from what is actually happening inside the system. Unless that reality becomes visible again, meaningful improvement is mostly guesswork. This guide should help you restore that contact with reality.


Improvement Is a Cost Before It Becomes a Capability

One of the hardest things to accept in struggling organizations is that systemic improvement is expensive. It consumes time, focus, political capital and delivery capacity before it produces visible stability improvements. Teams already under pressure rarely want to hear that the solution requires slowing down parts of the system temporarily in order to stop the long-term bleeding.

This is one of the reasons firefighting becomes so seductive. Reactive solutions usually feel cheaper at the moment.

Skipping deeper improvement work allows the organization to preserve short-term output. But the source of the problems does not disappear. Over time, new fires will get kindled and they will only get worse than the ones before.

At some point, the operational drag becomes so large that even maintaining the current pace starts consuming extraordinary effort. In last ditch effort, more new policies are applied that only exacerbate the problem. Where processes should be simplified, new steps are added and ultimately delivery slows down anyway, but now with significant damages also present. 

This is why especially organizations under survival pressure cannot afford to postpone systemic improvement indefinitely. Delaying the investment does not remove the cost. It only allows the cost to compound invisibly until it becomes unavoidable.


Two Solution Tracks Must Exist in Parallel

Many organizations make the mistake of treating firefighting and improvement as sequential activities. We fight the fires today, fix the system tomorrow. I explained why that is not a good idea in the previous article, the core of the argument is that new fires keep coming and “fix later” never happens.

At the same time, organizations cannot ignore immediate operational instability either. Severe incidents still need response. Customers still need support. Revenue pressure still exists.

This means two tracks need to exist simultaneously.

The first track is tactical stabilization, to prevent catastrophic failures, limit blast radius and prevent reputation from deteriorating. This needs to be focused, tactical firefighting. There will be fires not attended to, only the really major ones get any resources. Not attending every fire in existence opens up capacity for the second parallel track.

The second track is systemic self-understanding. Identifying what the organization is actually optimizing for, understanding how decisions are really made, exposing unchallenged assumptions, and simplifying processes. Only with a good understanding of the system can we hope to fix its weaknesses.

The first track is far more tangible and the results are immediately visible. That is why most organizations over-invest into it. However without the second track, firefighting simply becomes more organized firefighting.

This transition is not comfortable. Short-term throughput may temporarily decrease. Existing reporting structures may lose usefulness. Some teams will resist additional visibility because hidden adaptation often protects them from unrealistic expectations elsewhere in the system. Managers may feel loss of control when informal workarounds become visible. Long-standing assumptions about productivity and ownership may need to be challenged publicly.

None of this means the improvement effort is failing. In many organizations, discomfort is simply the first sign that operational reality is becoming visible again instead of being continuously absorbed by heroic intervention and organizational coping mechanisms.


Identify Who the System Actually Serves

“The customer” is often not a single entity. The person paying for the product may not be the person using it. The user may not be the person deciding whether the contract gets renewed. Regulatory bodies may impose constraints that customers themselves do not care about directly. Shareholders care about sustainability and profitability. Engineers care about maintainability and operational sanity. Leadership cares about predictability and growth. Competition shapes expectations even without interacting with your product directly.

All of these groups exert pressure on the system simultaneously.

The problem is that many organizations never explicitly map these pressures. When those competing pressures remain implicit, organizations often drift into accidental prioritization. Product may optimize feature throughput while support absorbs customer frustration. Engineering may optimize maintainability while sales continuously introduces custom commitments. Leadership may optimize reporting predictability while operational instability quietly grows underneath. Priorities emerge implicitly through urgency, politics, habit or whoever is shouting the loudest during the current escalation.

This is one of the reasons firefighting organizations drift into reactive decision making. They optimize locally for whichever pressure feels most immediate while losing visibility of the larger system they are trying to balance.

The goal here is not creating a giant stakeholder matrix nobody will ever read again.

The goal is forcing the organization to explicitly understand who this system is actually trying to serve and what their problems and needs are. Only once those become visible can they start being transformed into actionable objectives that can be prioritized and solved.

If this is not clear, organizations tend to optimize proxies instead of outcomes. Feature throughput becomes a proxy for customer value. Utilization becomes a proxy for productivity. Bug count becomes a proxy for quality. Delivery speed becomes a proxy for adaptability. Over time, the organization slowly loses contact with whether those proxies still reflect reality at all.

Before improving the system, you need to understand what the system is actually trying to achieve, for whom, and at what cost.


Map Reality, Not Process Diagrams

Once you understand who the system is supposed to serve and what objectives emerge from those pressures, the next step is understanding how work actually moves through the organization today. Not theoretically. Not according to Confluence. Not according to the onboarding presentation. How it moves in reality.

Deadlines exerted pressure, a change request came mid-sprint, teams bypass parts of the process to hit commitments. Hand-offs create misunderstandings, engineers compensate for missing clarity with assumptions, product introduces additional oversight to deal with delivery uncertainty. Over time, the actual delivery system slowly diverges from the documented one. This divergence is often difficult to acknowledge because the documented process does not merely describe how the organization works, it is aspirational and also represents how the organization believes it works, it represents its identity and understanding of itself. 

A process diagram might say: “Requirements are clarified during refinement.”

Reality might say: “Requirements stabilize three days after implementation has already started.”

A process diagram might say: “Teams own services independently.”

Reality might say: “Three senior engineers informally coordinate most critical decisions across the organization.”

A process diagram might say: “Quality is ensured through automated validation.”

Reality might say: “Production stability depends mostly on tribal knowledge and heroic intervention.”

This is why mapping reality requires observation, not documentation review.

The point of this exercise is not blaming teams for adapting. The point is understanding what your organization actually optimized itself into over time, because you cannot improve a system you are not looking at truthfully.

It is also important to understand that many dysfunctional patterns survive because they solve some local problem successfully.

A reporting layer may exist because leadership lost trust in delivery predictability. Excessive approvals may exist because previous failures created fear around autonomous decision making. Hidden dependencies may persist because certain individuals became the only reliable coordination mechanism inside an unstable system. These structures often remain in place not because organizations enjoy inefficiency, but because removing them without addressing the underlying fear or constraint can initially make the system feel even less stable.

This is why it is important to understand what a process is protecting the organization from, before attempting to improve or remove it.


Trace Where Learning Arrives Too Late

Most recurring failures are not isolated events. The production incident is usually not the moment the problem was created. It is the moment the organization could not ignore the problem anymore. The useful question is not: “How did this incident happen?” it is: “When could we realistically have discovered this earlier?”

The answer is usually way sooner than people think. The moment an idea is conceived is the moment when quality and operational risk should start being challenged, because every assumption discovered late becomes significantly more expensive to correct. Organizations trapped in the firefighting loop often discover problems only after the cost of correction has already become expensive and improving late-stage detection does not really help here.

When you diagnose an incident, chronologically follow the mapped real process and ask yourself if at any given point you could have learned about the problem and prevented it. Moreover, figure out what would need to be different to learn about the problem sooner. Be on a lookout for missing feedback loops, lack of communication and siloed processes.


Reduce the Cost of Being Wrong

Brittle systems are not necessarily evidence of incompetence. Complex systems naturally accumulate constraints over time. What separates mature organizations from the rest is they tend to surface brittleness earlier, understand its implications sooner and they respond before the cost becomes catastrophic.

Immature organizations often normalize fragility for too long because the surrounding organizational system continuously rewards short-term tradeoffs while delaying visibility of long-term consequences. Fixing the technical symptoms matters. But if the organizational conditions continuously reproduce the same technical outcomes, the instability eventually returns in a different form.

High performing systems usually focus more heavily on reducing the cost of discovering incorrect assumptions. This is why smaller delivery increments matter so much. Smaller batches help expose misunderstandings earlier, they shorten feedback cycles and reduce blast radius. Rollbacks and recovery is easier and customers get their value more often. All of this is not moving faster for its own sake. It is about making learning cheaper. That in itself does not magically remove technical constraints, staffing limitations or market pressure. But delayed learning makes all of those problems dramatically more expensive to solve because incorrect assumptions survive longer before the organization reacts to them. 


Remove Structures That Hide Reality

Some organizational structures exist primarily to improve operational understanding. Others mainly improve the feeling of control. Organizations under pressure often accumulate layers, ritualize behaviours, and obscures reality. 

These structures are not necessarily malicious. Most were introduced with good intentions. But over time, many organizations unintentionally optimize perceived accountability more than actual system understanding.

Most of these mechanisms were introduced for understandable reasons. Previous incidents created fear. Leadership lost visibility. Trust deteriorated. Teams started compensating for instability by introducing more oversight.

The problem is that many organizations eventually optimize for perceived accountability more than actual system understanding.

A dashboard may create visibility into metrics while hiding operational uncertainty entirely. A detailed delivery process may create the appearance of predictability while work continues moving chaotically underneath. Extensive reporting may communicate activity volume without improving understanding of customer outcomes or systemic risk.

This is how the process slowly becomes theater. The process itself is not the issue here. Mature organizations often rely on extremely disciplined operational systems. The difference is that healthy controls improve visibility and decision quality, while unhealthy controls primarily provide psychological reassurance. This is the reason metrics become dangerous.

Bug counts, test execution totals, approval compliance and throughput numbers can create the illusion of improvement while the underlying system remains unstable.

Many organizations use heavy separation of responsibilities and divide ownership cleanly. Product owns requirements, engineering owns implementation, QA owns quality. But complex failures rarely respect organizational boundaries. When dealing with failures, hand-offs fragment context, delayed communication undermines accountability and silos destroy ownership.

Do not eliminate specialization. However, teams should own delivery together and care about it from conception to deployment. Make sure that the people shaping decisions remain connected to the operational consequences of those decisions.

Shared ownership does not mean everybody does everything. It means the system does not allow responsibility to disappear into departmental borders.


Protect Capacity for Improvement

One of the most common failure patterns is that organizations consume every efficiency gain immediately. A team becomes slightly more stable, so additional delivery pressure appears instantly. Some operational breathing room emerges, so new initiatives get added immediately.

The organization improves just enough to become overloaded again.

Overloaded systems lose the ability to reason clearly about themselves. Every available cognitive resource gets redirected toward immediate operational survival.

Organizations trapped in constant firefighting often believe they cannot afford to stop fighting fires. In reality, they usually cannot afford to keep fighting them indefinitely. Improvement capacity is not organizational luxury. It is an ongoing maintenance cost for preventing operational debt from compounding faster than the business can absorb it. 

Protect capacity for continuous improvement. Stable organizations are not organizations without incidents, mistakes or technical constraints. They are organizations that continuously reduce the time to learn and adapt.

That process never ends.

There is no final state where the organization becomes “done” improving. The goal is not perfection. The goal is preventing the organization from drifting so far away from operational reality that firefighting becomes its primary mode of existence.


The next steps

If you want to start somewhere next week, do not start with a transformation program.

Pick one recent incident, delayed feature or failed initiative and trace it honestly from beginning to end.

Map who made the important decisions. Identify where assumptions appeared. Find where work stopped moving. Look for where feedback arrived too late to matter cheaply. Observe where ownership fragmented and where the organization compensated with process instead of understanding.

Do not optimize anything yet.

First, make reality transparently visible.

Organizations rarely escape firefighting through one massive intervention. Most begin improving when they finally stop treating recurring instability as isolated bad luck and start understanding the system continuously reproducing it.

 
 
bottom of page