top of page

[1/2] Firefighting Fallacy: When Fighting Fires becomes the System

1 day ago
6 min read

You have been there before. Another production incident happened. The customer reached out and this time they were not even angry anymore. Just tired of it all. A little resigned. At least it was nothing catastrophic this time. Still bad enough to trigger the usual meeting. Product managers are frustrated. Engineers are exhausted. Leadership wants confidence restored before the next release cycle.


Photo by Jay Heike
Photo by Jay Heike

The discussion quickly converges on familiar solutions.

We need stricter deployment gates.

We need broader regression coverage.

We need more automated tests.

We need another approval step before release.

Nobody in the room is irrational. Nobody is shouting. Nobody is throwing tantrums. The proposals are sensible. Responsible, even. And yet many organizations unknowingly reinforce the exact dynamics that created the instability in the first place. Because most quality problems do not originate where organizations usually try to solve them.


Too little, too late

When organizations experience rising defect counts, they often interpret the problem as insufficient verification. The assumption is straightforward: if more issues escape into production, then the organization must improve its ability to catch them before release.

So the response becomes reactive. They create more validation gates, ask for more approvals. Another reporting layer is added. Maybe a bigger regression suite will help. Anything that feels like additional control.

The problem in my experience is that reactive controls rarely address the conditions that created the defects in the first place. In many cases, the defect itself is only the final visible symptom of a much earlier organizational failure:


  • unclear priorities

  • disconnected decision making

  • ambiguous ownership

  • fragmented communication

  • oversized delivery batches

  • slow feedback cycles

  • assumptions staying invisible for too long


These problems often feel relatively vague and hard to pinpoint. They are not easily measurable, and frankly, often attack the ego, as they can feel like personal failure, even if they are in fact systemic. That is why it is easy to miss them, or in the worst case to have the tendency to dismiss them altogether. By the time testing discovers the issue, the organization has often already spent weeks reinforcing the wrong assumptions and it is not willing to let go.


The Firefighting loop

This is where many organizations eventually enter a firefighting loop. It goes something like this:


  1. A production issue appears. 

  2. Controls increase.

  3. Delivery slows down.

  4. Feedback arrives later.

  5. Learning becomes slower and more expensive.

  6. Coordination overhead grows.

  7. The organization becomes less adaptable.

  8. Another production issue appears.


Repeat ad nauseam.

The intention here is risk reduction. The actual outcome is bureaucracy and quality theater. The most dangerous part is that at first glance the process feels responsible the entire time.


Verification is not enough

This article is not an argument against testing. In fact automated testing is valuable. End to end regression matters. Validation matters. Exploratory testing remains one of the best ways to generate knowledge about a system. It is an argument against pretending testing can compensate for a dysfunctional delivery system.

Late verification cannot compensate for early systemic dysfunction. You cannot sustainably improve quality in a continuously misaligned environment. In order to scale quality successfully, you need to be able to learn about your product and systems faster than their instability can spread. You need to shorten feedback loops. You need to reduce ambiguity earlier. Engineers get faster answers. PMs spend less time clarifying tickets halfway through implementation. Smaller releases mean rollbacks stop becoming multi-team recovery efforts.

Shortly put, organizations that scale quality successfully optimize for learning speed instead of perceived simplicity.


Signs of improvement

In my career, one of the clearest examples of this came from an initiative focused primarily on delivery performance, not testing.

At the time, our company was under increasing pressure to deliver faster while maintaining stability. The intuitive response from a quality perspective would have been straightforward: add more controls around releases so fewer issues can escape into production. Manual approval gates would have slowed delivery too much, so naturally the expectation was to expand automated coverage and strengthen deployment verification.

I did not do any of that, instead I decided to completely remove our test management tool. Sounds crazy, right? But I did not do that because testing did not matter, I did it because the tool optimized documentation and reporting of testing activity far better than it improved learning, feedback quality, or shared understanding.

Then, we focused heavily on reducing lead time. Teams started splitting work into smaller increments. Delivery batches became smaller. Prioritization became more aggressive. People were grumpy at first. Some openly said we were just gaming arbitrary delivery metrics.

At first glance, none of this sounds like a quality initiative, but the effects appeared almost immediately.

There were fewer misunderstandings between product managers and engineers. Tasks described intended behavior more clearly. There was less back-and-forth clarification. The time between grooming and implementation became shorter. Changes became smaller and more focused. Incorrect assumptions became visible earlier. Smaller releases reduced the blast radius of mistakes. Simply put, we started catching bad assumptions while they were still cheap. Over time, deployments scaled significantly while incident levels remained stable and severe incidents disappeared entirely.

Importantly, this was not achieved by adding more approvals or expanding bureaucratic oversight. In fact, many traditional “quality control” mechanisms became less important over time because the organization was reducing instability at the source instead of compensating for it downstream.


The Fix it later illusion

One of the more counterintuitive lessons from this experience was realizing how misleading many quality metrics can become.

For example, organizations often interpret lower bug counts as evidence of higher quality. But bug counts are observational metrics, not outcome metrics. Finding fewer bugs does not automatically mean the product improved. It may simply mean the organization became worse at discovering problems.

Even when bug counts are accurate, they still measure only a narrow category of failure: deviations from expected behavior. They tell us very little about whether the product meaningfully helps customers, whether the right problem is being solved, whether the product direction is strategically sound or if the organization is learning effectively. Many expensive product failures technically work exactly as designed. Verification passed. The problem is that the wrong thing is being built.

This is also why purely reactive organizations struggle to transform themselves. Firefighting consumes capacity. Teams become trapped maintaining increasingly heavy coordination structures:


  • more approval chains

  • excessive regression burden

  • reporting overhead

  • long stabilization phases

  • never ending hand-offs

  • fragmented ownership

  • expanding process layers

  • policies and processes everywhere


Because the organization spends most of its energy responding to instability, it rarely creates enough space to improve the underlying system producing the instability. This creates a dangerous illusion: “We will improve the process later once things calm down.” The argument is that they do not have the capacity now and once they have it, they can fix things for real. But later eventually becomes today again. And the organization is still fighting fires.


Escaping the loop

So how do organizations escape this loop?

Quality and productivity are not opposing forces. Poor systemic quality eventually destroys productivity through rework, coordination overhead, delayed releases, unclear ownership, and slower learning. Eventually the friction inside the organization leaks directly into the product. Product quality is deeply influenced by the quality of the organizational system producing it. Communication structures shape software behavior. Incentives shape decision quality. Feedback timing shapes adaptability. A healthy system produces good quality products at pace.

The practical response is not to abandon testing or to remove all controls, it is to simplify the system enough to create room for learning again.

Identify the smallest possible number of activities in your organization that can prevent the majority of recurring problems. Make sure they are owned by the whole team. Is it having a good Definition of Ready? Have it. Is it Mob Testing session before deployment? Do it. Is it getting rid of the Once-a-month release structure? Release at will. Make sure everybody is invested in the change. Reduce the unnecessary coordination overhead. Focus on smaller delivery increments. Improve visibility of intent and risk earlier in the lifecycle.

The goal is to move testing closer to where decisions are made.

Because stable quality is rarely created by organizations that become exceptionally good at reacting to failure. It emerges from organizations that stop continuously recreating the conditions producing the next fire.

 
 
bottom of page