DOP 320: Why Dashboards Alone Are Not Enough for Incident Response

Episode 320

Show Notes

#320: In this episode, Darin and Viktor are joined by Jim Hirschauer, Head of Product Marketing at Xurrent, for a deep dive into the realities of incident management in today’s complex IT environments. While dashboards and monitoring tools have become ubiquitous in operations centers, the panel discusses why these visualizations alone often fall short when it comes to actually resolving incidents.

Drawing on decades of experience, they share stories of war rooms, recurring outages, and the persistent challenges that technology alone can’t solve. The conversation highlights the critical role of human expertise, communication, and organizational culture in bridging the gap between raw data and effective action.

Whether you’re an IT leader, SRE, or anyone responsible for uptime, this episode offers practical insights into what it really takes to keep systems running smoothly.

Frequently Asked Questions

Has incident management actually improved in twenty years?

Jim Hirschauer says on DevOps Paradox episode 320 that the stories he hears from people in IT operations today match his own from fifteen to twenty years ago: war rooms, disconnects, and silos nobody has managed to break down. When he asks conference audiences whether the same incidents keep recurring, around three quarters of the room raises a hand. Viktor Farcic counters that the old problems were solved and the new ones are simply harder.

Why do the same incidents keep happening?

Jim Hirschauer explains on DevOps Paradox episode 320 that after service is restored and everyone celebrates, the team returns to an email queue and a backlog that grew during the war room. The postmortem identified real causes and follow-up work, and that work competes with catching up. He describes it as neglect driven by being overworked rather than by not knowing what to fix.

What are the four phases of incident management?

Jim Hirschauer breaks it down on DevOps Paradox episode 320 into pre-incident work such as runbooks, detection through observability and alerting, resolution, and post-incident follow-up. His argument is that the industry has spent years getting good at the two middle steps and neglected both bookends. Runbooks start sparse because development runs late, then go stale as the application changes.

Why are dashboards not enough for incident response?

Viktor Farcic argues on DevOps Paradox episode 320 that a dashboard represents things you already knew could happen, so it fails precisely on the incident that has never occurred before. Jim Hirschauer describes hitting this himself: he had CPU, disk and memory charts he felt good about, and during a major incident realised he could not say what normal looked like without digging back through history.

Should you use static thresholds for alerting?

Jim Hirschauer says on DevOps Paradox episode 320 that he dislikes static thresholds in most cases, with disk utilisation as the exception since filling up is genuinely predictable. He implemented dynamic thresholding instead, building bands of normal behaviour and alerting when something moves far enough outside. Crossing that line means pay attention rather than something is broken, and combinations of such signals carry more meaning than any one.

Why is it hard to keep executives informed during an incident?

Viktor Farcic diagnoses it on DevOps Paradox episode 320 as a language problem. Tools are built for the person who buys them and speak only that role’s language, so anyone else has no way to read what they show and comes to ask instead. He once wrote in an incident report that repeated requests for status were the reason resolution took so long.

What is the DevOps Paradox podcast?

DevOps Paradox is a weekly podcast co-hosted by Darin Pope and Viktor Farcic, covering DevOps, platform engineering, and modern software delivery. Episode 320 brings in Jim Hirschauer of Xurrent to discuss why incident management looks much as it did twenty years ago, which phases teams neglect, and why a wall of dashboards gives false confidence. Every episode page carries the audio, the video, and a full transcript.

Topics

Share and Download

Guests

Jim Hirschauer

Jim Hirschauer

Jim Hirschauer is the Head of Product Marketing for Xurrent, the modern service management platform. Jim spent 15 years in IT operations in various roles ranging from Systems Administrator to IT Architect. For the past few years, Jim has been focused on improving process workflows and automations that deliver consistent and reliable results. With the recent advances in AI technology, Jim has been exploring how enterprises can gain real-world efficiencies using secure AI technologies. Jim is an award-winning speaker having been the recipient of the J. William Mullen Award for both technical excellence and an engaging presentation style.

Hosts

Viktor Farcic

Viktor Farcic

Viktor Farcic is a member of the Google Developer Experts and Docker Captains groups, and published author.

His big passions are DevOps, Containers, Kubernetes, Microservices, Continuous Integration, Delivery and Deployment (CI/CD) and Test-Driven Development (TDD).

He often speaks at community gatherings and conferences.

He has published DevOps Paradox and Test-Driven Java Development.

His random thoughts and tutorials can be found in his blog The DevOps Toolkit.