The quick download
Modern incident response depends on giving responders the context to diagnose accurately, not simply reducing the time to close a ticket.
-
MTTR can obscure the diagnostic work that determines whether teams understand the failure and prevent recurrence
-
Mature incident management brings together telemetry, change history, dependencies, prior incidents, and operational knowledge at the point of triage
-
Observability becomes more valuable when it supplies that context directly into incident workflows, giving Edwin AI a stronger foundation for correlation, diagnosis, and guided action
Incident response has spent years compressing time.
Monitoring detects a problem sooner. Paging reaches the right engineer sooner. Automation removes another handoff. MTTR falls, and the program looks healthier.
Forrester’s Incident Management Has Outgrown Its Playbook exposes a weakness in that model. A lower resolution time says little about the quality of the diagnosis that produced it. It cannot tell you whether responders understood the failure, whether they restored service through a durable fix, or whether the same condition will return under a different ticket number. Forrester argues that combining diagnosis and resolution into one metric can obscure the part of an incident where teams struggle most: establishing what happened and why.
For modern ITOps teams, that distinction matters. The scarce resource during a complex incident is often usable context.
We optimized the wrong variable
Mean Time to Resolution (MTTR) became useful because time is easy to measure and expensive to waste. The problem begins when the metric becomes a proxy for response quality.
Consider two incidents closed in 40 minutes. In one, the team identifies the affected dependency, connects it to a recent change, restores service, and records the failure pattern for future use. In the other, engineers restart a service, symptoms disappear, and the ticket closes while the initiating condition remains poorly understood.
The metric treats those outcomes as equivalent.
Forrester recommends a broader view of incident performance, including recurrence, recovery capability, business impact, and the quality of AI-assisted decisions. Those measures ask whether the response improved the system rather than simply ending the event.
That makes diagnosis more important than a single MTTR number suggests. Every minute spent searching for the relevant change record, tracing an upstream dependency, or reconstructing a similar incident from six months ago belongs to the response timeline, yet each delay has a different remedy.
Faster paging cannot solve a missing dependency map.
Incident management is a knowledge problem
Forrester describes mature incident management as increasingly knowledge-centric because modern incidents are more likely to cross systems and require historical or architectural context. Responders need enough information to form a sound hypothesis before they act, and the quality of that information affects both resolution accuracy and recurrence.
The knowledge frequently exists. Access is the problem. For example, a senior product executive interviewed for the report described records with substantial gaps because teams lack the time or incentive to document everything they learn. Another practitioner described knowledge distributed across ServiceNow, Teams folders, and personal laptops.
That fragmentation becomes expensive during an incident. A responder may need telemetry from one system, a service relationship from another, a recent deployment record from somewhere else, and a previous incident buried in the ITSM platform. Each source contributes part of the answer while leaving the engineer responsible for assembling the whole.
The records have holes
The obvious response would be better documentation. The report suggests a more demanding standard.
“Our records have a lot of holes. We have been asking humans to capture data, and they don’t want to do it and don’t have time to do it.”
Knowledge has to be available during the incident, in context, without relying on a responder to know where every useful artifact lives. Forrester describes knowledge orchestration as the ability to bring together service desk records, observability data, change history, previous incidents, and architectural documentation, then retrieve relevant material as part of triage.
Taken together, it all moves knowledge management closer to the mechanics of incident response.
A runbook written six months ago has limited value if the responder cannot find it. A previous incident becomes more useful when the system recognizes the same pattern and surfaces it automatically. A topology map matters most when it shows which upstream change could explain several downstream symptoms at once.
AI can help with this retrieval and synthesis because the task involves finding relationships across large volumes of operational data. Its value depends on the quality of the context underneath it.
Resilience requires planning for partial failure
Forrester connects the knowledge problem to a broader principle: mature incident management assumes that some failures will exceed the organization’s ability to prevent them.
Dependency chains now extend through cloud providers, SaaS platforms, networks, software components, and infrastructure owned by other organizations. Perfect availability becomes an increasingly poor design assumption under those conditions.
The report recommends graceful degradation: constrain the blast radius, preserve critical functionality where possible, and build recovery procedures around plausible failure scenarios. That can include service decomposition, redundancy across regional failures, manual fallback procedures, and recovery exercises that account for automation becoming unavailable or contributing to the incident.
This changes the role of incident context again. Recovery depends on knowing what is affected, which dependencies matter, and which parts of the service can be restored safely while the underlying failure remains active.
Detection tells you something has changed. Recovery requires a model of the system.
Observability becomes the incident context engine
That model increasingly lives in observability.
Forrester argues that observability data should move directly into incident-response workflows so responders can see system health, recent changes, and dependency information while they investigate. The report describes this convergence as a significant change in the role observability plays within operations.
An observability platform that produces another alert contributes additional evidence. An observability platform that connects that alert to the affected service, a recent change, related signals, and the relevant dependency chain contributes to diagnosis.
Forrester gets concrete in its recommendations. When an incident opens, the responding team should receive current health information, recent change activity, upstream and downstream dependency status, and historical matches from previous incidents.
That is also the clearest connection to Edwin AI.
Edwin AI correlates signals across hybrid environments and applies topology, change history, and service context to help teams understand how an incident fits together. Rather than leaving responders to reconcile related alerts manually, it can group signals around a likely incident, identify probable cause, surface affected dependencies, and provide a more complete starting point for investigation.
The benefit is greater diagnostic leverage. An engineer begins with a stronger hypothesis and spends less of the incident reconstructing information the organization already possesses.
AI-assisted response becomes substantially more useful when the underlying operational context is connected. Forrester makes the same sequencing argument: integrating observability into incident workflows creates a foundation for more sophisticated AI augmentation.
What to change over the next 12 months
Forrester recommends closing the gap between observability and incident response as one of the most useful moves organizations can make over the next year.
A practical assessment can start with a recent major incident. Examine how long the team spent establishing cause and impact, then identify where the required information lived. Look closely at the handoffs between monitoring, change records, dependency data, incident history, and the tools responders used to coordinate.
Those gaps tell you more than the final MTTR number.
The next step is to bring the relevant context into the incident workflow automatically, so responders begin with evidence that would otherwise take minutes or hours to assemble. That foundation improves human investigation today and gives AI systems a richer operating context for correlation, diagnosis, and eventually governed action.
The question for mature teams is therefore narrower than “How do we respond faster?” They need to know how much of the response timeline is consumed by finding and connecting information that already exists.
That is where incident management becomes a knowledge discipline, and where observability starts carrying more of the response itself.
See how IT leaders are adapting for the AI era
Hear from Forrester analyst Julie Mohr on how modern failures are changing incident response, where agentic AI fits, and how teams can define safe boundaries for autonomous action.




