The quick download:
The infrastructure looked healthy at first, but synthetic tests showed the public API slowing and failing from several regions. Putting both data sets on the same timeline revealed the backend problem behind it.
-
Internal network telemetry from LogicMonitor Envision identified traffic drops and routing instability across backbone links, but could not confirm customer impact on its own.
-
Synthetic tests confirmed response time spikes of several seconds, with server errors across three global regions, before hard downtime occurred.
-
Aligning both signals showed that backend queuing, caused by internal network instability, was producing the elevated API latency customers were experiencing.
-
Monitor your most important customer-facing services from inside and outside the environment, so teams can quickly distinguish an infrastructure issue from a business-impacting customer problem.
When a customer-facing service degrades, infrastructure teams need answers to two questions: what changed inside the environment, and what are customers experiencing outside it?
LogicMonitor’s acquisition of Catchpoint brought these two perspectives together, extending its infrastructure monitoring capabilities with Internet Performance Monitoring, including synthetic testing. The world’s largest global payments platform relied on both LogicMonitor and Catchpoint to monitor its environment. When its internal dashboards did not initially show an obvious service failure, external testing told a different story. Its public authentication API was slowing down and intermittently returning errors across multiple regions as customer complaints rose.
For a payments business, issues like this can disrupt authentication and transaction journeys, increase support demand, and put customer trust at risk. By comparing internal network telemetry with external synthetic monitoring, the team identified the source of the degradation, confirmed that users were affected, and focused its response on the right problem. Here’s how they did it.
Reading the Internal and External Signals Together
LogicMonitor Envision’s network data showed abrupt drops and rebounds across core and backbone links. The pattern looked like instability rather than normal congestion.
That data made it clear something was wrong with the network. On its own, it did not confirm whether the issue was reaching customers or whether it required escalation beyond routine network remediation.

LogicMonitor Envision’s backbone utilization view during the incident window. The shift from steady congestion to sudden drops and rebounds, visible across most links at once, is what flagged instability rather than a routine traffic pattern.
Internet Performance Monitoring data told a different part of the story. Tests against the public OAuth endpoint recorded server wait times and total response times several seconds above baseline, followed by intermittent server errors.

Internet Performance Monitoring view of the same window: response time and wait time climbing well above baseline, an Experience Score drop to 63, and downtime confirmed from three global test locations, independent evidence that real users were affected.
The external test also ruled things out. DNS, connect, and SSL timings stayed normal throughout the incident, which meant the delay wasn’t a client-side or ISP problem. That single detail narrowed the investigation considerably. Whatever was happening, it was happening on the backend, not in the delivery path to the user.
The network data explained where to look. The OAuth tests established that the disruption had reached a public-facing service. Once the two incident windows were compared, the investigation narrowed quickly to the backend.
A Practical Triage Workflow
The path from first alert to root cause followed a workflow that applies well beyond this one incident:
1. Establish the internal baseline. Confirm whether infrastructure and network telemetry show anomalous behavior, and characterize it: sudden shifts, gradual degradation, or isolated spikes.
2. Check external experience independently. Run or review synthetic tests against the customer-facing service to see whether real users are affected, and how severely.
3. Rule out the delivery path. Use DNS, connect, and SSL timing data to determine whether the problem originates outside the organization’s infrastructure, such as with an ISP or CDN, or inside it.
4. Compare the timelines. Look for overlap between the infrastructure event and the change in API or transaction performance.
5. Use the combined evidence to narrow the investigation and decide how seriously to treat the incident.

Every individual test result, plotted by timestamp across the test window. The cluster of failed tests (red) inside the highlighted window is what separates a genuine incident from ordinary test-to-test noise.
For the payments platform, step five was the turning point. Aligning the network instability window against the API latency window, roughly 90 minutes of overlapping activity, showed that backend queuing, caused by internal network instability, was producing delayed service-to-service calls, which reached customers as elevated latency and intermittent server errors.

One representative test run from inside the incident window. The wait-time breakdown is what ruled out DNS, connect, and SSL as the cause, and pointed the investigation straight at the backend.
One representative test endpoint was enough to confirm platform-level degradation. The team could understand the incident without instrumenting every API. A representative endpoint, paired with the network telemetry, was enough to show that the infrastructure instability and customer-facing degradation were occurring at the same time.
How the Monitoring Pieces Fit Together
Infrastructure monitoring gave the team the network context behind the incident. Internet Performance Monitoring showed what was happening at the public OAuth endpoint from outside the company’s environment, including the response-time increase, server errors, and geographic spread. Taken together, the two capabilities connected an unstable network condition to a degraded customer-facing service.
Today, these capabilities are becoming easier to use together. Customers can now move between the two platforms with a more seamless authentication experience. Synthetics and Internet Performance Monitoring alert data can be routed to Edwin AI through REST APIs and webhooks. Further integration work is under way, with the aim of making this kind of correlation easier to do in one place.
Apply the Same Approach to Your Own Services
Internal telemetry can tell you that the environment has changed. Internet Performance Monitoring can show whether that change is degrading the experience customers receive. By correlating both views on one incident timeline, teams can eliminate false leads, prioritize the right escalation, and reduce the time it takes to move from alert to action.
Start with one API or transaction that customers rely on. If it slows down or fails, can your team see the infrastructure event behind it and the effect on the service at the same time?
If not, your investigation will begin with only part of the picture.
See the full picture, inside and out
Connect infrastructure and digital experience monitoring to detect issues sooner, troubleshoot faster, and reduce their impact on customers.



