The countdown to Elevate 2026 is on. Join us in Chicago, London, or Sydney.

Register here

Partners

Docs

LM Academy

LM Community

Platform

Solutions

Pricing

Resources

Company

Platform
  • Infrastructure
  • Cloud & Multi-Cloud
  • Log Management
  • Edwin AI
Solution
  • Automation
  • Tool Consolidation
  • Reduce MTTR
  • Cost Optimization
Industry
  • Healthcare
  • Financial Services
  • Public Sector
  • MSP
Role
  • CIO
  • ITOps
  • CloudOps
  • AIOps
There is no result.
Try it free

14-day access to the full LogicMonitor platform

Explore Platform

One platform, one system for observability, intelligence, and action.

Agentic AIOps

Infrastructure Observability

Cloud Observability

Internet Performance Monitoring

Digital Experience Monitoring

Log Management

3000+ Integrations
3000+ Integrations

Agentic AIOps Overview

Autonomously detect, diagnose, and resolve issues across your environment.

Meet Edwin AI

Turn fragmented cross-domain event noise into explainable, guided action.

AI Agent

Deploy specialized AI agents to handle investigation across the incident lifecycle.

Event Intelligence

Compress raw alert storms into high-fidelity, prioritized insights.

AI Automation

Execute governed, closed-loop remediation across automation playbooks.

ITOps Context Graph

NEW

Unify topology, telemetry, and changes into an AI-ready context layer.

MCP

NEW

Establish traceable, secure governance boundaries for AI tool integrations.

Infrastructure Observability Overview

Full visibility across your entire hybrid estate to eliminate tool sprawl.

Network Monitoring

Accelerate time to innocence with deep network path and device visibility.

Server Monitoring

Track server health, OS metrics, and resource utilization across environments.

Remote Monitoring

Monitor distributed endpoints, branch networks, and remote facility health.

VM Monitoring

Maximize hypervisor performance and streamline compute capacity planning.

SD-WAN Monitoring

Keep multi-site cloud networks connected with real-time edge visibility.

Database Monitoring

Pinpoint database query bottlenecks to keep business applications fast.

Configuration Monitoring

Minimize change failure rates by tracking device configuration drift.

Storage Monitoring

Track SAN/NAS arrays, IOPS bottlenecks, and storage capacity trends.

Cloud Observability Overview

Multi-cloud and hybrid environments unified into a single operational pane.

Container Monitoring

Automated, real-time visibility for Kubernetes and ephemeral microservices.

AWS Monitoring

Track AWS services, scaling, and costs alongside on-premises data.

Google Cloud Monitoring

Monitor native GCP infrastructure, compute, and serverless resources.

Azure Monitoring

Comprehensive visibility into Azure environments, gateways, and workloads.

AI Monitoring

Track LLM infrastructure, GPU utilization, and AI application stack health.

Oracle Cloud Monitoring

Track OCI native compute, enterprise databases, and cloud storage.

SaaS Monitoring

Validate availability and workforce productivity for critical SaaS apps.

Cloud Cost Optimization

Optimize cloud spend, maintain performance, and control budgets.

Internet Performance Monitoring Overview

Understand performance across the full stack wherever users depend on it.

Internet Health

NEW

Use global vantage points to independently validate internet outages.

Real User Monitoring

NEW

Capture actual customer journeys and frontend performance in real time.

Synthetic Monitoring

NEW

Emulate user transactions and SaaS workflows to catch problems early.

Endpoint Monitoring

NEW

Diagnose remote workforce digital experience across devices and networks.

Digital Experience Monitoring

See every dependency, regardless of ownership or location.

Website Monitoring

Protect revenue journeys with proactive synthetic checks and uptime tracking.

CDN Monitoring

NEW

Audit edge performance and latency variance across your CDN providers.

API Monitoring

NEW

Test endpoints and third-party API reliability for critical app integrations.

Application Performance Monitoring

Connect code execution and traces directly to infrastructure health.

DNS Monitoring

NEW

Speed up time-to-innocence by tracking global nameserver resolution times.

DevOps Lifecycle Monitoring

NEW

Protect release velocity by validating dependencies during deployments.

BGP Monitoring

NEW

Trace global routing changes and path leaks to secure internet reachability.

Log Management Overview

Centralize and correlate log data to resolve incidents before they escalate.

Log Analytics & Intelligence

Correlate contextual log data with metrics to speed up root-cause analysis.

WebPageTest Web Performance

Test, compare, and optimize website speed, Core Web Vitals, and performance across real devices and global locations.

Learn more
Explore Solutions

Proactively manage modern hybrid environments with predictive insights, intelligent automation, and full-stack observability.

By Business Outcome

By Role

By Industry

Professional Services

Autonomous IT

Predictive, autonomous IT built

for resilience.

Automation

Eliminate operational toil with safe, policy-governed remediation workflows.

Modernization and Transformation

Accelerate complex technology transitions while protecting core enterprise resilience.

Cloud Migration

Maintain workload performance throughout migration.

Tool Consolidation

Reduce licensing costs and silos by replacing fragmented monitoring tools.

Cost Optimization

Lower your total cost-to-serve by finding cloud waste and underused resources.

Operational Efficiency

Maximize team capacity by reducing alert storms and shift-handoff friction.

Reduce MTTR

Shorten war-rooms by surfacing topology-aware probable cause in mins.

Network Reachability

NEW

Independently audit external BGP, ISP, and SaaS provider connectivity boundaries.

Edge Deployment Optimization

NEW

Monitor SLOs, compare providers, and validate cloud and edge delivery.

Web Performance Optimization

NEW

Maximize digital checkout conversions by tracking global frontend latency metrics.

Application Resilience

NEW

Safeguard business services against transaction failures and costly downtime.

Workforce Productivity

NEW

Troubleshoot remote hardware and network issues to protect productivity.

CIO

Maximize enterprise resilience and align AI investments to measurable business ROI.

AIOps

Compress cross-domain event noise into explainable, automated ops leverage.

DevOps

Speed up releases by protecting engineering roadmaps from toil.

ITOps

Standardize incident response to reduce alert fatigue and after-hours work.

CloudOps

Unify multi-cloud visibility to optimize costs and track hybrid blast radius.

Healthcare

Protect continuity of care and EHR availability across clinical workflows.

Public Sector

Ensure mission continuity and audit readiness for citizen-facing services.

MSP

Protect service margins and scale ops using multi-tenant, AI-assisted triage.

Retail & E-commerce

Safeguard peak retail campaigns, POS uptime, and digital customer journeys.

Technology

Protect customer trust and engineering velocity with SLA-driven visibility.

Hospitality

Deliver frictionless guest experiences and keep booking engines online.

Education

Maintain always-on student portals, learning platforms, and campus networks.

Manufacturing

Prevent production downtime by unifying IT, OT-adjacent, and edge systems.

Financial Services

Secure transaction trust and meet strict resilience compliance requirements.

Why LogicMonitor?

Discover why leading IT teams trust us to unify hybrid observability and eliminate tool sprawl.

Learn more
Explore Resources

Check out our resource library for IT pros, featuring expert guides, strategies, and insights for smarter, AI-driven operations.

Resources

Upcoming Events

Platform Help

Blog

Insights and advice from the experts on all things observability and AI.

Case Studies

See what real users have to say about the LogicMonitor platform.

Webinars

Live and on-demand learning, all in one place.

IT Guides

Learn from expert guides on the topics that matter most to IT teams.

How We Compare

See how our platform stacks up against other solutions.

Viee of a bridge over a river leading to Cologne cathedral rising against the skyline and a blue sky
CONFERENCE

Digital X Cologne

September 8, 2026

Cologne

CONFERENCE

SWORD Day

September 17, 2026

Geneva

View all events

Join us at innovation-focused conferences, tech talks, webinars, and other events.

Support Docs

Access product docs, release notes, and support resources.

LM Community

Join the community to learn from peers, ask questions, and connect with experts.

Customer Education

Learn more about our platform through resources and live trainings.

2026 The Year of Autonomous IT

NEW

Discover the trends, benchmarks, and strategies driving the industry shift to Autonomous IT.

Read the report
About LogicMonitor

Our observability platform proactively delivers the insights and automation CIOs need to accelerate innovation.

Leadership

Meet the leaders building the future of observability and AI.

Our Customers

See the proof of how IT teams win with LogicMonitor.

Careers

Find job openings and learn about our employee benefits.

Newsroom

Stay current with our latest mentions, press releases, and events.

Culture

NEW

Join a collaborative, values-driven culture built on innovation and growth.

Security

Purpose-built security for the hybrid observability and AI era.

Contact & Locations

Connect with our experts to explore AI-powered observability solutions.

Sustainability

Our commitment to the environment and the people in it.

The countdown to Elevate 2026 is on. Join us in Chicago, London, or Sydney.

Register here
Try it free

Platform

Explore Platform

One platform, one system for observability, intelligence, and action.

Agentic AIOps

Infrastructure Observability

Cloud Observability

Internet Performance Monitoring

Digital Experience Monitoring

Log Management

3000+ Integrations

WebPageTest Web Performance

Test, compare, and optimize website speed, Core Web Vitals, and performance across real devices and global locations.

Solutions

Explore Solutions

Proactively manage modern hybrid environments with predictive insights, intelligent automation, and full-stack observability.

By Business Outcome

By Role

By Industry

Professional Services

Why LogicMonitor?

Discover why leading IT teams trust us to unify hybrid observability and eliminate tool sprawl.

Pricing

Resources

Explore Resources

Check out our resource library for IT pros, featuring expert guides, strategies, and insights for smarter, AI-driven operations.

Resources

Upcoming Events

Platform Help

NEW

2026 The Year of Autonomous IT

Discover the trends, benchmarks, and strategies driving the industry shift to Autonomous IT.

Company

About LogicMonitor

Our observability platform proactively delivers the insights and automation CIOs need to accelerate innovation.

Leadership

Meet the leaders building the future of observability and AI.

Careers

Find job openings and learn about our employee benefits.

Culture

NEW

Join a collaborative, values-driven culture built on innovation and growth.

Contact & Locations

Connect with our experts to explore AI-powered observability solutions.

Our Customers

See the proof of how IT teams win with LogicMonitor.

Newsroom

Stay current with our latest mentions, press releases, and events.

Security

Purpose-built security for the hybrid observability and AI era.

Sustainability

Our commitment to the environment and the people in it.

Partners

Docs

LM Academy

LM Community

Agentic AIOps

Agentic AIOps Overview

Autonomously detect, diagnose, and resolve issues across your environment.

Meet Edwin AI

Turn fragmented cross-domain event noise into explainable, guided action.

AI Agent

Deploy specialized AI agents to handle investigation across the incident lifecycle.

Event Intelligence

Compress raw alert storms into high-fidelity, prioritized insights.

AI Automation

Execute governed, closed-loop remediation across automation playbooks.

ITOps Context Graph

NEW

Unify topology, telemetry, and changes into an AI-ready context layer.

MCP

NEW

Establish traceable, secure governance boundaries for AI tool integrations.

Infrastructure Observability

Infrastructure Observability Overview

Full visibility across your entire hybrid estate to eliminate tool sprawl.

Network Monitoring

Accelerate time to innocence with deep network path and device visibility.

Server Monitoring

Track server health, OS metrics, and resource utilization across environments.

Remote Monitoring

Monitor distributed endpoints, branch networks, and remote facility health.

VM Monitoring

Maximize hypervisor performance and streamline compute capacity planning.

SD-WAN Monitoring

Keep multi-site cloud networks connected with real-time edge visibility.

Database Monitoring

Pinpoint database query bottlenecks to keep business applications fast.

Configuration Monitoring

Minimize change failure rates by tracking device configuration drift.

Storage Monitoring

Track SAN/NAS arrays, IOPS bottlenecks, and storage capacity trends.

Cloud Observability

Cloud Observability Overview

Multi-cloud and hybrid environments unified into a single operational pane.

Container Monitoring

Automated, real-time visibility for Kubernetes and ephemeral microservices.

AWS Monitoring

Track AWS services, scaling, and costs alongside on-premises data.

Google Cloud Monitoring

Monitor native GCP infrastructure, compute, and serverless resources.

Azure Monitoring

Comprehensive visibility into Azure environments, gateways, and workloads.

AI Monitoring

Track LLM infrastructure, GPU utilization, and AI application stack health.

Oracle Cloud Monitoring

Track OCI native compute, enterprise databases, and cloud storage.

SaaS Monitoring

Validate availability and workforce productivity for critical SaaS apps.

Cloud Cost Optimization

Optimize cloud spend, maintain performance, and control budgets.

Internet Performance Monitoring

Internet Performance Monitoring Overview

Understand performance across the full stack wherever users depend on it.

Internet Health

NEW

Use global vantage points for independent validation of internet outages.

Real User Monitoring

NEW

Capture actual customer journeys and frontend performance in real time.

Synthetic Monitoring

NEW

Emulate user transactions and SaaS workflows to catch problems early.

Endpoint Monitoring

NEW

Diagnose remote workforce digital experience across devices and networks.

Digital Experience Monitoring

Digital Experience Monitoring

See every dependency, regardless of ownership or location.

Website Monitoring

Protect revenue journeys with proactive synthetic checks and uptime tracking.

CDN Monitoring

NEW

Audit edge performance and latency variance across your CDN providers.

API Monitoring

NEW

Test endpoints and third-party API reliability for critical app integrations.

Application Performance Monitoring

Connect code execution and traces directly to infrastructure health.

DNS Monitoring

NEW

Speed up time to innocence by tracking global nameserver resolution times.

DevOps Lifecycle Monitoring

NEW

Protect release velocity by validating dependencies during deployments.

BGP Monitoring

NEW

Trace global routing changes and path leaks to secure internet reachability.

Logs

Log Management Overview

Centralize and correlate log data to resolve incidents before they escalate.

Log Analytics & Intelligence

Correlate contextual log data with metrics to speed up root-cause analysis.

By Business Outcome

Autonomous IT

Predictive, autonomous IT built for resilience.

Automation

Eliminate repetitive operational toil with safe, policy-governed remediation workflows.

Modernization and Transformation

Accelerate complex technology transitions while protecting core enterprise resilience.

Cloud Migration

Maintain workload performance throughout migration.

Tool Consolidation

Reduce licensing costs and data silos by replacing fragmented monitoring tools.

Cost Optimization

Lower your total cost-to-serve by finding cloud waste and underused resources.

Operational Efficiency

Maximize team capacity by reducing alert storms and shift-handoff friction.

Reduce MTTR

Shorten war-room by surfacing topology-aware probable cause in mins.

Network Reachability

NEW

Independently audit external BGP, ISP, and SaaS provider connectivity boundaries.

Edge Deployment Optimization

NEW

Monitor SLOs, compare providers, and validate cloud and edge delivery.

Web Performance Optimization

NEW

Maximize digital checkout conversions by tracking global frontend latency metrics.

Application Resilience

NEW

Safeguard business services against transaction failures and costly downtime.

Workforce Productivity

NEW

Troubleshoot remote hardware and network issues to protect productivity.

By Role

CIO

Maximize enterprise resilience and align AI investments to measurable business ROI.

AIOps

Compress cross-domain event noise into explainable, automated ops leverage.

DevOps

Speed up releases by protecting engineering roadmaps from toil.

ITOps

Standardize incident response to reduce alert fatigue and after-hours work.

CloudOps

Unify multi-cloud visibility to optimize costs and track hybrid blast radius.

By Industry

Healthcare

Protect continuity of care and EHR availability across clinical workflows.

Public Sector

Ensure mission continuity and audit readiness for citizen-facing services.

MSP

Protect service margins and scale ops using multi-tenant, AI-assisted triage.

Retail & E-commerce

Safeguard peak retail campaigns, POS uptime, and digital customer journeys.

Technology

Protect customer trust and engineering velocity with SLA-driven visibility.

Hospitality

Deliver frictionless guest experiences and keep booking engines online.

Education

Maintain always-on student portals, learning platforms, and campus networks.

Manufacturing

Prevent production downtime by unifying IT, OT-adjacent, and edge systems.

Financial Services

Secure transaction trust and meet strict operational resilience compliance requirements.

Resources

Blog

Insights and advice from the experts on all things observability and AI.

Case Studies

See what real users have to say about the LogicMonitor platform.

Webinars

Live and on-demand learning, all in one place.

IT Guides

Learn from expert guides on the topics that matter most to IT teams.

How We Compare

See how our platform stacks up against other solutions.

Upcoming Events

Viee of a bridge over a river leading to Cologne cathedral rising against the skyline and a blue sky

CONFERENCE

Digital X Cologne

September 8, 2026

CONFERENCE

SWORD Day

September 17, 2026

View all events

Join us at innovation-focused conferences, tech talks, webinars, and other events.

Platform Help

Support Docs

Access product docs, release notes, and support resources.

LM Community

Join the community to learn from peers, ask questions, and connect with experts.

Customer Education

Learn more about our platform through resources and live trainings.

OBSERVABILITY VS MONITORING

Distributed Tracing

Trace requests across distributed services to find bottlenecks fast. Covers concepts, a Python tutorial with Jaeger, and production best practices.

9–13 minutes
June 24, 2026
Denton Chikura

IN THIS DEEP DIVE

CHAPTERS

    NEWSLETTER

    Subscribe to our newsletter

    Get the latest blogs, whitepapers, eGuides, and more straight into your inbox.

    SHARE

    The quick download:

    Distributed tracing turns complex multi-service request flows into a clear, followable path from root cause to resolution.

    • Traces assign unique identifiers to every request, linking spans across services so you can pinpoint exactly where bottlenecks occur without guessing which team owns the problem.

    • Head-based sampling keeps things simple for smaller environments; tail-based sampling gives you fine-grained control over what trace data you keep, at the cost of added infrastructure.

    • Open source tools like OpenTelemetry, Jaeger, and Zipkin make distributed tracing accessible without vendor lock-in, and the ecosystem is mature enough for production workloads.

    • Start with a clear sampling strategy and consistent instrumentation standards before scaling tracing across your organization.

    Distributed tracing

    Modern software teams manage dozens or hundreds of independent services, each built and deployed separately.

    Serving modern workloads, including richer pages, larger data volumes, and more concurrent users, pushes teams toward distributed architectures. The result is higher system reliability and improved scalability, both of which are common priorities for teams managing services at scale.

    Although microservices, containerization, cloud computing, and serverless have facilitated application development, they’ve also introduced blind spots that traditional monitoring alone can’t cover. Application components and services are now being built and maintained by different teams. As a result, there’s less insight into the broader user journey across an entire application, from the end user through the Internet path to the underlying infrastructure.

    The solution to effectively observe and monitor applications in a distributed world is to design a distributed tracing strategy. Distributed tracing extends your existing monitoring by adding an outside-in perspective: once you find a slowdown in the end-user experience, you can follow the user’s journey (the trace) to pinpoint the exact technical issue causing the delay. This ability to follow tracing across thousands of independent services lets engineering teams diagnose and resolve incidents faster.

    Traces, metrics, and logs each contribute distinct signals to root cause analysis (RCA), and you need all three working together with your infrastructure, cloud, and network monitoring to get a complete picture.

    What is distributed tracing?

    Distributed tracing is a method for observing system requests as they flow through an application, from the front end to back-end services and data stores. The concept implies that your software architecture is implemented in a distributed computing environment.

    What may start as a request to fetch a user’s cart total on a checkout page might involve a long chain of requests that include validating item availability, fetching payment options, and retrieving item prices before returning to the user. Many of these operations occur asynchronously and are abstracted from the end user. However, from an engineer’s perspective, they’re difficult to observe in log data because it’s impossible to know where a bottleneck occurred along the path.

    Distributed tracing works by following an origin request and tracking it along its path until it reaches the end. The tracing framework assigns a unique identifier to each request along the path and labels each individual piece (e.g., an API call or a query to a data store) as a span, with the origin API call labeled as the root span.

    Whenever a request enters a new service call, a child span is created, and each of its operations is identified and labeled. The root contains the entire execution path of the originating request, and each child can also serve as a root for its own nested spans. Each child’s unique identifier will include the original trace identifier, along with other relevant metadata, including error information and user information.

    Using distributed tracing, support engineers and DevOps teams can get a clear picture of where a bottleneck happened, irrespective of who owns which service. There’s no longer a need to figure out where it happened because the trace shows it clearly.

    It’s possible to drill down into span data in real time, giving support team members direct visibility into the end-user experience. When combined with infrastructure and network monitoring, these insights into the broader user journey, from the application layer to the underlying infrastructure, allow you to find, diagnose, and solve issues faster.

    The following diagram from OpenTracing can serve as an example.

    Figure 1: Illustration of a distributed trace (source)

    This visual representation shows how a request can have many child spans and various service calls along the way. With distributed tracing, you get a clear picture of the parent request’s journey, giving you visibility into which segment along the way you should dig deeper to find relevant information.

    One last important concept in distributed tracing is sampling, the process of filtering trace data according to a rule to decide which data to keep and which to discard.

    There are two main sampling approaches:

    1. Head-based sampling: Starting at the root, randomly choose which subsequent trace data to collect and persist to storage.
    2. Tail-based sampling: After the span has completed, retroactively decide which data to persist to storage.

    Both have advantages and disadvantages to consider, but at a high level, head-based sampling is simpler and can be implemented more quickly. It works well for companies that don’t operate in a highly distributed environment.

    Tail-based sampling gives more control over which data you retain because you decide after the trace completes. It does add operational complexity, and you’ll need infrastructure to buffer and process trace data before it’s stored.

    Netflix has documented how they built their distributed tracing infrastructure, which is a useful reference if you’re planning at scale.

    Key terms:

    • Request: Communication at the API layer enabling applications, services, and machines in general to share information with each other.
    • Trace: Relevant data about the requests in a system along some execution path.
    • Span: API calls that belong to a trace.
    • Root Span: The top span in a distributed trace.
    • Child span: Any nested span relative to the root.
    • Sampling: The process of deciding which trace data to store and which to discard. The two most popular are head-based and tail-based.

    Sample use case

    Let’s look at a simple implementation of OpenTracing’s tracing library in Python. We’ll follow its Python tutorial, which is hosted on GitHub. We’ll use a Dockerfile to run a visualization backend in Jaeger.

    Prerequisites

    1. This tutorial uses Python, pip, and virtualenv. Make sure you have these required dependencies before proceeding.
    2. You’ll need the Docker CLI set up locally to complete this tutorial. Docker is free and easy to set up for your local machine here, or if on a Mac, simply run:

    brew install –cask docker

    The following tutorial will use several dependency packages, namely Jaeger, to serve as the tracing back-end. Once you have Docker installed, run the following command to open a port at https://localhost:16686 that runs Jaeger’s all-in-one binary.

    docker run \
      --rm \
      -p 6831:6831/udp \
      -p 6832:6832/udp \
      -p 16686:16686 \
      jaegertracing/all-in-one:1.7 \
      --log-level=debug

    Now that that’s running, openhttps://localhost:16686.

    Figure 2: Screenshot of the Jaeger UI, Dashboard Overview

    You can see that Jaeger provides a clean UI for finding traces (left menu) and a search bar in the top menu. Once we start collecting some trace data, we’ll revisit this dashboard.

    5. Next, you’ll want to clone the OpenTracing tutorial repo on GitHub and then navigate to the Python directory. Then execute the commands below.

    git clone https://github.com/yurishkuro/opentracing-tutorial.git
    cd opentracing-tutorial/python
    virtualenv env
    source env/bin/activate
    pip install -r requirements.txt

    6. Now, let’s walk through creating a Tracer and implementing a simple trace. We’ll create a simple “hello world” service that takes an argument and prints “Hello, {arg}!”. Create two files, one called hello.py and one called _init_.py.

    git clone https://github.com/yurishkuro/opentracing-tutorial.git
    cd opentracing-tutorial/python
    virtualenv env
    source env/bin/activate
    pip install -r requirements.txt

    7. Open hello.py and first create the service.

    import sys
    import time
    
    def say_hello(hello_to):
           hello_str = 'Hello, %s!' % hello_to
           print(hello_str)
    
    assert len(sys.argv) == 2
    
    hello_to = sys.argv[1]
    say_hello(hello_to)

    8. This simple service will take a command-line argument, format it, and print the statement. To add tracing to it, use the tracing.py file from the python/lib directory to instantiate a Tracer. It contains code to interact with the Jaeger client. Recall that a trace is just a directed acyclic graph (DAG) of a bunch of spans. Each span in OpenTracing must contain, at a minimum, the name of the operation you’re tracing and the duration (start and end times). For simplicity, create a trace that consists of just one span.

    import sys
    import time
    from lib.tracing import init_tracer
    
    def say_hello(hello_to):
       with tracer.start_span('say-hello') as span:
           span.set_tag('hello-to', hello_to)
    
           hello_str = 'Hello, %s!' % hello_to
           span.log_kv({'event': 'string-format', 'value': hello_str})
    
           print(hello_str)
           span.log_kv({'event': 'println'})
    
    assert len(sys.argv) == 2
    
    tracer = init_tracer('hello-world') # initialize the tracer using the Jaeger client
    
    hello_to = sys.argv[1]
    say_hello(hello_to)
    
    # yield to IOLoop to flush the spans
    time.sleep(2)
    tracer.close()

    Let’s walk through what’s going on:

    • start_span() on the Tracer instance starts the span and takes the operation name as an argument. Each span must be finished by calling finish(), and the start and end timestamps will be captured for you. Note that we’re using Python’s context manager (the with keyword), which is the same thing as writing this:
    def say_hello(hello_to):
       span = tracer.start_span('say-hello')
       hello_str = 'Hello, %s!' % hello_to
       print(hello_str)
       span.finish()
    • set_tag() on the span instance will add a tag to the span.
    • log_kv() on the span instance will log values in a structured format. It’s important to be consistent in logging, as log aggregation systems need to process this information.
    • init_tracer() is the function defined in python/lib/tracer.py that allows us to connect to the Jaeger client. It takes an identifier in and will mark all child spans as originating from our service.
    • Lastly, we must introduce a delay into the program to allow enough time for the spans to flush to the Jaeger back-end before closing the tracer.
    pytho n -m lesson01.tutorial.hello I❤️tracing

    You’ll notice the following output:

    ...
    Initializing Jaeger Tracer with UDP reporter
    Using selector: KqueueSelector
    Using sampler ConstSampler(True)
    opentracing.tracer initialized to <jaeger_client.tracer.Tracer object at 0x10f619fd0>[app_name=hello-world]
    Hello, I❤️tracing!
    Reporting span 3c0e89bb739bc46:5f288b31bba91815:0:1 hello-world.say-hello
    Using selector: KqueueSelector

    Open the Jaeger UI to see the results.

    Figure 3: Screenshot of the Jaeger UI, Traces matching search query

    You’ll get a time-series graph and all the traces that match your query. Click on the trace to drill down into it.

    Figure 4: Screenshot of the Jaeger UI, an example trace

    This will show you how many services this trace interacted with, the total depth, and total spans. It will also show you the associated logs and metadata.

    If you want to challenge yourself further, follow the remaining tutorials on OpenTracing’s GitHub.

    Distributed tracing best practices

    A few practices make distributed tracing more reliable at scale.

    1. Remember Sampling

    Sampling is important because with hundreds to thousands of microservices, a large volume of trace data will be produced. To tackle this data storage problem and mitigate high costs and added complexity, choose a sampling strategy based on your environment’s scale, traffic patterns, and cost constraints.

    1. Follow Well-Established Standards

    Adopt established interoperability standards for trace context propagation. The W3C Trace Context specification defines how trace identifiers are passed across services and frameworks. Develop good practices early, so your tracing data is collected and organized in a way that scales and is easy to communicate across teams.

    1. Consult with Professionals Across Your Team

    Have discussions with your technical leaders and software architects to determine which path is feasible and makes sense for your use case. Some companies operating on a smaller scale don’t require many microservices or advanced observability techniques. Before adding distributed tracing infrastructure, confirm your traffic volume and operational needs justify the added complexity.

    Available tools

    Many tools, in addition to Jaeger and OpenTracing, are at your disposal for distributed tracing. There are great options in both the paid and unpaid categories.

    Several widely used open source tools cover the range of distributed tracing needs:

    1. Jaeger: As we’ve seen, Jaeger is an open source tool that provides distributed transaction monitoring, performance and latency optimization, RCA, service dependency analysis, and distributed context propagation. It provides an all-in-one executable to get you started fast with quick local testing.
    2. OpenCensus: A framework that originated with Google. It supports many major back-end service providers and provides metrics and tracing solutions.
    3. OpenTelemetry: OpenTelemetry is the modern successor to OpenTracing and OpenCensus, combining them into a single framework. It provides tools, SDKs, and APIs for gathering telemetry data across your application environment.
    4. OpenTracing: OpenTracing provides APIs in many different popular programming languages. You can use its SDK as the tracing back-end or integrate it with Jaeger via Docker, but it’s mainly known as a tracing solution. Note that any new implementations of a distributed tracing library should use OpenTelemetry.
    5. OpenZipkin: OpenZipkin is easy to use and provides features for both collection and data lookup.
    Figure 5: Overview of distributed tracing (source)

    Explore some of these open-source options and determine which best fits your use case. Gauge their capabilities across observability signals: logs, metrics, and traces.

    Conclusion

    Distributed tracing is most effective when you capture the right telemetry data and build a consistent instrumentation strategy from the start. It doesn’t replace your existing monitoring; it strengthens it by connecting end-user experience data to infrastructure, network, and application signals. Focus on traces, metrics, and logs together, and pair them with a sampling approach that keeps costs manageable without losing signal on critical paths.

    See how LogicMonitor connects traces, metrics, and logs across your entire stack in one platform.

    Distributed tracing gives you the request path. LogicMonitor connects that path to your infrastructure, cloud, network, and Internet performance data, so your team can trace a problem from the end user’s experience through the Internet path down to the service that caused it.

    Request a demo

    FAQs

    What’s the difference between distributed tracing and traditional logging?

    Traditional logging captures events within individual services, but it can’t connect those events across service boundaries. Distributed tracing assigns a unique identifier to each request and follows it across every service it touches, giving you a complete picture of the request’s journey. Logs and traces are complementary: logs tell you what happened within a service, and traces connect those events into a full end-to-end path.

    Do I need to rewrite my application to add distributed tracing?

    No. Most distributed tracing frameworks support instrumentation that lets you add tracing code to your existing application without a full rewrite. Tools like OpenTelemetry provide SDKs in many programming languages and support both automatic and manual instrumentation, so you can start with minimal code changes.

    How do I decide between head-based and tail-based sampling?

    Head-based sampling is simpler to implement and works well for smaller environments where you don’t need fine-grained control over which traces you keep. Tail-based sampling gives you more control because you can make decisions after the trace completes, but it requires additional infrastructure to buffer and process trace data. Start with head-based sampling if you’re just getting started with tracing, and move to tail-based when you need to optimize storage costs at scale.

    Can distributed tracing work across different programming languages and frameworks?

    Yes. Standards like the W3C Trace Context and tools like OpenTelemetry are designed to propagate trace context across services regardless of the language or framework they’re built with. A request can start in a Python service, pass through a Java microservice, and end in a Go service, and the trace will remain intact.

    By Denton Chikura

    Technical Writer

    Denton Chikura is a technical writer and longtime observability advocate focused on helping site reliability engineers and engineering teams discover the tools and capabilities that strengthen internet resilience. He works at the intersection of monitoring, performance, and infrastructure to make complex systems more understandable and usable, bridging the gap between deep technical detail and real‑world operations. His goal is to help teams build faster, detect issues earlier, and recover smarter, ultimately making the internet a better, more reliable place for everyone.

    Disclaimer: The views expressed on this blog are those of the author and do not necessarily reflect the views of LogicMonitor or its affiliates.

    © LogicMonitor 2026 | All rights reserved. | All trademarks, trade names, service marks, and logos referenced herein belong to their respective companies.

    Product

    Platform

    Infrastructure

    Cloud & Multi-Cloud

    Log Management

    Edwin AI

    Enterprise

    Demo

    Pricing

    WebPageTest Pricing

    RUM Monitoring

    IPM Monitoring

    Synthetic Monitoring

    How We Compare

    Datadog

    Dynatrace

    Virtana

    Solarwinds

    PRTG

    ManageEngine

    ScienceLogic

    SiteScope

    BigPanda

    About

    Careers

    Our Partners

    Leadership

    Newsroom

    Security

    AI Governance

    Sustainability

    Legal

    Documentation

    Docs Hub

    Release Notes

    Security

    Support Center

    Resources

    Autonomous IT in 2026

    Resource Library

    LM Academy

    Blog

    Case Studies

    Customer Education

    Connect

    Contact & Locations

    Submit a Ticket

    Events

    LM Community

    Careers


    Product

    Platform

    Infrastructure

    Cloud & Multi-Cloud

    Log Management

    Edwin AI

    Enterprise

    Demo

    Pricing

    WebPageTest Pricing

    RUM Monitoring

    IPM Monitoring

    Synthetic Monitoring


    How We Compare

    Datadog

    Dynatrace

    Virtana

    Zenoss

    Solarwinds

    PRTG

    ManageEngine

    ScienceLogic

    SiteScope

    BigPanda


    About

    Careers

    Our Partners

    Leadership

    Newsroom

    Security

    AI Governance

    Sustainability

    Legal


    Documentation

    Docs Hub

    Release Notes

    Security

    Support Center


    Resources

    Autonomous IT in 2026

    Resource Library

    LM Academy

    Blog

    Case Studies

    Customer Education


    Connect

    Contact & Locations

    Submit a Ticket

    Events

    LM Community

    Careers


    Privacy Policy

    Terms of Use

    Preference Center

    Do Not Sell My Information

    © 2026 LogicMonitor