Back to blog

Closing the Observability Maturity Gap in Datadog

A clearer path to trustworthy signals

X

min read

October 7, 2026

Ryan Kelly

The Trust Problem Hiding in Your Dashboards

Ask any engineering leader if their team has observability, and the answer is almost always yes. Metrics are flowing, logs are indexed, traces are captured, dashboards cover every wall of the virtual war room. Ask that same leader how long it took to find the root cause of last quarter's worst incident, and the story gets more complicated. Somewhere between "we collect everything" and "we understood what happened in minutes" sits the observability maturity gap, the space between having telemetry and being able to trust it.

This gap doesn't show up because teams lack data. It shows up because volume and maturity are different axes entirely. A team can ingest terabytes of logs, thousands of custom metrics, and full distributed traces, and still get paged at 2 a.m. with three dashboards open and no clear idea which service actually broke. More telemetry, on its own, doesn't produce more truth. It just produces more noise to sort through under pressure.

Where the Gap Comes From 

The gap tends to grow the same way technical debt does: one reasonable decision at a time. A team instruments a service in a hurry before a launch, without anyone deciding who owns its alerts. Someone builds a dashboard for one incident and never updates it as the architecture changes underneath it. Custom metrics accumulate tags nobody remembers adding, because tagging felt free until the bill said otherwise. None of these were mistakes in isolation. They're just what happens when instrumentation outpaces the discipline of maintaining it.

The economics make the gap easy to miss until it's expensive. Metrics sent through OpenTelemetry, for example, bypass a vendor's native integration pipeline, so the vendor bills every one of them as a custom metric, and a single metric with a few unbounded tags, like a container ID crossed with a user ID, can multiply into millions of unique time series before anyone notices. Observability spend now runs 15–25% of total cloud cost at mid-scale organizations, and a meaningful share of that pays for metrics nobody has queried in months. Cost is the symptom; the underlying condition is telemetry nobody ever curated, owned, or connected to anything a responder can act on.

What Immature Observability Actually Looks Like 

In my experience, a few patterns show up again and again in organizations that have plenty of telemetry but little maturity:

  • Alerts with no clear owner: a monitor fires, and the first ten minutes of an incident go to figuring out whose service it belongs to, not fixing it
  • Dashboards that don't survive architecture changes: technically live, practically ignored, because nobody trusts them to reflect the current system
  • Unbounded custom metrics:  high-cardinality tags that multiply cost without multiplying insight
  • No SLOs tied to business impact:  teams track CPU and latency but can't say whether users were actually affected
  • Telemetry with no ownership metadata: logs, traces, and metrics that exist in isolation from the service catalog, so nobody can trace a signal back to a team

Any one of these is manageable. All five together can turn a routine deployment into a multi-hour incident review.

Why the Gap Is Expensive to Ignore 

Cloud bills aren't the only place the cost shows up, though that number gets attention fast. It also shows up in MTTR that doesn't improve no matter how much data teams add, in on-call engineers who stop trusting alerts because too many of them are noise, and in postmortems that keep surfacing the same root cause: nobody could see the whole picture fast enough. Maturity is what turns telemetry into a shared, trustworthy source of truth instead of a pile of signals everyone interprets differently under stress.

Closing the Gap with Datadog 

Closing this gap is less about adding more instrumentation and more about giving the instrumentation you already have structure, ownership, and judgment. 

Ownership: Software Catalog

Datadog's Software Catalog addresses the ownership half of the problem directly. It maps services to the teams that own them, attaches monitoring data, dependencies, and security posture to that ownership record, and turns "whose alert is this?" into a lookup instead of a Slack thread.

Noise: Watchdog

Watchdog, Datadog's AI engine, now powered by the Toto timeseries foundation model, learns the normal operating baseline for a system, including seasonality and inter-service dependencies, so a predictable traffic spike doesn't page anyone at midnight. Datadog reports this kind of contextual suppression can cut alert noise by up to 40%, which is the difference between an on-call rotation people trust and one they've learned to tune out.

Cost: Custom Metrics Governance

The cost side of the maturity gap has its own fix. Custom metrics governance practices, like using Metrics without Limits to trim indexed cardinality, reviewing the Metrics Summary page for metrics unqueried in 30+ days, and setting cost monitors on week-over-week growth, turn an unbounded bill into a managed one. Infinite Cardinality Metrics go a step further, pricing by metric name rather than unique tag combinations, so teams can capture the dimensions that matter without cardinality becoming a budget risk on its own.

And when an incident does happen, SLO alerts reframe the conversation around user impact rather than raw infrastructure metrics, while Bits AI's investigation agent is built specifically for the 2 a.m. page, pinpointing likely root causes across the telemetry that already exists, rather than asking a half-awake engineer to correlate five dashboards by hand. Datadog cites root-cause investigations resolving up to 90% faster with that kind of assistance in the loop.

A Clearer Path to Trustworthy Telemetry 

Most organizations already collect plenty, so closing the observability maturity gap comes down to deciding which signals actually matter, what metrics and events can (and should!) be tossed before they reach the observability platform, and building the habits that keep telemetry honest as systems change. Datadog makes the data you already have ownable, quiet when it should be, and fast to act on when it counts. That turns telemetry from a pile of numbers into something your team can actually trust.

Want to see where your Datadog setup lands on the maturity map? Talk to our team.