Good Monitors Aren't Enough
When engineering teams implement Datadog, they do the right things: they instrument their services, configure monitors for known failure modes, and build dashboards for the metrics that matter. But applications are living systems: services change, configurations drift, new endpoints get added, and old ones quietly break. Error codes get misclassified, sensitive data starts leaking through traces, and no one writes a monitor for the thing they haven't imagined yet. The result is an observability gap: a growing distance between what your monitors believe is true about your application and what is actually happening inside it, widening slowly, one undetected failure at a time.
The most dangerous failures are the ones that look fine in Datadog until a customer calls, an audit happens, or the error volume becomes impossible to ignore.
The Reliability and Visibility (RV) Investigation
The investigation was developed internally at RapDev to surface the gap between what monitoring thinks is happening and what is actually occurring in production. RapDev Managed Datadog engineers run a structured RV Investigation, a proactive, hands-on deep-dive across a customer's entire Datadog environment. The goal is simple: find every observability problem in the environment, regardless of which product or layer it lives in, and surface it with enough evidence that the right team can act immediately.
The RV Investigation is an open-ended, exhaustive investigation that follows the evidence wherever it leads. Our engineers examine every layer of a customer's Datadog environment: APM, logs, infrastructure, dashboards, monitors, RUM, database monitoring, security signals, tagging, pipelines, and beyond, until they have a complete picture of what is working, what is not, and what is degrading silently.
What do we look for?
- APM trace health: Error rates, span volume anomalies, mislabeled HTTP status codes, and service-level failure patterns across the full trace lifecycle.
- Database-layer failures: Query-level failures, partition issues, and DB errors that surface as upstream API errors with no database context visible.
- Log quality and classification: Whether logs are correctly classified at source; error-level messages emitted as info silently disabling entire categories of log-based alerting.
- Monitor and visibility gaps: Production services with no monitors, no dashboards, no logs, or no environment tags gaps that no automated alert will ever surface on its own.
- RUM and frontend instrumentation: Session quota configuration, application tagging, sampling rule hygiene, and platform-level issues that silently drop frontend visibility.
- Security and compliance signals: PII flowing through traces, stale API keys, public dashboards, and Sensitive Data Scanner gaps that create compliance exposure.
- Infrastructure and agent health: Hosts missing the Datadog agent, containers without standard tags, network-level visibility gaps, and agent configuration drift.
- Cost and pipeline governance: Runaway log ingestion, unnecessary Sensitive Data Scanner processing, and quota misconfigurations that drive cost without improving coverage.
Real Investigations, Real Outcomes
These findings aren't hypothetical; they come from real customer environments RapDev has investigated. Below are two recent examples of what the RV Investigation uncovered.
1. Financial services customer database-layer failure: A ten-hour notification outage caused by a missing database partition
At a financial services customer, a notification service was silently failing for ten hours from 3 AM to 1 PM with no monitor firing and no alert reaching any on-call team. The investigation traced the failures through APM trace data to a PostgreSQL partition gap: insert operations were failing because the notifications table had no partition covering the incoming date range. Every insert was rejected at the database layer; the errors propagated upstream as generic 500 responses with no database context visible in the traces.
Because Database Monitoring was not enabled, the recovery event itself left no trace in Datadog. The investigation confirmed that the issue was resolved, but could not identify what corrective action was taken or when, a direct consequence of the DBM coverage gap.
Outcome: Recommendations delivered for partition pre-creation automation, APM error rate monitors with specific thresholds, and DBM enablement for the affected database instance.
2. Media and publishing customer log classification and PII: 30% error rate, silent log monitors, and PII flowing through traces
At a media customer, an APM spike investigation into a content subscription service revealed a 30% error rate in production; real users were actively unable to complete subscription actions, and no alert had fired. A second finding during log analysis uncovered error-level log messages being emitted at info severity across multiple services simultaneously, meaning the customer's log-based monitors would never fire for this entire class of failure. During trace analysis, sensitive user data was found being passed as plain-text API parameters and captured inside Datadog traces.
Outcome: Follow-up with the application team confirmed that 2 out of 3 investigation findings mapped directly to real application issues the team then acted on. The third was confirmed as expected behaviour.
Why Expertise Matters More Than Tooling
Every observability team is close to their own environment. That proximity is a strength they know their services, their history, their deployment patterns. But it also creates blind spots. Teams build monitors for failures they have seen before. They do not monitor for the failure modes they have not imagined yet.
RapDev engineers bring cross-industry experience across dozens of Datadog environments, and a systematic investigation methodology designed specifically to surface the gap between what your monitoring thinks is true and what your application is actually doing. The RV Investigation is not a one-time audit. It is a continuous practice, a regular exercise in honest assessment of your observability posture, delivered by engineers whose job is to find what was missed.
Datadog is a powerful platform. Whether you are extracting its full value depends less on the tool and more on the expertise applied to it. That is what Managed Datadog at RapDev is built to provide.
Want an RV Investigation for Your Environment?
RapDev Managed Datadog engineers can run a Reliability and Visibility investigation on your Datadog account, surfacing degradation, coverage gaps, and compliance risks before your customers find them. Reach out to RapDev to get started.


