Back to Case Studies

Enabling SAIT to Reduce MTTR with Modern Observability

How SAIT reduced MTTR by deploying Datadog observability and improving correlation across key services

Download PDF

RapDev partnered with SAIT to deploy Datadog APM, infrastructure monitoring, and logging, cutting mean time to resolution (MTTR) by 20% within six months and delivering full visibility into 70+ servers

Customer Challenges

  • Disparate monitoring tools owned by separate teams
  • Limited visibility across the technology stack
  • SQL servers with no monitoring at all
  • Manual configuration delaying response times

RapDev
Solutions

  • APM, infrastructure, and logging deployed together
  • Ansible automation for agent rollout
  • Tagging standardization across on-prem hosts
  • Synthetic testing simulating real user journeys

Business
Outcomes

  • 20% reduction in MTTR within six months
  • 70+ servers with complete visibility
  • 30% increase in proactive incident detection
  • Under a week to scale licenses and new servers

The Southern Alberta Institute of Technology (SAIT) is a polytechnic institute in Calgary, Alberta, Canada.

Established in 1916, it is Calgary's second-oldest post-secondary institution and Canada's first publicly funded technical institute.SAIT offers more than 110 career programs in technology, trades, and business.

3K+

Employees

12,000+

Students/year

30,000+

Avg. Registrations/season

The Challenge

Siloed Visibility Impacting User Experiences

Core Challenge

Disparate observability tools were each managed independently by different teams, and the siloed approach hindered collaboration, scalability, and efficiency. Detecting and resolving issues was hardest during heavy-load events like open enrollment, and SQL servers lacked monitoring altogether.

Heavy manual workflows and no configuration management practice raised both overhead and the perceived risk of migration.

The Solution

Automating Operations for Proactive Monitoring

VISIBILITY INTO ERP SYSTEMS

SAIT's primary concern was its Ellucian Banner ERP. RapDev delivered comprehensive tracking of Banner application performance, from infrastructure metrics through to APM requests, closing the visibility gap that mattered most to students, educators, and administrators.

‍

ANSIBLE AUTOMATION

RapDev built an automated deployment platform for provisioning monitoring agents across SAIT's infrastructure. A strong Ansible inventory covers Windows and Linux on-premises hosts, installing Datadog agents and APM components without manual effort.

‍

SYNTHETIC TESTING

Synthetic tests simulate real user interactions, letting SAIT find performance bottlenecks and fix them before students or staff feel the impact. It replaced a posture that had relied mostly on Nagios alerts with little visibility.

‍

DASHBOARDS & ALERTING

Comprehensive dashboards consolidated views that had been scattered across Nagios and Uptime Robot, with customized alerting built to match. SAIT gained a unified picture of infrastructure, applications, and user experience.

‍

TAGGING & LOGGING

RapDev's engineers standardized automated tagging across on-premises hosts. Comprehensive logging allowed SAIT to collect, search, and analyze log data efficiently. Integration with existing Terraform and Ansible tooling kept the whole setup cohesive.

‍

COLLABORATIVE WORKFLOW

RapDev fostered collaboration and knowledge-sharing across SAIT's engineering teams, with guidance on best practices for monitoring configuration, alerting, and incident response. Internal resources came away able to support the platform themselves.

The Results

MTTR Down, Visibility Up

Zero Sev 1 Incidents

Unexpected shutdowns dropped from four a month to none post deployment

250+ Dashboards

Filterable by region, country, and role for every audience

1,300+ Servers Covered

Full Datadog deployment across a globally distributed estate

Automated Health Checks

Proactive alerts and synthetic transactions catch issues first

“What's really exciting is we have built, with RapDev, an automated deployment platform so that when we want to purchase more licenses, change our existing agreement, or scale out to those 200 servers, we can do it in less than a week.”

- Ross Henderson | Principal Consultant, SAIT

What's Next

Continual Innovation

SAIT plans to expand monitoring coverage, refine dashboards for deeper insights, and implement advanced analytics for proactive issue resolution. The robust logging foundation also lays the groundwork for defining and monitoring service level objectives.

Let’s get started

Ready to maximize your observability investment?

Get In Touch