Back to Case Studies
Enabling SAIT to Reduce MTTR with Modern Observability
How SAIT reduced MTTR by deploying Datadog observability and improving correlation across key services
RapDev partnered with SAIT to deploy Datadog APM, infrastructure monitoring, and logging, cutting mean time to resolution (MTTR) by 20% within six months and delivering full visibility into 70+ servers
Customer Challenges
- Disparate monitoring tools owned by separate teams
- Limited visibility across the technology stack
- SQL servers with no monitoring at all
- Manual configuration delaying response times
RapDev
Solutions
- APM, infrastructure, and logging deployed together
- Ansible automation for agent rollout
- Tagging standardization across on-prem hosts
- Synthetic testing simulating real user journeys
Business
Outcomes
- 20% reduction in MTTR within six months
- 70+ servers with complete visibility
- 30% increase in proactive incident detection
- Under a week to scale licenses and new servers
The Southern Alberta Institute of Technology (SAIT) is a polytechnic institute in Calgary, Alberta, Canada.
Established in 1916, it is Calgary's second-oldest post-secondary institution and Canada's first publicly funded technical institute.SAIT offers more than 110 career programs in technology, trades, and business.
3K+
Employees
12,000+
Students/year
30,000+
Avg. Registrations/season
The Challenge
Siloed Visibility Impacting User Experiences
Core Challenge
Disparate observability tools were each managed independently by different teams, and the siloed approach hindered collaboration, scalability, and efficiency. Detecting and resolving issues was hardest during heavy-load events like open enrollment, and SQL servers lacked monitoring altogether.
Heavy manual workflows and no configuration management practice raised both overhead and the perceived risk of migration.
The Solution
Automating Operations for Proactive Monitoring
VISIBILITY INTO ERP SYSTEMS
SAIT's primary concern was its Ellucian Banner ERP. RapDev delivered comprehensive tracking of Banner application performance, from infrastructure metrics through to APM requests, closing the visibility gap that mattered most to students, educators, and administrators.
ANSIBLE AUTOMATION
RapDev built an automated deployment platform for provisioning monitoring agents across SAIT's infrastructure. A strong Ansible inventory covers Windows and Linux on-premises hosts, installing Datadog agents and APM components without manual effort.
SYNTHETIC TESTING
Synthetic tests simulate real user interactions, letting SAIT find performance bottlenecks and fix them before students or staff feel the impact. It replaced a posture that had relied mostly on Nagios alerts with little visibility.
DASHBOARDS & ALERTING
Comprehensive dashboards consolidated views that had been scattered across Nagios and Uptime Robot, with customized alerting built to match. SAIT gained a unified picture of infrastructure, applications, and user experience.
TAGGING & LOGGING
RapDev's engineers standardized automated tagging across on-premises hosts. Comprehensive logging allowed SAIT to collect, search, and analyze log data efficiently. Integration with existing Terraform and Ansible tooling kept the whole setup cohesive.
COLLABORATIVE WORKFLOW
RapDev fostered collaboration and knowledge-sharing across SAIT's engineering teams, with guidance on best practices for monitoring configuration, alerting, and incident response. Internal resources came away able to support the platform themselves.
The Results
MTTR Down, Visibility Up
Zero Sev 1 Incidents
Unexpected shutdowns dropped from four a month to none post deployment
250+ Dashboards
Filterable by region, country, and role for every audience
1,300+ Servers Covered
Full Datadog deployment across a globally distributed estate
Automated Health Checks
Proactive alerts and synthetic transactions catch issues first
“What's really exciting is we have built, with RapDev, an automated deployment platform so that when we want to purchase more licenses, change our existing agreement, or scale out to those 200 servers, we can do it in less than a week.”
- Ross Henderson | Principal Consultant, SAIT
What's Next
Continual Innovation
SAIT plans to expand monitoring coverage, refine dashboards for deeper insights, and implement advanced analytics for proactive issue resolution. The robust logging foundation also lays the groundwork for defining and monitoring service level objectives.
We don’t believe in hoarding knowledge
We go further and faster when we collaborate. Geek out with our team of engineers on our learnings, insights, and best practices.
Blog Posts
Resources
Let’s get started
Ready to maximize your observability investment?
