Skip to content
BritonOne Technology
Cloud & DevOpsTelemedicine

24/7 SRE for a telemedicine platform

Ran a telemedicine platform to a 99.98% SLO with a follow-the-sun on-call.

99.98%
KubernetesPrometheusPagerDutyTerraform
24/7 SRE for a telemedicine platform
IndustryTelemedicine
DisciplineManaged Cloud Ops
CountryUnited Kingdom
Headline result99.98%
The story

Problem, approach, and the outcome

About the client

The client is a UK telemedicine provider whose patients rely on the platform around the clock, including out-of-hours consultations. High availability here is a matter of patient access to care.

Their in-house team was lean and could not sustainably cover 24/7 reliability without burning out or over-hiring.

The challenge

A lean in-house team could not safely cover the out-of-hours consultations the service depended on, and burnout risk was real. Round-the-clock cover was beyond what a small team could sustain.

Reliability targets were high because patients rely on the platform at all hours, but the team was too small to hold them alone. The gap between the obligation and the capacity was widening.

The provider needed genuine round-the-clock coverage without over-hiring. Sustainable reliability, not heroics, was the goal.

Our approach

We took on a follow-the-sun SRE rota, so coverage passes between regions and no one carries an unsustainable night shift. Distributing on-call across time zones is what makes 24/7 humane and sustainable.

We ran the platform to explicit SLOs with runbooks and error budgets, making reliability measurable rather than heroic. Clear targets and playbooks turned firefighting into managed operations.

Every incident fed back into hardening, so the same failure did not recur, and on-call was blameless by design, keeping the retrospective focused on the system, not the individual. Reliability improved steadily as the platform learned from each incident.

Results
  • 99.98% SLO sustained
  • Follow-the-sun, blameless on-call
  • Error budget driving reliability work
  • Every incident fed back into permanent hardening
Next step

Get a senior architect on the call, first time, every time.

No SDR gauntlet. 30 minutes with an engineer who can scope the problem, name the risks, and give you an honest feasibility call.