Skip to content
BritonOne Technology
Quality Assurance & TestingGovernment & Public

Load and soak testing for an enterprise cloud platform

Validated a 99.9% uptime SLA under peak load, so the cloud portal and provisioning APIs held through demand spikes without incident.

99.9%
k6JMeterGrafanaPrometheus
Load and soak testing for an enterprise cloud platform
IndustryGovernment & Public
DisciplinePerformance Testing
CountryUnited States
Headline result99.9%
The story

Problem, approach, and the outcome

About the client

The client operates a large public-sector cloud platform through which agencies provision and manage shared digital services. For this kind of platform, availability is not a marketing number; a portal outage stalls the work of thousands of downstream users at once.

The service carried a contractual 99.9% uptime SLA, but nobody could prove the platform would actually hold that target when demand peaked. The commitment had been signed on hope rather than evidence.

The challenge

The portal and its provisioning APIs had never been exercised at production-like load, so the team had no idea where the platform would bend or break. Every busy period was effectively a live experiment.

Uptime commitments were measured after the fact from incident logs, not proven ahead of time. A breached SLA on a public service carries penalties and reputational cost, and the platform team could not afford to find its ceiling during a real peak.

What they needed was hard evidence, under sustained realistic load, that the SLA would hold, plus a clear picture of where the first constraint sat. Guesswork had to be replaced with measured proof.

Our approach

We modelled the real traffic mix from the platform's own telemetry, ramps, sustained peaks, and long quiet tails, then drove it against a production-parity environment with k6 and JMeter. Matching the real shape of demand is what makes the results transfer to live operation.

We layered soak testing over the load runs, holding peak concurrency for hours to expose the slow leaks and connection-pool exhaustion that only surface over time. Short tests hide the failures that actually cause outages.

Grafana and Prometheus correlated every run with CPU, memory, and connection-pool telemetry, so a wobble in the SLA pointed straight at its cause rather than a vague suspicion. We handed over reproducible scripts and a prioritised tuning list, and re-ran until the 99.9% target held with margin to spare. The commitment became evidenced, not hoped for.

Results
  • 99.9% uptime SLA validated under peak load
  • Multi-hour soak runs exposed a connection-pool ceiling before launch
  • Every run mapped to server-side telemetry for root cause
  • Reproducible load scripts handed to the platform team
Next step

Get a senior architect on the call, first time, every time.

No SDR gauntlet. 30 minutes with an engineer who can scope the problem, name the risks, and give you an honest feasibility call.