Performance engineering

When a system is slow or expensive to run, we measure first: which requests are slow, where the time goes and what it costs. Then we fix the parts that matter and measure again, so every change comes with a number.

p95 latency over time, dropping after a measured change A line chart of 95th percentile latency over time. The line is high and jagged, then a dashed vertical marker labelled change deployed, after which the line drops and stays low. Measurements before and after are labelled. p95 time slow fast one change deployed before: measured after: measured Measure, fix one thing, measure again

What we do

  • Database and query tuning Slow queries, missing or unused indexes, locking and contention, and data that should come from a replica or a cache instead of the primary database.
  • Profiling slow services and endpoints Finding the calls, serialization, waits and retries that add up to a slow response.
  • Load tests and capacity planning How the system behaves at two, five and ten times today's traffic, where it breaks, and what it will cost to serve.
  • Cost tuning Right-sized instances, cached repeated work, batching, and the other changes that lower the bill without lowering reliability.
  • Fixing the slow parts From a rewritten query to a change in how a service is split or how data is stored.

When you need this

  • Pages or API calls that take seconds instead of milliseconds, especially at peak times.
  • Database load that keeps climbing, with a bigger server as the only fix so far.
  • A cloud bill growing faster than your traffic.
  • A launch, campaign or season ahead, and you need to know the system will hold.

How an engagement works

  1. Measure We add the instrumentation that is missing, record a baseline and profile the slow paths.
  2. Rank We list the bottlenecks by how much they cost you and how much effort each fix takes.
  3. Fix and verify One change at a time, each measured against the baseline and reviewed by your team.
  4. Report What changed, with before-and-after numbers, what remains and what to watch.

What you get

  • Baseline measurements and a ranked list of bottlenecks
  • Fixes, each with a before-and-after measurement
  • Load test results and capacity estimates
  • Dashboards and alerts for the metrics that matter
  • A short report of what changed and what to watch

Related work

Common questions

How quickly can you find the problem?

Usually within the first days. Most slow systems have a few bottlenecks that dominate, and profiling finds them fast. Fixing them can take longer, depending on how deep the change goes.

Will you need access to production?

Read access to metrics and logs, and a copy of the data, is usually enough to find the problems. Changes go through your normal review and deployment process; we do not make unreviewed changes to a live system.

Can you fix performance without a rewrite?

In most cases, yes. The biggest gains usually come from queries, indexes, caching and how work is batched, not from rewriting the system. If a rewrite of some part is the right answer, we say so and show the measurements behind it.

Do you do load testing on its own?

Yes. Before a launch or a busy season, we load-test the system against the traffic you expect and report where its limits are, with or without a tuning phase afterwards.

Tell us about your system

Send a few lines about what you are building or what needs fixing, and we will set up a call.