See where your architecture breaks and what it costs, before you build it.
Drag load balancers, services, caches and databases onto a canvas, press play and watch throughput, p99 latency, errors, bottlenecks and monthly cost live. It runs in your browser, without an account.
Triple the traffic, watch the database tip over
This is the real simulation engine, running the Web app with database template in a Web Worker on this page: 400 requests/s through a load balancer to 4 API servers and one Postgres primary. The traffic dial moves from 1× to 3× and back. At 3× the database runs at about 130 % load, the API runs out of threads waiting for it, p99 goes from under 100 ms to about 1 s and about a quarter of the requests fail.
In the full playground (early access) you can go one step further: put a Redis cache in front of the reads and the database drops to about 36 % at 3×, p99 to 120–150 ms, for about $192/month more.
What you get
-
Live bottlenecks
Every node shows utilization, queue, latency and errors while the simulation runs. When something overloads, Stackrig points at the root cause, not at every node that waits on it, and says what would help.
-
p99 and cost, side by side
Tail latency comes from sampled request paths, not from averages. Monthly cost comes from AWS, GCP, Azure and Hetzner price tables, each with its source and date.
-
Chaos on demand
Switch off a node or some of its instances, add latency or errors, then recover it. Timeouts and retries are set per connection, so a retry storm shows up as one.
-
Templates to learn from
Eight designs, from a URL shortener to event-driven order processing, each with one lesson you can reproduce in a minute: what breaks first, and the cheapest fix.
How it works
-
A fluid simulation
Rates, queues, utilization, errors and retries are computed per node and connection, so the cost of a step depends on the size of the design, not on the load: 10 requests/s and 10 million requests/s cost the same. Latency percentiles come from sampled request paths.
-
Checked against an exact reference
A discrete-event simulation that follows every single request serves as the reference. In 15 validation scenarios the fluid model has to match it within 5 % for throughput and utilization and within 15 % for p50, p95 and p99. For the template in the demo above it matches within 0.3 %. Where it misses, for example overload tails, correlated retries and retry-storm tipping points, the deviation is measured and listed, not hidden.
-
Calibration in progress
Calibration against real cloud VMs is in progress. Until those results are published, the numbers are model results that agree with an exact reference, not measurements of your system.
Join the waitlist
Stackrig is in private preview. Leave your email address and we tell you when it opens. We may also ask whether you would like to talk to us about how you design systems today. No newsletter, no tracking.
The full playground at /app is invite-only while Stackrig is in private preview.
Share links (stackrig.dev/app#…, and older ones to stackrig.dev/#…) now
open a sign-in page first: enter an invited email address and the one-time code we send you.
If the design does not open after signing in, open the link again. Join the waitlist to get an
invite.
The waitlist opens soon.