App, database, files, and slow work together
Interactive systems lab / 05
System Evolution Lab.
Start simple. Add one boundary at a time when the workload gives you evidence that it helps.
Grow the system when evidence asks you to.
First move: Pick “Growing product.” Compare stages 2 and 3 at the same load, then fail one app instance.
Scroll to see all seven stages →
Stress the design
No added disruption.
All capacities below are teaching budgets. Measure a real workload before setting limits.
One box.
Ship a simple product with one deployment surface.
The current scenario fits these teaching ceilings. Probe a failure or higher peak before adding another component.
Demand against capacity
One failed machine or noisy query affects every part of the product.
Earliest stage within this model’s budgets: 1. One box.Change the part that measurements show is constrained. Stage 3 deliberately adds visibility without adding throughput. A cache helps reads; durable jobs protect slow work; replicas and partitions add their own operating costs.
The stage budgets are illustrative, not hardware ratings or an availability promise. App work treats each slow task as three request-equivalent units until it is moved to workers. Cache hits are fixed at 70% after stage 5 unless the cache outage scenario is selected. This model omits database locks, shared CPU, connection pools, network delay, and deployment cost.
Under the model
A stage solves a named problem.
Separating a database changes isolation. Instrumentation makes a bottleneck visible without adding capacity. Stateless application replicas spread request work. A shared cache reduces repeated reads. Durable jobs move slow work out of request lifetimes. Replicas and partitions extend data capacity at a higher operating cost.
The budgets in the simulator are illustrative. Real scale decisions need query plans, load tests, pool budgets, retry semantics, and incident evidence. Multiple changes may be needed at once, but each should have a reason.
- Hold the workload constant and compare stages 2 and 3.
- Lose one app instance before and after load balancing.
- Trigger a cache outage at stage 5 and inspect database reads.
Lesson 05 / Check your understanding
Can you explain the result?
Stage 3 adds metrics and tracing but keeps the same hardware. What changes immediately?