Start with the promise to the user
Imagine class bookings. A customer sees available places, chooses a class, pays and receives confirmation. The crucial guarantee is that two people cannot receive confirmed bookings for the last place.
Describe available, held, paid and cancelled states. Who owns each transition? What if payment arrives after the hold expires? Should the system confirm without a place or refund? These rules determine transaction and integration boundaries.
The class catalogue may tolerate delayed updates while the last-place decision needs a separate check. Different operations can have different guarantees. Choosing one consistency label for the whole product hides the important distinctions.
We define the system boundary with the workflow participants. A payment provider owns money movement while our service owns seat confirmation. One local call cannot promise an atomic transaction across both. We describe how outcomes are reconciled and who handles exceptions instead of leaving that boundary hidden behind an API arrow.
Each transition gets a condition and owner. A hold requires a free seat; confirmation requires a valid hold and matching payment. Cancellation may require a refund. Rules scattered across the UI, event handler and background job can diverge, so ownership needs to be explicit even when implementation spans several components.
We examine two simultaneous requests. Checking availability and then separately writing a booking can fail if another request does the same in between. The invariant needs protection under real concurrency. The mechanism follows the data model and storage guarantees; passing a sequential happy-path test does not establish it.
Administrative actions belong in the design. Staff may close a class after sales begin or add capacity. These changes need the same guarantees and a traceable outcome. An admin panel must not bypass the rules, or the system will be correct only when every operator follows assumptions the software does not enforce.
We also state what is acceptable: waiting for confirmation, briefly stale catalog availability or a delayed report. This removes unnecessary strictness from peripheral paths and lets the design spend complexity on the booking guarantee. Clear tolerance is as useful as a list of things that must never happen.
Estimate system load and capacity limits
Suppose there are 100000 bookings a day. The average is about 1.16 per second, yet opening a popular class can attract hundreds of attempts within seconds. The average does not describe capacity needed for that event.
At 2 KB per main record, new data is approximately 200 MB daily or 73 GB yearly. Indexes, logs, history, backups and replicas add physical storage. Estimate them separately and distinguish measurements from assumptions.
The estimate should lead to concrete questions: how many connections we keep, where a hot row forms, how long customers wait and what storage costs. Multiplying average traffic by ten is not a peak model. That margin might be excessive for the whole service and still insufficient for one popular class.
Operation classes have different costs. Catalog reads can greatly outnumber bookings, and payment confirmations arrive through another flow. Counting bookings alone drops read pressure and external calls from the model. Each path gets frequency, data size and dependencies rather than inheriting the average cost of a business transaction.
A hot object is tested separately from overall RPS. A thousand attempts spread across a thousand classes behave differently from a thousand attempts on one last seat. More application processes may increase contention for one record without increasing useful confirmations. Distribution matters as much as total arrivals.
Storage estimates include lifecycle. A detailed event log may stay for a month while financial records need a much longer period. Keeping everything for the longest period makes cost depend on an accidental blanket policy. Access, deletion and restoration need consideration for each data class.
External capacity is part of the estimate. If a provider limits requests, local servers do not remove that constraint. We need a queue, rate limit and waiting-time expectation or an agreed quota increase. The user promise must fit these conditions rather than only the capacity of infrastructure we control.
Numbers test feasibility rather than create false precision. For an unknown parameter, we use a range and check sensitivity. If doubling response size changes little, the estimate is robust there. If a small write-rate change breaks the design, we should measure that parameter before investing in less consequential detail.
Translate consistency into observable behavior
If two copies cannot communicate, independently confirming the final place on both risks breaking the promise. One option is to pause confirmation where authority over the place cannot be established. The catalogue can remain available.
This is a decision about a network partition, not a reason to switch off the whole product during any failure. For each operation, we decide whether stale data is acceptable, whether a request can wait for processing or whether the result needs immediate confirmation. Those answers define the availability boundary.
Users also need to observe their own changes. "Booking created" followed by "no booking found" damages trust even if replicas converge later. Specify the guarantees required by these sequences rather than leaving them behind a technical label.
Stale availability has a user cost. Someone can see a seat and be refused at confirmation. That may be acceptable if the product explains it and the frequency is tolerable. Eventual consistency names a model; it does not decide whether the experience meets the product promise.
Seeing one's own write can use several approaches: route sensitive reads to current data, wait for a required version or carry the confirmed result into the next screen. Each has latency and failure implications. We choose based on the promised sequence rather than apply one rule blindly to every read.
Caches add copies, while search and analytics often update through separate streams. Their lag need not block the booking, but the UI must understand it. Searching for a just-created reservation may require explanation or another lookup path if the normal index has not caught up.
Last-write-wins is not a universal conflict solution. Competing description edits might tolerate it; competing confirmations of one seat cannot. Financial and limited-resource operations need semantic rules. Their protocol follows the invariant rather than the convenience of a generic merge function that happens to fit the storage layer.
A network failure includes reconnection. Queued events may arrive in a different order after connectivity returns. Components need to identify valid changes and superseded ones. Restored communication does not automatically remove conflicts, so the end of the failure is part of the design too.
Resolve the uncertainty left by a timeout
A payment call times out. The provider may not have processed it, or may have processed it while the reply disappeared. Repeating with a new identifier can create another operation. The customer needs a pending state and the system needs reconciliation.
Define operation identifiers, durable results and conflicting payloads under the same identifier. An idempotent receiver can return the earlier result, but does not create atomicity between your database and an external payment provider.
Handle late payment after cancellation, duplicate events and reordered delivery explicitly. Each needs a path to confirmation, refund, manual review or another status check. An arrow joining two diagram boxes does not capture those decisions.
An unknown result becomes a state we can query. The operation has an ID, start time and current status. A client can look it up again and a reconciliation job can contact the provider. If that state is not stored, a restart leaves the team guessing from incomplete logs.
An idempotency key needs to survive possible repeats. Deleting it immediately after an answer can permit a duplicate from a client that never received confirmation. Keeping it forever has a cost too. Retention and behavior after expiry follow the protocol and operation rather than an arbitrary cache setting.
Publishing after a local write creates another failure window: the reservation is saved, but the process dies before sending the event. Storing publication intent with the business change can support later delivery. Receivers still handle duplicates, and stuck publication work needs detection. The mechanism solves a specific gap, not every delivery problem.
Compensation may not undo a physical outcome immediately. Refunds can take time, emails have been read and an external system may be unavailable. We define what completion means and what the customer sees in between. Naming a compensation pattern does not supply those business rules.
Unhandled exceptions need an owner and a safe manual path. An employee might verify a disputed payment and record a decision. The action needs duplicate protection and an audit trail. Otherwise the recovery tool becomes a new way to violate the invariant it was intended to repair.
Separate availability, data loss and recovery time
Normal availability, acceptable data loss and recovery time are different requirements. RPO expresses the tolerated loss window; RTO expresses a recovery target. Measure actual recovery under the chosen failure scenario.
A one-hour recovery promise includes detection, deciding to switch, startup, data verification and restoring traffic. A quick database command proves only one part. Bookings may also need reconciliation with payments preserved independently by the provider.
RPO and RTO follow a specific failure. Losing one process and losing the entire data store have different effects. Compute may return quickly while records take much longer. One number for every scenario usually hides either expensive excess capacity or a promise the actual recovery process cannot meet.
Recovery dependencies include access, configuration, keys, network and people. A backup can be intact while its read credentials are unavailable. Waiting for permission or a specialist contributes to downtime. The exercise should follow the path the on-call team will really use, not an expert's private shortcuts.
Restored data needs business checks. Bookings and payments can diverge, events repeat and holds expire during downtime. Traffic should resume with reconciliation and safe continuation. Returning requests immediately after database startup can be riskier than spending time verifying the promises that matter.
The exercise limits harm and has an observable result. We do not need to break production to find a broken instruction. But an overly artificial environment misses realistic data sizes and dependencies. We distinguish fully tested steps from assumptions, so a successful exercise is not credited with proving more than it did.
Recovery expectations change with the product. More data can double restoration time and a new integration can add reconciliation. An old test does not establish the new system's RTO. The plan is reviewed with data growth and substantial workflow changes rather than treated as a document finished at launch.
Choose a design the team can verify and operate
If one database and a few processes meet the load and guarantees, another service needs a concrete purpose. Every component brings deployments, contracts, monitoring and partial failures to investigate. Team time and operational capacity belong in the constraints too.
We record the decision, alternatives, assumptions and a trigger for revisiting it. Then we test the risk: the last-seat race, a lost payment response, a booking spike or recovery. This gives us reasons why the design fits and when it might stop fitting. Without those checks, the diagram remains an assumption.
We compare options using the same requirements: peak capacity, invariant protection, failure detection and recovery. One service, several services and an added queue should all answer those questions. This reveals where complexity buys a needed property instead of making the most detailed diagram look strongest by default.
Service boundaries follow data ownership and independent change. Components constantly needing one shared transaction acquire a protocol when separated. Independent load or teams can justify that cost, but the decision should state why. A process boundary changes failure behavior even when the business feature stays the same.
Operational complexity shows up in everyday work: deployments per change, places to inspect during failure and people who can recover a component at night. A design understood by one person has an organizational failure point despite multiple technical replicas. Team knowledge belongs in the availability discussion.
A small prototype checks the disputed risk, such as last-seat contention or upstream latency. It need not reproduce the whole future product. Its value comes from an experiment able to change the architecture before broad implementation, rather than a demonstration built to confirm the option already preferred.
A revisit trigger is observable: tail latency, conflict rate, recovery time or operational workload. "When we grow" is too vague to guide action. Specific triggers make a simple current design an intentional stage with limits, instead of an unspoken promise that it will fit forever.
Testing the last-seat design
We begin with two simultaneous customers. Only one receives a valid hold; the other gets an understandable refusal. The first pays and the response is lost. A repeat with the same key must not create another charge or reservation. The UI exposes a pending state whose outcome can be checked.
Next comes a late payment after expiry, when another person owns the seat. The agreed refund or alternative path runs explicitly. Repeated events are safe, the refund outcome is stored and disputed cases have an owner. Delivery order does not decide the business policy by accident.
Load tests distinguish many separate classes from one popular class. We inspect waiting, record contention and provider limits. During a partition, confirmation respects the chosen boundary while the catalog keeps its permitted behavior. Reconnection includes reconciliation of queued changes and final states, rather than assuming restored networking resolved the outcome.
Finally we restore data, reconcile payments and return confirmation traffic. Measured time is compared with RTO, the recovered data point with RPO. The decision record states tested cases and remaining assumptions. These results describe the design more precisely than an architecture label or the number of boxes on a diagram.