Three System Design Principles That Help You Build Better Software
An online store looks simple from the outside: you choose a product, pay, and receive an order confirmation. Behind that experience, several pieces of software must cooperate—and keep working when traffic rises, a network slows down, or a payment service stops responding.
System design is the work of deciding how those pieces fit together. You don’t need a huge application to benefit from it. Three principles give you a practical starting point: separate responsibilities, expect failures, and improve the part that actually limits performance.
1. Separation of Concerns: Give Each Part a Clear Job
Separation of concerns means dividing a system into components with distinct responsibilities. A component is a part of the software that performs a particular job.
Think of a restaurant. The kitchen prepares food, the servers look after customers, and the cashier handles payment. They coordinate, but you wouldn’t want every person to perform every task at once.
Software benefits from similar boundaries.
What this looks like in an online store
You might divide the store into these areas:
| Component | Main responsibility |
|---|---|
| Product catalog | Describe products and display their prices |
| Inventory management | Track available stock and reserve items for orders |
| Payment processing | Request payments and record their outcomes |
| Order management | Track an order from placement through fulfillment |
Inventory management, for example, answers questions such as “How many units are available?” and “Have these units already been reserved for another order?” Payment processing answers a different question: “Has the customer paid?”
Those responsibilities interact, but they are not the same.
Why the boundaries matter
Suppose you change payment providers. With clear separation, most of the changes should stay within the payment component rather than spread through stock tracking and product descriptions.
That makes the system easier to:
- Understand: You know where to look for a particular behavior.
- Test: You can check stock calculations without making real payments.
- Change: Updating one responsibility is less likely to disturb unrelated behavior.
The important detail is that separation does not require separate servers. A small application can have clearly organized components while still running as one program. Splitting everything into independent services introduces its own coordination and operational work.
Give each component a clear responsibility before deciding where it should run.
A useful design question is: If this business rule changes, how many unrelated parts of the system will I need to edit? If the answer is “almost everything,” your boundaries may need attention.
2. Design for Failure: Assume Something Will Go Wrong
Even well-built systems fail occasionally. Computers restart, networks lose connections, and outside services become unavailable.
Designing for failure means planning how the system should behave under those conditions instead of treating them as unimaginable exceptions.
The aim isn’t to prevent every possible failure. It’s to limit the damage and make recovery manageable.
Put a limit on waiting
A timeout is a limit on how long one component waits for another.
Imagine your store requests payment approval, but the payment provider never responds. Without a timeout, the request could wait indefinitely, tying up resources and leaving the customer staring at a spinner.
With a timeout, the store can stop waiting and take a deliberate next step.
But there is an important catch: a timeout does not prove that the payment failed. The provider may have processed the payment while its response was lost.
That uncertainty shapes how you handle retries.
Retry carefully—especially when money is involved
A retry repeats an operation that failed or did not receive a response. It can help with temporary problems, but careless repetition can create new ones.
If you retry a payment request, you must avoid charging the customer twice.
One common safeguard is idempotency: designing an operation so that repeating the same logical request does not apply its effect multiple times. For payments, this often means attaching a unique identifier to the payment attempt. A system that supports this behavior recognizes a repeated identifier and avoids creating another charge.
You should also limit retries and leave pauses between them. Constantly repeating requests can add pressure to a service that is already struggling.
Prepare for recovery
Different failures need different safeguards:
- Timeouts prevent indefinite waiting.
- Limited, carefully spaced retries help with temporary failures.
- Backups provide a way to recover lost or damaged data.
- Graceful degradation keeps useful parts working when optional features fail.
For example, if product recommendations are unavailable, the store may still let you browse and buy. If payment processing is unavailable, it should not pretend that payment succeeded.
A resilient system does not make every failure disappear. It gives failures a controlled outcome.
For each important dependency—a service or component your system relies on—ask: What happens if it is slow, unavailable, or gives us an uncertain result?
3. Scale the Bottleneck: Improve What Actually Limits Performance
When an application gets slow, adding more servers can feel like the obvious answer. Sometimes it helps. Sometimes it changes almost nothing.
A bottleneck is the part of a system that limits its overall performance. Think of a supermarket with plenty of aisles but only one open checkout. Making the aisles wider won’t shorten the payment queue.
Scaling means increasing a system’s ability to handle work. Effective scaling starts by finding the bottleneck.
Measure before choosing a solution
Look for evidence about where time and resources are being spent:
- Which requests are slow?
- How much time goes into database work?
- Are application servers fully occupied?
- Are requests waiting on an outside service?
A database stores and retrieves structured information, such as products, orders, and stock levels. If retrieving that information is the slowest step, adding application servers may simply send more work to the same overloaded database.
Match the remedy to the problem
| Measured problem | Possible response |
|---|---|
| Repeatedly reading the same product details is expensive | Use a cache |
| One application server cannot handle incoming requests | Add servers and distribute traffic |
| A particular database query is slow | Improve how the query retrieves data |
A cache stores a reusable copy of information so it can be retrieved more quickly. It can reduce repeated database reads, but the copy may become outdated. That matters more for available stock than for a rarely changing product description.
A load balancer distributes incoming requests across multiple servers. This can help when application-server capacity is the constraint, but it does not automatically fix a slow shared database.
The practical loop is simple:
- Measure performance.
- Identify the limiting component.
- Make a targeted improvement.
- Measure again.
Once you relieve one bottleneck, another may become visible.
Use the Three Principles Together
These principles reinforce one another. Clear responsibilities help you locate performance problems and decide where failure handling belongs. Failure safeguards keep slow or unavailable components from causing uncontrolled damage. Measurement helps you spend effort where it matters.
When reviewing a design, start with three questions:
- Responsibility: Does each part have a clear job?
- Failure: What happens when a part stops working?
- Performance: What evidence shows where the system is constrained?
You don’t need the most elaborate architecture. You need a system whose parts are understandable, whose failures are manageable, and whose improvements address real problems.