← Engineering Notes

Reliability Engineering

Retry vs Circuit Breaker vs Fallback

When backend services depend on external systems, failure handling is not just about trying again. Retries, circuit breakers, and fallbacks solve different problems — and using the wrong one can make a failure worse.

The problem

One thing I have been paying more attention to in backend systems is how they behave when something they depend on stops behaving normally.

An API call to an external service can fail for many reasons. The dependency might be temporarily unavailable, slow to respond, returning errors, or failing consistently.

The first instinct is often simple:

“Try again.”

Sometimes that is exactly what the system needs.

But retrying is not a complete failure strategy.

If the dependency is already struggling, repeatedly sending the same request can create even more traffic. A temporary failure can turn into a larger problem if every caller keeps retrying without any limits.

That is where I started thinking about retries, circuit breakers, and fallbacks as three different tools rather than three interchangeable patterns.

A retry asks:

“Should I try this operation again?”

A circuit breaker asks:

“Should I keep calling this dependency at all right now?”

A fallback asks:

“If the dependency cannot complete this operation, what can I do instead?”

Those are different questions.

Retry: dealing with transient failures

Retries are useful when a failure is likely to be temporary.

For example, a network request might fail because of a short-lived connection problem or a temporary interruption in the dependency.

Instead of immediately failing the request, the application can try the operation again.

Conceptually:

The important part is that retries need boundaries.

A retry policy should consider things such as:

  • How many times should the operation be retried?
  • Which failures are actually retryable?
  • Should there be a delay between attempts?
  • Should the delay increase between attempts?
  • Is the operation safe to repeat?

The last question is particularly important.

Retrying a read is generally a different problem from retrying an operation that changes state. If an operation creates something, charges something, sends something, or otherwise has side effects, blindly repeating it can create a second operation rather than simply recovering the first one.

So the lesson for me is not “add retries.”

It is:

Retry when the failure looks transient and the operation can be safely retried.

Why retries alone are not enough

Imagine an external dependency is consistently failing.

Without a limit, every incoming request can trigger multiple attempts.

One request becomes several requests to the failing service.

At scale, that can look something like:

The application is spending more work trying to reach something that is already unavailable.

This is where retries can become part of the problem.

A retry mechanism should therefore be designed with the possibility that the dependency may remain unavailable for longer than expected.

That led me to the next question:

What happens when continuing to call the dependency is no longer useful?

Circuit breaker: stopping repeated failure

A circuit breaker introduces another decision.

Instead of allowing every request to keep calling a failing dependency, the application can temporarily stop making those calls after failures cross a defined threshold.

Conceptually:

The useful part is not the name “circuit breaker.”

The useful part is the change in behavior.

Instead of repeatedly discovering that the dependency is unavailable, the application can recognize the failure pattern and stop sending unnecessary traffic for a period of time.

That gives the dependency some breathing room and prevents the application from spending resources repeatedly waiting for failures.

Eventually, the circuit can allow requests through again to determine whether the dependency has recovered.

The important idea is that a circuit breaker is about controlling calls to an unhealthy dependency.

A retry is about giving an individual operation another chance.

They solve different problems.

Retry vs Circuit Breaker

I think about the distinction like this:

Retry

“Maybe this particular failure is temporary.”

The system makes another attempt.

Circuit breaker

“This dependency appears unhealthy right now.”

The system stops repeatedly calling it.

That means they can work together.

A request might retry a small number of times for a transient failure, while the circuit breaker prevents the service from continuing to hammer the dependency when failures become persistent.

The exact thresholds and timing depend on the system. There is no universal number that makes a circuit breaker correct.

The important part is understanding what decision the mechanism is making.

Fallback: doing something useful when the dependency fails

Retries and circuit breakers answer questions about calling the dependency.

A fallback answers a different question:

“What should the application do if that dependency cannot provide the result?”

Sometimes the answer is to return a controlled error.

Sometimes the application can use previously available data.

Sometimes a default response is acceptable.

Sometimes the operation simply cannot continue and the system should communicate that clearly to the caller.

Conceptually:

A fallback should not pretend that the dependency succeeded when it did not.

That distinction matters.

Returning stale or partial information might be acceptable for one type of operation and completely wrong for another.

The fallback has to match the business meaning of the operation.

Putting the three together

The most useful mental model for me is not:

“Which one should I use?”

It is:

“What kind of failure am I dealing with, and what should happen next?”

A simplified flow looks like this:

This is only a conceptual model. Real systems can have more states and more detailed policies.

The important part is that each mechanism has a different responsibility.

  • Retry: Give a potentially transient operation another chance.
  • Circuit breaker: Stop repeatedly calling a dependency that appears unhealthy.
  • Fallback: Provide an alternative behavior when the dependency cannot complete the operation.

The part I find easy to get wrong

It is tempting to think of reliability patterns as things you simply add to a service.

Add retries.

Add a circuit breaker.

Add a fallback.

Done.

But each one changes system behavior.

A retry changes how many times an operation can execute.

A circuit breaker changes whether a dependency is called at all.

A fallback changes what the user or calling service receives when something fails.

Those decisions should therefore be connected to the actual operation.

For example, retrying a request that is safe to repeat is a different problem from retrying an operation with side effects.

Similarly, returning cached data as a fallback might be useful for a read-heavy endpoint but unacceptable when the caller requires the latest authoritative state.

Reliability is therefore not only about keeping requests alive.

It is also about deciding which behavior is safe when the normal path stops working.

Failure handling is also about protecting the system

One thing I increasingly appreciate about these patterns is that they protect more than the individual request.

A retry can help recover from a short-lived problem.

A circuit breaker can prevent a struggling dependency from receiving an increasing number of repeated requests.

A fallback can allow the rest of the system to degrade in a controlled way instead of failing unpredictably.

That means failure handling is partly about protecting the dependency and partly about protecting the application itself.

The goal is not necessarily:

“Never show an error.”

Sometimes the correct behavior is a clear, controlled failure.

A reliable system is not one where failures never happen.

It is one where failures have predictable behavior.

Trade-offs

Retries

Useful for:

  • Transient failures
  • Short-lived network problems
  • Operations that are safe to repeat

Trade-offs:

  • Additional latency
  • Additional traffic
  • Risk of duplicate side effects if the operation is not safe to repeat
  • Can amplify an existing outage if poorly bounded

Circuit breakers

Useful for:

  • Consistently failing dependencies
  • Preventing repeated calls to an unhealthy service
  • Failing fast when continuing to call is unlikely to help

Trade-offs:

  • Additional state and configuration
  • Temporary failures can cause calls to be rejected while the circuit is open
  • Thresholds and recovery behavior need to match the system

Fallbacks

Useful for:

  • Graceful degradation
  • Returning alternative or previously available information
  • Providing controlled behavior when a dependency is unavailable

Trade-offs:

  • The fallback may provide less complete information
  • Stale or partial data can be misleading if used incorrectly
  • Not every operation has a meaningful fallback

None of these patterns is universally required.

The architecture should reflect the dependency, the operation being performed, and what failure means to the application.

What I took away

The biggest shift for me was moving away from thinking about retries as the default answer to every external-service failure.

Sometimes a retry is exactly right.

Sometimes continuing to retry is the wrong thing to do.

Sometimes the system should stop calling the dependency for a while.

Sometimes the right answer is to return something different.

Retries, circuit breakers, and fallbacks are therefore not competing solutions. They are different decisions in the failure path.

The useful question is not:

“Which pattern should I add?”

It is:

“What should this system do when the normal path stops working?”

That question leads to much better backend designs.