Skip to main content

Command Palette

Search for a command to run...

When Your Learning Becomes Your Debugging Toolkit

How the Thundering Herd Problem went from something I wrote about to something I recognised and applied in a real system

Updated
•10 min read•View as Markdown
When Your Learning Becomes Your Debugging Toolkit
A
Experienced QA Engineer with a strong background in test automation using modern day tools. Currently expanding into AI-driven testing, web development, and system design to grow into a well-rounded software engineer. I write about testing, development, and scalable system architecture based on hands-on learning and real-world experience.

Earlier this year, I wrote about the Thundering Herd Problem as part of my journey of learning more about distributed systems.

At the time, it was a concept I was studying.

I understood the basic idea: when many clients or processes perform the same operation at nearly the same time, a shared resource can suddenly experience a burst of traffic that it wasn't designed to handle.

I wrote about things like staggering, backoff, jitter, request coalescing, and rate limiting.

What I didn't expect was that, months later, I would encounter a real-world problem where that knowledge would become useful.

And that turned out to be a much more valuable learning experience than simply reading about the concept.

The Problem

I was working with a parallel E2E test suite. Every now and then, some tests would fail unexpectedly.

There wasn't a particular test that consistently failed. One run might fail in one place. Another run might fail somewhere completely different. And when one of those tests was run in isolation, it would pass.

That pattern immediately made me suspicious.

If the same test passes by itself but intermittently fails when the entire suite runs in parallel, the problem may not actually be inside that test. Something about the environment, timing, concurrency, or shared dependencies may be involved. The tests were running concurrently, and several of them were performing similar authentication-related operations around the same time.

That made me ask a question I'd encountered before:

What happens when multiple workers suddenly do the same thing at the same time?

- That sounded familiar.

Recognising the Pattern

This was the moment when a concept stopped being theoretical.

The Thundering Herd Problem is fundamentally about synchronisation. It's not necessarily that one request is bad. It's that many otherwise legitimate requests arrive at nearly the same time.

Think about the difference between:

Request → Request → Request → Request

and:

         ┌── Request
         ├── Request
         ├── Request
         ├── Request
         ├── Request
         └── Request
              |
         Shared Resource

Parallel execution was creating a concentrated burst of similar activity, which was then triggering a security mechanism in the environment.

What initially looked like random test failures started to look much more like a pattern.

The Engineering Question I Asked

My first instinct wasn't to change the infrastructure.

Instead, I started with a simpler question: what could I change from my side to reduce the problem without changing the underlying environment?

That distinction became important because infrastructure changes can have broader implications, while a change within my area of control could be tested quickly and safely. When a problem crosses into infrastructure territory, there may be several possible solutions.

  • You could change application behaviour.

  • You could change infrastructure configuration.

  • You could modify security controls.

  • You could change rate limits.

  • You could add exceptions or allowlists.

Some of those may ultimately be the right solution.

But they can also introduce other considerations: security, cost, operational complexity, coordination, or unintended side effects.

So before changing something outside my immediate area of control, I asked myself:

What can I improve from my side?

That question led me back to something I had already learned. Applying What I Had Learned and shared previously.

One of the techniques I had discussed in my original article was staggering.

Instead of allowing every parallel worker to perform the same operation at exactly the same time, I could introduce a small delay based on the worker's identity.

Conceptually:

Worker 0 → immediately
Worker 1 → wait a little
Worker 2 → wait a little more
Worker 3 → wait a little more
...

The goal wasn't to remove parallelism.

Parallel execution was still important. The goal was to reduce synchronisation.

In other words:

Keep the concurrency, but avoid the herd.

The first iteration helped. The intermittent failures decreased. But they didn't disappear. And that led to the next part of the investigation.

Why Staggering Alone Wasn't Enough?

A deterministic stagger is useful, particularly at the beginning of a run. But a parallel test suite isn't a single event. It runs for a significant amount of time. Different workers finish different tests at different speeds.

  • Some tests take longer.

  • Some take less time.

  • Retries can happen.

  • Execution paths can differ.

Eventually, workers that started at slightly different times can naturally drift back into synchronisation.

Something like this:

Initial execution

W1 ────────────────
W2    ────────────────
W3       ────────────────


Later...

W1 ────────────────┐
W2 ────────────────┤
W3 ────────────────┘
                   ↑
             collision again

This is where jitter becomes useful

Instead of applying exactly the same deterministic delay every time, introduce a small random component.

Conceptually:

delay = worker-based delay + random jitter

The purpose of the randomness isn't randomness for its own sake. It's to break synchronisation.

Two workers are less likely to repeatedly fall into the same rhythm.

The Result:

The second iteration improved the situation further. The intermittent failures became less frequent. It wasn't a magical fix that eliminated every environmental factor. But it significantly reduced the impact of the synchronised workload. More importantly, I now had a much better explanation for the behaviour.

And that distinction matters.

A flaky test suite can easily lead you toward solutions like:

  • increase the timeout

  • add another retry

  • rerun the failed test

  • increase the number of retries

  • ignore the occasional failure

Sometimes those are legitimate steps but they can also hide symptoms.

The more useful question was:

Why are these operations happening together in the first place?

Once I started thinking about the problem in terms of concurrency and synchronisation, the behaviour made much more sense.

  • The Interesting Part Wasn't the Code Change

  • The most interesting part of this experience wasn't adding a delay or jitter.

I'd already written about those techniques.

The interesting part was recognising the pattern.

Before this happened, the Thundering Herd Problem was something I had learned about. After this happened, it became something I had experienced.

There's a big difference.

You can read about a distributed-systems concept and understand its definition. But when you encounter a real system behaving that way, the concept becomes much more tangible.

That's one of the things I enjoy about engineering.

Sometimes something you learned months ago suddenly becomes the tool you need today.

Engineering Beyond the Tool

There's another reason this experience stayed with me.

The problem happened while working with E2E tests, but the underlying reasoning wasn't really about E2E testing.

It was about:

  • concurrency

  • synchronisation

  • request patterns

  • shared resources

  • retries

  • backoff

  • system behaviour

The test framework was simply the environment in which I encountered the problem. That's an important distinction for me.

I don't want to think of automation simply as a collection of tools, frameworks, selectors, assertions, and test scripts.

Good automation is also an opportunity to understand the system you're testing.

And sometimes, the problems you encounter while building or running tests are actually software engineering problems in disguise.

What I Took Away

  1. Parallelism changes system behaviour

Running multiple workers isn't simply a faster version of running one worker. Concurrency changes the shape of the traffic reaching everything downstream.

A system may behave perfectly when requests are distributed over time and behave very differently when the same requests arrive almost simultaneously.

2. Flakiness isn't always a test problem

When different tests fail on different runs, don't immediately assume that the failing tests are broken.

  • Look at what they share.

  • Look at timing.

  • Look at concurrency.

  • Look at dependencies.

Look at what happens when multiple workers perform similar operations simultaneously.

3. Retries can amplify problems

Retries are useful, but they're not free. If the underlying problem is a burst of requests, an immediate retry can simply create another burst. Sometimes the right retry strategy requires backoff and jitter, rather than simply trying again immediately.

4. Small changes can change system behavior

A relatively small delay can have a meaningful impact when it changes the timing of a large number of operations. The objective isn't always to make something faster. Sometimes it's to make the workload more predictable and less synchronised.

5. Fix the part you can control

There may be multiple layers at which a problem can be addressed. But not every problem requires changing every layer. Before reaching for an infrastructure change, I find it useful to ask:

What can I improve within my area of control without compromising an existing design or introducing unnecessary complexity?

Sometimes the smallest change is the most appropriate first step.

6. Learning compounds when you apply it

This is probably my biggest takeaway.

I originally wrote about the Thundering Herd Problem because I was learning about distributed systems.Months later, I recognized a similar pattern during a real debugging exercise.

The knowledge went through a simple progression:

Learn → Understand → Recognise → Apply → Learn more

That's when technical learning becomes much more valuable. You're no longer just collecting concepts.

You're building a mental toolbox.

And occasionally, a real engineering problem gives you the opportunity to reach into that toolbox and use something you learned months ago.

Final Thought

One thing this experience reinforced for me is that there is rarely one absolute fix for a problem like this.

What worked in my situation was a mitigation I could make within my own area of control. By changing how the parallel workload was generated, I was able to reduce the problem without immediately requiring changes to the underlying environment.

But that doesn't mean every problem can—or should—be solved at the code or test level.

Some problems may ultimately require intervention from infrastructure, security, platform, or operations teams. There may be cases where the correct solution is a configuration change, a security-policy adjustment, a capacity change, or something else entirely.

The important part, at least for me, is what happens before we reach that point.

Instead of immediately saying:

“This is an infrastructure problem.”

or

“Someone else needs to fix this.”

I think we should first try to understand what is actually happening, why it is happening, and what part of the problem we can influence from our side.

Sometimes we can solve it ourselves.

Sometimes we can reduce the impact.

And sometimes, after understanding the problem properly, we can provide the right context to the team that needs to intervene.

For me, that's the bigger lesson from this experience:

Understand the problem first. Try to solve what you can control. And when the solution belongs to another layer, escalate with understanding rather than assumptions.

I wrote about the Thundering Herd Problem as a concept.

Months later, I encountered a real situation where that knowledge helped me reason about a problem and make a meaningful improvement.

That experience tied everything together for me—it showed that learning a concept is only the first step. The real value comes later, when you recognise a familiar pattern in a different context and are able to apply that understanding to make a practical difference.

That's the kind of learning I want to keep pursuing—not just learning more tools, but becoming better at understanding systems, recognising patterns, making informed trade-offs, and solving problems wherever I can contribute.

Sometimes, your best debugging tool isn't another tool at all.It's something you learned a few months ago.

System Design Fundamentals for Engineers

Part 3 of 3

A beginner-friendly system design series documenting my journey as an SDET learning distributed systems. Topics include the thundering herd problem, caching strategies, traffic spikes, and scalability patterns used in modern systems.

Start from the beginning

⚡ Understanding the Thundering Herd Problem in Distributed Systems

Introduction: A Simple Real-World Analogy Before e-commerce became mainstream, people would queue outside stores during special sales like Black Friday. The moment the doors opened, everyone rushed in

More from this blog

A

Anuvrat.Dev-Test

3 posts

This publication is about learning, engineering, and solving real-world problems.

Expect insights on software engineering, system design, distributed systems, automation, AI, debugging, reliability, and trade-offs—based on real problems and ongoing learning.

Great engineers don’t just use tools; they understand how systems work, why they fail, and how to reason about solutions.

I’ll share what I learn, build, break, and the lessons behind it.