When Your Learning Becomes Your Debugging Toolkit
How the Thundering Herd Problem went from something I wrote about to something I recognised and applied in a real system

Earlier this year, I wrote about the Thundering Herd Problem as part of my journey of learning more about distributed systems.
At the time, it was a concept I was studying.
I understood the basic idea: when many clients or processes perform the same operation at nearly the same time, a shared resource can suddenly experience a burst of traffic that it wasn't designed to handle.
I wrote about things like staggering, backoff, jitter, request coalescing, and rate limiting.
What I didn't expect was that, months later, I would encounter a real-world problem where that knowledge would become useful.
And that turned out to be a much more valuable learning experience than simply reading about the concept.
The Problem
I was working with a parallel E2E test suite. Every now and then, some tests would fail unexpectedly.
There wasn't a particular test that consistently failed. One run might fail in one place. Another run might fail somewhere completely different. And when one of those tests was run in isolation, it would pass.
That pattern immediately made me suspicious.
If the same test passes by itself but intermittently fails when the entire suite runs in parallel, the problem may not actually be inside that test. Something about the environment, timing, concurrency, or shared dependencies may be involved. The tests were running concurrently, and several of them were performing similar authentication-related operations around the same time.
That made me ask a question I'd encountered before:
What happens when multiple workers suddenly do the same thing at the same time?
- That sounded familiar.
Recognising the Pattern
This was the moment when a concept stopped being theoretical.
The Thundering Herd Problem is fundamentally about synchronisation. It's not necessarily that one request is bad. It's that many otherwise legitimate requests arrive at nearly the same time.
Think about the difference between:
Request → Request → Request → Request
and:
┌── Request
├── Request
├── Request
├── Request
├── Request
└── Request
|
Shared Resource
Parallel execution was creating a concentrated burst of similar activity, which was then triggering a security mechanism in the environment.
What initially looked like random test failures started to look much more like a pattern.
The Engineering Question I Asked
My first instinct wasn't to change the infrastructure.
Instead, I started with a simpler question: what could I change from my side to reduce the problem without changing the underlying environment?
That distinction became important because infrastructure changes can have broader implications, while a change within my area of control could be tested quickly and safely. When a problem crosses into infrastructure territory, there may be several possible solutions.
You could change application behaviour.
You could change infrastructure configuration.
You could modify security controls.
You could change rate limits.
You could add exceptions or allowlists.
Some of those may ultimately be the right solution.
But they can also introduce other considerations: security, cost, operational complexity, coordination, or unintended side effects.
So before changing something outside my immediate area of control, I asked myself:
What can I improve from my side?
That question led me back to something I had already learned. Applying What I Had Learned and shared previously.
One of the techniques I had discussed in my original article was staggering.
Instead of allowing every parallel worker to perform the same operation at exactly the same time, I could introduce a small delay based on the worker's identity.
Conceptually:
Worker 0 → immediately
Worker 1 → wait a little
Worker 2 → wait a little more
Worker 3 → wait a little more
...
The goal wasn't to remove parallelism.
Parallel execution was still important. The goal was to reduce synchronisation.
In other words:
Keep the concurrency, but avoid the herd.
The first iteration helped. The intermittent failures decreased. But they didn't disappear. And that led to the next part of the investigation.
Why Staggering Alone Wasn't Enough?
A deterministic stagger is useful, particularly at the beginning of a run. But a parallel test suite isn't a single event. It runs for a significant amount of time. Different workers finish different tests at different speeds.
Some tests take longer.
Some take less time.
Retries can happen.
Execution paths can differ.
Eventually, workers that started at slightly different times can naturally drift back into synchronisation.
Something like this:
Initial execution
W1 ────────────────
W2 ────────────────
W3 ────────────────
Later...
W1 ────────────────┐
W2 ────────────────┤
W3 ────────────────┘
↑
collision again
This is where jitter becomes useful
Instead of applying exactly the same deterministic delay every time, introduce a small random component.
Conceptually:
delay = worker-based delay + random jitter
The purpose of the randomness isn't randomness for its own sake. It's to break synchronisation.
Two workers are less likely to repeatedly fall into the same rhythm.
The Result:
The second iteration improved the situation further. The intermittent failures became less frequent. It wasn't a magical fix that eliminated every environmental factor. But it significantly reduced the impact of the synchronised workload. More importantly, I now had a much better explanation for the behaviour.
And that distinction matters.
A flaky test suite can easily lead you toward solutions like:
increase the timeout
add another retry
rerun the failed test
increase the number of retries
ignore the occasional failure
Sometimes those are legitimate steps but they can also hide symptoms.
The more useful question was:
Why are these operations happening together in the first place?
Once I started thinking about the problem in terms of concurrency and synchronisation, the behaviour made much more sense.
The Interesting Part Wasn't the Code Change
The most interesting part of this experience wasn't adding a delay or jitter.
I'd already written about those techniques.
The interesting part was recognising the pattern.
Before this happened, the Thundering Herd Problem was something I had learned about. After this happened, it became something I had experienced.
There's a big difference.
You can read about a distributed-systems concept and understand its definition. But when you encounter a real system behaving that way, the concept becomes much more tangible.
That's one of the things I enjoy about engineering.
Sometimes something you learned months ago suddenly becomes the tool you need today.
Engineering Beyond the Tool
There's another reason this experience stayed with me.
The problem happened while working with E2E tests, but the underlying reasoning wasn't really about E2E testing.
It was about:
concurrency
synchronisation
request patterns
shared resources
retries
backoff
system behaviour
The test framework was simply the environment in which I encountered the problem. That's an important distinction for me.
I don't want to think of automation simply as a collection of tools, frameworks, selectors, assertions, and test scripts.
Good automation is also an opportunity to understand the system you're testing.
And sometimes, the problems you encounter while building or running tests are actually software engineering problems in disguise.
What I Took Away
- Parallelism changes system behaviour
Running multiple workers isn't simply a faster version of running one worker. Concurrency changes the shape of the traffic reaching everything downstream.
A system may behave perfectly when requests are distributed over time and behave very differently when the same requests arrive almost simultaneously.
2. Flakiness isn't always a test problem
When different tests fail on different runs, don't immediately assume that the failing tests are broken.
Look at what they share.
Look at timing.
Look at concurrency.
Look at dependencies.
Look at what happens when multiple workers perform similar operations simultaneously.
3. Retries can amplify problems
Retries are useful, but they're not free. If the underlying problem is a burst of requests, an immediate retry can simply create another burst. Sometimes the right retry strategy requires backoff and jitter, rather than simply trying again immediately.
4. Small changes can change system behavior
A relatively small delay can have a meaningful impact when it changes the timing of a large number of operations. The objective isn't always to make something faster. Sometimes it's to make the workload more predictable and less synchronised.
5. Fix the part you can control
There may be multiple layers at which a problem can be addressed. But not every problem requires changing every layer. Before reaching for an infrastructure change, I find it useful to ask:
What can I improve within my area of control without compromising an existing design or introducing unnecessary complexity?
Sometimes the smallest change is the most appropriate first step.
6. Learning compounds when you apply it
This is probably my biggest takeaway.
I originally wrote about the Thundering Herd Problem because I was learning about distributed systems.Months later, I recognized a similar pattern during a real debugging exercise.
The knowledge went through a simple progression:
Learn → Understand → Recognise → Apply → Learn more
That's when technical learning becomes much more valuable. You're no longer just collecting concepts.
You're building a mental toolbox.
And occasionally, a real engineering problem gives you the opportunity to reach into that toolbox and use something you learned months ago.
Final Thought
One thing this experience reinforced for me is that there is rarely one absolute fix for a problem like this.
What worked in my situation was a mitigation I could make within my own area of control. By changing how the parallel workload was generated, I was able to reduce the problem without immediately requiring changes to the underlying environment.
But that doesn't mean every problem can—or should—be solved at the code or test level.
Some problems may ultimately require intervention from infrastructure, security, platform, or operations teams. There may be cases where the correct solution is a configuration change, a security-policy adjustment, a capacity change, or something else entirely.
The important part, at least for me, is what happens before we reach that point.
Instead of immediately saying:
“This is an infrastructure problem.”
or
“Someone else needs to fix this.”
I think we should first try to understand what is actually happening, why it is happening, and what part of the problem we can influence from our side.
Sometimes we can solve it ourselves.
Sometimes we can reduce the impact.
And sometimes, after understanding the problem properly, we can provide the right context to the team that needs to intervene.
For me, that's the bigger lesson from this experience:
Understand the problem first. Try to solve what you can control. And when the solution belongs to another layer, escalate with understanding rather than assumptions.
I wrote about the Thundering Herd Problem as a concept.
Months later, I encountered a real situation where that knowledge helped me reason about a problem and make a meaningful improvement.
That experience tied everything together for me—it showed that learning a concept is only the first step. The real value comes later, when you recognise a familiar pattern in a different context and are able to apply that understanding to make a practical difference.
That's the kind of learning I want to keep pursuing—not just learning more tools, but becoming better at understanding systems, recognising patterns, making informed trade-offs, and solving problems wherever I can contribute.
Sometimes, your best debugging tool isn't another tool at all.It's something you learned a few months ago.


