The best and the worst kinds of tests for microservices
Which kinds of tests should we avoid in microservices, and which should we adopt? This article explores a question that matters a lot for a healthy project.
There’s no doubt that good tests are indispensable for ensuring the quality of a microservice. What many people don’t realize is that some categories of tests can hurt a team’s productivity and should be avoided. There are also categories of tests with an absurdly high return on investment. In this article, I explore this landscape to show which kinds of tests work best for microservices.
For the last three years I’ve been working with microservices at AmbevTech, and one thing became clear along the way: the classic test pyramid is a poor fit for microservices.

When it was conceived, in the early 2000s, the idea was that unit tests were easier to write and faster to run than other kinds, such as integration tests. That was before Docker, before microservices. It was a different reality, still ruled by large applications.
But when a microservice has just a few thousand lines of code, it’s so small that unit tests inevitably end up coupled to implementation details. They become anemic tests.
Anemic tests
Anemic tests are tests that rely too heavily on mock setup to work. Overusing mocks is an anti-pattern known as Mockery.
Anemic tests usually have the following traits:
- They’re hard to maintain
- They can be hard to understand
- They tend to break often
- They discourage refactoring
That last point matters a lot: anemic tests break at the slightest thing. A small, well-intentioned refactoring can break tests simply because some of them cared more about a class’s inner workings than about the result of an operation.
Good tests should support and encourage refactoring. After all, refactoring is desirable in any project, since it’s vital for keeping the codebase healthy over the project’s lifetime.
So if tests break even under refactorings that change neither the behavior nor the results of the processes, those tests are in the way, hurting more than helping.
A good test should give the developer confidence that their refactoring isn’t accidentally changing the business rules. If the test breaks during a refactoring, it becomes useless, and deleting it may be better in the long run.
Apparently, at Spotify they avoid calling these tests “unit tests”, preferring the term “implementation detail tests”.
Quite often, the coupling between the test suite and the application will look like this:

What is the best unit of testing?
In projects that use object-oriented design, there’s a tendency to treat classes as the units to be tested. In procedural or functional projects, the tendency is to treat functions as the unit of testing.
But as this article by Fabio Pereira points out, the real units are behaviors. That’s an important statement, because it points to BDD (Behavior Driven Development) as a potential practice for breaking this deadlock.
In Kent Beck’s book Test Driven Development: By Example, we find a good definition of unit tests. He defines them as “tests that are independent of each other – meaning that running one shouldn’t affect the others”.
In other words, they have nothing to do with testing classes or functions, and everything to do with testing isolated units deterministically: every time we run a test, the result must be consistently the same.
Using BDD naturally leads to a way out of this deadlock: testing each behavior in isolation from the others, focusing on inputs and outputs, leads us to tests that are independent of implementation details.
Behavior-focused tests usually have the following traits:
- They’re easy to maintain
- They’re easy to understand
- They’re resilient to refactoring
- They encourage refactoring
They’re basically the antithesis of anemic tests, with exactly the traits we look for in an agile development environment that aims for quality and productivity.
In this Spotify article on testing microservices, they state that integration tests offer the best cost-benefit ratio for microservices:

These two articles by Kent C. Dodds also discuss this topic in the world of JavaScript applications and are very interesting reads:
Tests like these will tend to be what Martin Fowler likes to call sociable tests 1:

Integration tests as the preferred way to test
Integration tests that apply BDD properly are usually quite easy to understand and cheap to maintain. They also stay true to the business rules for longer.
Still according to the same Spotify article, their tests usually aren’t any more complex than this:

On several projects at AmbevTech, we use the MediatR library to standardize how interactions with the application are invoked, which makes it easier to write tests with these traits. Our tests usually look like this:
public class ListagemBonificacoesTests : ApplicationTestBase
{
[Fact]
public async Task Deve_ordenar_pelo_id_decrescente_quando_nao_for_especificada_nenhuma_ordenacao()
{
InsertMany(ListaBonificacoesComIdsSequencias());
ListarBonificacoesPaginadasRequest requestSemOrdenacao = new();
ListaBonificacoesPaginadasResponse response = await Handle<ListarBonificacoesPaginadasRequest>(requestSemOrdenacao);
var bonificacoes = response.Items;
bonificacoes.Should().BeInDescendingOrder();
}
} Either way, it pays to standardize how we interact with the application, so the tests can focus on those interactions with well-defined inputs and limit themselves to checking the outputs and any side effects, such as records inserted into the database or messages published to a broker.
Following these steps, we end up with a reliable test suite that’s well decoupled from implementation details, as illustrated in the figure below:

Conclusions
As we’ve seen, unit tests in microservices tend to create coupling with implementation details, so they should be reserved for specific, highly complex components, or ones that are hard to test any other way, such as a call to an external service.
On the other hand, the preferred way to test microservices is behavior-focused integration tests. Applying BDD techniques brings countless benefits: above all, ease of maintenance and reliability, supporting refactoring and the evolution of the services.
These conclusions come from my studies and reflections over the last few years on a subject I’m very fond of, and I hope they’re useful to you. Now, dear reader, I invite you to leave a comment below. Thanks for reading!
Footnotes
-
The terms “sociable tests” and “solitary tests” were coined by Jay Fields, in the book “Working Effectively with Unit Tests” ↩