Tricks for better developability Part I - Testing
When Atlas Principle was born, I was roughly 15 years into the world of software. During this time I met with many approaches that didn't work, and a couple that were genius, but were way above my level at the time.
Since the purpose of Atlas Principle is to help you put your software in front of customers at a consistent pace and quality, I believe what I picked up over the years will be helpful to you.
Disclaimer: The code you will find here is pseudo-code, the syntax might be a mix of javascript, rust, and python. The implementations are for demonstration purposes only.
Let's build this topic up from the point of view of an organization that's scrambling to get their failing deployments under control. Every time a new build goes live developers take shelter from the inevitable urgent bug reports filed by customers who were in the middle of their work and now face a broken application.
First instinct: end-to-end tests
Naturally, many companies reach for end-to-end (E2E) tests. I'm sure you're already aware that these tests are one of the most complex and time-consuming ones to maintain. They rely on certain text elements on the website to be found, visible, and/or even clickable. A slower (or faster) network during testing can break your tests, and if all is well, sometimes the headless browser engine you use will act differently enough that you'll chase a hard to reproduce issue for hours.
Later on I'll offer some insight on what you can do to have the confidence in releasing with less moving parts, but let's assume that E2E tests are really your best option for now.
Common mistakes I've seen with E2E tests
Tying tests to copy
Let's assume you have a save button, and you'd like to make sure there's some sort of feedback if it worked:
it('should save the form', () => {
const saveButton = findElement(By.text('Save'));
saveButton.click();
expect(findElement(By.text('Save successful'))).toBeDefined();
});
As soon as you throw internationalization (i18n) in the mix, you have a potential of these tests breaking simply because the environment uses a different language by default than what's expected in the tests. Or if instead of "Save successful" you output "Your form has been submitted". These are the wrong reasons a test should fail.
It's an improvement if you write all your copy in English, and you make sure there's some configuration for these tests that will force that language.
It's even better if you use the data-testid attribute in your HTML. Your test will be similarly obvious, but now you can freely change your copy without it ever breaking your tests
it('should save the form', () => {
const saveButton = findElement(By.testId('saveButton'));
saveButton.click();
expect(findElement(By.testId('successFeedback'))).toBeDefined();
});
Mocking and simulating
In some cases E2E tests use watchers and mocks to avoid sending real data and waiting for the actual response from a backend. I think this mostly devolves E2E tests into automated UI tests, which can still be valuable, but not nearly as much. The general problem I have with mocking, is you can make it give you anything you want, the application's existing logic and guardrails don't apply. This makes it trivial to invent ideal world scenarios that test nothing in reality.
I'd suggest setting up a complete system in your testing pipeline with comprehensive test data, so that you can meaningfully simulate what would happen if the current batch of changes was to run on production.
You might need to resort to listening to some network requests completing, to make sure you don't wait based on a flat timer, but assert a certain thing has happened once a given response has finished. This is also very brittle. Sometimes you have a setup such as Finish(A) -> Finish (B) -> assert; if the reality is Finish (B) -> Finish (A) you might end up in a stuck state, as from the tests point of view there is no request B completing after A.
What's better than E2E tests then?
We've established what E2E tests are, bringing up your entire stack, checking it through a browser, and making sure everything works as you'd expect.
Let's keep pushing in a direction that will result in faster, more reliable tests. It can be argued they will never provide you the proof that if you click it in the browser it will work. The point of automated tests is: confidence in releasing. Whatever worked yesterday will continue to work, and whatever I'm changing now will work as expected.
Component tests
If you develop a Single-Page Application (SPA), you probably can rely on component tests instead of E2E ones.
A component test renders only the "unit under test" in JavaScript. You don't need a backend, you don't need networking, and you don't render it in a browser. That's already 3 entire system parts you no longer rely on.
Let's contrast it to the E2E test for save confirmation:
it('should save the form', () => {
const storage = new InMemoryStorage<Profile, ProfileFilters>(); // explained in the "Fakes, not mocks section"
const component = ProfileForm.render({ storage, data: sentinel.formData });
const saveButton = component.getByTestId('saveButton');
saveButton.click();
expect(storage).toContain(sentinel.formData);
expect(component.getByTestId('feedback')).toHaveTextContent('forms.saveSuccess');
});
This should test the exact same logic as your E2E test, and by a rough estimation I'd say it's almost a second faster. A second might not sound much, but once you have around a hundred tests, the difference is running your entire suite under 2 minutes versus 2 seconds. Most of the mature projects I've worked with had around 2-5 thousand tests. A proper testing strategy and attention to these design choices will be the difference between a 5-minute test run and a 45-minute one.
It's important to clarify that these two types of tests are not equal. E2E tests without any mocks really ensure close to 100% that your application is working as expected. E2E with mocks in my opinion is testing ~90% that your frontend is working (as with mocks you won't see if you're sending the wrong data, using the wrong endpoint, or if you're getting an unexpected response). Component tests however are almost identical to mocked E2E tests in the confidence they give you about the health of your release.
Given how often these tests should run, I strongly suggest investing the time upfront to make sure your tests stay fast.
For comparison, Atlas Principle's current version has these numbers (with ~84% coverage):
Test Files 717 passed (717)
Tests 4537 passed (4537)
Start at 14:02:38
Duration 13.01s (environment 93%, tests 4%, transform 2%, import 1%)
Better than component tests?
Maybe you work with Server-Side Rendered (SSR) applications and don't have access to component tests. Maybe you just want to have tests that are even faster than component tests?
In that case I'd strongly suggest (and this will be mentioned again later in this same post) investing into types and interfaces. Let's keep abusing the same example scenario
enum FormFeedback {
Saved = 'forms.saveSuccess',
Failed = 'forms.saveFailed',
}
it('should store the profile', () => {
const storage = new InMemoryStorage<Profile, ProfileFilters>();
const form = new Form({ storage, data: makeProfile({ name: 'Ada' }) });
form.save();
expect(storage.items).toContainEqual(form.data);
});
it('should report a successful save', () => {
const form = new Form({ storage: new InMemoryStorage(), data: makeProfile() });
const result: FormResult = form.save();
expect(result.feedback).toBe(FormFeedback.Saved);
});
What this tries to demonstrate, is that you can move away from some logic in your templates: move the logic into the class/function that gives you an outcome, a state, or a snapshot if you will. Your templates will go from
<div>
<p v-if="form.feedback === FormFeedback.Saved">
{{ t("forms.saveSuccess") }}
</p>
<p v-else>
{{ t("forms.saveFailed") }}
</p>
</div>
to be more like
<p data-testid="feedback">{{ t(form.feedback) }}</p>
to keep your tests quick, independent, and repeatable.
Side note about TDD
It's no secret I'm a fan. TDD is one of the things that I tried (well was guided through really), and was immediately sold on the workflow and the benefits of it. If you don't do TDD you can still write testable code, it's just way easier with following the red-green-refactor cycle.
Anyway, the only danger I wanted to warn you about when not following TDD, is when sometimes your tests pass for the wrong reasons, and effectively test nothing of importance.
class Invoice {
// ... omitted for brevity
update() {
if (this.amount == null) {
this.status = InvoiceStatus.Draft
}
}
}
it('should set the invoice status to DRAFT without an amount', () => {
const invoice = factory.makeInvoice();
invoice.amount = null;
invoice.update();
expect(invoice.status).toBe(InvoiceStatus.Draft);
});
If you write your implementation first or at the same time with your tests, you see it's passing, and you move on... but wait! Remove the change we just added and the test still passes, how?
You couldn't have known (and it's kind of the point), but the default status in the invoice factory is draft. This is also a common pattern you might encounter in medium to large size applications, there are some defaults or some assumptions that will slip by you.
Had you written the test first, and ran it before any of your production code was altered, you would've seen the classic "wtf" moments in TDD: when your test is unexpectedly passing. In my own anecdotal experience this happens 5-10% of the time, which is not a lot, but certainly more than never.
What's even better than tests?
This won't work in every programming language, but what I really enjoy doing is modeling states with types. I'm sure you've heard about "making invalid states unrepresentable", and that's exactly what I'm talking about!
Let's assume you have several places where you check whether a user is a staff member to determine if they're allowed to perform an action:
// these tests are repeated across every operation
it('should raise an error if the account is not of a staff member', () => {
const account = factory.makeAccount({isStaff: false});
const result = () => performAction(account, ...inputs);
expect(result).toThrow(OperationForbidden);
});
A better way in my opinion is to express this constraint via types.
class Customer {
readonly isStaff = false;
// ... omitted for brevity
}
class StaffMember {
readonly isStaff = true;
// ... omitted for brevity
}
type Account = Customer | StaffMember;
function toAccount(row: AccountRow): Account {
if (row.isStaff) {
return new StaffMember(row);
}
return new Customer(row);
}
and in your views
function performAction(account: StaffMember, ...inputs) { ... }
the compiler will make sure you can't perform an action anymore with a regular account, only staff is accepted. This might perform worse in languages where it's easy to cast to any other type, or if the type system is weaker. TypeScript would allow customer as unknown as StaffMember to be passed in for example.
It's not just for guards though, if you move away from primitive obsession, you can also model values with the help of Type-Driven Development, and they can each know about how to present themselves in your application.
// instead of
const amount: number = -5;
<div>
<p v-if="change > 0">+{{ amount }}</p>
<p v-else-if="amount === 0">0</p>
<p v-else>{{ amount }}</p>
</div>
// you could
class Income {
constructor(readonly amount: number) {}
toRepresentation() {
if (this.amount === 0) {
return '0';
}
return `+${this.amount}`;
}
}
class Expense {
constructor(readonly amount: number) {}
toRepresentation() {
return `-${this.amount}`;
}
}
type Amount = Income | Expense;
const change: Amount = new Income(20);
<div>
{{ change.toRepresentation() }}
</div>
Again, instead of relying on UI tests, the code itself can be tested in isolation, faster, and more easily.
Closing thoughts on tests
In addition to the above, I'd like to reiterate two things we already touched on briefly:
Fakes, not mocks & Dependency Injection
Mocks are lies, I can mock my authentication function to respond with a picture of a kitten. While amusing, I don't think it's particularly useful.
The way I prefer this, and I frequently do it with persistence/databases, is to have a common interface for the "concept" you're building with concrete implementations for every variation you support:
interface Persistence<E, F> {
load(filters: F): E[];
save(entity: E): E;
}
class InMemoryStorage<E, F> implements Persistence<E, F> {
readonly items: E[] = [];
load(filters: F): E[] {
// ... find in the array
}
save(entity: E): E {
this.items.push(entity);
return entity;
}
}
class PostgresStorage<E, F> implements Persistence<E, F> {
// ... omitted for brevity
}
function saveAccount(account: Account, storage: Persistence<Account, AccountFilters>) {
// ... omitted for brevity
}
// production
saveAccount(account, new PostgresStorage(db));
// tests
saveAccount(account, new InMemoryStorage());
What this enables you to do is to have a production implementation saving and loading to and from a Postgres database for example, while your tests can inject a much faster, dependency and infrastructure free in-memory storage class to test your logic.
If your first thought is that now database specific queries and operations aren't tested, you would be right! Those would deserve their own integration tests. Then again since we don't need to test every operation with a real database, a handful of integration tests for retrieving and storing data are sufficient.
So what do you start with?
Some of the advice here requires massive effort or tedious groundwork. My suggestion would be to find what's easiest for you to get started with, avoid the pitfalls mentioned, and have a clear plan on which level you'd like to arrive at, and how to get there.
Most of the advice here is based on what I've seen, and what pain points I detected in various teams, projects, and companies. It still isn't a one-size-fits-all solution for everyone.
This concludes Part I of this series, in future episodes we'll touch on:
- Screaming architecture
- Trunk-based development and in general why PRs are not always required
- Pushing mutability to both ends of your pipelines while keeping the core functional
- Coupling and Cohesion
- Painfully strict linters
- CQRS-lite
- Optimizing for code deletion