Appearance
Terraform Automated Testing
Why Terraform tests center on the state changes a module goes through, and how to tell a useful test from a weak one
Related Concepts: Terraform Change Control | Terraform Module Philosophy | Unit Testing | Implements: Terraform Testing
Modules as Units of Code
A Terraform module is a unit of code, and we test it the way we would test a function or a class in any other language. A module takes inputs, processes them, creates resources from them, and produces outputs. A function does the same with its arguments and its return value.
The closest comparison is a small group of classes that share their data. A single class keeps its state to itself. The resources in a module read the same variables and locals and pass attributes to each other, so the data belongs to the module as a whole. Martin Fowler describes this kind of unit in Unit Test. He calls the choice of unit "a situational thing" and says he often treats a group of closely related classes as one unit. A module fits that description, and most of what Unit Testing says about application code applies to it.
State Changes
The highest priority in a Terraform test suite is the set of state changes a module has to go through. A module is healthy and working as designed when it moves through each of them cleanly, so those transitions are what the tests confirm first.
Two kinds of tests cover these state changes:
- Integration tests spawn real infrastructure. They confirm behavior only the provider and the cloud API can prove, such as IAM policy evaluation or provider-side validation.
- Unit tests use mock providers. They assert logic that happens inside the Terraform engine, such as conditionals,
for_eachexpansion, and string interpolation, and they need no real infrastructure.
The line between the two follows the speed-based definition in Unit Testing. In A Set of Unit Testing Rules, Michael Feathers says a test that "communicates across the network" is an integration test. A test against a mock provider makes no network calls and needs no credentials, so it is a unit test whether it plans or applies. A test with real credentials talks to the cloud API, so it is an integration test.
A module's test suite uses both kinds. Terraform Change Control covers why the cost of real infrastructure pushes most of the suite toward unit tests.
Mock Providers
Terraform unit tests lean heavily on mock providers, and at first glance that breaks the rule in Why Mocking is Usually a Smell. In application code, heavy mocking means business logic is tangled up with infrastructure. The fix is to pull the logic out into pure functions, or put a port in front of the infrastructure and swap in a simple adapter for tests. Terraform gives us neither option, so the reasoning has to change.
The reason is state. A module's state lives in an external system, the cloud provider's API, and nearly every expression in a module ends up feeding a resource the provider has to create. Application code can move its logic away from the database. A module has nowhere to move its logic, because Terraform has no general-purpose functions. Logic that lives outside a resource needs either a custom provider with its own functions, or a child module that holds only variables, locals, and outputs and has no provider at all. We have not adopted the second approach. Every piece of logic would need its own child module, and a module tree built that way gets messy and hard to follow fast.
So the mock provider is a means to an end. It lets us test logic that never depended on the provider without being forced to create real infrastructure, and it does that for two reasons:
- Cost. Real infrastructure costs real money, and a test of a
for_eachexpansion or a naming convention gets nothing for it. - Speed. Many resources take minutes to provision and tear down. A mocked run finishes in seconds.
A mock provider is also a different kind of mock from the ones the Unit Testing article warns about. A jest.fn() mock returns whatever the test author tells it to, and a test can pass on those made-up values while the real code is broken. A mock provider is mostly dumb. It reads the provider's schema and fills in computed attributes with values that only match the type, per the mocking documentation: a random 8 character string for a string, 0 for a number, false for a boolean, and empty collections. Nothing about the format or structure of the real provider's output survives, which is why Terraform Testing notes that mocked ARNs fail AWS validation. A test can still abuse a mock provider through override blocks. In normal use, though, the only thing left to assert against is the module's own logic, which is what a unit test should cover.
The mock provider also sits where Unit Testing allows a mock, at the I/O boundary. The resources inside the module still work with each other for real, so Terraform unit tests stay sociable tests in Fowler's sense.
Plan, Apply, Update, Delete
Every module must handle four states:
- Plan. The module produces a plan without errors for the inputs it supports.
- Apply. The module applies the plan and creates the infrastructure it describes.
- Update. An already applied module accepts a change to its inputs and applies it with no manual overhead. Manual overhead is anything a person does by hand to make the change go through, such as running
terraform state mv, applying with-targetor-replacefirst, or editing a resource in the cloud console. - Delete. The module destroys everything it created.
Kent C. Dodds puts the reason for this focus in one line, quoted in Frontend Testing: "The more your tests resemble the way your software is used, the more confidence they can give you." Engineers use a module by planning it, applying it, changing its inputs, and destroying it. A test that walks through those states resembles real use. A test that checks one attribute in isolation covers a small slice of it.
The native test framework maps onto the four states directly. A run block with command = plan covers the plan. A run block with command = apply covers the apply. Run blocks in the same test file share state, so a second apply run with changed variables exercises an update against the infrastructure the first run created. When the file finishes, the framework destroys the resources in reverse order and reports any failure, which covers the delete.
Weak Tests
It is easy to write Terraform tests that look thorough and prove little. Two patterns come up often.
Testing the Provider
A test that asserts behavior the provider defines entirely tests the provider's implementation. Asserting that an S3 bucket created with versioning enabled reports versioning as enabled checks the AWS provider's code. Application code follows the same rule. A test of a third-party library's behavior belongs in the library's own suite, and Unit Testing asks us to test the behavior our own code adds.
The exception is a module that contains its own logic to handle a provider edge case. Some provider behavior is rigid or surprising, and a module sometimes carries a workaround for it. That workaround is module logic. A provider upgrade or a refactor can break it without warning, so a test that pins it down is a strong test.
Echo Tests
A test that passes a string into a module and asserts the same string comes back out of an output or attribute tests no behavior of the module. We call these vanity tests. They add a few points of coverage and prove nothing about how the module works.
Fowler names the underlying problem in Test Coverage: "high coverage numbers are too easy to reach with low quality testing." When a coverage number becomes the goal, people write the tests that move the number, and echo tests are the cheapest way to move it in Terraform.
An echo test becomes a real test once the module transforms the input. A name built from interpolation, a tag map merged with defaults, or a value chosen by a conditional all carry module logic worth asserting.
Prioritizing Update Tests
Update tests need the most judgment. A module can have dozens of inputs, testing every change to every one of them is expensive, and some inputs matter far more than others. An input earns a high priority for update tests when it meets all three of these conditions:
- It is likely to change often.
- It is important to the core behavior of the module and its intended use case.
- It has cascading effects on the module's internals, such as driving
for_eachkeys, toggling dynamic blocks, or feeding the names of other resources.
Inputs that change rarely, or that pass through to a single attribute with no effect on anything else, are lower priority. We cannot predict every input that will change. Before refactoring how an input flows through a module's internals, though, we can judge whether that input's implementation is settled or likely to move. An input that changes often and ripples through the module is the one most likely to force a replacement or need manual steps on update, so it gets update tests first.
Limitations
Automated tests cover a lot of a module's behavior, but parts of Terraform still work against them. This section records the ones we have run into.
Import Blocks
At the time of writing, a unit test cannot plan a configuration that contains an import block. The mock provider rejects every import on purpose and fails the run with this error:
Cannot import resources from mock providers. Use an `override_resource` block
to targeting the specific resource being imported instead.The override_resource block the error suggests does work, and we don't use it. An import block is temporary. It gets removed once it has applied, so an override written for it is a bespoke test change for code that is about to disappear. Every new import would mean editing every test that loads the configuration, which is not practical. Integration tests have it worse. Against a real provider, the import fails when the object it points at does not exist in the test account, and hashicorp/terraform#34150, which asks for a way to skip imports during tests, is still open.
So imports are not part of what we test. Every import block in a root module lives in imports.tf, and CI deletes that file from its working copy before it runs terraform test. Terraform has no flag to leave a file out of a test run, so removing it is the mechanism. The tests stay the same whether an import is pending or not.
It is an annoyance, and it is also easy to see how it happened. Terraform has no concept of migrations outside the CLI. An import block is syntax for the same engine operation as terraform import, the same way a moved block stands in for terraform state mv. Both exist to deal with state, and a test run starts from an empty state with nothing in it to import. The blocks are a workaround for Terraform being stateful. We think it is an acceptable workaround given what the maintainers are dealing with, and it still makes testing harder. moved blocks have no such problem, since moves happen offline and the mock provider passes them straight through.
The closest comparison in application code is a database migration. Picture two changes in flight that both depend on a change to the shape of a table. Neither change can reconcile shape A with shape B on its own, so a migration runs once and moves the real data from one shape to the other. An import block does that job for Terraform state.
We treat import blocks the way we treat data migrations. Each one is an intentional, one-time change, usually with much lower risk than a data migration. Once it has applied everywhere it needs to, we delete it from imports.tf. We hope the maintainers add a way to run tests with these blocks in place.
Further Reading
- Unit Test by Martin Fowler covers what counts as a unit and the difference between sociable and solitary tests.
- Test Coverage by Martin Fowler covers what coverage numbers can and cannot tell you.
- A Set of Unit Testing Rules by Michael Feathers is the source of the speed-based unit test definition.
- Testing Implementation Details by Kent C. Dodds covers testing from the point of view of the people who use the code.