Example-Based Tests
Given X, expect Y. The fundamental behavioral sensor.
Fundamentally different from coverage: coverage says “this code executed.” A behavioral assertion says “this code produced the right result.” That is a massive distinction.
The weakness
Example-based tests are only as good as the examples chosen. They test what the author thought to test. Property-based testing, mutation testing and metamorphic testing exist to find the gaps.
In practice
A reading is an assertion failure with both sides of the comparison spelled out:
FAILED tests/test_pricing.py::test_discount_applies_above_threshold
tests/test_pricing.py:42: in test_discount_applies_above_threshold
assert price_with_discount(100, tier="gold") == 90.0
E AssertionError: assert 95.0 == 90.0
E + where 95.0 = price_with_discount(100, tier='gold')
================== 1 failed, 213 passed in 4.21s ==================
Reading it well takes three habits:
- Read the assertion before the code. The failure shows two values; the first question is which one is wrong. The assertion encodes the intended behavior, and sometimes it is the test that needs fixing, not the code.
- One failure is a signal, not a ratio. “1 failed, 213 passed” is not 99.5% healthy. The pass count is context; the failure is the reading.
- Separate the verdict from the harness. An assertion error is a behavioral reading. A collection error, import error, or fixture timeout is the harness breaking, which means the suite produced no reading at all.
How it gets gamed
The suite belongs to the same people as the code, so it can be bent:
- Delete or skip the failure.
skip,xfail, and “I will fix it tomorrow” turn red to green without touching the code. A rising skip count is the suite telling you it is being silenced. - Weaken the assertion. Replacing
== 90.0withis not Nonemakes the test pass and the sensor blind. The test still runs, so the suite looks healthy while detecting less. - Pin the symptom. Hard-coding the current output into the expectation turns the test into a tautology that passes for any implementation.
- Sample only the happy path. Choosing examples that avoid the buggy branch keeps the suite green and the bug alive. Mutation testing is the sensor for this gap.
The meta-signal is the skip count plus the share of test-file diffs that weaken an assertion. Neither shows up in coverage.
Response playbook
When a test fails:
- Reproduce it locally. A failure you cannot reproduce is not yet a finding; it is a question. Get the exact input and the exact assertion before anything else.
- Decide which side is wrong. The assertion encodes intended behavior. If the code violated it, fix the code. If the expectation was wrong, fix the test and say why, because a silently rewritten expectation is a deleted sensor.
- Fix the cause, then re-run the full suite. A fix that breaks a sibling test has revealed a second assumption you did not know you had.
- Add the neighbor cases. If
100failed, the values just above and below the boundary are now suspects. The failure is a free tour of the edge of the behavior. - Never delete the test to unblock a merge. If the test is wrong, rewrite it with the reason in the commit. If it is right, the merge is not done.
What it cannot detect
Missing behavior, untested edge cases, and integration failures that emerge only when components are connected.
Related sensors
References
Publications
- Techniques for Improving Regression Testing in Continuous Integration Development Environments —
- On the Effectiveness of the Test-First Approach to Programming —
- Software Testing Techniques —
Tooling
- pytestPython testing framework
- JestJavaScript testing framework
- JUnitJava testing framework
- Go testingGo's built-in testing package