An industry resource
What independent observations
would cause us to believe
this software is correct?
Software is increasingly an opaque artifact. We cannot — and often do not want to — fully understand every implementation. The Software Observatory is a catalog of epistemic sensors: the observable signals that reduce uncertainty about whether a system is correct, maintainable, and behaving as intended. Not "code quality metrics." Measurement instruments pointed at different failure modes.
No single sensor measures correctness. Coverage measures execution. Mutation testing measures test sensitivity. Types measure a particular class of structural inconsistency. Contracts measure boundary assumptions. Observability measures what actually happened and preserves enough dimensionality to investigate unknown unknowns.
They are all measurement instruments pointed at different failure modes.
10 sensor families
The catalog is organized into 10 families, each asking a different question about the system. Each family rules out a different way the software could be wrong. No single sensor is sufficient, and neither is a pile of them: a sensor earns its place by eliminating a doubt the others leave open, not by adding another green check.
Structural
"Is this artifact internally coherent?"
Compiler, type checker, linter, formatter, schema validator
02Behavioral
"Does it do what we expect?"
Unit, integration, E2E, contract, snapshot tests
03Test Effectiveness
"Do our tests actually detect failures?"
Coverage, diff coverage, mutation testing
04Invariants
"What must always be true?"
Balance >= 0, every FK valid, every request has one ID
05Adversarial
"Can we make our evidence of correctness fail?"
Fuzzing, fault injection, chaos, metamorphic testing
06Runtime
"What is it actually doing?"
Logs, traces, metrics, profiles, high-cardinality events
07Change
"What did this change actually affect?"
API compatibility, canary, shadow traffic, error budget, A/B testing
08Architecture
"Is the system becoming harder to reason about?"
Dependency graphs, coupling, fitness functions, hotspots
09Evolution
"Does this look like changes that caused trouble before?"
Revert rate, regression rate, churn, incident correlation
10Human Comprehension
"Can another observer understand and challenge this?"
Review, explainability tests, documentation drift, onboarding
The confidence landscape
No single sensor is sufficient, so there is no total ordering across sensors — no "best" sensor. But each dimension (oracle strength, latency, scope) is a partial order, and the atlas's left-to-right axis is time, not quality. Each one trades feedback latency — how long you wait for the signal — against efficacy: how much the signal can actually tell you. Compilation is instant and definitive about validity; user outcomes are slow and definitive about everything that matters. Most sensors live somewhere in between.
Hover a point to name it. Click to open the entry.
Sensor list (latency × efficacy)
| Sensor | Family | Feedback latency | Efficacy (oracle) |
|---|---|---|---|
| A/B Testing | Change | days | high |
| API Compatibility | Change | seconds | high |
| Boundary Sensors | Architecture | seconds | high |
| Branch Coverage | Test Effectiveness | seconds | low |
| Build Provenance & SBOM | Structural | minutes | high |
| Business Invariants | Invariants | hours | high |
| Canary Analysis | Change | minutes | high |
| Change Coupling | Evolution | days | medium |
| Compiler | Structural | milliseconds | maximum |
| Continuous Profiling | Runtime | seconds | medium |
| Contract & Refinement Types | Structural | milliseconds | high |
| Contract Tests | Behavioral | minutes | high |
| Database Invariants | Invariants | minutes | high |
| Decision Provenance | Human Comprehension | days | low |
| Dependency Graph | Architecture | seconds | low |
| Diff Coverage | Test Effectiveness | minutes | low |
| Differential Testing | Adversarial | minutes-hours | high |
| Distributed Traces | Runtime | seconds | medium |
| Documentation Drift | Human Comprehension | days | low |
| DORA Metrics | Evolution | days | medium |
| Error-Budget Impact | Change | hours | medium |
| Escaped Defect Rate | Test Effectiveness | months | medium |
| Example-Based Tests | Behavioral | seconds | high |
| Fault Injection | Adversarial | minutes | medium |
| Feature Flag Exposure Telemetry | Change | seconds | high |
| Architecture Fitness Functions | Architecture | seconds | high |
| Fuzzing | Adversarial | minutes-hours | high |
| Hotspot Analysis | Architecture | minutes | low |
| Incident Correlation | Evolution | weeks | medium |
| Incremental Build Correctness | Change | minutes | medium |
| Independent Review | Human Comprehension | hours | medium |
| Integration Tests | Behavioral | minutes | high |
| Line Coverage | Test Effectiveness | seconds | low |
| Linter | Structural | milliseconds | medium |
| Live Chaos Experiments | Adversarial | hours | high |
| Live Service Graph Discovery | Architecture | minutes | high |
| Load Testing | Runtime | minutes | medium |
| Metamorphic Testing | Adversarial | minutes | high |
| Model Checking | Structural | minutes | maximum |
| Mutation Testing | Test Effectiveness | minutes-hours | high |
| Observability Events | Runtime | seconds | medium |
| Onboarding Experiment | Human Comprehension | weeks | low |
| Pre-Promotion Invariant Gates | Invariants | minutes | high |
| Property-Based Testing | Adversarial | seconds | high |
| Resource Telemetry | Runtime | seconds | low |
| Revert Rate | Evolution | days | medium |
| Runtime Invariants | Invariants | seconds-hours | high |
| Schema Validator | Structural | seconds | high |
| Second-Agent Review | Human Comprehension | minutes | medium |
| Shadow Traffic | Change | minutes | high |
| Smoke Tests | Behavioral | seconds | medium |
| Snapshot Tests | Behavioral | seconds | medium |
| Static Analysis | Structural | seconds | medium |
| Static Security Analysis | Adversarial | minutes | medium |
| Statically Checked Invariants | Invariants | milliseconds | high |
| Synthetic Monitoring | Behavioral | minutes | high |
| Theorem Proving | Structural | hours | maximum |
| Time-to-Repair | Evolution | weeks | medium |
| Type Checker | Structural | milliseconds | maximum |
Mutation Testing
Take if user.is_admin: allow() and mutate it to
if not user.is_admin: allow(). If all your tests still
pass, your tests did not actually establish the behavior you thought
you established. Mutation testing is a sensor of test
sensitivity rather than test presence.
Recently reviewed
- Human Comprehension Decision Provenance reviewed 2026-09 Can you answer “why is this weird thing here?” A sensor of archaeological accessibility — can you determine why code exists, not just what…
- Invariants Database Invariants reviewed 2026-09 Every foreign key refers to an existing object. Every request has exactly one request_id. created_at <= updated_at.
- Behavioral Contract Tests reviewed 2026-09 Does service A continue satisfying the assumptions of service B? Contract tests are a sensor of boundary assumptions — the implicit…
- Runtime Continuous Profiling reviewed 2026-09 Where did computation actually go? Not “CPU is 82%” but “this function consumed 40% of the time in these specific requests.” A sensor of…
- Adversarial Static Security Analysis reviewed 2026-08 Attacking the code before it runs. Taint tracking, dataflow analysis, and pattern-based scanners (Semgrep, CodeQL) ask: “is there any path…
- Adversarial Property-Based Testing reviewed 2026-08 You state a property that should hold for every input, and the tool generates inputs trying to break it.