An industry resource

What independent observations
would cause us to believe
this software is correct?

Software is increasingly an opaque artifact. We cannot — and often do not want to — fully understand every implementation. The Software Observatory is a catalog of epistemic sensors: the observable signals that reduce uncertainty about whether a system is correct, maintainable, and behaving as intended. Not "code quality metrics." Measurement instruments pointed at different failure modes.

No single sensor measures correctness. Coverage measures execution. Mutation testing measures test sensitivity. Types measure a particular class of structural inconsistency. Contracts measure boundary assumptions. Observability measures what actually happened and preserves enough dimensionality to investigate unknown unknowns.

They are all measurement instruments pointed at different failure modes.

10 sensor families

The catalog is organized into 10 families, each asking a different question about the system. Each family rules out a different way the software could be wrong. No single sensor is sufficient, and neither is a pile of them: a sensor earns its place by eliminating a doubt the others leave open, not by adding another green check.

The confidence landscape

No single sensor is sufficient, so there is no total ordering across sensors — no "best" sensor. But each dimension (oracle strength, latency, scope) is a partial order, and the atlas's left-to-right axis is time, not quality. Each one trades feedback latency — how long you wait for the signal — against efficacy: how much the signal can actually tell you. Compilation is instant and definitive about validity; user outcomes are slow and definitive about everything that matters. Most sensors live somewhere in between.

efficacy of the signal
definitive suggestive
User outcome Production behavior Canary / shadow Integration tests Behavioral tests Property / metamorphic Mutation testing Static analysis / types Compilation A/B Testing API Compatibility Boundary Sensors Example-Based Tests Feature Flag Exposure Telemetry Architecture Fitness Functions Property-Based Testing Schema Validator Branch Coverage Dependency Graph Line Coverage Resource Telemetry Build Provenance & SBOM Canary Analysis Contract Tests Database Invariants Integration Tests Live Service Graph Discovery Metamorphic Testing Pre-Promotion Invariant Gates Shadow Traffic Synthetic Monitoring Business Invariants Live Chaos Experiments Change Coupling DORA Metrics Revert Rate Compiler Type Checker Continuous Profiling Distributed Traces Observability Events Smoke Tests Snapshot Tests Static Analysis Contract & Refinement Types Statically Checked Invariants Decision Provenance Documentation Drift Diff Coverage Hotspot Analysis Differential Testing Fuzzing Mutation Testing Error-Budget Impact Independent Review Escaped Defect Rate Fault Injection Incremental Build Correctness Load Testing Second-Agent Review Static Security Analysis Incident Correlation Time-to-Repair Linter Model Checking Onboarding Experiment Runtime Invariants Theorem Proving
instant feedback latency → slow

Hover a point to name it. Click to open the entry.

Structural Behavioral Test Effectiveness Invariants Adversarial Runtime Change Architecture Evolution Human Comprehension
Sensor list (latency × efficacy)
SensorFamilyFeedback latencyEfficacy (oracle)
A/B TestingChangedayshigh
API CompatibilityChangesecondshigh
Boundary SensorsArchitecturesecondshigh
Branch CoverageTest Effectivenesssecondslow
Build Provenance & SBOMStructuralminuteshigh
Business InvariantsInvariantshourshigh
Canary AnalysisChangeminuteshigh
Change CouplingEvolutiondaysmedium
CompilerStructuralmillisecondsmaximum
Continuous ProfilingRuntimesecondsmedium
Contract & Refinement TypesStructuralmillisecondshigh
Contract TestsBehavioralminuteshigh
Database InvariantsInvariantsminuteshigh
Decision ProvenanceHuman Comprehensiondayslow
Dependency GraphArchitecturesecondslow
Diff CoverageTest Effectivenessminuteslow
Differential TestingAdversarialminutes-hourshigh
Distributed TracesRuntimesecondsmedium
Documentation DriftHuman Comprehensiondayslow
DORA MetricsEvolutiondaysmedium
Error-Budget ImpactChangehoursmedium
Escaped Defect RateTest Effectivenessmonthsmedium
Example-Based TestsBehavioralsecondshigh
Fault InjectionAdversarialminutesmedium
Feature Flag Exposure TelemetryChangesecondshigh
Architecture Fitness FunctionsArchitecturesecondshigh
FuzzingAdversarialminutes-hourshigh
Hotspot AnalysisArchitectureminuteslow
Incident CorrelationEvolutionweeksmedium
Incremental Build CorrectnessChangeminutesmedium
Independent ReviewHuman Comprehensionhoursmedium
Integration TestsBehavioralminuteshigh
Line CoverageTest Effectivenesssecondslow
LinterStructuralmillisecondsmedium
Live Chaos ExperimentsAdversarialhourshigh
Live Service Graph DiscoveryArchitectureminuteshigh
Load TestingRuntimeminutesmedium
Metamorphic TestingAdversarialminuteshigh
Model CheckingStructuralminutesmaximum
Mutation TestingTest Effectivenessminutes-hourshigh
Observability EventsRuntimesecondsmedium
Onboarding ExperimentHuman Comprehensionweekslow
Pre-Promotion Invariant GatesInvariantsminuteshigh
Property-Based TestingAdversarialsecondshigh
Resource TelemetryRuntimesecondslow
Revert RateEvolutiondaysmedium
Runtime InvariantsInvariantsseconds-hourshigh
Schema ValidatorStructuralsecondshigh
Second-Agent ReviewHuman Comprehensionminutesmedium
Shadow TrafficChangeminuteshigh
Smoke TestsBehavioralsecondsmedium
Snapshot TestsBehavioralsecondsmedium
Static AnalysisStructuralsecondsmedium
Static Security AnalysisAdversarialminutesmedium
Statically Checked InvariantsInvariantsmillisecondshigh
Synthetic MonitoringBehavioralminuteshigh
Theorem ProvingStructuralhoursmaximum
Time-to-RepairEvolutionweeksmedium
Type CheckerStructuralmillisecondsmaximum

Recently reviewed