History

All state transitions must be preserved. Learn how to build systems where every change is stored as an immutable fact that can be replayed, inspected, and verified.

Historical integrity means that every state transition is stored as an immutable fact that can be replayed, inspected, and independently verified. State changes are treated as first-class values.

In a basic CRUD world with only periodic snapshots (H0), history vanishes in an instant, leaving you at 3am digging through WALs, nightly backups, and random log fragments during an outage. As you level up to H1, every change becomes an immutable, durably written event. At peak maturity, your system’s historical integrity means you can reconstruct its exact state at any moment, intact and reliable, even after major refactorings.

To evaluate this, we define five levels of historical integrity:

  • H0: Basic Archival: Periodic snapshots capture moments in time, but with gaps
  • H1: Durable History: All state changes are recorded as immutable events
  • H2: Deterministic Replay: State can be reconstructed exactly from event history
  • H3: Evolvable History: Historical data survives schema changes and system evolution
  • H4: Time-Travel Provenance: Every event is immutable, causally linked, and queryable across time

H0: Basic Archival

At primordial stages of historical integrity, systems periodically take full or partial snapshots of system state, like nightly dumps or hourly exports. These allow operators to restore to a recent checkpoint, but any changes made between snapshots are typically lost. At this maturity level, the goal is archival durability, not precise point-in-time recovery. Snapshots serve as a fallback against catastrophic failure, but restoring from them almost always involves some degree of data loss.

You can:

  • Restore the system to a previous archive
  • Inspect state at specific intervals

This level supports:

  • Manual rollbacks and recovery
  • Basic debugging of gross state changes

While data recovery is now possible, snapshots still miss the transitions between states and remain prone to significant data loss. To move beyond vague terms like “nightly” or “hourly”, let’s be more precise about acceptable risk at H0 maturity.

A practical heuristic is to keep snapshot data under 15% of the raw event log size per entity or stream, or take a snapshot nightly, whichever comes first. This strikes a workable balance between historical fidelity and replay performance, especially in legacy environments that still rely on maintenance windows.

That said, snapshotting is ultimately an operational concern (speaking as a former SRE manager). Frequency and retention policies should be tuned based on your system’s load characteristics, recovery time objectives, and infrastructure cost. There’s no one-size-fits-all rule, but there should always be a rule of some kind, otherwise history remains ephemeral.

H1: Durable History

Once each state transition is captured as an immutable event in an append-only journal at C1, H1 adds the durability guarantee, preserving every insert, essentially forever. It builds on C1’s foundation, where a single writer per stream maintains strict causal order. At C1 ∩ H1, you now have a replayable event log that lets you scan a single writer’s timeline and reconstruct the full history of an entity.

H1 asserts that history should never be discarded, and that journals must be preserved forever.

With this assertion you can:

  • Track what happened and when
  • Store events as a sequential, time-ordered log (with single-writer per journal wall clocks)
  • Build a durable audit trail

This level enables straightforward historical inspection and supports confident long-term storage strategies. Updates and deletes are not permitted, only inserts and queries, ensuring the integrity of the recorded history.

This is easier said than done. Keeping data forever requires operational discipline: log rotation, compaction, journal archival, and retrieval automation. But here’s what makes it worthwhile: H1 transforms the role of H0. At H0, snapshots are your recovery mechanism, gaps and all. At H1, snapshots become a performance optimization backed by complete, immutable logs. Complete journals become the archive, and snapshots simply accelerate replay. With cursor-based log reading, realtime consumers track their position and always read from the head. This means there’s no tradeoff between long-term completeness and immediate performance, as long as you never archive unprocessed events.

It’s worth pointing out an exception to append-only and immutable data. Some data will need to become inaccessible over time due to privacy regulations like GDPR. Cryptographic tombstones or key-encrypted shredding allow data to remain physically present but cryptographically inaccessible. Instead of overwriting or erasing, systems destroy the encryption key, rendering the underlying data unreadable. This approach preserves the fact that data once existed (maintaining audit trails or chain-of-custody) without retaining sensitive content, aligning with privacy regulations while keeping the historical timeline intact.

H2: Deterministic Replay

H2 asserts that you must be able to discard read models and regenerate them from the journal at will. At this stage of maturity the system embraces full event sourcing with deterministic replays, allowing current state to be derived from the complete event log, and also enabling reliable point‑in‑time restoration.

With read-side projections, you can:

  • Discard and regenerate state at will
  • Reproject any view or model from first principles
  • Eliminate the need for mutable state in the system of record
Tip · What's in the box?

Many event sourcing frameworks include this type of snapshotting out of the box:

  • Akka/Pekko lets you snapshot based on event count, custom logic, or time.
  • Rust libraries like eventually-rs support this pattern with explicit hooks.
  • EventStoreDB consumers can store their projection progress as offsets or versioned checkpoints.
  • ObzenFlow supports -⁠-⁠replay-from to replay from archived source journals, with future support for replaying from any vector clock position.

Replaying state from the beginning of time will cause very long restarts and missed SLAs, potentially hours of latency if replaying billions of events. This is where replay checkpoints come in. H2 also introduces a requirement for “operational snapshots”, automatically-created restart anchors that limit how far back the system must replay during recovery.

A practical policy is:

  • Snapshot after N events (e.g, 1,000–10,000)
  • Or after T time (e.g, every 1–4 hours of wall clock)

This ensures your system doesn’t always start from zero, gaining all the benefits of read-side projections while avoiding cold-start problems.

H3: Evolvable History

Past events remain compatible as the system evolves. The journal supports schema changes, domain refactors, and feature growth without invalidating history or requiring destructive migrations.

Tip · ML and evolvable history
Evolvable history turns your journal into a feature store. You can re-run feature logic over raw historical events for reproducible ML pipelines, backfill new features without ETL, and maintain train/serve parity with shared logic across batch and streaming inference. It’s also ideal for fine-tuning domain-specific micro models (LoRA, adapters) on clean, versioned event data.

At this level of maturity, you must be able to:

  • Upcast safely. Each (event_type@version) is translated in place by a pure, side-effect-free function. Upcasters are chained and must be total. CI fails if coverage is incomplete.
  • Deploy schema-first. Consumers must be upgraded before producers emit new versions. Unknown event types are quarantined or rejected.
  • Monitor version drift. Metrics on unknown events, upcaster failures, and projection errors provide early detection of compatibility regressions.
  • Reprocess events. Old events can be replayed through current logic to rebuild projections or regenerate derived state.

At this maturity level, expect to be implementing client-side wrappers over legacy schemas to upcast into new events when schemas change. What we don’t do is mutate old events to fit the new shape, as this would break lower levels of maturity guarantees.

H3 emphasizes compatibility over completeness. It ensures that the past does not break under new code and that long-lived systems can evolve without discarding history.

H4: Time-Travel Provenance

History becomes a query surface. Events are immutable. Each one is linked to its predecessor by a cryptographic hash and tagged with the causal context that produced it. The system supports point-in-time reconstruction, branching timelines, and automatic verification of past behaviour.

This is a significant leap, and few systems will ever reach it. But for those that do, the journal becomes a truth machine: a verifiable, tamper-evident record that can answer questions about the past with the same confidence as the present. At this level of maturity, you can:

  • Query any entity at a moment in time. Periodic snapshots act as indexed checkpoints. The engine loads the nearest snapshot, replays only trailing events, and returns state in milliseconds.
  • Fork history for simulation or backtesting. Branch from any snapshot, append hypothetical events, and compare outcomes without touching the main timeline.
  • Detect tampering. Each event stores the hash of its predecessor; any insertion, deletion, or shuffle breaks the chain and raises an alert.
  • Navigate causality. Metadata records what triggered each event. A graph index answers “why” questions and traces knock-on effects across aggregates.
  • Verify replay. Every branch replay runs state transitions through invariant checks. Violations fail CI and trigger production alerts.

The leap from H3 to H4 is substantial:

FeatureH3 Evolvable HistoryH4 Time-Travel Provenance
Replay scopeFull log or checkpoint-plus-tailSnapshot-indexed point-in-time query
BranchingManual copy to a test DB/streamNative branch or tag with isolation
IntegrityTrust storage and checksumsHash-link every event and verify on read

This is what makes the journal a truth machine: indexed snapshots for instant recall, branching APIs for safe experimentation, and cryptographic integrity so tampering is impossible to hide. Few systems need this level of rigour, but for those that do, like regulated finance, healthcare, legal discovery, and high-stakes simulation, H4 is the target.

Concrete pay-offs:

  • A regulator requests the state of an account on 12 Apr at 14:03; the system answers in near real time.
  • Data science forks last year’s trades, tests a new pricing model, and measures impact within minutes.
  • Audit replays historic decisions under new policy and produces a signed report.
  • Nightly replay catches an invariant breach from two months ago and flags it before production suffers.
Tip · Hard problems await, but the benefits are real

H4 leans on formal modelling more than storage tricks. Define business invariants as executable rules and halt replays the instant a transition breaks one. Roll out format changes safely: upgrade readers first, quarantine snapshots new code cannot parse, and monitor metrics such as snapshot_verify_failed and hash_chain_gap. Test rollouts under partition to keep branches consistent.

Time-travel provenance turns the past into a strategic asset. But it also unlocks the future. Imagine running game day exercises on alternate timelines: what would your investment portfolio look like if you’d made a different trade two years ago? That sounds simple enough, but now imagine an institutional investor running that analysis across millions of accounts in minutes. Or a hospital network simulating how different treatment protocols would have affected patient outcomes across a decade of records. Or an airline replaying a month of operations to test whether a new maintenance schedule would have prevented delays. H4 makes these scenarios queryable.