A system exhibits causal integrity when the current state of any resource can be traced back through a sequence of prior events that explain how it came to be. These events must record what happened, why it happened, and what triggered it.
This is the foundation of reflective systems. With a verifiable causal chain, you gain evidence of cause and effect. Unlike logs or secondary artefacts, which are often incomplete or decoupled from domain meaning, a system with causal integrity treats causality as a first-class concern.
In CHAIN, all decisions, state changes, and outputs must be causally derivable. To evaluate this, we define five levels of causal integrity:
- C0: Opaque Causality: State is stored directly, overwriting history with no traceable context
- C1: Single-Writer Replay: Immutable events are stored in a way that enables sequential replay and state reconstruction
- C2: Causal Linking: Events include causal metadata for direct (but unverified) cause-effect linkage
- C3: Graph Navigation: Causality forms a navigable graph; impact and root cause can be queried
- C4: Enforced Causality: Causal structure is validated, simulated, and enforced across workflows
ClientOpenedPortfolio"]:::event C2001["command-2001
SendAcknowledgementEmail"]:::command C2002["command-2002
PerformCreditCheck"]:::command C2003["command-2003
ScheduleAdvisorCall"]:::command E1237["event-1237
AcknowledgementEmailSent"]:::event E1235["event-1235
CreditCheckPerformed"]:::event E1236["event-1236
AdvisorCallScheduled"]:::event C2004["command-2004
NotifyComplianceOfficer"]:::command %% Links E1234 --> C2001 E1234 --> C2002 E1234 --> C2003 C2001 --> E1237 C2002 --> E1235 C2003 --> E1236 E1235 --> C2004 %% Styles classDef event stroke:#f97316,stroke-width:2px,color:#f4f4f5; classDef command stroke:#4f46e5,stroke-width:2px,color:#f4f4f5; %% Assign classes class E1234,E1235,E1236,E1237 event; class C2001,C2002,C2003,C2004 command;
C0: Opaque Causality
State is freely mutated and discarded with no causal metadata being preserved alongside data. Limited causal linking may be possible using secondary sources of context (like logs), if available.
- No way to explain how the latest state was reached without stitching together external context
- Secondary sources of context like logs are often incomplete and difficult to reason about
- Typical causal maturity of most CRUD systems with mutable rows
No trail or chain, only current state is persisted directly within the domain.
C1: Single‑Writer Journals
At the C1 maturity level, each identifiable entity maintains its own append-only journal. One entity, one writer, one process. This enforces a parallelization factor of one, which can limit horizontal scalability under extreme load, but it guarantees causal ordering. Every state-changing command results in a new immutable event, appended in timestamp order, forming a reliable, replayable sequence of events.
Because each journal has exactly one writer:
- There are no write-write conflicts or race conditions
- The sequence of events is reasonably deterministic
- Causal relationships within the stream are trivially preserved
With single-writer journals, causality becomes implied based on order rather than links. Events are written in a clear sequence, so it’s always obvious what happened next. Each command produces exactly one event, making cause and effect trivial to trace, but still requires manual effort. Retries are safe because conflicting commands are caught before they reach the log, and since the writer sees the latest state before appending, it can enforce business rules like “balance never negative” with confidence.
This creates a causally consistent event log scoped to a single entity. It may seem simple, but this structure forms the first solid foundation for understanding why things happened, and it’s what makes most event-sourced systems reliable in practice.
Worth noting is that each event in a C1 journal must be tagged with a monotonic ULID: a 128-bit identifier whose first 48 bits are the millisecond timestamp and whose remaining randomness is forced to tick upward if two events land in the same millisecond. Because ULIDs are lexicographically sortable, the writer’s append order is baked right into the ID itself: no two events from the same writer can share an identifier, and every new ULID is guaranteed to compare greater than the one before it.
C1 is therefore your first real checkpoint on the CHAIN path: durability, replayability, and a clear internal narrative, all delivered by keeping a journal and not letting anyone else write in it.
C2: Causal Linking
Entities encode causal metadata, essentially a reference to what triggered the event, and what broader operation it participated in. This can be explicit (linking to a command, saga, or policy) or implicit (via a singly linked causal_parent_id field).
Causal linking works when expectations are reasonable and scoped intentionally. In this series, we scope C2 at the aggregate level (in DDD terms; more on this in later units), which refers to a single-writer process operating within the same VM. To ensure that causal metadata is trustworthy, C2 must lean directly on C1’s guarantees: a strictly ordered, append-only journal with monotonic identifiers like ULIDs. These IDs make it safe to include references such as causal_parent_id, since every referenced event must already have been observed and persisted by the writer.
At C2, your process-local journal is the sole source of truth, and within a single writer’s process you can be confident that causal metadata reflects an observed, local truth.
Now you can trace:
causal_parent_ulid: the upstream event that caused this onecorrelation_ulid: the enclosing workflow or transaction
Reaching C2 is the first true level of locally linked causality, making your state evolution reconstructable and chainable within the single-writer boundary, e.g, reconstructing the state of an aggregate from a journal.
C2 only holds when confined to a well-defined boundary, typically a single aggregate in DDD terms. Without that discipline, a causal_parent_ulid can point at events that were never observed, or even at the wrong stream. C2 is your first step toward verifiable causality; true cross-aggregate guarantees demand stricter coordination (vector clocks, HLCs, or similar), which we’ll explore in later units.
C3: Graph Navigation
C3 builds on the maturity achieved at C2 by adding comprehension of causality by turning sequential traces into structural graphs.
At this level, causality becomes a queryable graph, no longer just a singly linked list of events as in C2. Events and commands are connected via explicit causal edges that allow you to traverse what happened, why it happened, and what it affected, but only within the bounded scope defined at C2 (for example, a workflow, saga, or aggregate boundary).
You can now ask:
“What are all the events downstream of
event-1234within this aggregate boundary?”
This unlocks:
- Root cause analysis
- Blast radius inspection (what might this command affect?)
- Workflow debugging and timeline reconstruction
- Lightweight observability without enforcing distributed guarantees
C3 improves traceability, not correctness. The system records causal structure, but it does not verify that an event’s dependencies were satisfied at write time.
Most systems never reach this level of causal reasoning. They jump straight to global lineage and burn out on the complexity. (Adding full graph semantics across a system with distributed writers and no clock discipline feels like chasing the distributed systems dragon.)
But if you hold the scope boundary, C3 becomes very achievable and quite powerful. You’ll notice a shift from incident response to impact forecasting, and from reactive debugging to proactive safety checks.
Social and causal graphs follow a power law: most nodes have few edges, but a few have very many. When a “celebrity” node emits an event, naïve fan-out writes or on-demand graph traversals can swamp your indexes.
Mitigation strategies:
- Hybrid push/pull: push writes for low-fan-out nodes; pull (read-time assembly) for mega-nodes
- Cache traversals: store results of common impact queries (e.g. “all effects of event X”)
- Shard by scope: partition edges so hotspots only hit their own shards
- Materialize views: precompute and store subgraphs for known heavy hitters
Event sourcing advocates tend to materialize views, but the main point is not dictating a solution. There are numerous strategies to deal with “celebrity nodes”, pick the right one for your use case.
It’s not because of magic, it’s a two-fold effect:
- The structure itself: implementing causal edges and correlation graphs gives you direct leverage. You can trace effects, anticipate problems, and explain system behaviour clearly.
- The design discipline required to get here: reaching this level of maturity forces teams to think in terms of boundaries, intent, and traceability. That mindset compounds. Once in place, it improves code quality, testability, operability, even onboarding.
This kind of maturity functions like architectural compound interest. You pay a little up front and it keeps paying you back.
C4: Enforced Causality
C4 is where causality becomes enforced across distributed boundaries. At this level, your system records causal relationships and it verifies them using formal approaches to partial distributed ordering. Events carry vector clocks or hybrid logical clocks (HLCs) that encode their known dependencies.
Logical or hybrid clocks (Lamport timestamps, vector clocks, or HLCs) do not enforce causal order by themselves. They merely record each event’s view of its dependencies. Your event processor must:
- Read those timestamps or vector entries
- Verify that all declared predecessors have been persisted (or stash the event until they are)
- Only then accept or act on the event
Only the combination of accurate clock metadata and strict consumer-side enforcement delivers true causal maturity. With both observation and enforcement in place, the system can now assert:
“This domain event cannot be valid, as it occurred before its own parent.”
This enables systems to reason not only about what happened, but what should have happened first, even across services, writers, or domains.
Each event asserts a partial causal order:
“This event occurred after these known versions of all relevant writers.”
With this information, a system can:
- Reject or defer events with unmet causal dependencies
- Detect and prevent causality violations across services
- Simulate workflows or validate invariants like, “Every
TradeSettledmust causally follow aTradeExecuted”
In practice, few teams will ever fully reach C4, it’s a big jump from C3. Beyond the technical hurdles, you need organizational buy-in, rigorous domain modeling, and infrastructure that can capture and enforce distributed lineage. Remember, C4 isn’t about creating a single global timeline or linearizability, it’s about ensuring each event only fires after its declared dependencies. You’ll still process independent events in parallel; the nuance at this level is in handling the subtle cases where causality really matters.
C4 brings real-world costs, as well as rigour. Managing vector clocks, dependency checks, out-of-order buffering, and replay logic all add complexity and measurable latency. In practice, you should expect an enforcement overhead on the order of tens of milliseconds on commodity SSD clusters; call this your causal-enforcement SLA so you’re not guilty of hand-waving.
Equally important are the failure semantics at this level of maturity. E.g, what happens when a dependency never arrives? Do you dead-letter the event, issue an alert after T seconds, or back off and retry forever? Define a clear liveness policy, e.g, “any event with unmet causal prerequisites moves to a dead-letter queue after 30s and raises an incident”, so your system fails loudly, not silently.
Strive for C4 only within tight boundaries where causal correctness across services is essential for safety, compliance, or resilience, and understand the performance tradeoffs.
When you do need it, usually in safety-critical, financial, or operationally regulated systems, C4 becomes non-negotiable. It can prevent entire classes of bugs and failures that C3 can only observe after the fact. But you should reach for it after the cost of causal ambiguity is greater than the cost of enforcing this level of correctness within a boundary.