Most maturity models assume a clean separation between “the software” and “operations.” But production systems are holistic. Operators SSH into nodes. DBAs run migrations. Scripts patch data. Incident responders flip feature flags. If these mutations don’t flow through the same channels as application code, your view of state is fractured.
Intent captures anything and everything that mutates state, regardless of origin. And state itself is a combination of intent and outcome. Without captured intent, state is just “what is” with no explanation of “what was meant to be.” You’re left reverse-engineering decisions from consequences, a fragile and lossy exercise usually performed during postmortems under duress.
To preserve and understand intent, we define five levels of intent integrity:
- I0: Untraced: State changes occur through side channels; no record exists of what happened or why
- I1: Described: Human-authored goals are recorded externally (e.g. in tickets, emails, or Slack) but disconnected from actual mutations
- I2: Attached: All state changes flow through the journal with intent embedded in the event itself
- I3: Modeled: Intent is a first-class construct; it is modeled, versioned, and drives system behavior
- I4: Validated: The system can assert that outcomes align with declared goals, and surface divergences
I0: Untraced
The most common symptom of I0 is that current state is the source of truth, and state change information is treated as operational noise. A relational database holds the “real” data. The WAL exists for crash recovery, not for understanding. Audit logs, if they exist, are a compliance checkbox that nobody queries. The system knows what is but has no authoritative record of how it got there or why.
If your primary data store is a mutable row that gets UPDATEd in place, and the only change record is something like a WAL or a trigger-based audit table, you’re at I0. The change information is a side channel, not the source of truth. When that WAL rotates out or that audit table gets truncated for space, the history vanishes.
At I0:
- Current state is the only authoritative record
- Change history lives in secondary artifacts (WALs, audit logs, CDC streams) that are treated as operational concerns
- Those artifacts don’t capture why something changed, only that it changed
- Intent is nowhere: not in the current state, not in the change record
The more extreme form of I0 is when state changes bypass the system entirely. An operator SSHs into a production node and mutates data directly. A DBA runs a one-off UPDATE to fix corrupted records. There may or may not be a Jira ticket somewhere, but even if there is, it exists in a parallel universe, completely disconnected from the actual mutation. The system’s view of state and actual state have diverged, and you have no way to detect it.
This is common in legacy systems, but also in modern systems that haven’t thought carefully about operational boundaries. Every manual intervention that bypasses the application is a hole in your causal chain. Every CRUD mutation that overwrites history is a lost explanation.
You can’t:
- Reconstruct what changed or when, beyond what current state implies
- Attribute any outcome to intent, because intent was never captured
- Distinguish deliberate changes from accidents, bugs, or malice
I1: Described
Intent exists, but it lives outside the system. A Jira ticket says “allow this operator to delete these accounts.” An operator runs a script. Maybe the script logs something. Maybe it doesn’t. Either way, the connection between the ticket and the system’s state is manual and fragile.
At this level, you have parallel tracks:
- The external record: tickets, Slack threads, emails, runbooks
- The system’s record: logs, events, database state
Correlating them requires tribal knowledge, timestamp matching, and heroic grep sessions. The link is never authoritative. You can reconstruct a plausible narrative, but you can’t prove it.
At I1, observability is fragmented across the “three pillars”: logs, metrics, and traces. You’re hopping between systems, correlating by hand. The intent lives in one place (a ticket), the symptoms live in another (metrics), and the details live in a third (logs). Nothing is connected.
Even with effort, you often can’t:
- Verify the system followed intent without “expert testimony”
- Detect when an outcome diverged from the original goal
- Prove that a specific ticket corresponds to a specific state change
At I1, most teams rely on third-party tooling to stitch together plausible intent from fragmented signals. Dashboards ingest logs. Analytics platforms slice and dice metrics. Trace viewers render spans. The tooling is often impressive, but it’s reconstructing a narrative from side-channeled data that was never designed to tell a coherent story. The result is plausible, not authoritative. Useful for debugging, but not auditable. You can build a convincing timeline, but you can’t prove it.
I2: Attached
At I2, the journal becomes the complete observability substrate. All state changes, regardless of origin, flow through it. No more side channels. No more hopping between logs, metrics, and traces. The journal is the single source of truth.
This is the shift from fragmented observability to what Charity Majors and the Honeycomb team call “wide events”: arbitrarily-wide structured events that capture all context about a state change. Actor, intent, timestamp, correlation ID, payload, outcome, everything lives in one place. Metrics and traces are derived from the journal, not maintained as separate pillars.
{
"event_id": "01HY3...",
"event_type": "AccountClosed",
"ts": "2025-05-15T12:07:13.044Z",
"actor_id": "operator:jsmith",
"intent": "customer_requested_closure",
"correlation_id": "ticket-OPS-4421",
"source": "admin_cli",
"payload": { "account_id": "acct-9931", "reason_code": "GDPR_REQUEST" }
}At I2, operational actions flow through the same channel as application actions. When an operator fixes something, that fix emits an event. The CLI tool, the admin interface, the migration script. They all write to the journal. There are no back doors. The journal isn’t a view of what the application did; it’s a view of what happened to state, period.
Traditional observability treats logs, metrics, and traces as separate systems. You see something scary on a dashboard, hop to logs, grep for timestamps, then hop to traces. Wide events eliminate this fragmentation. One source of truth, arbitrarily wide, with all context preserved. You derive metrics and SLOs from the events themselves.
At this level:
- Every mutation flows through the journal, including operational interventions
- Events embed intent directly: why this action was taken, not just what happened
- External references (ticket IDs, correlation IDs) are attached, but the journal is authoritative
- Observability is unified, not fragmented across pillars
To go deeper, we distinguish between human-driven intent and intent derived by the system itself.
I2a: Human-Looped Intent
Human-looped intent means explicitly preserving the reason a person caused an event, even when the resulting effect is triggered by automation. A stop-loss order expresses an investor’s intent to limit downside risk. When that stop-loss triggers later, the system embeds the original intent in the resulting event. The loop closes between human reasoning and system action.
This matters for operational actions too. When an operator runs a migration, the event captures not just “migration ran” but “migration ran to fix data corruption from incident INC-2234.” The intent travels with the mutation.
I2b: Agentic Intent
Agentic intent arises when the system makes decisions on the user’s behalf, based on encoded strategies, configurations, or policies. The originating intent may be human, such as “maintain a 60/40 risk profile”, but the specific actions are derived by the system. This is not artificial intelligence in the buzzword sense, more like your architecture acting as a decision-making agent.
At I2, both human and agentic intent are captured in the same journal, with the same structure. You can query across them uniformly.
At this level, your system can:
- Attribute actions to stated goals, whether human or system-derived
- Audit intent across layers and across time
- Begin to reason about what was intended versus what actually occurred
- Answer “why did this happen?” from a single query, not a scavenger hunt
Most systems stop here and call it good enough. But if we want reflectivity, we need to go further. We need systems that can not only remember what they did and why, but evaluate that reasoning and improve.
I3: Modeled
At I2, intent is captured in wide events. At I3, intent becomes a first-class concept in the domain model itself. It is structured, versioned data that drives behaviour across the system, not just metadata attached to events.
Examples:
- Declarative goals like
"intent": "maximize_yield"shape system behaviour, such as an AI agent dynamically adjusting portfolio risk - Services express intent in executable plans, with intent flowing from initiation to completion
The shift here is profound. Intent changes from “passive” to “active”, becoming the input to machine intelligence and learning. This is the threshold where:
- Reinforcement loops become possible
- Retrospective analysis can factor in goals
- You shift from debugging code to debugging plans and models
Once you’ve reached I2a and I2b, this becomes a natural next step, even though the step up is fairly massive. Intent metadata like thresholds, risk profiles, or target allocations become inputs to learning systems. Raw events like StopLossTriggered or PortfolioRebalance are transformed into structured features that models can train on and reason about.
At I3, intent becomes the connective tissue between domain modeling and machine learning.
I4: Validated
At I4, intent becomes enforceable, testable, and traceable. This is where systems begin to understand how their actions perform, using deliberate and measurable forms of machine intelligence.
You can:
- Simulate plans to validate if goals can be met
- Assert “this outcome satisfied the original intent”
- Flag when actions deviate from their intended purpose (e.g. due to slippage, drift, or constraints)
- Audit why a system failed to achieve its goal, and whether it was due to bad input or bad logic
This level enables:
- Self-reflective, adaptive behaviour
- Safe delegation to AI agents or autonomous subsystems
- Postmortems that distinguish between alignment failure and execution failure
Examples of adaptive behaviour are numerous. Here are a few to frame the possibilities without becoming too prescriptive.
| Capability | Description |
|---|---|
| Goal satisfaction | A validation job compares intent and outcome events, writing results to an intent_validations table with fields like was_satisfied, deviation, and reason. This becomes your source of truth for goal alignment. |
| Drift and data quality | Tools like Great Expectations, WhyLabs, or Monte Carlo monitor feature distributions and trigger alerts when live data diverges from historical patterns, hinting at misalignment with original intent. |
| Experiment tracking | Systems like MLflow or DVC log feature sets, model versions, and training metadata. Combined with validation results, this enables traceability between goals, inputs, decisions, and outcomes. |
At I3, systems use structured intent to train and serve machine learning models. Tools like Feast help bridge that gap, turning intent into features for inference.
At I4, the system begins to validate whether those inferences and decisions actually satisfied the original intent. This requires building capabilities such as validation engines, outcome audits, and intent–result comparisons stored as first-class records.
What really changes here is accountability. Intent is more stringently verified at this level of maturity. Systems that act must be able to prove whether they upheld that contract, and if not, why.
This is the threshold where outcome monitoring becomes goal-oriented. Systems assess how well actions fulfill intent, and operators work to understand when and why alignment breaks down. The specific tools you use do not matter as much as ensuring that your team, within a clearly defined scope, can demonstrate confidence and traceability between intent and outcome.