The Agent Changed Its Plan Mid-Run: Reconciling AI Decisions With Completed Temporal Activities
Building a Product Recommendation Engine With Neo4j — No ML Library Required
Code Review Core Practices
Getting Started With DevSecOps
It has been such a joy getting to know our DZone contributors beyond reading their incredible articles. And, pretends to be shocked, developers do, in fact, have lives outside of working on their computers. From building open-source projects to spending quality time with their families, DZone’s community is full of not only subject-matter experts, but passionate individuals with interesting stories to tell. One of those individuals happens to be our Member Spotlight of the week. Mayowa Fajobi may be a newer face on DZone, but he’s already an established leader in tech spaces like platform engineering, AI-driven solutions, and open source strategy. But don’t just take my word for it! Learn more about Mayowa, his background, and the career journey that brought him to where he is today. 1. What first got you interested in technology? I’ve always been fascinated by the idea that technology can turn complex problems into something simple and repeatable. My interest really took shape when I started working with Linux and infrastructure and realised that I could automate processes that would otherwise require hours of manual work. That eventually led me into DevOps, cloud-native technologies and Kubernetes, and then into open source, where I could contribute to the tools I was using rather than simply consuming them. 2. If you could only keep three tools in your developer toolkit, what would they be and why? Linux, Git, and Kubernetes. Linux gives me the foundation for understanding what is actually happening beneath the applications and platforms I build. Git is indispensable for collaboration, experimentation, and open-source contribution. Kubernetes brings the two together at scale and has become one of the most interesting platforms for solving distributed infrastructure problems. If I had to pick a fourth, I’d probably struggle not to say a terminal! 3. What advice would you give someone just getting started in your field? Don’t wait until you feel like an expert before you start building. Learn the fundamentals, build things that break, investigate why they broke, and repeat the process. I would also encourage people to contribute to open source early; even documentation fixes, tests, or small bug fixes can teach you how real-world software is designed and maintained. Most importantly, focus on solving problems rather than collecting technologies. 4. As someone working across platform engineering, AI, and open source, what change do you think will have the biggest impact on how engineering teams build and operate software over the next few years, and what do you think technical leaders should be doing now to prepare for it? I think the biggest shift will be the move from software built around long-running services to software built around intelligent, autonomous, and ephemeral workloads. AI is accelerating this transition. Instead of predictable compute shaped by static deployments, we’re entering a world where agents appear, execute work, pause, resume, and scale dynamically based on context. This will force infrastructure to become far more adaptive, capable of allocating, reusing, and governing compute in real time rather than through fixed-capacity models. Because of this, platform engineering must evolve. It will shift from simply providing clusters and pipelines to providing a unified abstraction layer that manages workload lifecycle, security, observability, and cost for increasingly dynamic systems. Technical leaders should prepare now by: Investing in strong platform foundations and treating infrastructure as a product.Designing for automation, portability, and policy-driven governance rather than rigid setups.Creating safe pathways for engineers to experiment with AI while maintaining strict security and visibility boundaries. Finally, open source will be essential. No single organization can solve these challenges alone. The teams that adapt well will be those that combine solid engineering fundamentals with active participation in the wider ecosystem. 5. What’s your ideal way to spend a weekend? A good balance of technology, music, and time away from technology. I enjoy spending time with family, exploring somewhere new, and playing the piano. I also try to switch off for a while, although I’ll inevitably end up opening my laptop at some point, usually because an interesting open-source problem has been sitting in the back of my mind all week! Check out Mayowa's content here.
Every production incident starts with a simple question: "Has this happened before?" I've lost count of how many incident bridges I've joined where that question came up within the first few minutes. Before anyone proposes restarting a service or rolling back a deployment, someone inevitably starts searching. They look through Slack conversations from previous outages, browse old postmortems, compare dashboards with similar incidents, or dig through runbooks to see whether another team has already solved the same problem. Do you notice what's happening here? The engineers aren't trying to demonstrate how much they know about distributed systems. They're trying to remember. That observation has become increasingly important as AI assistants find their way into engineering organizations. Today's large language models are remarkably good at explaining Kubernetes concepts, debugging stack traces, writing SQL queries, or summarizing log files. Those capabilities are valuable, but they only solve part of the problem. During an incident, reasoning is rarely the bottleneck. Finding the right context is. The engineer who resolves an outage the fastest isn't always the one with the deepest theoretical knowledge. More often, it's the engineer who remembers that a similar issue occurred eight months ago after a database failover, or who knows that a particular service has historically exhibited the same failure pattern after specific deployment changes. Experience is a form of memory. If we want AI to become a trusted operational partner instead of just another chatbot, we need to think about memory as carefully as we think about intelligence. Intelligence Answers Questions. Memory Solves Problems. Large language models are exceptional at answering questions because they've been trained on an enormous amount of public knowledge. Ask an LLM to explain consensus algorithms, Kubernetes scheduling, or distributed tracing, and you'll probably receive a detailed, technically accurate explanation within seconds. Production incidents, however, ask very different questions. Instead of asking: What is Kubernetes? Engineers ask: Why did our Kubernetes cluster start failing after yesterday's deployment?Has this service failed in the same way before?Which runbook actually worked the last time?Who owns this dependency today?What changed in the last hour that could explain this behavior? Those aren't questions about computer science. They're questions about organizational memory. The answers don't exist in a foundation model's training data because they're unique to every organization. They live in deployment histories, incident timelines, internal documentation, architecture decisions, monitoring dashboards, chat conversations, and postmortems accumulated over years of operating software. An AI that understands distributed systems but lacks access to this operational history is like an experienced consultant joining your incident bridge for the first time. It may offer useful suggestions, but it doesn't know your environment, your systems, or your team's accumulated experience. That's why intelligence alone isn't enough. Every Incident Is a Search Problem One pattern I've noticed is that incident response often looks less like debugging and more like information retrieval. Think about what engineers actually do during the first fifteen minutes of a major incident. One person opens dashboards to identify where the failure started. Another compares the current deployment with the previous version. Someone searches Slack for keywords that resemble the current symptoms. Another engineer opens the last postmortem for the affected service. Meanwhile, the incident commander tries to understand which teams need to be involved. None of these activities involve writing complex algorithms. They're all attempts to reconstruct context. If you mapped the engineer's workflow, it might look something like this: Alert → Metrics → Logs → Deployment History → Previous Incidents → Runbooks → Slack Discussions → Architecture Documentation → Decision The common thread is that engineers are constantly retrieving information before making decisions. That retrieval process is exactly where AI can provide the most value—not by replacing engineering judgment, but by dramatically reducing the time required to gather relevant context. Not All Memory Is the Same When we talk about memory in AI, it's easy to think only about conversation history or a vector database. In practice, incident response depends on several different kinds of memory, each answering a different set of questions. Incident memory This is the collective history of operational failures. Previous incidents, timelines, root causes, postmortems, and lessons learned all fall into this category. During an outage, one of the first questions engineers ask is whether they've seen the problem before. An AI that can retrieve similar incidents and explain how they were resolved immediately provides value because it shortens the investigation. Operational memory Runbooks, playbooks, escalation procedures, and service ownership represent another form of memory. These artifacts capture how an organization expects engineers to respond under different circumstances. Instead of generating a generic remediation plan, an AI can recommend the procedure that has already been validated by the organization. Infrastructure memory Production systems change constantly. Deployments, feature flags, infrastructure updates, configuration changes, and dependency upgrades all influence system behavior. Understanding what changed recently is often more useful than understanding how a technology works in theory. Organizational memory Some of the most valuable operational knowledge never reaches formal documentation. Engineers discuss recurring issues in Slack, record architectural decisions in design documents, and exchange troubleshooting tips during retrospectives. Over time, this becomes institutional knowledge that experienced engineers rely on instinctively. AI should be able to surface that knowledge instead of forcing every engineer to rediscover it. Memory Changes the Quality of Recommendations Imagine two AI assistants responding to the same latency alert. The first assistant says: CPU utilization is high. Consider restarting the service. It's not necessarily wrong, but it's also not particularly helpful. Now imagine a second assistant with access to organizational memory says: A similar incident occurred three months ago after deployment version 6.4. During that incident, restarting the service temporarily reduced latency, but the underlying cause was an inefficient database query introduced by the deployment. The query was reverted, and latency returned to normal within six minutes. A deployment with similar changes occurred eighteen minutes before the current alert. I recommend validating query performance before restarting the service. Neither assistant is more intelligent in the traditional sense. The second assistant is simply making better use of memory. That additional context changes the recommendation from a generic suggestion into operational guidance grounded in the organization's own experience. Building Memory Into AI Systems Memory isn't a single database or a single technology. It's an architectural capability that combines multiple sources of operational knowledge into a coherent context for reasoning. A production-ready incident assistant might continuously ingest information from observability platforms, deployment pipelines, service catalogs, incident management systems, internal documentation, and communication channels. Rather than asking engineers to manually gather information from each source, the AI assembles the relevant context before generating a recommendation. The language model is still responsible for reasoning, summarization, and communication. The memory layer ensures that reasoning is grounded in facts that are specific to the organization rather than generic patterns learned during training. In many ways, this mirrors how experienced engineers work. They don't solve incidents by relying only on theoretical knowledge. They combine technical understanding with years of accumulated operational experience. That's exactly the capability our AI systems should emulate. Final Thoughts As large language models continue to improve, it's tempting to believe that more intelligence alone will solve the challenges of operational AI. My experience suggests otherwise. The most effective incident response systems aren't necessarily the ones with the largest models or the most sophisticated prompts. They're the ones that help engineers remember. They surface the right runbook, identify the last time a service failed in the same way, highlight the deployment that introduced the problem, and connect today's symptoms with yesterday's lessons. In other words, they make organizational experience accessible when it's needed most. Incident response has always been a combination of reasoning and memory. AI has made remarkable progress on the first half of that equation. The next step isn't simply building smarter models—it's building systems that remember.
If your team distributes AI development skills (as a Claude Code or Cursor plugin or a shared rules file), you own a catalog. That catalog covers things like: How to structure a serviceWhat must pass before a commitWhich internal library to use instead of rolling your own Skills load cheaply, thanks to progressive disclosure. Only the name and one-line description sit in context until the description matches the work. So when token spend climbs, the intuitive read is context bloat, and the obvious lever is tuning descriptions to fire less. That work is worthwhile as it improves routing quality. But it addresses only one of the two ways a skill costs money. The other is procedural amplification: a skill that encodes a fixed procedure as prose, so the model reconstructs it, deliberating, calling tools, checking results, on every invocation. That doesn’t show up as a large context payload. It shows up as extra inference round trips, and an aggregate cost dashboard cannot see it. Predictability does not prove a step should be automated. It identifies places where you should ask whether you’re paying a model to make a decision the system has effectively already made. A few terms come up more than once below, defined here so they don’t slow you down later: TermMeaningSkillA packaged set of instructions Claude Code loads only when the task matches it.OTelShort for OpenTelemetry, the open standard Claude Code uses to report what it did and how much it cost.API request / round tripOne call to the model and its response. This article counts these to measure cost.HookA small script Claude Code runs automatically at a specific moment, such as right after every tool call.p50 / p90 / p95 / p99Percentile timing. p50 is the typical (median) call; p99 is close to the worst case you’ll see.AsyncA setting that lets a hook run in the background instead of making the agent wait for it to finish.NDJSONOne JSON record per line in a file. Simple to append to and to stream. Use the Vendor’s Telemetry With Small Customization I’ll correct a claim I believed myself and have seen repeated: OpenTelemetry only provides aggregate counters, so you have to build your own event pipeline. That’s false, and starting from scratch costs you a week. Claude Code’s OTel export includes events, and they’re richer than I expected: EventCarriesclaude_code.api_requestskill.name, cost_usd, input_tokens, output_tokens, duration_ms,event.sequence, prompt.idclaude_code.tool_resulttool_use_id, tool_name, success, duration_ms skill.name on the API request is the important one: inference round trips are already attributed to the skill that was active. cost_usd is documented as an estimate, not billing data. The One Field to Add Getting command detail out of the native exporter requires OTEL_LOG_TOOL_DETAILS=1, which attaches full_command, bash_command, and file_path to tool events. That’s exactly the payload you cannot centralize off a fleet of developer laptops: branch names, commit messages, customer identifiers, the occasional pasted token. So a small local hook fills one gap: a privacy-preserving shape of the command, joined to native telemetry on tool_use_id, which the docs describe as matching the tool_use_id passed to hooks, allowing correlation between OTel events and hook-captured data. Plain Text native OTel ──┐ │ api_request: skill.name, cost, tokens, round trips ├── join on tool_use_id ──▶ analysis │ tool_result: tool_use_id, tool_name, success local hook ──┘ The hook doesn’t rebuild the trace. It emits the join key plus the one thing native telemetry can’t safely give you. Anything OTel already reports (success, duration_ms, skill.name, token counts) is deliberately not duplicated. Two sources of truth for one field will eventually disagree, and you’ll trust the wrong one. The command-to-shape transform itself is a per-binary allowlist, not a secret detector, and it has a sharp edge worth knowing: The transform: git commit -m "fix auth for acme" becomes git commit -m <ARG>.The edge case: a value-taking flag like --token abc123 leaks its argument as a bare token unless you explicitly track which flags consume the next word. Get the allowlist wrong, and the safety argument for the whole pipeline goes with it. What the Hook Costs (Measured) Tool hooks run on the critical path, so I benchmarked one: 1,000 synthetic PostToolUse payloads through the real script, mixed across Bash/Read/Edit/Write/Grep/Skill. Plain Text p50 p90 p95 p99 max mean subprocess 24.28ms 26.91ms 27.91ms 32.31ms 52.42ms 24.93ms in-process 0.08ms 0.15ms 0.16ms 0.20ms 0.35ms 0.10ms macOS, arm64, 10 cores, Python 3.9.6, n=1000 after 25 discarded warmups. The interesting part isn’t the total; it’s the split. About 24.2 ms of that 24.3 ms p50 is Python interpreter startup. The actual work — scrubbing, serializing, appending — costs 0.08 ms. That points the fix somewhere I would not have guessed: Don’t optimize the scrubber. It’s already three orders of magnitude below the launch cost.Do fix how the hook is launched. Register it async: true so nobody waits, or make it a thin client to a long-lived local collector so you pay startup once per session instead of per tool call. Had I estimated instead of measured, I’d have guessed low single-digit milliseconds and gone tuning the scrubbing code. Wrong target entirely, which is the argument for measuring rather than estimating, in miniature. Event records came in at 364 bytes each. One engineer at roughly 40 sessions a week and 60 tool calls a session generates about 2,400 records, under 1 MB of raw NDJSON weekly. It stays small-data at any plausible team size, because the enrichment record only carries what native OTel doesn’t. Measuring Your Own Catalog Skills have two token costs that behave differently, and knowing which one dominates for you tells you where to look first: Cost componentPaid whenScales withDescriptionEvery request, every session, whether the skill fires or notCatalog sizeBody (procedural amplification)On invocation, then re-sent on each subsequent request in the sessionHow often the skill fires, and how long sessions run Measuring the 25 skills in Claude Code’s official plugin marketplace, a catalog anyone can install and re-measure, gives: Plain Text body chars: min 989 median 11,395 max 32,625 (33x spread) desc chars: min 105 median 357 max 906 always-resident total (all 25 descriptions): 9,703 chars Read those two numbers against each other. Carrying every description in the catalog costs less than a third of one large body. The single biggest skill is roughly 3.4 times the entire always-resident footprint, and it’s charged again on every request for the rest of any session that invokes it. That’s the shape that makes procedural amplification worth hunting. If your always-resident total instead dwarfs your median body, the lever really is catalog size and description tuning. You can stop there. (These are characters, not tokens. The ratio shifts with how much code versus prose a body contains. Count with the real tokenizer, messages.count_tokens, which is free and rate-limited only by requests per minute; don’t reuse a chars/4 estimate or another vendor’s tokenizer, and don’t reuse an old Claude count either: the tokenizer used by Claude 4.7 and later yields roughly 30% more tokens for the same text than earlier models.) Why This Is Worth Measuring at All The subject isn’t token optimization. It’s finding misplaced probabilistic computation: places where the system already knows the next operation and is paying a model to rediscover it. If the next operation is…It belongs in…Genuinely uncertainA skill, policy and judgment, model reasoningAlready determinedA script, deterministic execution A skill that spells out a fixed procedure in prose has put that boundary in the wrong place. The industry is converging on the same line from the runtime side: programmatic tool calling exists precisely to keep deterministic multi-tool sequences out of the inference loop. Key Takeaways Aggregate cost accounting answers the question you already knew to ask. Event-level behavioral traces let you find the one you didn’t, and in Claude Code, most of that stream already exists, attributed to the skill, waiting on one privacy-preserving field to become useful. Procedural amplification leaves no trace in a token counter, raises no error, and turns no dashboard red. It’s visible in the order of the calls, and in how many round trips it takes to get through them. Measured, not estimated: Hook overhead is about 24 ms per call, almost entirely interpreter startup, not the redaction logic itself.Measured, not estimated: In one real catalog, a single skill body costs 3.4× what the entire always-resident description set costs.Next step: Measure by the actual stretch of work, not by an arbitrary number of calls. Group everything that happens while one skill is active, however long that turns out to be, rather than picking a fixed count upfront. When one model turn fires off several tool calls at once, count that as the single decision it was, not several. And don't trim the long, expensive stretches out of the data before analyzing it. Those are usually the exact cases worth finding.
The Batch Processing Problem Batch processing isn't inherently a disadvantage. It becomes a problem when the business needs a decision now, but the architecture was designed to make that information available later. Picture an enterprise system processing millions of customer interactions. Transactions land across multiple systems throughout the day. Every few hours, a scheduled job extracts the data, transforms it, updates another system, and eventually makes it available downstream. This works fine — until the business asks: "Why can't we react to this the moment it happens?" A suspicious transaction. A changed preference. A completed payment. A signed document. A submitted service request. An account state change. In large enterprise environments, I've watched a fairly consistent pattern play out. Teams initially focus on throughput and infrastructure capacity — can the pipeline handle the volume, can it finish the batch window in time? As the systems mature, the harder questions shift elsewhere entirely: who owns a given event, how failures get recovered, how schemas evolve without breaking consumers nobody remembers exist, and — most importantly — what the actual business impact is when a consumer falls behind. Running the batch more frequently doesn't answer any of those questions. Eventually, the architecture itself has to change. Event-Driven Doesn't Mean "Install Kafka" This is where most transformations quietly stall. A common pattern looks like: batch system → add Kafka → the same tightly coupled design underneath. The organization now calls itself event-driven, but nothing structural has actually changed. Real event-driven architecture requires rethinking state ownership, service boundaries, data contracts, failure handling, consistency assumptions, observability, and operational responsibility — not just swapping the transport layer. Plain Text Business Action | v Producer Service | v Event Backbone / | \ v v v Risk Customer Analytics Svc Svc Svc The producer shouldn't need to know or care who's downstream. That's the real architectural benefit — a producer that stays ignorant of its consumers is what genuine decoupling looks like. If the producer still has to know which five systems need to be updated and in what order, you haven't built an event-driven system — you've built a batch job that happens to run on Kafka. Business Events Are Not Technical Messages There's an important distinction between commands and events. A command — UpdateCustomerProfile, SendNotification — says do something. An event — PaymentAuthorized, DocumentSigned — says something happened. Well-designed events represent durable business facts, not implementation instructions. Publish PaymentAuthorized, and Fraud Detection, Notifications, Analytics, Accounting, and Audit can all react independently, without the producer orchestrating any of them. That's the difference between an event-driven system and a batch system wearing a streaming costume. The Hard Problems Start After the First Event Duplicates Most messaging systems guarantee at-least-once delivery, so PaymentAuthorized may legitimately arrive twice. The customer shouldn't be charged twice. Idempotency — via event IDs, business transaction IDs, or a processed-event store — isn't optional polish. Duplicate delivery is a normal condition in a distributed system, not an edge case you occasionally trip over. Ordering If AccountClosed is processed before AccountCreated ever arrives, the consumer ends up holding a state that shouldn't be able to exist — an account that's closed but was never opened. The instinct is to enforce global ordering everywhere, but that kills scalability. The better question is narrower: what actually needs to be ordered? Usually it's events for the same business entity — the same customer, the same account — not the entire enterprise-wide stream. Schema Evolution An event schema gains a new field six months after a consumer was deployed against the old one. Does it break? Backward compatibility, schema registries, and contract testing matter here because events tend to outlive the applications that created them. Treat event contracts like governed APIs, not like internal implementation details nobody needs to track. Failure Don't retry forever. A sane strategy escalates in stages: an initial attempt, then a short retry, then backoff, then a dedicated retry queue, then a dead-letter queue for anything that still hasn't succeeded, then manual investigation or replay. A malformed or logically invalid "poison" event shouldn't be allowed to block the pipeline indefinitely just because it keeps failing the same way. Worth watching closely: retry count, dead-letter volume, consumer failure rate, and the age of the oldest unprocessed event. Eventual Consistency Changes How Teams Think In a synchronous system, an update and its visibility happen together — you write, you read back the new value, done. In an asynchronous architecture, that guarantee disappears. One consumer might reflect a change in twenty milliseconds; another might take two seconds; a third might be temporarily unavailable and catch up later. Different systems can legitimately hold different states for a period of time, and that isn't automatically a defect. The real architectural question is: how stale can this information safely become? Fraud decisioning tolerates almost none — a few hundred milliseconds of staleness can be the difference between catching and missing something. Marketing analytics can tolerate a great deal more. Audit cares more about completeness than about speed. This needs to be decided per business function, not applied as one blanket policy across the platform. It's also worth being honest about what "real-time" actually means in practice. A system that processes an event in milliseconds isn't meaningfully real-time if the downstream systems that act on that event take minutes to reflect the result. I've seen teams celebrate a fast event pipeline while the actual customer-facing decision — the offer shown, the risk flag raised — still lagged well behind because a downstream dependency hadn't caught up. Real-time decisioning has to be measured end-to-end, at the point where the business decision is made, not just at the point where the event was published. Migrating Off Legacy Without a Big-Bang Cutover Ripping out a legacy system in one motion rarely goes well. A more workable path is incremental: capture changes from the legacy system as events, route them through the event backbone, and let new and existing systems consume from the same stream during the transition. Plain Text Legacy System | v Change / Event Capture | v Event Backbone / | \ v v v New New Existing Svc Svc Systems The transactional outbox pattern is useful here: write the event to an outbox table in the same database transaction as the business update, then publish from the outbox separately. That avoids the classic dual-write problem, where the database commit succeeds but the event publish fails, silently leaving downstream systems out of sync. Change Data Capture can also help expose changes from a legacy system as a migration bridge. But it's worth being deliberate about this: a raw database row change is not automatically a well-designed business event. CDC tells you a row changed; it doesn't tell you why, or whether that change represents something a downstream consumer should actually care about. Treating every CDC record as a business event is one of the more common ways these migrations end up producing noisy, low-value streams. Observability Has to Be Designed In, Not Added Later A customer says: "My transaction disappeared." Where do you look, across a chain of services and events? You need correlation IDs, trace IDs, event IDs, business transaction IDs, timestamps, producer identity, and schema versions threaded through everything — plus the standard infrastructure metrics: consumer lag, event age, processing latency, retry rates, dead-letter volume, error rate. But infrastructure telemetry on its own isn't enough. Knowing "consumer lag is 12,000" is far less useful than knowing "12,000 customer transactions are currently delayed." That translation — from technical signal to business impact — is what tends to separate a platform that's merely instrumented from one that's genuinely observable. It's also usually the gap that shows up first when something goes wrong in production: the engineering team sees a metric, and it takes real effort to connect that metric to what a customer or a business stakeholder is actually experiencing. Security and Governance Have to Follow the Data Event-driven architecture multiplies how much data moves around a system, so security has to travel with the data rather than sit only at the application perimeter. That means clear authentication and authorization for who can publish and consume which topics, encryption both in transit and at rest, discipline about not routinely copying sensitive or personal information into every event just because it's convenient, defined retention policies, and clear auditability of who produced what and when. When Not to Use Event-Driven Architecture Don't adopt EDA because it's fashionable. A synchronous API is often the better choice when immediate request-response is required, the workflow is simple, only one system needs the result, or strong immediate consistency is essential. Batch remains entirely appropriate for monthly statements, historical reporting, bulk reconciliation, archival, and much regulatory reporting. The mature position isn't "everything must become event-driven." It's choosing synchronous, asynchronous, and batch patterns based on what the business actually requires — and having a clear answer for why. A Practical Decision Framework Before converting a workload, it's worth asking a short set of questions: Does the business genuinely require lower latency — or would nobody notice the difference between seconds and hours?Do multiple independent consumers need the same business change? If so, event-driven design becomes attractive.Can the business tolerate eventual consistency? If not, the workflow needs closer examination before proceeding.Can the organization actually operate distributed, asynchronous systems — with the observability, on-call practices, and schema governance that requires?What happens when one component fails? If the design can't answer that clearly before production, it isn't ready for production. A Reference Architecture Plain Text ┌──────────────┐ │ Channels │ └───────┬──────┘ │ v ┌──────────────┐ │ API / Domain │ │ Services │ └───────┬──────┘ │ Business Events │ v ┌────────────────────────┐ │ Event Backbone │ └────────────────────────┘ │ │ │ ┌────┘ │ └────┐ v v v Decisioning Notifications Analytics │ │ │ v v v Data Store Data Store Data Store ──── Observability ──── ────── Security ─────── ───── Governance ────── Observability, security, and governance aren't downstream services bolted onto the diagram — they span the whole architecture, or they don't really work. Conclusion The real transformation isn't batch → Kafka. It's delayed processing → continuous business awareness, and central orchestration → autonomous consumers responding to business facts. That shift comes at a cost: more distribution, more asynchronous behavior, more operational complexity, more governance overhead. So the goal was never to produce more events. The goal is systems capable of making timely, reliable decisions — while staying understandable and operable when, inevitably, something fails.
I went looking for a clean way to show what prompt caching actually saves an agent, and the first thing I found was a fact that's easy to miss if you only read the "up to 90% savings" headline. Caching a fresh conversation's first turn costs more than not caching it. There's no cache to read from yet, so you pay the input price on the content plus a 25% premium to write to the cache, and get nothing back. The savings arrive starting on turn two, once there's something to read. Where This Lives in deepagents Deep Agents ships AnthropicPromptCachingMiddleware from langchain-anthropic in its default middleware stack. It doesn't decide what to cache by guessing. It tags exactly two things on every model call: Python # langchain_anthropic/middleware/prompt_caching.py (trimmed) def _apply_caching(self, request: ModelRequest) -> ModelRequest: overrides: dict[str, Any] = {} cache_control = self._cache_control # {"type": "ephemeral", "ttl": self.ttl} overrides["model_settings"] = {**request.model_settings, "cache_control": cache_control} system_message = _tag_system_message(request.system_message, cache_control) if system_message is not request.system_message: overrides["system_message"] = system_message tools = _tag_tools(request.tools, cache_control) if tools is not request.tools: overrides["tools"] = tools return request.override(**overrides) The system prompt gets one breakpoint on its last content block, and the tool list gets one breakpoint on its last tool. Tool definitions are sent as one contiguous block; a single trailing breakpoint caches the entire tool set. On top of that, the top-level cache_control on model_settings gets translated by Anthropic's own auto-caching behavior into a breakpoint at the end of whatever's cacheable in the request. This is what lets the cached prefix grow to cover prior conversation turns as an agent session gets longer, not just the fixed system+tools prefix. That's the whole mechanism. No decision logic, no cost estimation, no adaptive behavior. It tags the stable parts and lets the API's own prefix-match caching do the rest. The Actual Cost Shape, Measured Live Here's the part that isn't obvious from "caching saves money": it saves money on a schedule, not uniformly. Prompt caching is a prefix match. Turn 2 can only read from cache what turn 1 wrote, turn 3 reads what's accumulated through turn 2, and so on. That gives every agent conversation the same two-phase cost curve: write-only first turn, then reads that get proportionally cheaper as the stable prefix grows relative to what's new each turn. To measure this for real rather than project it, I ran the same 8-turn engineering conversation (a realistic back-and-forth about refactoring a blocking-I/O call inside an async FastAPI handler) through claude-opus-5 twice - once with no caching, once with top-level auto-caching (cache_control={"type": "ephemeral"} on every request, the exact mechanism AnthropicPromptCachingMiddleware uses via model_settings["cache_control"]). The system prompt was a real ~1,345-token engineering-assistant prompt, counted with count_tokens before running anything billed: Python """python run_experiment.py (needs ANTHROPIC_API_KEY)""" import anthropic MODEL = "claude-opus-5" INPUT_PRICE = 5.00 / 1_000_000 OUTPUT_PRICE = 25.00 / 1_000_000 CACHE_WRITE_MULT = 1.25 # 5-minute TTL CACHE_READ_MULT = 0.10 SYSTEM_PROMPT = open("system_prompt.txt").read() # a real ~1,345-token prompt TURNS = [...] # 8 real follow-up questions in one coherent conversation client = anthropic.Anthropic() def run(*, use_caching: bool) -> list[dict]: messages: list[dict] = [] turn_records = [] for user_text in TURNS: messages.append({"role": "user", "content": user_text}) kwargs = { "model": MODEL, "max_tokens": 400, "system": SYSTEM_PROMPT, "messages": messages, } if use_caching: kwargs["cache_control"] = {"type": "ephemeral"} resp = client.messages.create(**kwargs) usage = resp.usage messages.append({ "role": "assistant", "content": "".join(b.text for b in resp.content if b.type == "text"), }) turn_records.append({ "input_tokens": usage.input_tokens, "output_tokens": usage.output_tokens, "cache_creation_input_tokens": getattr(usage, "cache_creation_input_tokens", 0) or 0, "cache_read_input_tokens": getattr(usage, "cache_read_input_tokens", 0) or 0, }) return turn_records def cost_for_turn(r: dict) -> float: return ( r["input_tokens"] * INPUT_PRICE + r["cache_creation_input_tokens"] * INPUT_PRICE * CACHE_WRITE_MULT + r["cache_read_input_tokens"] * INPUT_PRICE * CACHE_READ_MULT + r["output_tokens"] * OUTPUT_PRICE ) The real output, 16 live API calls, $0.24 total: TurnNo-cache $Cached $Savingscache_read/creation (cached run)10.0170$0.0187$-10.2%read=0 write=138920.0173$0.0111$35.7%read=1389 write=6430.0175$0.0110$37.1%read=1453 write=4040.0177$0.0110$37.7%read=1493 write=4050.0179$0.0110$38.4%read=1533 write=3760.0182$0.0112$38.5%read=1570 write=6170.0184$0.0111$39.6%read=1631 write=4680.0185$0.0110$40.5%read=1677 write=27Total0.1423$0.0961$32.5% Turn 1 really is negative: 10.2% more expensive than no caching at all, exactly as the documented economics predict. A 5-minute-TTL cache write needs at least two requests to break even (1.25x write + 0.1x read ≈ 1.35x, against 2x for two uncached requests). After that, savings climb every single turn, because the thing growing is the cached portion of the prompt (the cache_read_input_tokens column climbing from 1,389 to 1,677 across the run), while the thing staying flat is the new portion (the write column). By turn 8, caching is saving 40.5% on that turn alone. This is also why a one-shot script or a short-lived Lambda almost never benefits from caching. If the conversation ends after one or two turns, you're stuck in the loss zone the math shows on turn 1. Caching pays for itself on sustained, multi-turn agent sessions, which is exactly the shape of a deepagents CLI session or a long-running subagent loop, not a single classification call. How Much the Stable Prefix Matters The 32.5% overall figure here is real, but it's specific to this run's cacheable prefix being a fairly modest ~1,345-token system prompt. That number moves with the size of the stable part of the prompt relative to what's genuinely new each turn. Swapping SYSTEM_PROMPT in the script above for a larger, more tool-heavy prefix (closer to what a deepagents agent with 15+ MCP tool schemas actually carries) and rerunning would show the savings percentage climb further. The shape of the curve doesn't change, only how much of the bill it's discounting. If you want to know your own number, this script is a five-minute run against your own agent's actual system prompt. Two Levers, and They Stack Caching cuts the price per token of everything that repeats turn over turn. It says nothing about how many tokens you're sending in the first place. That's a second, independent lever: tool-selection filtering, which cuts the token count itself by not sending schemas irrelevant to the current turn. Measured earlier on a 61-tool registry, a lexical-overlap scorer at top_k=10 gets 67% recall at an 84% reduction in tool-schema payload per turn. These aren't competing techniques — they compose. Caching makes the tokens you do send cheaper turn over turn; tool selection reduces how many tokens are in the first place, in the one part of the prompt (the tool list) caching's own breakpoint sits right in front of. A 6,000-token stable prefix that's 40% tool schemas doesn't stay 6,000 tokens once you're only sending the top-10 relevant tools instead of all 61: it shrinks, and then gets cached at the smaller size. Neither technique substitutes for the other: one is about the price of a token, the other is about whether you send it at all. Takeaways The counterintuitive part is the one worth remembering: caching a conversation that only runs one or two turns can cost more than not caching it, because the write premium has nothing to amortize against. It only pays off on sessions that actually run long enough to read back what turn one wrote. Second, if you're trying to actually cut what an agent costs, price-per-token and token-count are separate dials: caching turns one, tool selection turns the other, and a real cost reduction effort turns both.
Software quality is about habits. The quality of our releases depends on the quality of our habits. Nobody ships a defect-riddled release because they didn't want quality. They ship it probably because the small daily behaviors that would have prevented it — the small habits — quietly stopped happening. One skipped test, one PR approved without being read, one “we'll fix it later” at a time. This article makes a quick introduction to those habits, small and large. It explains how small habits become large over time and how good and bad habits compound. I also discuss a specific danger I think we're underestimating: that AI, used without understanding and judgment, doesn't just fail to build good-quality habits. It can actively erode the ones that we may already have. Small Habits Small (or atomic) habits are small enough to survive a bad day and cheap enough to repeat without thinking about it. None of them is impressive in isolation. That's exactly why they're the unit that everything else is built from. Writing and testing: Writing the failing test before writing the fix, not afterChecking the null case, the empty list, the empty string, the zero, the timeoutWriting one assertion that actually checks behaviorDeleting a test that no longer tests anything meaningful instead of leaving it as decorationRunning the full suite locally before pushing, not just the file you touchedAdding a regression test the same day a bug is fixed, not “when there's time” Code review and collaboration: Actually reading a diff line by line before approving itLeaving a comment when you don't understand something, instead of rubber-stamping itAsking “what happens if this call fails” on every PR that adds a network or disk callRequesting a second reviewer on anything touching auth, payments, or data migration, without being told toReviewing your own diff once before asking anyone else to Naming, structure, and documentation: Naming things so the next person doesn't need you to explain themUpdating the doc or README in the same PR as the code change, not in a follow-up that never comesWriting a one-line commit message that says why, not just whatDeleting dead code the moment you notice it, instead of leaving it “in case”Logging enough context at the point of failure that you won't need to reproduce it just to read a log Everyday discipline: Reading the error message fully before searching for itReproducing a bug before claiming to have fixed itFlagging a flaky test the first time you see it, instead of re-running until it passesSaying “I don't know why this works” out loud instead of merging it anyway Large Habits Large habits need sustained effort and organizational will, not just individual discipline. They're expensive to install and easy to let quietly decay. I haven't worked with many teams that actually have the large habits that they think they have. A test suite that's trusted enough that a red build actually stops a merge, every time, with no manual override cultureA blameless postmortem process that produces real changes, not just a document nobody rereadsA quality gate in the release pipeline that blocks a release rather than getting routed around under deadline pressureAn onboarding program that transmits testing and review culture to new hires, not just repo access and a Slack inviteArchitecture reviews that ask “how will this fail” and “what happens at 10x load” before “does this work”A living catalog of architectural decisions and the reasoning behind them, kept current enough that people actually consult itA production monitoring and alerting setup tuned enough that on-call trusts the alerts instead of muting themA deprecation and technical-debt process with real budget and real deadlinesChaos engineering or systematic failure-injection practiced routinely, not once after a bad incidentA culture where raising a quality concern is rewarded even when it slows a release down How Small Habits Become Large A large habit is what a small habit looks like after a few hundred repetitions and a bit of organizational scaffolding around it. A few concrete examples: Writing one unit test with every change, consistently, for a year, is what a real unit test suite is made of. Unit test suites often accumulate, PR by PR, from the atomic habit of not merging untested logic. A test suite guards against regressions once many individual tests are maintained and are trusted enough so that a failure blocks a merge.Leaving one honest review comment when something is unclear, repeated across hundreds of PRs, is what eventually produces the right review culture. A culture where junior engineers feel safe asking questions and senior engineers feel obligated to answer them carefully. No policy document creates that; it's the residue of a habit practiced long enough to become the norm.Writing down why a bug happened, every time, in a shared place, is what a real blameless postmortem process is made of. The first few write-ups are just notes. After enough of them accumulate, cross-referenced and occasionally reread, they become an institutional memory the team consults before making the same category of mistake twice.Updating a doc or writing a short note about a decision each time a decision is made is what eventually becomes a living architectural record. This is something that new hires can read to understand not just what the system does, but why it's shaped the way it is.Flagging one flaky test instead of re-running it is what, multiplied across a team and a year, is the difference between a CI pipeline people trust and one people route around. The large habit — a quality gate that actually informs us about releases — cannot exist unless enough individuals have consistently practiced the small habit. The direction of causation matters here. You cannot batch-install a large habit. You can write the policy, buy the tool, mandate the process — but if the atomic habits underneath it aren't being practiced by individuals, the large habit is just a costume. The postmortem process without honest small write-ups is meaningless. The quality gate without individuals who trust failing checks is a formality people learn to bypass. Large habits are the compounded, aggregated shape of thousands of small habits; you build them from the bottom up, or you don't really have them at all. Habits Compound — In Both Directions A team that writes one more edge-case test per PR than it used to, for two years straight, ends up somewhere completely different from a team that writes one fewer. Neither team can point to the day quality became a strength or a liability. It's a rounding error every day and a number of dissatisfied customers over a year. The bad-habit version compounds just as reliably. “We'll add the test later” becomes later never comes. That's how test suites that nobody fully trusts are built. That's also how red CI that gets ignored becomes a quality gate that's causing problems instead of solving them. The organization now believes it has quality controls it doesn't actually have — and that belief gap is dangerous. Where AI Genuinely Helps Currently, AI is legitimately good at collapsing friction. Writing code and drafting test scaffolding or suggesting the edge case a tired engineer misses at 6 p.m. Summarizing a diff so a reviewer actually reads it instead of skimming, flagging the null case nobody handled — these are atomic-habit amplifiers. They make good habits cheaper to perform, and cheap good habits are often the ones that survive. As AI keeps improving, I expect it to be good at more and more aspects of the SDLC. Where It Becomes Dangerous “AI is good and bad” is a cliché I don't want to hide behind, so here's the specific mechanism. The danger isn't that AI writes bad code or hallucinates an API. It's that when AI fully closes a loop — writes the fix, writes a test that passes, resolves the incident — it can quietly remove the exact process that used to build the judgment we rely on to know when something is fragile. Before AI: an engineer hits a bug, doesn't understand it immediately, tries something, fails, forms a theory, tries again, and eventually understands the system well enough to fix it. More importantly, well enough to recognize the next three places the same class of bug will show up. The fix is important, but it is not the most valuable output of that process. The mental model built while failing is the most valuable output. When AI closes the loop opaquely, even if the bug is really fixed, the mental model may never get built. Engineers will lose knowledge while their judgment and critical thinking weaken. The ability to map incidents to root causes, to understand a system in more depth as we analyze it again and again — all that is gone if we simply accept an AI output without understanding. Terence Tao on AI in Mathematics Fields medalist Terence Tao raised a similar argument about mathematics. His claim isn't that machine-generated proofs are worthless but that the field of mathematics is stagnating because it is entering a crisis of value and understanding. Good, fruitful open problems are mined in a non-renewable fashion, Tao mentions. He goes on to explain: "In short, the indiscriminate use of powerful solution-extraction tools can achieve the immediate short-term goal of solving problems at hand, but at the cost of sustaining the ecosystem for the next wave of progress, or in understanding the progress already obtained." In essence, Tao isn’t objecting to AI’s capability; he is objecting to AI’s opacity. Using the Navier–Stokes regularity problem as his example, Tao described a scenario where an AI system runs an entire search process privately and simply hands back a finished proof. The problem may get solved, but the field gains almost nothing. And this is because the value of a hard open problem is not in the answer itself. The value lies in the decades of people's failed attempts: the tools that we had to invent along the way, the adjacent structure that we mapped while stuck, the wrong turns that became new subfields on their own. Solve a problem transparently, and the field grows by understanding and learning. Solve it as a sealed black box, and you collect the answer but skip the part that was actually producing the mathematics. Let's think about that for a moment from an engineering perspective. Aren't open problems, and how well we understand them, one of the main reasons for engineering innovation? The struggle to keep distributed systems consistent and available produced the CAP theorem. By understanding the CAP theorem, consensus protocols like Paxos and Raft now exist in reliable infrastructure. Trying to run software at a scale that no team could reason about manually gave rise to container orchestration and service meshes. Such open problems gave rise to site reliability engineering as a discipline in its own right. The difficulty of trusting code that could never be exhaustively tested gave birth to property-based testing, fuzzying, and formal verification tools like TLA+. Realizing that testing everything is often meaningless and inefficient (if possible at all) shaped software testing as a discipline. Chaos engineering exists because a team treated “our systems fail in ways we can not predict” as a problem worth living inside rather than a problem to be solved instantly without understanding the solution. Open Problems And Software Quality Our version of “open problems” is production incidents, hard-to-reproduce bugs, and the uncomfortable stretches where nobody's quite sure why the system behaves the way it does. Historically, those gaps close the slow way: someone traces the failure, builds a mental model, writes it up, and the team's collective judgment gets a little better. That's our version of the “adjacent structure mapped while stuck” — debugging notes, a workaround that became a design pattern, the engineer who six months later says “oh, I've seen this shape before.” An AI that resolves an incident by generating a fix nobody on the team fully understands solves the problem in the narrowest possible sense and gives back nothing else. Multiply that across a year of incidents, and you get an organization that ships, that's green on every dashboard, and that has steadily lost the tribal knowledge that used to make it resilient. The loss won't show up as a regression in this quarter's metrics. It shows up the next time a hard problem appears, and nobody in the room has the instinct for it — because the instinct was never built. It was skipped. To keep the loop intact, we could: Treat an AI-generated fix like a junior engineer's PR, not a senior engineer's judgment call — someone accountable still explains, in their own words, why it works before it mergesDon't let AI's speed quietly cancel the postmortem — if it was a real incident, someone still writes down what happened, even if the fix took ten minutes instead of two daysUse AI to compress the mechanical part of the habit loop — scaffolding, boilerplate — but keep the diagnostic “why did this actually fail” a human exercise, at least until an AI's explanation has been checked and not just found plausibleWatch for the belief gap: a green pipeline built partly on unexamined AI fixes can look identical to one built on understood fixes, until it doesn'tWhen onboarding, keep exposing new engineers to the failed attempts, not only the final AI-polished diff — the failed attempts are still where judgment gets transmitted Wrapping Up Identity is built by repeated evidence because in software we are what we repeat. We are our habits. We become someone who doesn't skip tests by repeatedly not skipping tests. Quality cultures work the same way at the team level. One repetition at a time, small habits accumulating into large ones. AI can make each of those repetitions cheaper, and that's genuinely valuable. What it must not be allowed to do is to remove our learnings from repetitions themselves and hand back only the outcome. As mathematics loses something specific and non-recoverable when a hard problem is solved opaquely, software quality is exposed to the same risk — on a much shorter feedback cycle, in production, right now.
TL;DR: The AI Workflow Inventory Finalizes the A3 Delegation System You probably know your own AI shortcuts: the colleague who drafts the stakeholder update, the nightly job somebody set up before they left, the interview notes that go through a model on demand. However, I want to challenge you: that perceived knowledge can create false confidence, because it feels like knowing the team’s way of working with AI, when it is still only a diffuse understanding of the practice. Ask the team to combine those individual accounts into one list, and you may discover how little of the whole anyone can see. The AI Workflow Inventory is the artifact for that list: one row per recurring task, refined into provisional task classes to prepare the team’s next AI delegation decisions. Disclaimer: I read Charniak/McDermott’s book on “Artificial Intelligence” decades ago; of course, I use AI for research, translation, proofreading, challenging story arcs and article structures, and summarization. It is a production tool, not a substitute for thinking. Your Own Shortcuts Are the Easy Part When I published the AI Delegation Lifecycle in June 2026, the point before stage 1 was deliberately left without an artifact: the team was supposed to know which work it does with a model, at what frequency, and at what stakes, which I call a forensic analysis of your own workflow. However, as I quickly discovered, this assumption that a team could produce the analysis from memory was wrong, and updating the student reference for the A3 Delegation System earlier this month showed me its design flaw: The Routing Policy assigns an AI model tier per task class,The AI Definition of Done is written per task class (never per task), andThe Delegation Audit opens with a walkthrough of what the team has handed over. All three artifacts consume task classes, but none of the six stages of the A3 Delegation System produces them. The AI Working Agreement is the exception on the other side of the lifecycle: it records the team’s defaults and boundaries and can exist before the inventory is complete. The gap is easy to miss from the inside, because every individual on the team can answer for their own use, and nobody can answer for the sum: who else drafts with a model, which outputs leave the team that way, and which jobs run unattended. That sum, useful, undocumented, and unowned, is what I call “AI Debt,” and it works right up to the moment someone asks who is responsible for it. What One Row Looks Like in the AI Workflow Inventory The AI Workflow Inventory is a one-page canvas with an identity strip (team, owner, date, review date), a panel on where to look, a panel with five refinement rules, and the inventory table itself. Let me show you an example from a Scrum team: Number: I-01Task: Transcribe photos of Retrospective sticky notesProcess: Retrospective facilitationTask class: Transcriptions for internal useHow often: Per SprintWho runs it: Scrum MasterTool: GPT-5.6 SolOutput goes to: Stays in the teamData that enters the model: Photos of handwritten notes and names. The task is written as verb and object, the class is named by output and audience, and “output goes to” has three values: stays in the team, leaves the team, or customer-facing. The number travels to the Workflow Card later, so the row and the respective decision can be matched months apart. Two design choices carry most of the weight: The person and the tool sit in separate columns because people change and tools change while the task stays; a field that says “Anna, ChatGPT” makes it harder to update either one when Anna moves teams or the license changes.The last two fields ask different questions that teams routinely conflate: where the output goes tells you about the stakes on the way out, whereas what enters the model tells you about the exposure on the way in, and an output that “stays in the team” can still have been produced from photographs with names on them. The entry in the data field is a fact to review; it is not permission to keep uploading the names. What the canvas does not have is a life cycle stage or status column, and the omission is deliberate. The inventory records what happens today; the poster and the Workflow Cards show where a workflow stands in the A3 Delegation Life Cycle, and a retired workflow is a row in the Re-classification Log, identified by the same inventory number. A stage column would turn the pre-decision list into a second, competing status board, and agile practitioners already know how that ends from the Product Backlog in Jira and the “real” one in a spreadsheet debate. The AI Workflow Inventory takes work to maintain, but it reduces repeated documentation and decision-making: The team captures each recurring task once.Compatible tasks share one review standard and one routing decision instead of getting one each.The inventory number connects a task to every later decision about it.The life cycle status stays on the poster and in the records that already exist. Update the inventory when the work changes or when the team discovers, corrects, or refines what it has recorded. (I know, it sounds like a database schema, and I started preliminary work on how to design an application for the A3 Delegation System.) Where to Look The canvas names three sources for AI workflows, and the 30-day window matters because memory is a poor inventory tool, while last month’s outputs are evidence: Your meetings come first: which outputs of the events that repeat (planning, pipeline or campaign reviews, stand-ups, monthly reporting, Retrospectives, etc.) did a model draft, summarize, transcribe, or translate?Then your tools: each colleague looks at their own chat history from the last 30 days and brings the prompts, applied Skills, or agent tasks that come back every week or every Sprint, who wrote them, and how often they were rewritten; this is a self-report, and it has to be, since nobody should be reading anyone else’s conversations.Then your outbox: the emails, reports, tickets, and documents that left the team, and the question of which of them started as a model draft. Thirty days is the starting window. Once the obvious rows are on the page, ask about the work no chat history shows: the jobs that run unattended, the quarterly report, the thing that only happens at release time. A nightly test run, such as the one in the example below, may leave no trace in anyone’s recent chat history. The line that decides the session is the last one in that panel: list the shortcuts nobody approved as well, because that is where the AI debt sits. If people read the inventory as a check on whether they followed the rules, those shortcuts stay invisible, and the team ends up with a list that looks tidy and describes a team that does not exist. Therefore, going first with a shortcut of your own helps, as does explaining, before the first row, that recording a task does not approve its use: the session starts with understanding what happens today, and concerns that need immediate attention under the team’s existing boundaries will still be addressed. Provisional Task Classes and the Grouping Test The grouping rule creates a puzzle that a careful reader will spot: it asks whether two tasks share the same AI Definition of Done, the same reviewer, and the same model tier, and those are decisions for later stages of the AI Delegation lifecycle, which the team has not reached. The resolution is that classes in the first session are provisional. Name them by output and audience, test the grouping against whatever review standards and routing decisions the team already has, and where those are missing, leave the grouping provisional and revisit it as the team works through the stages. An AI workflow inventory can carry uncertainty; what it must not carry is a settled-looking decision nobody made. The five rules on the AI Workflow Inventory canvas are short: Capture first, judge later: The A3 decision comes in stage 1, after the row exists, and a row approves nothing.One row per recurring task: Four prompt revisions used for the same weekly update are one task, and the tool goes in its own column.Name the class by output and audience: “Transcriptions for internal use,” “status communication leaving the company.” A class name that fits any task (“AI support”) is too wide to tell anyone which standard applies. (That naming approach is not different from coding.)Same standard, same reviewer, same tier, one class: When one of the three differs, split the class.Every class goes to stage 1; every audit refines the list: A class without an A3 decision is AI debt with a date on it. Three failure patterns are worth watching for: the sanctioned-use list, where only approved use gets written down, and your existing AI debt stays invisible; the prompt or Skill catalog, with a row per prompt/Skill instead of per task, so the classes never form; and the one-time census, filled once and never refined, so that six months later the audit walks a stale list and everyone nods at rows that no longer run. Example: Four Rows From a Scrum Team Here is a full first pass from a fictitious Scrum team of seven: Transcribe photos of Retrospective sticky notes (Scrum Master, GPT-5.6 Sol, stays in the team; input: handwritten notes with names): Internal output can still involve sensitive input.Rewrite customer feedback tickets into Product Backlog item drafts (class: requirement drafts for internal use): A concrete task inside a reusable class.Draft acceptance criteria for new items (same class): A second task that may share that class, as long as reviewer, standard, and tier stay the same.Rewrite interview notes into the hiring feedback template (on demand, Scrum Master, leaves the team; input: candidate names): Informal use that an approved-use-only list would have missed. The last row is on the list only because capture came first; under a rule that lists only approved use, the person running it would have left it out, and the team would have gone into A3’s stage 1 without knowing candidate names were entering a model on demand. From Row to Workflow Card Each decided row of the AI Workflow Inventory becomes a Workflow Card within the A3 Delegation System: workflow, task class, the A3 category, and later the tier, the owner, and the audit dates; the card goes on the poster where the workflow stands today. The word “decided” matters: the row exists before the A3 “Assist-Automate-Avoid” decision; the card exists after it. A row without a card means one of two things, and the remedy differs: either nobody has decided, in which case the class goes to the A3 Framework next, or somebody decided and never recorded it, in which case the team confirms the decision and writes it down. Ask which it is before doing either. For the transcription row, that means: the AI Workflow Inventory gives it a number, a task, and a provisional class; A3’s stage 1 decides whether it is Assist, Automate, or Avoid, and knowing that GPT-5.6 Sol runs it today settles nothing; stage 3 settles the owner, and the Scrum Master who runs it this Sprint does not become the owner by default. Both decisions go on the card. Your First 60 Minutes with the AI Workflow Inventory Book the session, aim for eight rows, and treat it as a first pass. One way to split the time: State the purpose and the boundary (capture, not approval): 10 minutesWalk the three sources, including unapproved shortcuts and unattended jobs: 20 minutesForm provisional task classes with the grouping test: 20 minutesName the inventory owner and the review date, and agree on the next step: 10 minutes The next step is A3’s stage 1: take every class to the A3 decision with the people who run the tasks in the room. If a row needs attention under the team’s existing boundaries before that (candidate names in a model, say), address it; “capture first” gives the decision its own step and does not postpone it. At each Delegation Audit, the list gets two kinds of maintenance. Add the rows that appeared since the last audit, and reconcile the existing ones: does the task still run, does the same person run it, with the same tool, on the same inputs, for the same audience? When you retire a task, record that decision in the Re-classification Log under the same inventory number, and keep the link to its original inventory record. Conclusion Which recurring use would your team’s current list miss? If your team cannot describe where AI already contributes to recurring work, start with a 60-minute AI Workflow Inventory session. The list gives you a shared basis for A3 decisions and helps you find the workflows that individual recollection might miss; then choose where in the A3 Delegation System to apply the method first.
Editor’s Note: The following is an article written for and published in DZone’s 2026 Trend Report, Cloud-Native Foundations: Kubernetes, Platform Engineering, and Distributed Operations at Scale. Every engineering organization that I have worked with eventually faces the same issue, which is that each team ships services differently. One team used Helm, another wrote raw manifests, and a third would have built a custom Bash script. As these different approaches accumulate, the supporting deployment steps often end up scattered across multiple Wiki pages that quickly go stale. New engineers then spend their first two weeks copying configuration values from an old repository and hoping they still work. A golden path fixes this without turning the platform team into a gatekeeper. It provides users with a standardized workflow for the shortest and most obvious route from a fresh repo to a production workload. This guide walks you through designing a minimum viable golden path, where guardrails belong, and how to keep it useful after v1. Choose the First Golden Path Start with one workflow to standardize first; the strongest candidate is usually the workflow your teams ship most often, or one that teams experience the most friction with. In many organizations, that workflow is a stateless HTTP service exposing a REST or gRPC API endpoint, deployed to Kubernetes and owned by one application team. For this walkthrough, we will use orders-api, a stateless HTTP service on Kubernetes, as our reference throughout this article. The intended users are application developers, not platform engineers — those who create the golden path itself. The path starts with a create-service command in a CLI or a form in an internal developer portal. It should end when the service is running in production with logs, metrics, ownership, and on-call rotation attached. Keep the first version deliberately narrow. A workload that needs GPU nodes, a queue-driven scaling model, or a stateful sidecar can wait. Trying to capture every exception at the beginning turns a practical delivery path into a long platform program. A golden path’s success criteria are qualitative, not quantitative. Analyze the first release by user adoption and experience. Are teams using standardized workflows instead of copying an old repository? Can a new engineer understand the end-to-end deployment process without asking around? Are on-call handoffs easier because services have the same operational shape? The answers to these questions matter more than looking at any adoption numbers displayed on a dashboard in the first few months. Define What the Path Standardizes A golden path is a curated set of decisions that are made once and reused consistently across services: The workload template should provide a Dockerfile, fully maintained base image, Kubernetes manifests, probes, resource requests and limits, a Pod Disruption Budget (PDB), autoscaling defaults, and consistent labels.The delivery pipeline should build, test, scan, sign, and publish the image.The platform defaults should include namespace rules, quotas, network policies, ingress, TLS, logging, metrics, tracing, and basic alerts. The path should not own product decisions; teams will still choose their language, framework, business logic, schema, feature flags, test strategy, and service-specific objectives. This boundary is very important. If we over-standardize, developers will work around the platform, and if we under-standardize, every instance will start with a different set of commands and dashboards. Also make sure the path is easy to find. One internal documentation page, one command, and one entry in the developer portal are enough. If a developer has to ask which template to use, the path has already failed and created friction. The table below shows the differences between shared standards the path owns and decisions each service team owns. Shared Standards vs. Team-Owned Decisions shared standard team decision Dockerfile, base image, patching cadence Language and framework choiceDeployment manifests, probes, resource requests/limits, PDB, Horizontal Pod Autoscaler Business logic, schema, feature flags Build, test, scan, sign, and publish pipeline Test suites specific to the service Namespaces, quotas, network policies, ingress, and TLS defaults Non-standard scaling (queue-driven consumers, GPU jobs) Logging, metrics, tracing, and alerting defaults Business-specific dashboards and SLOs Turn Common Requests Into Self-Service Actions Once the path is created and available to users, review the top 10 tickets your platform team receives. Look for repeated requests such as creating namespaces, adding a database, registering a DNS name, rotating a secret, or creating another environment. These are all good candidates because the desired outcome is already understood, and the steps are mostly predictable. For the Orders API golden path, the platform team can provide the following self-service actions and apply guardrails based on the risk from each change: Fully automated. These actions are reversible and have a limited blast radius. Creating a development namespace for orders-api, spinning up a preview environment on a PR, or rotating a non-production secret happens on demand without a human involved to review.Light review. Actions that change cost, security exposure, or shared infrastructure should require a light review. Provisioning production Postgres for orders-api opens a pre-filled change request that needs one approval. A new public DNS record on a shared domain is reviewed through a one-click approval on a pre-filled PR.Approval mechanism. Every self-service action generates a PR against a config repo, pre-fills the values, tags the reviewer, and merges on approval. The change flows through the same pipeline as code, and every action leaves an audit trail because it’s a git commit. The self-service interface should offer supported choices instead of exposing raw cloud APIs. For example, allowing every team to choose any PostgreSQL version, instance class, or backup schedule can leave the platform team operating 30 different database configurations. A better approach is to provide a small, opinionated set of options such as small, medium, and large. This gives developers enough flexibility while keeping the operational model understandable. For our Orders API, the developer-facing configuration can stay small: YAML # svc.yaml name: orders-api owner: team-orders tier: standard # small | standard | high runtime: http dependencies: - kind: postgres size: small # opinionated preset, not raw config on_call: orders-oncall The configuration captures the developer’s intent, while the golden path translates each request into an approved action with the right guardrail and a clear record of what happened. The table below shows how this works for the Orders API. Orders API Self-Service Actions, Guardrails, and Evidence Step Self-Service Action Guardrail Evidence Create service Run svc new via CLI or submit a portal form Template pinned to current version; namespace quotas applied Repository created with owner metadata; entry in service catalog Add dependency Pick from opinionated list (small/medium/large DB) One-click PR review for prod-tier resources Merged PR against config repo with reviewer name Deploy to prod Merge to main triggers promotion Progressive rollout with auto-rollback on error/latency signals Deployment record with canary metrics and rollback status Rotate secret Run svc rotate-secret New version issued; old version revoked after grace window Audit log entry linked to requester Create a Consistent Path From Code to Deployment Every service on the golden path should move through the same basic stages: pull request → merge to main → staging → production. The exact tooling can vary, but the meaning of each stage should not. At the PR stage, CI runs unit tests, linting, the container build, and security checks. Produce an immutable image tagged with the commit identifier, but do not deploy it to production.On merge to main, the same image is promoted to staging automatically. Rebuilding at each stage creates uncertainty because the artifact tested is no longer guaranteed to be the artifact released. Run integration and smoke tests in this stage.Promoting the image to production reveals the delivery guardrails. Start with a small percentage of traffic (5-10%), monitor health signals, and continue increasing traffic to 25%, then 100%. Roll back automatically when error rate, latency, or probe failures cross agreed thresholds. A developer should not have to recreate this logic in every repository — it should be baked into the deployment tooling. A failed orders-api canary would look like this end to end: The pipeline promotes the new image to 5% of production pods.The error rate for the /orders endpoint rises sharply during the observation window.The deployment controller restores the previous image and drains the new pods based on the rollback threshold.The pipeline posts a message in the orders-oncall service channel with a link to the failing dashboard and offending commit identifier (SHA).An incident record is created automatically only when rollback fails, or the service remains unhealthy. Teams may skip a stage for a documented case (e.g., configuration-only change), but the exception should be an explicit setting with an owner, not an informal workaround. Plain Text # pipeline stages (pseudo) on_pr: [test, lint, build, scan, sign] on_merge: [promote_to_staging, integration-tests] on_green: [canary-5, wait-signals, canary-25, wait-signals, full-rollout] On_regress: [auto-rollback, notify-oncall, record-failure, open-incident] Observability and Day-1 Operational Defaults Even if its pods are running, a service is not ready until the owning team can determine whether it is healthy and knows what action to take when it is not. The golden path should therefore create the minimum operational surface at the same time as the service. The template includes the following list on day one: Structured logs to the central log store, with request ID and trace identifiersRequest rate, error rate, latency percentiles, and saturation metricsDistributed traces with a platform-managed sampling defaultA standard dashboard created from the service nameAlerts for high errors, high latency, restart loops, and resource pressureLiveness and readiness checks connected to a health endpoint Ownership should also be captured during service creation. Ask for the team, on-call rotation, and support channel, then reuse those values in alert routing, the service catalog, and the runbook. Generate a simple runbook with sections dedicated to common failures such as stalled deployments, elevated errors, and pod eviction. A partially completed runbook with a familiar structure is far more useful than a blank page, and consistency here pays off during an incident. Keep the Golden Path Useful Over Time Exceptions are inevitable, so record the failure reason, owner, and expiry date rather than letting the exception become a permanent member. At review time, either the service returns to the path or the platform team decides the pattern is common enough to support. Treat templates and defaults like product code: review changes, version them, and provide a propagation method. When a base image or manifest default changes, open a change against each service instead of relying on teams to notice a document update. Silent drift is one of the fastest ways to lose developer trust in the path. Track a small set of signals such as the time from service creation to first production deployment, template version distribution, open exceptions, and the percentage of new services created through the path. Pair those numbers with developer feedback. A slow step that teams repeatedly bypass tells you where the next path improvement belongs. A new template version without a propagation plan becomes a fork. Extend the path when a pattern is used by three or more teams, but keep it narrow while it is still one team’s edge case. Plain Text # template bump propagation (pseudo) on template_release(new_version): for svc in services_on_path(): open_pr(svc, bump_template = new_version, auto_merge = svc.opts.auto_bump, reviewer = svc.owner) Making the Golden Path Useful in Practice A golden path succeeds when it is easier to follow than to work around. Start with one common workflow, standardize what is shared, and leave product choices with the service team. Make routine actions self-service, place checks in the delivery flow, and include observability from the first deployment. Usage signals can then inform future improvements to the path. A small path that ships, earns trust, and changes steadily will have a greater impact on engineering speed than a broad platform program that remains unfinished. Resources: CNCF TAG App DeliveryOpenTelemetry General Semantic ConventionsKubernetes Pod Security StandardsBackstage Software Templates“Building a CI/CD Pipeline With Kubernetes” by Naga Santhosh Reddy VootukuriKubernetes Security Essentials, DZone Refcard by Yitaek HwangPlatform Engineering Essentials, DZone Refcard by Apostolos Giannakidis This is an excerpt from DZone’s 2026 Trend Report, Cloud-Native Foundations: Kubernetes, Platform Engineering, and Distributed Operations at Scale.Read the Free Report
Combining Temporal and LangGraph creates a deceptively simple question: which runtime owns the state of the agent? Both preserve execution progress, but they preserve different kinds of progress. Temporal reconstructs Workflow state from Event History and reuses recorded Activity results during replay. LangGraph persists thread-scoped graph state as checkpoints and resumes from super-step boundaries. Treating those mechanisms as interchangeable creates ambiguous recovery semantics. Production integration therefore needs explicit authority for business progress, agent working state, and the handoff between them. Deployment language, storage backend, model provider, and hosting topology remain unspecified assumptions. The current Temporal LangGraph integration narrows the problem. Its public-preview Python plugin can run LangGraph nodes as Temporal Activities or deterministic Workflow code, while Temporal provides durability; the documentation recommends an in-memory LangGraph checkpointer rather than a separate PostgreSQL or Redis checkpointer. Continue-As-New can carry cached task results into the next Workflow Run. The harder two-runtime problem appears when LangGraph retains an independent persistent checkpointer while Temporal separately orchestrates the business lifecycle. That case needs an application-level consistency contract. Ownership and Handoff The cleanest ownership rule is semantic. Temporal should own business lifecycle state: whether an execution is open, waiting for approval, canceled, timed out, compensated, or complete. LangGraph should own agent working state: messages, retrieved evidence, tentative plans, tool proposals, and graph position. External systems should remain authoritative for effects in their own domains. A payment processor owns whether a charge occurred; a deployment service owns whether a release exists. Temporal and LangGraph may retain receipts, but neither should invent contradictory domain truth. This matches their native models: Event History records Workflow progress, while LangGraph checkpointers persist thread state. The handoff should be modeled as a command protocol. A business execution ID identifies the long-lived application process and can map to a stable Temporal Workflow ID. A graph thread ID identifies checkpoint lineage. A command ID identifies one requested graph advance, while a monotonically increasing application revision identifies the state version against which that command was accepted. Temporal Run ID should not become the business identifier because Continue-As-New preserves Workflow ID while creating a new Run ID. LangGraph uses thread_id to load checkpoint history. Stable application identities must therefore outlive either runtime’s individual run instance. A graph-advance Activity can enforce that contract without exposing persistence details to Workflow code: Python def advance_agent(cmd): receipt = receipts.get(cmd.command_id) if receipt and receipt.status == "completed": return receipt.result state = threads.acquire( cmd.thread_id, fencing_token=cmd.revision, ) if state.revision != cmd.expected_revision: raise StaleCommand(cmd.command_id) result = graph.invoke( cmd.input, {"configurable": {"thread_id": cmd.thread_id}, durability="sync", ) return receipts.complete(cmd, result) The key property is the admission rule. The same command ID must never become new input merely because an Activity retried. A command accepted at revision 17 remains command 17 across worker crashes and timeouts. A genuinely new turn receives a new command ID and expected revision. LangGraph’s synchronous durability mode persists each checkpoint before the next step starts, reducing checkpoint-loss risk, but it does not create a transaction with Temporal Event History. Failure and Re-Execution The critical failure window begins after LangGraph commits a checkpoint and before Temporal records Activity completion. Temporal documents the analogous edge directly: an Activity can finish, the worker can crash before reporting completion, and the Activity can then execute again. Completed Activities are not re-executed during Workflow replay, but an unrecorded completion is indistinguishable from unfinished work at the orchestration boundary. Blindly injecting the same graph input on retry can therefore advance the agent twice for one logical transition. A durable command receipt closes that ambiguity. It records command ID, thread ID, expected revision, resulting revision, checkpoint reference, status, and result reference. When checkpoint and receipt records share a database, a custom persistence adapter can commit command acceptance and checkpoint metadata in one local transaction. When stores cannot share a transaction, recovery can reconstruct a missing receipt from checkpoint metadata containing the command ID and resulting revision. That reconstruction is an application-level inference based on LangGraph checkpoint metadata and lookup capabilities. A separate receipt written only after graph completion leaves another crash window. External effects require a second deduplication boundary. Temporal recommends idempotent Activities because Activity execution may happen more than once, and idempotency keys must ultimately be enforced by the called service. The graph should therefore propose an effect before performing it. An interrupt can expose the proposal, allowing Temporal to own approval, deadlines, and cancellation while a dedicated Activity performs the mutation with a stable operation ID. LangGraph documents that an interrupted node restarts from its beginning when resumed, so code before interrupt() executes again and pre-interrupt effects must be safe to repeat. Python def await_effect(state): receipt = interrupt({ "operation_id": state["operation_id"], "proposal": state["proposal"], "revision": state["revision"], }) if receipt["operation_id"] != state["operation_id"]: raise ValueError("effect receipt mismatch") return {"effect_receipt": receipt} This arrangement also clarifies replay. Temporal replay re-executes Workflow code while matching Commands against Event History; recorded Activity results are reused. LangGraph replay from an older checkpoint instead re-executes nodes after that checkpoint, including LLM calls, API requests, and interrupts. A LangGraph fork is therefore a new computational branch, not restoration of external reality. Previously completed business effects remain attached to their original operation receipts, while newly proposed effects require fresh authorization. Operational Semantics Concurrency control must prevent overlapping Activity attempts from advancing one thread simultaneously. A lease alone is insufficient if an expired holder can still write. A monotonically increasing fencing token tied to the accepted application revision provides a stronger rule: persistence rejects writes from an older token after a newer command is admitted. This is an integration pattern rather than a built-in Temporal or LangGraph guarantee. Administrative retries, manual resumes, and human approvals should pass through the same admission path. Temporal’s documented retry model establishes the underlying reason for such protection: Activity execution can occur more than once even though successful completion is observed once by the Workflow. Retry policy should remain layered. Temporal should own Activity retry delivery, while LangGraph-level retries should remain narrowly scoped to graph operations whose repetition is safe. Independent retry loops at both layers can multiply attempts and obscure the failure budget. Cancellation also needs explicit semantics. Temporal delivers cancellation to heartbeat-enabled Activities through heartbeats, but cancellation of orchestration does not prove that an already accepted remote effect was reversed. Ambiguous operations therefore require reconciliation with the authoritative downstream system. Continue-As-New changes run identity but not business identity. Temporal starts a fresh Event History with the same Workflow ID and a different Run ID, carrying selected state forward. In a dual-runtime design, the business execution ID, graph thread ID, latest revision, outstanding commands, and unresolved effect receipts must cross that boundary. Temporal scopes Update-ID deduplication to a Workflow Run, so deduplication that must survive Continue-As-New cannot rely solely on server-side Update identity. The official LangGraph integration similarly carries serialized cached results across Continue-As-New. Operational tests should target boundaries rather than happy paths. Worker termination immediately after checkpoint commit, after an external service accepts an operation, and before Activity acknowledgment exposes duplicate-execution defects. Concurrent retries should prove that stale fencing tokens cannot write. Stale approvals should prove that proposal revision checks reject obsolete decisions. Continue-As-New tests should prove that logical identities and deduplication records survive the run transition. Temporal provides test hooks for exercising Continue-As-New, while LangGraph checkpoint history and replay make recovery assertions observable. Conclusion A Temporal-and-LangGraph agent becomes reliable only when durability is subordinate to ownership. Temporal should decide business progress, LangGraph should preserve agent working state, and external systems should remain authoritative for real-world effects. The handoff needs stable business, thread, command, and revision identities; durable receipts; idempotent effect execution; explicit replay semantics; fenced concurrency; and continuity across Continue-As-New. The official Temporal LangGraph plugin can collapse much of this complexity by making Temporal the durable execution substrate. When independent persistence remains on both sides, however, two durable runtimes do not become one consistent runtime automatically. A precise state-ownership protocol is what turns duplicated durability into controlled recovery.
Unit testing business logic often requires surprisingly little business data. Suppose we want to test a program that loads an order, calculates its total, and rejects it when the amount exceeds a limit. The decision we want to verify is simple: The program loads the specified order.It asks for the order’s total.If the total is too high, it does not approve the order.It returns FALSE. Yet a conventional Java unit test may have to construct an Order. That order may require a customer, line items, currencies, prices, tax information, identifiers, and other objects that have nothing to do with the decision being tested. Builders, fixtures and mocking frameworks reduce the typing, but they do not eliminate the underlying problem: the test must participate in the internal representation of the domain model. BUBAS takes a different approach. Domain objects are opaque to a BUBAS program. Because the program cannot inspect them, a unit test does not need to construct them. It needs only a token. The Business Program Consider this BUBAS program: SQL PROGRAM ApproveOrder(orderId INTEGER, limit DECIMAL) RETURNS BOOLEAN DECLARE purchase Order DECLARE total DECIMAL purchase = LOAD_ORDER(orderId) IF NOT ORDER_WAS_FOUND(purchase) THEN LOG_EVENT "ERROR", "no such order: " + orderId RETURN FALSE END IF total = ORDER_TOTAL(purchase) IF total > limit THEN LOG_EVENT "INFO", "over limit: " + total RETURN FALSE END IF APPROVE purchase RETURN TRUE END. Order is a Java domain type registered by the application embedding BUBAS. The program can store an Order in a variable and pass it to operations that accept an Order, but it cannot access its fields or invoke its methods. There is no expression such as: SQL purchase.customer.account.balance If the program needs information about an order, the application must expose an operation for obtaining it: SQL total = ORDER_TOTAL(purchase) This restriction is primarily an encapsulation mechanism. The business program depends on the vocabulary of its domain rather than on the internal structure of Java objects. It also has an important consequence for testing. Replace the Object With Identity Here is a BUNIT test for the over-limit case: Gherkin PROGRAM OverLimitIsRejected "LOAD_ORDER" WITH ARGS(42) RETURNS "o1" "ORDER_TOTAL" WITH ARGS("o1") RETURNS 1500.00 "APPROVE _" IS MOCKED ARGUMENT "orderId" IS 42 ARGUMENT "limit" IS 1000.00 RUN RESULT IS FALSE "APPROVE _" WAS NOT CALLED END. The string "o1" is not an order serialized as text. It does not contain an order number, a total or any other property. It is a test token representing one opaque Order. The first mock says: Plain Text "LOAD_ORDER" WITH ARGS(42) RETURNS "o1" When the program calls LOAD_ORDER(42), BUNIT returns the token "o1" in place of the real Java object. The program stores it in purchase. Later it calls: Java ORDER_TOTAL(purchase) The second mock recognizes that same token and returns 1500.00. The program cannot tell that "o1" is not a real Order. It has no operation with which to inspect the object. It can only pass the value back through the vocabulary supplied by the host application. For this test, identity is all the domain object needs. We Are Testing the Conversation A BUBAS business program contains decisions and orchestration. Algorithms, persistence, infrastructure, and domain-object implementations remain in Java. Its unit test should therefore concentrate on questions such as: Which domain operations were invoked?With what arguments?What values did those operations return?Which branch did the program select?Which operations were deliberately not invoked?What result did the program produce? In the example, we do not test how ORDER_TOTAL calculates a total. That belongs in the Java test for the implementation of ORDER_TOTAL. We test what the business program does when ORDER_TOTAL reports 1500.00. This division gives us two focused tests rather than one oversized test: Java tests verify the individual domain operations.BUNIT tests verify how a business program coordinates them. The BUNIT test documents the business scenario directly. An order identified by 42 exists, its total is 1500.00, the approval limit is 1000.00, and the program must not approve it. The test does not explain how to manufacture an object graph that produces those facts. Opacity Buys Mockability Mocking domain objects in a general-purpose language is often difficult precisely because the production code can observe so much about them. It may call methods, inspect nested objects, compare values, serialize the object or pass it to code that expects a particular implementation. A substitute must reproduce every observable property used along the tested path. An opaque BUBAS value has only the observations provided by the registered vocabulary. If the vocabulary exposes ORDER_TOTAL, then the mock controls the answer to ORDER_TOTAL. If it does not expose the customer’s internal account object, neither the program nor the test needs to know that such an object exists. The object boundary and the testing boundary are the same boundary. This is stronger than merely saying that business programs should avoid inspecting domain objects. They cannot inspect them unless the embedder deliberately provides an operation that does so. Consequently, a token can stand in for any opaque value as long as the mocks define how the exposed operations respond to it. Multiple objects require only multiple identities: Plain Text "LOAD_ORDER" WITH ARGS(42) RETURNS "o1" "LOAD_ORDER" WITH ARGS(43) RETURNS "o2" "ORDER_TOTAL" WITH ARGS("o1") RETURNS 1500.00 "ORDER_TOTAL" WITH ARGS("o2") RETURNS 200.00 The test describes the distinctions that matter without constructing either order. The Test Uses the Real Language A dangerous form of mocking creates a second, simplified interface used only by tests. Eventually, the production vocabulary changes while the test vocabulary does not. BUNIT does not compile the business program against a parallel language. The program under test is compiled against the real sealed BUBAS language. Mocking happens later, at dispatch. Therefore, the test cannot silently keep using an operation that no longer exists in the production language. Nor can it casually return a value of the wrong BUBAS type. Before executing a test, BUNIT checks the mocks and the test configuration. It can report problems such as: A mock declared with the wrong number of arguments;A mock returning a value incompatible with the real operation;An argument supplied for a parameter the program does not accept;A mocked command that should initialize a variable but does not provide its value. The test reports these errors before the business program runs. The test remains artificial — as every unit test is — but it is artificial inside the actual language contract. Do Not Assert Everything A test becomes fragile when it records every interaction, whether or not that interaction matters to the scenario. BUNIT allows an expectation to specify only the relevant part of a call. For example: Plain Text "LOG_EVENT _, _" WAS CALLED WITH ARGS("INFO", CONTAINS("over limit")) The test requires an informational log message containing "over limit". It does not require the complete message to remain byte-for-byte identical. Similarly: Plain Text "APPROVE _" WAS NOT CALLED expresses the important negative requirement without inventing an Order merely to compare it with another Order. The purpose is not to reproduce the execution trace. It is to state the observable facts that define the business case. What This Does Not Test Opaque tokens do not prove that the Java implementation of LOAD_ORDER returns the right order. They do not prove that ORDER_TOTAL calculates taxes correctly or that APPROVE commits a transaction. Those operations require their own Java unit and integration tests. BUNIT tests the program at the orchestration boundary. This makes it possible to test business decisions without databases, service containers, or complete domain-object graphs, but it does not replace testing below or beyond that boundary. Nor does BUNIT make every Java application automatically testable. The application developer first has to expose a suitably designed vocabulary. If one enormous operation performs loading, calculation, approval and notification internally, BUNIT can mock that operation but cannot test the decisions hidden inside it. Testability therefore provides feedback about vocabulary design. Operations should represent meaningful domain capabilities at the level where business programs genuinely make choices. The Deeper Result Opaque domain types may initially look like a limitation. The program cannot examine its own values freely. It has to ask the vocabulary to interpret them. That limitation creates a clean separation: Java owns domain representation and implementation.BUBAS owns orchestration and decisions.BUNIT replaces domain capabilities at that same boundary.Tokens replace complex objects with identity when identity is all the test requires. The production program becomes independent of domain-object structure. The unit test inherits that independence. We do not need a fake Order with a fake customer containing fake line items whose prices happen to add up to 1500.00. For this business decision, we need only to say: Plain Text "ORDER_TOTAL" WITH ARGS("o1") RETURNS 1500.00 The business program never needed to know what was inside the order. Neither does its test. The detailed code and the BUBAS framework are available as open source at https://github.com/verhas/bubas.
Agile
Career Development
Methodologies
Team Management
Can Your Team Name the Work It Already Runs With AI?
September 25, 2026
by Stefan Wolpers
CORE
One Agent, Two Runtimes: Defining State Ownership Between Temporal and LangGraph
September 25, 2026
by Akhil Madineni
CORE
How to Build an Asynchronous AI-Content Review Workflow in C#
September 23, 2026
by Brian O'Neill
CORE
AI/ML
Big Data
Databases
IoT
Prompt Caching Doesn't Save Money on Turn One
September 25, 2026
by Ninaad Rao
CORE
Software Quality Habits and AI
September 25, 2026
by Stelios Manioudakis
CORE
Can Your Team Name the Work It Already Runs With AI?
September 25, 2026
by Stefan Wolpers
CORE
Cloud Architecture
Integration
Microservices
Performance
Beyond Batch: Engineering Enterprise Systems for Real-Time Decisioning
September 25, 2026 by Prem Kumar Gadhanki
September 25, 2026
by Naga Santhosh Reddy Vootukuri
CORE
How to Verify Response Data in API Testing With Playwright TypeScript
September 25, 2026
by Faisal Khatri
CORE
Frameworks
Java
JavaScript
Languages
Tools
September 25, 2026
by Naga Santhosh Reddy Vootukuri
CORE
Testing Business Programs Without Constructing Domain Objects
September 25, 2026
by Peter Verhas
CORE
How to Verify Response Data in API Testing With Playwright TypeScript
September 25, 2026
by Faisal Khatri
CORE
Deployment
DevOps and CI/CD
Maintenance
Monitoring and Observability
Software Quality Habits and AI
September 25, 2026
by Stelios Manioudakis
CORE
September 25, 2026
by Naga Santhosh Reddy Vootukuri
CORE
Testing Business Programs Without Constructing Domain Objects
September 25, 2026
by Peter Verhas
CORE
AI/ML
Java
JavaScript
Open Source
Software Quality Habits and AI
September 25, 2026
by Stelios Manioudakis
CORE
Can Your Team Name the Work It Already Runs With AI?
September 25, 2026
by Stefan Wolpers
CORE
Member Spotlight: Mayowa Fajobi
September 25, 2026 by Dominique Roller