Why Your GPU Fleet Is Both Full and Idle: Building a Capacity Orchestration Layer
Member Spotlight: Raghava Dittakavi
Agentic AI Threat Intelligence Essentials
Getting Started With Agentic AI for SecOps
Conventional unit tests remain essential around an LLM application, but they do not answer the most important production question about whether a change preserved useful behavior. A parser can still return valid JSON while the answer becomes less grounded. A tool call can satisfy its schema while selecting the wrong tool. A retrieval pipeline can return documents successfully while omitting the evidence required for the final answer. This gap exists because an LLM application is not a deterministic function whose correctness can always be represented as actual == expected. Empirical work on code generation has shown substantial output variation across repeated calls, including at temperature zero, so single-run assertions can mistake sampling noise for either success or failure. Regression testing therefore has to evaluate behavior statistically and semantically, not only execution paths. Unit Tests Prove Contracts, Not Behavior The first layer should still look familiar. Deterministic properties deserve deterministic tests like JSON must parse, required fields must exist, tool arguments must conform to a schema, forbidden operations must remain blocked, and latency or token budgets can be checked numerically. Anthropic’s evaluation guidance explicitly separates success criteria from the mechanisms used to measure them and includes exact-match, code-based, human, and model-based grading as different options rather than treating one metric as universal. A contract test for a support-answering endpoint can remain intentionally narrow: Python def validate_contract(result): assert result["status"] in {"answered", "abstained"} assert isinstance(result["answer"], str) assert len(result["answer"]) <= 2_000 assert all(c["source_id"] for c in result["citations"]) Passing this function proves that the response is consumable by downstream software. It does not prove that the answer is correct, that every material claim is supported, or that an abstention occurred when evidence was missing. Those are behavioral properties and need evaluators matched to the application’s actual failure modes. Research on LLM-based judging also shows why a single generic "quality" score is insufficient as judges can exhibit position, verbosity, and self-enhancement biases, even though strong judges can correlate well with human preferences under controlled evaluation. Build the Regression Corpus From Failures A useful regression suite begins with representative tasks rather than prompts invented solely for testing. Production traces, manually verified examples, support escalations, retrieval misses, tool selection failures, malformed outputs, and previously fixed incidents should become durable evaluation cases. LangSmith’s current evaluation documentation describes this feedback pattern where production traces can be added to datasets so that a failure observed in live traffic becomes a repeatable offline test, while experiments compare application versions on the same dataset. Each case should preserve enough context to reproduce the behavior that matters. A question alone is often insufficient for RAG or agent systems because the retrieved documents, tool state, permissions, conversation history, and expected policy can affect the result. The expected value also should not always be a canonical sentence. A stronger case stores assertions about required facts, forbidden claims, acceptable citations, expected tool choices, and whether abstention is mandatory. Python case = { "input": "Can an expired license be renewed online?", "required_facts": {"renewal_window": "30 days"}, "forbidden_claims": {"automatic_extension"}, "must_cite": {"policy-17"}, "expected_action": "answer" } The corpus also needs slices. A global average can improve while an important category regresses. Long queries, multilingual requests, ambiguous requests, high-risk tool actions, low-retrieval-confidence cases, and specific product domains should retain explicit tags so candidate performance can be compared within those populations. Google’s production ML guidance similarly recommends monitoring real-world metrics and inspecting data slices because aggregate quality can conceal skew or deterioration in subgroups. Score Properties Instead of Strings Exact matching works for classifications, IDs, tool names, and other canonical outputs. Open-ended language requires property-based evaluation. For RAG, retrieval and generation should be measured separately. Ragas formalized this decomposition with metrics targeting retrieval quality, answer relevance, and faithfulness rather than collapsing the entire pipeline into one score. That separation matters operationally because an unsupported answer and a retrieval miss require different fixes. A production-like evaluator can combine deterministic checks with semantic scoring: Python def evaluate(case, result): return { "contract": contract_score(result), "groundedness": groundedness_score( result["answer"], result["evidence"] ), "task_quality": rubric_score( case["input"], result["answer"], case ), "tool_correctness": tool_score(case, result), } The evaluator should expose dimensions rather than immediately averaging them. A groundedness score of zero must not be hidden by excellent style. A wrong financial action must not pass because the explanation is fluent. Hard safety and contract requirements should act as vetoes, and softer qualities such as completeness or tone can be aggregated only after those constraints pass. LLM as a judge is useful when the desired property cannot be expressed with code, but the judge itself becomes part of the test infrastructure. G-Eval demonstrated that rubric-driven LLM evaluation can align better with human judgments than older reference-based metrics for some generation tasks, while later judge research documented systematic biases and sensitivity to evaluation setup. The practical consequence is straightforward with the judge model, rubric, prompt, sampling settings, and parser should be versioned like any other dependency. Borderline cases should be sampled repeatedly or routed to human review rather than converted into false precision by a single score. Gate Changes Against a Baseline, Then Close the Production Loop Regression testing is most informative when a candidate is compared with a pinned baseline on identical cases. Absolute thresholds such as "quality must exceed 0.85" can hide a meaningful drop from 0.94 to 0.86. A gate should therefore combine non-negotiable case failures with relative movement in quality, cost, and latency. Python def release_allowed(baseline, candidate): if candidate["critical_failures"] > 0: return False quality_delta = candidate["quality"] - baseline["quality"] latency_ratio = candidate["p95_ms"] / baseline["p95_ms"] return quality_delta >= -0.01 and latency_ratio <= 1.10 The tolerances in this example are application policy, not universal constants. More important is the comparison model where the baseline and candidate run against the same versioned corpus, results remain inspectable by case and slice, and repeated samples are used when output variance is material. Recent research on judge reproducibility reinforces this requirement, finding that temperature control can reduce variability without eliminating it and arguing that grader disagreement should be treated as an evaluation signal rather than ignored. CI can then apply evaluation at different depths. Pull requests can execute a compact, high-signal corpus covering critical behaviors, while scheduled or pre-release runs can execute broader datasets and repeated samples. Official LangSmith guidance supports offline evaluation for comparing versions before deployment and provides CI/CD integration patterns around evaluation runs. The important design principle is not the specific platform, as evaluation results must be release evidence rather than a dashboard consulted only after a failure. Offline evaluation still cannot fully represent production traffic. Query distributions change, documents change, tools return new states, and upstream services evolve. Production monitoring therefore completes the regression loop. Google’s production ML guidance emphasizes live quality measurement, drift detection, and ongoing monitoring rather than treating validation as a one-time pre-deployment event. Sampling real traces, reviewing low-confidence or high-impact outcomes, and promoting confirmed failures back into the regression corpus turns operational incidents into permanent test coverage. AI regression testing becomes reliable when deterministic software tests and behavioral evaluations are treated as complementary rather than interchangeable. Unit tests should enforce contracts and invariants, curated datasets should preserve real failure modes, evaluators should measure task-specific properties, baseline comparisons should detect meaningful degradation, and production traces should continuously refresh the suite. An LLM application is ready for release not when every generated sentence matches a fixture, but when measured behavior remains within explicit quality, safety, cost, and latency tolerances under representative conditions. That shift turns evaluation from an ad hoc prompt-checking exercise into a disciplined software engineering control for probabilistic systems.
Every engineering team scaling AI systems in enterprise software eventually hits the same roadblock. On one side, developers are deploying retrieval-augmented generation (RAG) pipelines and LLM microservices. On the other side, risk and compliance committees present a long checklist derived from frameworks like NIST AI RMF or ISO 42001. In my experience leading enterprise technology controls and risk frameworks across regulated environments, I have repeatedly seen everything from small AI use cases to multimillion dollar AI initiatives stall for months simply because compliance teams could not verify risk controls through static spreadsheet reviews. The standard industry response has been manual post hoc reviews. Teams fill out algorithmic impact spreadsheets at the end of a sprint cycle, schedule sign-off meetings and delay release cycles by weeks. This manual approach fails in production. When applied to non-deterministic, fast-evolving AI models, spreadsheet governance creates an illusion of control while missing runtime failure modes. To build safe, compliant, and scalable AI systems, we must shift governance out of meeting rooms and embed it directly into CI/CD pipelines as Architectural Fitness Functions. The Flaw: AI Breaks Deterministic Non-Functional Testing In classic microservice architectures, Non-Functional Requirements (NFRs), such as memory footprint, response latency, and auth checks are deterministic. A test passes or fails based on binary, predictable logic. AI components break this paradigm in three distinct ways: Non-Deterministic Behavior: A minor update to prompt construction or model parameters can change system outputs across thousands of edge cases without throwing an error or crashing a service.Silent Failure Modes: System failures in AI rarely manifest as HTTP 500 errors. Instead, they appear as subtle hallucinations, context degradation, or unhandled prompt injections that quietly erode user trust.Dynamic Dependencies: Logic is defined not only by application code, but also by vector embeddings, model weights, and external API responses. Because these risks are runtime behaviors rather than static syntax bugs, pre-release security reviews inevitably miss them. We need a way to translate high-level Risk Appetite Statements (RAS) into automated build gates and continuous runtime assertions. The Concept: AI Governance as Architectural Fitness Functions In software architecture, a fitness function provides an objective, automated assessment of a system's architectural characteristics. To govern AI effectively, we must write AI Governance Fitness Functions, which are automated code assertions integrated into CI/CD pipelines and observability stacks that continuously test system behavior against predefined risk thresholds. Instead of asking engineers, "Did you complete the AI risk checklist?", the pipeline automatically asks, "Did this build satisfy our governance fitness assertions before reaching staging?" Implementing 3 Tiers of Pipeline Controls A robust governance automation strategy operates across three distinct stages of the delivery lifecycle. 1. Shift-Left Policy-as-Code Before code reaches staging, static analysis tools and policy engines like Open Policy Agent (OPA) evaluate configuration files, system prompts, and API parameters. For instance, an OPA Rego policy can enforce mandatory safety boundaries and PII masking rules directly on prompt templates: Plain Text package software.governance.ai default allow = false allow { input.component_type == "llm_prompt_wrapper" input.pii_masking_enabled == true input.max_output_tokens <= 2048 contains(input.system_prompt, "DO NOT override system safety boundaries") } If an engineer modifies a prompt template and removes safety boundaries, the pull request build fails instantly. 2. Automated Evaluation Gates in CI/CD Unit tests cannot verify whether a retrieval model's accuracy has degraded. During continuous integration, automated evaluation gates must execute standardized 'golden datasets' against the staging setup. Instead of checking exact string matches, the fitness function executes semantic assertions for hallucination thresholds and toxicity scores: Python import pytest from evaluation_framework import evaluate_rag_pipeline def test_ai_governance_hallucination_threshold(): eval_results = evaluate_rag_pipeline(dataset_path="tests/governance/golden_eval_set.json") hallucination_rate = eval_results.get_metric("faithfulness_score") toxicity_score = eval_results.get_metric("toxicity_score") assert hallucination_rate >= 0.98, f"Build failed: Hallucination rate reached {1 - hallucination_rate:.3f}" assert toxicity_score == 0.0, "Build failed: Toxic output detected" If a change degrades response quality or introduces bias, the pipeline blocks deployment automatically, treating risk violations with the same urgency as broken software builds. 3. Production Telemetry and Circuit Breakers Governance does not stop at deployment. Models suffer from data drift, and upstream API providers occasionally update models under the hood. In production, Key Risk Indicators (KRIs) must be monitored alongside standard Service Level Indicators (SLIs). If runtime telemetry detects elevated hallucination rates or prompt injection attacks, automated circuit breakers should instantly downgrade the feature to a deterministic fallback path without crashing the service. Mapping Risk Appetite to Engineering Metrics To bridge the gap between risk committees and engineering teams, qualitative compliance requirements must be mapped directly to automated assertions: mapping risk to fitness functionsQualitative Risk StatementNon-Functional Requirement (NFR)Architectural Fitness Function"Prevent inaccurate financial outputs."Faithfulness score >= 0.98 on golden evaluation datasets.CI build gate asserting context relevance."Protect customer PII."Pre-flight PII sanitization on all outbound payloads.Static analysis check enforcing client wrappers."Ensure service availability."99.9% uptime with graceful non-AI fallback.Production circuit breaker reverting to static rules on API timeout. Code Is the Only Source of Truth Treating AI governance as a manual, post-development compliance step introduces friction without delivering safety. By shifting governance left and expressing rules as Architectural Fitness Functions, engineering teams can automate compliance, protect production systems, and scale AI products with confidence.
As Large Language Models (LLMs) continue their rapid trajectory of development, software engineers and cloud architects regularly run into two systemic challenges: The Fixed Knowledge Cutoff: Model intelligence is inherently restricted to its training data window, making it blind to real-time changes. The "Air-Gap" Limitation: Out of the box, LLMs cannot securely interact with external systems or private APIs on their own. Historically, developers bypassed these hurdles by writing fragile, ad-hoc API wrappers or custom orchestrators. Enter the Model Context Protocol (MCP): an open standard designed to standardize how AI applications safely connect to external data sources and execution environments.In this article, we will explore the core architecture of MCP, look at why the Azure MCP Server is a game-changer for cloud engineers, and walk through a step-by-step guide to configuring it inside Visual Studio Code. The Core Architecture of MCP At its heart, MCP establishes a uniform "language" that allows AI applications (Hosts) to talk to external resources (Servers). Rather than building custom integrations for every new LLM or tool, developers can rely on a single, clean architecture: MCP Architecture The standard is built around four fundamental building blocks: MCP Host: The runtime environment or user interface where the AI agent operates (e.g., VS Code, Claude Desktop, Cursor).MCP Client: The architectural component within the host that initiates and maintains the active connection.MCP Server: A lightweight, modular helper service that exposes specific resources, prompts, and tools.Transport Layer: The underlying protocol facilitating communication, typically utilizing JSON-RPC 2.0 over standard input/output (stdio) or HTTPS. MCP Components How the MCP Handshake Works Instead of executing raw, unpredictable bash scripts, the interaction is highly structured: Initialization: The client connects to the server and queries its capabilities. Declaration: The server returns a structured schema listing the specific tools it supports. Execution Request: When an LLM determines it needs external data, the host requests a specific tool execution from the server. Context Injection: The server executes the local process, gathers the result, and returns a semantic JSON response to the host, which is then cleanly formatted for the user. MCP in Action Why Use the Azure MCP Server? If you are managing infrastructure on Microsoft Azure, the Azure MCP Server bridges the gap between your local AI assistant and your active cloud resources. Operating as a secure local process, it natively integrates with the Azure command-line context. The server supports over 40 Azure services and more than 170 tools out of the box, spanning critical cloud primitives: Compute & Containers: Azure Container Apps and Azure Kubernetes Service (AKS). Storage & Resource Management: Azure Storage (blobs and containers) and Azure Resource Groups. Infrastructure as Code: Integrated Azure Terraform Best Practices. Key Advantages Over Raw CLI Executions While you could technically let an AI agent execute arbitrary commands in a terminal, using a dedicated MCP server provides several major structural benefits: Rich Semantics: Instead of parsing messy, unstructured terminal stdout text, the MCP server passes rich semantic data blocks directly back to the LLM. Strict Governance & Scope: You can explicitly configure the server to run in read-only mode or expose only a select subset of tools, preventing the AI from accidentally deleting production infrastructure. Interactive Guardrails: The protocol requires explicit user confirmation before executing tools that touch sensitive data or perform mutative actions. Azure MCP Server Setup and Configuration Modes You can run the Azure MCP server locally across multiple development setups using stdio transport. Below are the three most common configuration schemas. NuGet Configuration For .NET teams, the server can be dynamically fetched and started using the dnx toolchain: JSON { "mcpServers": { "Azure MCP Server": { "command": "dnx", "args": [ "Azure.Mcp", "--source", "https://api.nuget.org/v3/index.json", "--yes", "--", "azmcp", "server", "start" ], "type": "stdio" } } } Node.js Configuration If you are developing in a standard JavaScript/TypeScript ecosystem, you can spin up the server dynamically using the latest npm package via npx: JSON { "mcpServers": { "azure-mcp-server": { "command": "npx", "args": [ "-y", "@azure/mcp@latest", "server", "start" ] } } } Docker Configuration For isolated development environments, you can run the server in a container. Note that you must provide a local environment file (.env) containing your Azure Service Principal credentials: JSON { "mcpServers": { "Azure MCP Server": { "comman { d": "docker", "args": [ "run", "-i", "--rm", "--env-file", "/full/path/to/.env", "mcr.microsoft.com/azure-sdk/azure-mcp:latest" ] } } } Step-by-Step Visual Studio Code Integration A great feature for those who want to work within Visual Studio Code is that they can also manage the Azure MCP Server through an exclusive extension available directly in their editor. The following steps walk you through the setup. Step 1: Install the Extension Search for and install the Azure MCP Server extension directly from the Visual Studio Code marketplace. Install Azure MCP Server Extension Step 2: Initialize and Verify Open the Command Palette (Cmd + Shift + P on macOS or Ctrl + Shift + P on Windows) and search for the extension commands to verify the server is active and running. Select an MCP Server Step 3: Configure Your Tool Accessibility Open your integrated AI chat window and select the Tools icon. Here, you will see a list of all active tools provided by the Azure MCP Server. You can toggle specific permissions on or off—for instance, disabling write operations while keeping read operations active. Select Your Tools Step 4: Interact Natively Now, your AI chat assistant can securely call Azure tools in the background. You can ask complex queries like: "Are there any inactive containers running in my resource group?""Upload our local configuration file directly to our Azure storage container blob storage." Any further interactions that reference an Azure account would use the MCP tools. Confirm Tools and Resources Real-World Engineering Use Cases To see the power of this setup, let’s look at how this changes day-to-day operations: Use Case A: Automated AKS Incident Troubleshooting When an incident occurs in an Azure Kubernetes Service (AKS) cluster, engineers typically run dozens of diagnostic commands. With the Azure MCP server connected, you can simply ask the LLM: "Investigate why the pods in our production namespace are crash-looping." The agent will call the relevant AKS tools, inspect the logs, identify the misconfiguration, and suggest the fix - all in seconds. Use Case B: Continuous Terraform and Compliance Audits Before deploying infrastructure, you can point your local AI agent to your code directory. Because the server incorporates Azure Terraform Best Practices, the agent can audit your configuration files, cross-reference them against your live Azure Resource Groups, and warn you if you are violating security compliance rules or generating drift. Conclusion The Model Context Protocol represents a major step forward in AI-assisted development. By standardizing the communication layer, the Azure MCP Server enables software engineers to transform static, isolated LLMs into active, context-aware cloud operators. Whether you are monitoring active Kubernetes clusters, auditing Terraform configurations, or automating file uploads to blob storage, MCP gives your AI assistant safe and highly scoped access to the "live" Azure ecosystem.
If you have ever drawn a forecast band around a price series, you have probably reached for the same formula everyone reaches for: take a volatility estimate, multiply by the square root of the horizon, and call the result your interval. sigma * sqrt(t) is the closed-form move; it is one line of code, and for fat-tailed financial series it is wrong in a way that is worth understanding precisely, because the error does not even keep the same sign as you change the horizon. This is a short engineering write-up on what breaks, how we measured it, and the handful of things you actually have to compute to build an honest, reproductible interval. What Sqrt(t) Assumes, and Why Crypto Violates It The sigma * sqrt(t) scaling falls out of assuming i.i.d. Gaussian log-returns: variance adds linearly in time, so standard deviation scales with the root of time, and a fixed multiple of it gives a fixed-probability band. Two assumptions are doing all the work — normality and independence — and daily crypto returns honor neither. Realized kurtosis on daily BTC log-returns sits near 16 against the Gaussian's 3. Returns are close to uncorrelated, but their magnitudes are strongly autocorrelated (volatility clustering), so the independence the formula needs is not there either. We measured the damage against the full Binance history rather than a recent window — 3,261 daily bars for BTC back to 2017. The quantity of interest is the ratio of an empirically measured 80% band to the sigma * rt(t) band at each horizon: Shell horizon empirical80 / sqrt-t band (BTC) 1 day ~0.80 7 days ~0.88 30 days ~1.00 Read that carefully: at one day the parametric band is too wide (0.80), and by thirty days it is about right (1.00). The error changes sign with the horizon, so there is no single scale factor you can bake in to fix it. The reason the short-horizon band is too wide despite fat tails is that the kurtosis lives in the extreme tails, not in the 10th–90th-percentile shoulders — so the 80% interval is actually narrower than Gaussian while the 99% interval is much wider. Fat tails and a narrow 80% band coexist, which is exactly the kind of thing a parametric shortcut hides from you. Compute the Empirical Band Instead The fix is to stop parameterizing and read the interval straight off the empirical distribution of realized h-day log-returns: Python import numpy as np def empirical_band(prices, horizon, lo=0.10, hi=0.90): p = np.asarray(prices, dtype=float) # overlapping h-day log-returns r = np.log(p[horizon:] / p[:-horizon]) q_lo, q_hi = np.quantile(r, [lo, hi]) spot = p[-1] return spot * np.exp(q_lo), spot * np.exp(q_hi) That is the whole idea, and it already beats the parametric band because it makes no distributional assumption. But there is a trap in the line that builds r. The Overlapping-Windows Trap Those h-day returns overlap: consecutive 30-day windows share 29 days of data. Overlapping samples are heavily autocorrelated, so if you report len(r) as your sample size, you are overstating your evidence by roughly the horizon. For BTC's 30-day band, ~3,231 overlapping windows correspond to only about 107 independent months. Quantile estimates from overlapping windows are still usable, but their uncertainty is far larger than the raw count implies, and you must report the independent count, not the overlapping one: Python def independent_count(n_bars, horizon): return max(1, (n_bars - horizon) // horizon) We print band from N independent windows directly on every chart for this reason. A band from 107 independent months is a different epistemic object than one implying 3,231 samples, and collapsing the two is how backtests quietly manufacture confidence. Evaluate With Coverage, in Both Directions The metric for an interval forecast is coverage, and the failure is symmetric. A claimed 50% band should contain the outcome about half the time across many out-of-sample days. If it contains 90%, the band is padded — a failure that a one-sided "were we inside?" check will never catch, because padding always looks safe. So score both the 50% core and the 80% band against their targets, over many days, and treat over-coverage as a miss. Make It Reproducible and Tamper-Evident The last piece is provenance, and it is pure engineering. We serialize each forecast object, hash it with SHA-256, and anchor the hash to the Bitcoin blockchain via OpenTimestamps before publishing. One gotcha worth flagging because it cost us a real bug: hash the exact bytes you publish. We were writing the JSON with a trailing newline but hashing the object without it, so a reader running shasum on the published file got a different digest — which reads as fraud even though nothing was wrong. Write the file, hash the file, timestamp the file; byte-for-byte, no re-serialization in between. None of this is an edge, and the write-up would be dishonest if it implied one. An empirical band lowers the cost of being wrong about volatility; it does not tell you direction. No method reliably beats a liquid market, and anyone promising that is selling something. What the empirical quantile buys you is a band whose width means what it says — and a pipeline where any reader can recompute the number and check the timestamp themselves. The live version, scored in public with the misses kept, is at neuportal.ai/experiment. Educational content — not financial advice.
The Problem The Model Context Protocol has an initialize handshake. A client connects, the server hands back an Mcp-Session-Id header, and every request after that carries the same ID. That's fine with one server. It falls apart with three. A load balancer doesn't know or care about Mcp-Session-Id. It's an application-layer detail, and by default, the load balancer routes based on whatever algorithm it uses: round robin, least connections, or source IP hash. The client's second request can land on a completely different instance than the one that issued the session. That instance has never heard of the session ID, so it returns an error or, worse, silently starts a new session with no memory of the tools the client already listed or the state it had built up. Teams hit this the same way every time. Local development works because there's one process. Staging works because there's one pod. Production breaks the first time the deployment scales beyond a single replica, and the failure looks like a flaky client rather than an infrastructure problem, so it takes a while to trace. The standard fixes were the same ones every stateful HTTP service has used for 20 years: sticky sessions at the load balancer (route by session ID or client IP, accept the uneven load distribution), a shared session store like Redis so any instance can pick up any session, or deep packet inspection at the gateway to route on the Mcp-Session-Id header specifically. All three work. All three add an operational dependency to what should be a stateless API call. What Changed in the Spec The 2026-07-28 MCP specification release candidate removes protocol-level session management. The Mcp-Session-Id header is gone, and so is the session it represented. Connection metadata that used to be exchanged once at initialize time (protocol version, client info, client capabilities) now travels _meta on every request instead. The practical effect: a remote MCP server can sit behind a plain round-robin load balancer, route traffic on the Mcp-Method header if you want method-aware routing, and let clients cache tools/list responses for as long as the server's ttlMs value says they're valid. No sticky sessions. No shared session store just to keep the protocol working. This doesn't mean your server has to be stateless. It means the protocol stopped assuming state lives at the transport layer. If your server genuinely needs to remember something between calls (a shopping basket, an open browser tab, a half-finished multi-step operation), you handle it the way HTTP APIs have always handled it: mint an explicit handle from a tool call and have the model pass that handle back as an ordinary argument on the next call. Migrating a Stateful Server Say you have an MCP tool that opens a browser session and later needs to make calls to act on that same browser. Before the spec change, you'd have been tempted to key that off the transport-level session ID. After that, you do it explicitly: Python from fastmcp import FastMCP import uuid mcp = FastMCP("browser-tools") class BrowserHandle: """Minimal stand-in for a real browser automation session (e.g. Playwright).""" def __init__(self, start_url: str): self.start_url = start_url def click(self, selector: str) -> None: # Real implementation would drive an actual browser page here. pass # In-memory for the example; use Redis or a DB in production _sessions: dict[str, BrowserHandle] = {} @mcp.tool() def open_browser(start_url: str) -> dict: """Open a browser and return a handle for later calls.""" handle_id = str(uuid.uuid4()) _sessions[handle_id] = BrowserHandle(start_url) return {"browser_id": handle_id, "status": "opened"} @mcp.tool() def click_element(browser_id: str, selector: str) -> dict: """Click an element in a previously opened browser.""" session = _sessions.get(browser_id) if session is None: return {"error": f"No browser session for id {browser_id}. Call open_browser first."} session.click(selector) return {"status": "clicked", "selector": selector} The model now carries browser_id as an ordinary tool argument, the same way it would carry a basket_id or an order_id. That handle can live in Redis with a TTL, in a database row, wherever makes sense for your durability requirements. It's no longer the protocol's job to keep it alive. If you're running behind a load balancer today with sticky sessions configured specifically to work around the old MCP session model, this is your cue to check whether you can drop that configuration once your server and client both support the 2026-07-28 spec, or its final released version. Check the actual spec version your SDK negotiates before you rip out sticky sessions. Older clients still speaking the pre-07-28 protocol will still expect Mcp-Session-Id to work, and mixed-version fleets are exactly the kind of thing that turns a clean migration into a bad on-call week. What to Check Before You Migrate A few things worth confirming before you touch the load balancer config: Check your MCP SDK version and whether it implements the stateless spec or still assumes session pinning. Not every SDK moved at the same pace.Check whether any of your tools rely on server-side state that is implicitly tied to the session lifecycle (e.g., an open file handle, a database transaction, or a lock). Those need an explicit handle now, not an assumption that "the session" will still be around.Check your load balancer's health check and routing rules for anything referencing Mcp-Session-Id specifically. If your ops team added rules to route on that header, they can likely come out.Test with a client that is legitimately routed to different instances across consecutive calls, not just a local single-instance setup. That's the scenario the old model broke on, and it's the one you want to confirm the new model handles. References The 2026-07-28 MCP Specification Release CandidateSEP-2567: Sessionless MCP via Explicit State HandlesSEP-1442: Make MCP Stateless (by default)Scaling AI Agent Infrastructure with the MCP Stateless updates (Google Developers Blog)
This article is for platform engineers, data engineers, security architects, and AI application teams building enterprise agents that retrieve data, call tools, or trigger workflows on behalf of users. Imagine a support engineer asking an internal AI agent for a customer summary. The user is cleared to see support tickets, but the agent runs under a broad service account that can also reach contract terms, payment history, and escalation notes. The agent does not need malicious intent to create a breach. If it retrieves contract terms the user could not normally access, governance has already failed at the data boundary. That is why enterprise AI governance cannot stop at the model. Eval sets, content filters, and prompt rules are useful, but they sit above the point where the real risk lives: the moment the agent reaches into your systems and pulls something back. This article focuses on that exact moment: the boundary where an AI agent moves from reasoning about a request to touching enterprise data or executing a tool. That boundary deserves to be treated as a first-class security control. The Problem With the Account the Agent Runs As Traditional data access control assumes one of two callers. Either a human is behind the query, authenticated and carrying their own permissions, or a fixed service account is running a known, reviewed workload. Role-based access control was designed for both. An agent is neither. An agent composes queries at runtime. It decides which tool to call, which data to fetch, and how to chain those calls in sequences nobody reviewed in advance. If it runs under a broad service account, every user talking to that agent can inherit that account's reach. The fix is not to make the model sound more careful. The fix is to put a control layer between the agent and enterprise systems, then enforce that control at the moment each data or tool call is made. The Pattern: A Guardrail Gate on Every Data Access A guardrail gate is a runtime enforcement layer between the agent and the systems it wants to access. The agent must pass through this gate on every data or tool call. The gate performs four jobs in sequence: bind the caller's real identity, scope what can be retrieved, gate the action, and record the decision. Figure 1. The guardrail gate: four control points between the agent and your data. The same pattern is easier to operationalize as a decision flow, because the important behavior is in the branches: allow, deny, or require human approval. Figure 2. Request lifecycle, with allow, deny, and human-approval branches. 1. Bind the Caller's Real Identity, and Carry It All the Way Down This is the control most teams get wrong, so it is worth slowing down on. The agent should not act with standing privileges. It should act as the human it serves, resolved fresh on every request. The key is mechanical: the user's identity must travel from the chat box to the data access layer without being swapped for a service account. The clean way to do this is a token exchange. The agent forwards the user's token, and the guardrail trades it for a short-lived downstream credential that authorizes as that user, not as the agent. This aligns with the OAuth 2.0 Token Exchange pattern described in RFC 8693, which defines how a security token service can issue a new token for delegation or impersonation across security domains. Python def exchange_on_behalf_of(user_token): # Trade the user's token for a short-lived downstream credential # that carries THAT user's identity, not the agent's. if not verify_signature(user_token): return None return sts.exchange( subject_token=user_token, audience="data-plane", # who the credential is for # the resulting credential authorizes AS the user ) Now the data layer can enforce row-level security under the user's identity instead of trusting the agent. If this control is right, the data plane already refuses anything the user could not see directly. 2. What the Agent Can Retrieve Before the Query Runs Even a correctly identified user can ask a question whose honest answer would require data they should not see. Retrieval scoping narrows the searchable surface before the agent ever runs a query, rather than filtering results after the fact. The distinction matters. Post-filtering means the sensitive rows were fetched, sat in memory, possibly landed in a log, and only then got dropped. Pre-scoping means they were never reachable. Here is the crux made concrete: an agent doing retrieval-augmented generation over a vector store, with the guardrail enforcing per-user access at retrieval time. Python def handle_agent_retrieval(agent_request): # (1) Bind the real caller. The agent forwards the user's token, # never its own service credential. principal = exchange_on_behalf_of(agent_request.user_token) if principal is None: return audit_and_deny(agent_request, reason="no verifiable caller") # (2) Turn the user's clearances into a metadata filter the vector # search cannot escape. This is a PRE-filter, not a post-filter. allowed_filter = access.metadata_filter_for(principal) # e.g. {"region": principal.region, "dept": principal.dept} if not allowed_filter: return audit_and_deny(agent_request, reason="no in-policy corpus") hits = vector_store.search( embedding=agent_request.query_embedding, top_k=8, metadata_filter=allowed_filter, # out-of-policy chunks are never returned ) audit.record(principal, "retrieval", allowed_filter, len(hits)) return hits The one line that does the work is metadata_filter=allowed_filter. Out-of-policy chunks are never retrieved, so they never enter the prompt, never reach the model, and never show up in a trace. 3. Gate the Action, Not Just the Read Retrieval is only half of what an agent does. The other half is acting: writing a record, triggering a workflow, sending something outward. A read policy that is airtight does nothing if the agent can then take an action the user was never allowed to take. Every tool the agent can call needs an explicit policy that answers three questions: Is this user allowed to invoke this tool? Are the arguments within approved bounds? Does the action require a human approval step before it commits? Python def check_action(principal, tool_call): policy = action_policy_for(tool_call.name) if not policy.permits(principal, tool_call.args): return deny(f"{tool_call.name} not permitted for this caller") if policy.requires_confirmation(tool_call.args): return require_human_approval(tool_call) # high blast-radius path return allow(tool_call) The high-blast-radius actions — anything that writes, spends, sends, or deletes — are the ones that most deserve a confirmation step. Be conservative here early and loosen later, rather than the reverse. 4. Record Every Access as an Event The last control point does not block anything, which is why people skip it, and it is the one that saves you when something goes wrong. Every access decision the guardrail makes — allow or deny — with the resolved identity, the scope, and the tool call, should be written as an audit event. This is not logging for its own sake. When an agent produces a surprising result three weeks from now, the audit trail is the only thing that lets you answer: What did it actually touch, on whose behalf, and why was that allowed? Without it, you are guessing. Putting the Four Controls Together Each control is simple on its own. The payoff comes when they compose into a single gate that every agent call passes through, retrieval and tool execution alike. In practice, that is one function, or one middleware, wrapping the agent's access to the outside world: Python # The four controls, composed into one gate every agent call passes through. # Wrap the agent's access to data and tools with this single entry point. def guardrail(agent_request): # 1. Identity: bind the real user, never the agent's service account. principal = exchange_on_behalf_of(agent_request.user_token) if principal is None: return audit_and_deny(agent_request, reason="no verifiable caller") # 2. Retrieval: scope the searchable surface BEFORE any query runs. if agent_request.kind == "retrieval": allowed = access.metadata_filter_for(principal) if not allowed: return audit_and_deny(agent_request, reason="no in-policy corpus") hits = vector_store.search( embedding=agent_request.query_embedding, top_k=8, metadata_filter=allowed, # pre-filter, not post-filter ) audit.record(principal, "retrieval", allowed, len(hits)) return hits # 3. Action: gate tools, route high-blast-radius calls to a human. if agent_request.kind == "tool_call": decision = check_action(principal, agent_request.tool_call) audit.record(principal, "action", agent_request.tool_call, decision.status) return decision # allow, deny, or require_human_approval return audit_and_deny(agent_request, reason="unknown request type") Every path through the gate, including each denial, ends in an audit record, so nothing the agent does escapes review. The four controls are not four features to build separately; they are one checkpoint the agent cannot go around. Why This Belongs at Runtime These controls belong at runtime because an agent's behavior is not fixed in advance. A pipeline's access can be reviewed once; an agent's access depends on the user request, model decision, and tool chain. The boundary must be enforced where each call is made. None of this requires a new platform. It requires treating the space between your agent and your data as a first-class component with its own responsibilities. That idea is consistent with the broader Zero Trust principle that access decisions should be explicit and resource-centered rather than assumed from network location, as described in NIST SP 800-207. Where the Guardrail Gate Sits in the Stack In production, the guardrail gate sits between the agent runtime and every system the agent can touch, exactly as Figure 1 shows. Beyond the agent, the gate, and the data plane already in that diagram, a real deployment leans on three supporting services: Identity provider or token exchange service: issues the short-lived, user-scoped credentials the gate binds to each request.Policy engine: decides whether this user, resource, operation, and set of arguments are permitted.Audit sink: receives allow, deny, and approval events for logs, governance dashboards, or a SIEM. This placement keeps the model useful while preventing it from becoming the security boundary. The model proposes; the gate, backed by identity, policy, and audit, decides. Before and After: Broad Service Account vs. User-Bound Retrieval Before: The support agent from the opening runs on one broad service account and searches the entire customer corpus. The app tells the model to avoid sensitive content, but the vector store still returns contract, finance, and escalation chunks. The model may ignore some of them, but unauthorized data has already entered the runtime. After: The same request flows through the guardrail gate. It exchanges the user's token, builds a metadata filter from the user's role and scope, and applies it before vector search. Out-of-scope chunks are never fetched, and the audit trail records the user, scope, query class, and result count. That one shift changes the security model. Instead of trusting the model to behave after it has seen too much, the system ensures it never receives data outside the user's allowed scope. Production Implementation Checklist For teams turning this pattern into production controls, the checklist works best when it is grouped by control area. Identity and Delegation Do not let agents use broad standing service credentials for user-facing requests.Exchange the user's token at runtime and carry the resolved user identity to the data plane.Use short-lived downstream credentials that expire quickly and are scoped to the requested resource. Retrieval Controls Apply retrieval filters before search so out-of-policy chunks are never fetched.Keep metadata filters close to the search layer rather than relying on the model to ignore unauthorized context.Treat vector indexes, document stores, and SQL endpoints as policy-enforced data planes, not passive context providers. Action Controls Define explicit policies for every tool the agent can call, including argument bounds.Require human approval for write, send, delete, spend, or other high-impact actions.Start conservative for high-blast-radius actions and loosen policy only after reviewing real usage patterns. Auditability and Review Record allow, deny, and approval decisions as audit events with caller, scope, tool, arguments, and reason.Review guardrail decisions periodically against security guidance such as the OWASP GenAI LLM Top 10 and risk-management practices such as the NIST AI Risk Management Framework.Feed denied requests, approval patterns, and near misses back into policy tuning and threat modeling. What Not to Do A few anti-patterns show up repeatedly in early agent deployments. They are tempting because they make the first demo easier, but they also move the security boundary to the weakest possible place. Do not let the agent use one broad service account and hope the prompt keeps it honest.Do not retrieve everything first and filter sensitive results afterward.Do not treat prompt instructions as access control.Do not give write-capable tools to an agent without a policy layer and an approval path.Do not log only the final answer; log the access decision that produced it. Governance that stops at the model is theater. The real control point is the runtime boundary between the agent and the systems it can touch. Bind the user, scope retrieval before search, gate actions before execution, and record every decision. Then inventory every data source and tool your agent can reach, and force each path through one guardrail gate before production. References The references below provide the identity, Zero Trust, and AI risk-management standards behind the pattern. RFC 8693: OAuth 2.0 Token ExchangeNIST SP 800-207: Zero Trust ArchitectureOWASP GenAI LLM Top 10 2026NIST AI Risk Management Framework
An approved model does not make an AI system accountable. The consequential behavior emerges from the model's interaction with runtime context, retrieved data, tools, permissions, orchestration rules, and human approval controls. That changes the assurance target. Engineering teams still need model evaluation, dataset documentation, bias testing, red teaming, and release approval. They also need to prove that the deployed system operated within its delegated authority and that they can reconstruct the execution path that produced an outcome. The practical shift is from governing a model artifact to governing a socio-technical execution boundary. Move the Assurance Boundary Outward Production AI systems increasingly assemble decisions at runtime. The orchestrator selects context, the retriever introduces enterprise data, policy services constrain actions, tools change external state, and reviewers approve or challenge recommendations. The model participates in the decision; it does not determine the system boundary. A stale vector index can produce a harmful recommendation even when the model performs perfectly. An overprivileged tool can convert a hallucination into an unauthorized financial transaction. A poorly designed approval queue can reduce human oversight to rubber-stamping. None of these failure modes exist within model weights. In architecture reviews, evaluate system capability as a function of several interacting controls: Capability = f (Model, Context, Data, Tools, Permissions, Policy, Workflow) This formulation sets a boundary rather than generating a score. Changing any term alters the application's risk profile without altering the model. For instance, a single foundation model can power three distinct risk profiles: Internal Summarizer: Read-only access with zero external side effects.Customer Service Agent: Reads account history, issues messages, and processes limited refunds.Procurement Agent: Compares suppliers, negotiates contract terms, and executes purchases. Model governance applies to all three. System governance must distinguish their different authority levels, blast radiuses, and evidence requirements. Governing Execution Paths, Not Isolated Operations Traditional identity and access management (IAM) evaluates operations in isolation: May this principal read this record? May this role invoke this API? Agentic workflows introduce risk through path composition. If an agent can read sensitive customer records and send external emails, both permissions might pass individual IAM checks. However, combining them creates a data exfiltration risk. Static role-based access control (RBAC) cannot reliably detect path-based failures because enforcement depends on prior events, data classification, business purpose, and accumulated context. The execution path itself must become a first-class governance object. Path-aware enforcement can be implemented using deterministic runtime rules: Context labeling: Tag data with classification metadata as it enters the context window.Metadata propagation: Pass sensitivity and purpose tags through intermediate artifacts.Accumulated state evaluation: Evaluate proposed tool calls against accumulated context labels.Transaction binding: Bind high-impact actions to transaction-specific approvals.Approval invalidation: Invalidate pending approvals if arguments (recipient, amount, payload) change.Fail-closed policy: Deny execution when required provenance metadata is missing. Using deterministic controls to evaluate proposed model actions prevents enforcement responsibilities from falling on the probabilistic model itself. Separating Technical Access from Delegated Authority An IAM role specifies which resources an identity can read or invoke, but it does not define which business decisions an agent may make. A procurement agent may require read access to supplier catalogs, price histories, and contract schemas; these permissions do not imply authority to select a vendor, accept commercial terms, or release payments. Represent delegated business authority as an explicit, machine-enforceable contract: YAML agent_id: procurement-agent-17 purpose: supplier_comparison data_scope: allow: - approved_supplier_catalog - purchase_history tools: allow: - search_catalog - request_quotation deny: - create_purchase_order - release_payment decision_authority: allow: - recommend_supplier deny: - accept_commercial_terms execution_limits: autonomous_spend_gbp: 0 subagent_delegation: false approval: required_for: - supplier_selection - purchase_order expires_after_minutes: 15 revocation: mode: immediate An authority schema must address six core operational questions: Purpose: What business objective is permitted?Data and tools: Which resources and interfaces are authorized for that purpose?Decision scope: Which choices may the agent recommend versus decide?Execution scope: Which side effects may it trigger autonomously?Constraints: What financial limits, approvals, and segregation-of-duties rules apply?Ownership: Which accountable principal can revoke this authority? Fine-grained authority models introduce policy maintenance overhead, whereas coarse roles obscure authority inside technical permissions. Organizations should standardize reusable authority profiles for common risk tiers, applying transaction-level constraints for consequential actions. Operationalizing Human Oversight Inserting a manual approval gate does not guarantee meaningful oversight. A reviewer evaluating hundreds of decisions daily with seconds per case will default to confirming automated outputs—especially when the UI displays only the final recommendation. Effective human-in-the-loop controls require structured operational conditions: Contextual delivery: Present recommendations alongside supporting evidence, provenance tags, and explicit uncertainty flags.Scope-bound approvals: Bind approvals strictly to exact proposed parameters. An approval for a £500 adjustment must not authorize a £5,000 transaction or a modified payload.Actionable workflows: Provide clear mechanisms to challenge, override, or escalate recommendations.Telemetry monitoring: Track review velocity, queue pressure, override rates, and approval patterns.Adaptive capacity control: Suspend or re-route automation if reviewer capacity drops below design thresholds. Applying risk-tiered review avoids operational bottlenecks: fully automate low-impact, reversible tasks; sample medium-risk workflows for quality control; and mandate explicit approval for high-impact or irreversible side effects. Connecting Design Time Intent to Runtime Evidence Design documentation states how a system should operate; runtime evidence proves how it executed. This distinction is vital when agents select tools dynamically, retrieval contexts shift, and workflows execute asynchronously. Post-incident analysis cannot reconstruct a decision using only a model version and final output. A robust architecture segregates concerns into three distinct planes: Governance (purpose, policy, authority), Execution (agent, model, context, tools), and Evidence (provenance, audit traces, outcomes). Missing LayerOperational FailureGovernance Without EnforcementPolicies remain static documentation rather than active runtime controls.Execution Without EvidenceActions occur, but the system cannot reconstruct or defend decisions during audits.Evidence Without GovernanceSystems collect logs without the contextual rules needed to detect policy violations. For every consequential execution, record typed events with a stable correlation ID to capture: Initiation context: Initiating identity, business purpose, and active authority profile.Component metadata: Model, prompt template, policy engine, and tool versions used.Lineage and logic: Retrieved source items, retrieval timestamps, policy rules evaluated, and tool arguments generated.Oversight and state: Reviewer identity, exact approval scope, executed state changes, and downstream business outcomes. Use structured, typed event schemas rather than unstructured trace logs for audit compliance. Unstructured logs remain valuable for real-time debugging, but schema drift makes them unreliable for formal control testing. Preserving Audit Evidence Securely Maximum observability differs from maximum accountability. Raw agent traces often contain personally identifiable information (PII), proprietary documents, system prompts, credentials, and confidential tool responses. Logging all context indiscriminately introduces security and compliance risks. Design audit systems for maximum verifiability with minimum necessary disclosure: Cryptographic references: Store immutable content hashes (SHA-256SHA-256) and versioned source IDs instead of duplicating whole payload documents.Segmented telemetry: Separate short-term operational logs from long-term, restricted-access audit records.Data minimization: Redact credentials, API keys, and unneeded PII before writing events to storage.Policy claims: Record structured policy decision inputs and outputs instead of raw contextual prompts.Tiered payload retention: Reserve full payload persistence for high-risk transaction classes that explicitly require complete replay capabilities. Defining the Governed System Boundary Different engineering domains maintain distinct perspectives on the system boundary: Model teams: View prompts, weights, and inference outputs.Application teams: Focus on orchestration, context assembly, and retrieval chains.Security teams: Monitor identities, network gateways, microservices, and data stores.Compliance teams: Track business purpose, data processing, policy checks, and end-user impacts.Operations teams: Manage throughput, system dependencies, error rates, and incidents. System governance aligns these perspectives around the end-to-end business outcome. When documenting architecture boundaries, trace the entire flow: initiating identities, runtime data context, orchestration modules, tools, human approval queues, downstream state changes, and final evidence persistence. If an architecture diagram ends at the model response while a downstream service executes a transaction, the assurance boundary is incomplete. Lifecycle Integration Guide Integrate accountability controls directly into the software development lifecycle: Map business consequences: Categorize system actions by reversibility, financial limits, data sensitivity, and recovery time.Define authority contracts: Establish explicit schemas covering data access, tool invocation, decision scope, execution boundaries, and revocation owners.Threat-model execution paths: Evaluate multi-step attack vectors, including prompt injection via retrieval data, confused deputy scenarios, parameter substitution, and tool-chaining exploits.Enforce at the side-effect boundary: Validate tool arguments, policy state, and accumulated context immediately before executing external actions.Calibrate review capacity: Design human approval interfaces to display clear evidence, enforce time-on-task standards, and automatically throttle automation if queues overflow.Emit structured evidence: Log typed events linked by correlation IDs to enable independent audit reconstruction.Verify revocation controls: Test kill-switches to ensure teams can revoke agent authority, invalidate pending approvals, and cancel queued tasks instantly. Architecture Review Checklist What real-world state changes can this system initiate? What runtime data enters the context window via retrieval and tools? What external side effects can each tool produce?What explicit decision boundaries govern the agent's actions?What sequence of individually permitted actions could produce a prohibited outcome?Where do deterministic policy engines validate proposed operations?Does human oversight remain effective during peak transaction volumes?Is approval cryptographically bound to the exact execution parameters?Can authority be revoked instantly to stop queued and in-flight operations?Can an independent auditor reconstruct a historical execution path without developer intervention?Is audit evidence stored securely without unnecessarily duplicating sensitive context? Model governance verifies the safety and training lineage of a technical component. However, it cannot confirm whether a running system stayed within its delegated authority, executed valid context, maintained human oversight, or initiated authorized business actions. Complete system assurance requires governing the full execution path: purpose, context, policy, tools, human controls, actions, and audit evidence. Placing deterministic controls at the action boundary ensures AI systems remain accountable, observable, and defensible in production.
"Give the agent memory" sounds like a feature request. In production systems, it is an architecture decision about durable state. The phrase agent memory is often used to describe several different things: Conversation historyWorkflow stateCheckpointsUser preferencesLong-running investigation contextRetrieved documentsTool results Combining all of them into one persistent conversation object is convenient during prototyping. It is risky in enterprise environments. Different forms of state have different durability requirements, access rules, and retention periods. A checkpoint required to resume an interrupted workflow is not the same thing as a user preference that may be useful six months later. The first step toward safe agent memory is therefore separating these concepts. Layer One: Transient Conversation Context The shortest-lived form of memory is the context needed for the current inference request. For example: Python messages = [ system_message, recent_user_message, relevant_tool_result ] This context helps the model reason about the immediate task. It does not necessarily need permanent storage. If the content contains large tool outputs, sensitive information, or temporary debugging data, persisting everything by default creates unnecessary exposure. Transient context should be aggressively scoped. Layer Two: Durable Run State A multi-step agent has execution state separate from conversation history. Consider: Python state = { "task": "analyze deployment failure", "current_step": "awaiting_approval", "candidate_fix": {...}, "attempts": 2 } If the process crashes, that state determines whether the workflow can resume. Without persistence, the entire run may restart from the beginning. That can be inconvenient for analysis and dangerous for workflows with side effects. Suppose the agent has already performed steps one through five and is waiting for approval before step six. A restart should not execute steps one through five again. The correct design is checkpointed execution. Failure mode: restart replays completed work If a workflow restarts from the beginning instead of from a durable boundary, already-completed tool calls can be repeated. The result may be duplicate tickets, duplicate notifications, repeated writes, or approvals applied twice. Checkpoint Before Important Boundaries A checkpointer stores graph state after meaningful transitions. Conceptually: Python checkpoint.save( principal_id=current_user.id, thread_id=thread_id, state=workflow_state, node="human_approval" ) If the process restarts: Python state = checkpoint.load(thread_id) graph.resume(state) The agent returns to the same logical boundary. This is especially useful for human approval. A workflow may pause overnight while an engineer reviews a proposed operation. The process that originally created the request does not need to remain alive. The durable checkpoint becomes the source of truth. Architecture at a Glance A useful way to model the system is to keep inference context, resumable execution state, and long-term memory as separate concerns around a governed state layer. Development and Production Need Different Backends During local development, a lightweight store is often enough. For example: Python checkpointer = LocalCheckpointStore( path="./agent_state.db" ) This makes agent development easy because engineers can restart processes and inspect state without deploying distributed infrastructure. Production requirements are different. Agent runs may execute on multiple instances. Containers may restart. Users may reconnect through another server. Workflows may remain paused for hours or days. A production checkpointer therefore needs distributed durability. Python checkpointer = DistributedCheckpointStore( namespace="agent-runs" ) The application should depend on a storage interface rather than backend-specific behavior: Python class CheckpointStore: def save(self, principal_id, thread_id, state): ... def load(self, principal_id, thread_id): ... def delete(self, principal_id, thread_id): ... The same agent runtime can then use a local implementation during development and a distributed transactional datastore in production. For example, a LangGraph checkpointer can use SQLite locally and PostgreSQL in production while preserving the same identity-scoped persistence contract. A Thread Is a Security Boundary: Not a Bearer Capability Long-running agent workflows often use a thread identifier. That identifier must not become the only access control. A request such as: HTTP GET /threads/48291 cannot mean “return this thread to whoever knows the ID.” The lookup should be scoped to identity: Python checkpoint.load( user_id=current_user.id, thread_id=request.thread_id ) Storage keys can make the boundary explicit: Plain Text /principal/{principal_id}/thread/{thread_id} Authorization must still be enforced independently, but physical key separation reduces accidental cross-user access. For enterprise agents, per-user or per-principal isolation is fundamental. “Resume yesterday’s investigation” should mean resume that principal’s authorized investigation, not retrieve a globally accessible conversation object. Failure mode: thread ID becomes a capability If knowing a thread identifier is enough to load its state, the identifier behaves like a bearer token. Treat thread IDs as locators, not authorization. Enforce identity and policy checks on every read and write. Redact Before Persistence If sensitive information should not be stored, removing it later is weaker than never storing it. Redaction should happen before writing state. Python safe_state = redact_sensitive_fields( workflow_state ) checkpoint.save( thread_id=thread_id, state=safe_state ) The redaction layer may target: CredentialsAuthentication tokensPersonally identifiable informationPrivate keysSensitive tool responsesRestricted document content A useful design keeps references instead of raw values when possible. Instead of: JSON { "api_token": "secret-value" } persist: JSON { "credential_reference": "credential://service-x" } On resume, the runtime can reacquire the authorized credential. This avoids turning the memory store into a shadow secret-management system. Do Not Persist the Entire Prompt by Habit Debugging frameworks frequently store every prompt, completion, and tool result. That is convenient until prompts contain sensitive information. Production memory should distinguish observability from durable agent state. For example, the checkpoint might contain: JSON { "task_type": "incident_analysis", "current_node": "collect_logs", "artifact_ids": [ "log-ref-782" ], "decision_status": "pending" } It may not need to contain the complete raw logs. The actual artifact can remain in its governed source system, where existing retention and access policies apply. Memory should store enough information to resume the workflow, not automatically duplicate every byte the workflow encountered. Retention Is Part of the Data Model A memory record without a retention policy is an indefinite record. That is rarely the right default. Different memory types need different lifetimes: Temporary inference state: minutesFailed-run diagnostics: daysApproval checkpoints: until completion plus audit windowLong-term user preferences: policy dependentAudit decisions: regulated retention schedule Retention metadata should travel with the record: JSON { "thread_id": "thread-55", "memory_class": "workflow_checkpoint", "created_at": "2026-08-10T18:00:00Z", "expires_at": "2026-09-10T18:00:00Z" } A background lifecycle process can enforce expiration. The agent itself should not decide how long regulated records remain available. Memory Needs Versioning Long-running workflows introduce another problem: application code changes while old checkpoints still exist. A state written by version 3 of an agent may be resumed by version 5. Without state versioning, deserialization can fail or, worse, succeed with changed semantics. Persist a schema version: JSON { "state_version": 3, "thread_id": "thread-55", "node": "approval", "payload": {} } The runtime can then migrate older states explicitly: Python state = store.load(thread_id) if state.version < CURRENT_VERSION: state = migrate(state) This is the same discipline used in database schema evolution. Agent memory is application data and should be engineered accordingly. Resumption Must Not Duplicate Side Effects Checkpointing becomes especially important around tool execution. Suppose a workflow: Creates a ticketSaves stateCrashes If the save occurred after ticket creation but failed before recording success, the resumed workflow may create the ticket again. State transitions and side effects need idempotency. A tool call can include a stable operation key: Python tool.execute( operation_id="thread-55-step-8", payload=request ) If the same step is retried: Plain Text operation_id already completed return previous result This makes resumption safe. Durable memory without idempotent side effects can actually make failure recovery more dangerous because the system confidently resumes into duplicate operations. Audit Memory Access Memory reads should be observable. For sensitive workflows, the system should record: ActorThreadOperationTimestampPurposePolicy decision A user should not be able to silently inspect another user’s stored agent state. An administrator may have legitimate access for incident response, but that access should be auditable. Memory governance is not limited to protecting writes. Reads can expose every question, document, and decision associated with an agent run. Separate Long-Term Memory From Workflow Checkpoints Long-term memory deserves an especially strict boundary. A checkpoint answers: Where was this workflow when it stopped? Long-term memory answers: What information should this agent remember in future interactions? Those are very different questions. Do not promote checkpoint data into long-term memory automatically. A user may want an investigation to resume tomorrow without wanting every detail of that investigation retained as a permanent preference or profile. Long-term memory should require explicit policies about what qualifies, who owns it, how it is updated, and when it expires. Memory Is Infrastructure, Not Model Context Reliable enterprise agents need to remember enough to continue useful work. That requirement should not turn every interaction into an indefinitely retained transcript. The architecture should separate transient context, durable run state, checkpoints, and long-term memory. It should isolate threads by principal, redact sensitive values before storage, version persistent schemas, enforce retention policies, audit memory access, and make resumed side effects idempotent. Once an agent can say “I remember where we stopped,” the storage behind that sentence becomes part of the enterprise data architecture. That means it deserves the same rigor as any other stateful production system. Durable memory makes agents more useful. Governed memory makes them safe enough to resume.
Picture the dashboard on a good day. Cycle time is green. Pull request counts are climbing. The throughput chart bends upward, exactly as the vendor deck promised. Now picture the person the dashboard cannot show you: a senior engineer spending her afternoon working out whether a flawless-looking pull request is actually correct. That gap between what the chart says and what the reviewer feels is the story of AI-assisted engineering in 2026. Writing code is no longer the scarce resource. Knowing what the code does, what it cost, and whether anyone checked it is. Most organizations are still measuring the old scarcity. What the Independent Evidence Shows Start with the most rigorous public research on delivery performance. Google Cloud's 2025 DORA report drew on survey responses from nearly 5,000 technology professionals. It found that AI adoption now correlates positively with delivery throughput and product performance, but still correlates negatively with delivery stability. The authors' explanation is the technical heart of this whole topic: without robust control systems such as strong automated testing, mature version control practices, and fast feedback loops, a rise in change volume produces instability. The pipeline is receiving more input than its safeguards were sized for. Developer sentiment moves the same way. Stack Overflow's 2025 developer survey, fielded from May 29 to June 23, 2025, with 49,009 responses according to ADTmag's coverage, found that 84% of developers use or plan to use AI tools, while trust in the accuracy of the output fell to 29% from 40% the year before. Adoption normally builds confidence. Here it did the opposite. Then there is the perception problem. METR ran a randomized trial in 2025 with 16 experienced open source developers across 246 real tasks. Developers forecast a 24% speedup, and the measured result was that tasks took 19% longer, yet afterward they still believed AI had made them 20% faster. I want to be careful, because this result is routinely overstated. It covers one small group using early 2025 tools, METR itself now calls the results historical, and its 2026 follow-up was too compromised by selection effects to give a reliable estimate. The durable lesson is not that AI slows people down. It is that felt productivity and measured productivity can diverge, so self-report is a weak instrument for a measurement problem. Corroboration From the Vendors, With Caveats Two recent surveys come from companies that sell code verification tools, so read them as directional, not definitive. Sonar's State of Code survey, released January 8, 2026, covered over 1,100 developers worldwide. Respondents reported that AI accounts for 42% of their committed code, with an expected 65% by 2027. Also, 96% said they do not fully trust AI output, yet only 48% said they always verify it before committing. The Register's coverage adds that 38% said reviewing AI code takes more effort than reviewing a colleague's, against 27% who said the opposite. Qodo published its 2026 State of AI Code Quality Report on September 23. Censuswide surveyed 500 developers and 300 engineering leaders, all in the United States and all at organizations where AI already does meaningful work, between August 7 and August 14, 2026 (survey details). Developers and leaders, answering separate questionnaires, both named reviewing and validating AI-generated code as their top delivery constraint, at 26% each. The sample is not representative of all engineering organizations, and every figure is self-reported. Still, one result stands out. 90% of leaders said they can report AI's impact to executives or the board, while only 45% said they can trace AI activity to the code changes it produced. The same report found 89% of organizations had experienced an AI-related production incident. Every one of these surveys measures perception, at a specific moment, with tools that keep changing. That limitation shapes what follows. Why the Old Metrics Miss It Three mechanisms explain why familiar dashboards mislead once AI writes a large share of the code. The unit of measurement drifted from the unit of change. Tickets and story points record intent. They were a rough proxy for the amount of code that shipped, and that proxy held well enough when humans typed every line. When one engineer with an assistant can turn a multi-week refactor into an afternoon, the relationship between a ticket and the change it produced stops being stable. Verification cost never appears on a clock. Qodo's report describes why. AI-authored changes tend to look finished, with clean naming and passing tests, but nothing in the diff shows which alternatives were considered or which assumptions carried over. Reviewers spend more effort reaching the same confidence, and 36% of developers in Qodo's sample said review takes the same time but demands greater cognitive effort. Cycle time cannot register effort that does not lengthen the calendar. Feedback loops were sized for the old volume. This is DORA's finding restated. Review queues, test suites, and deployment safeguards were built for a certain rate of change. Raise the rate and the weakest safeguard becomes the constraint. These mechanisms open three distinct measurement gaps. Spend cannot be tied to what got built. Activity rises without a matching rise in shipped value. And real effort, such as careful verification and cleanup of generated code, never gets a ticket. Two Practitioner Perspectives Flux, a Boston company building a code-first engineering intelligence platform, supplied commentary for this piece. Flux sells the kind of analysis its executives recommend, so weigh their views accordingly, but each speaks to a real part of the problem. Ted Julian, Flux's founder and CEO, describes the pressure from the finance side: "Every engineering leader we talk to has been doubling down on AI: more tooling and more code moving through the pipeline. And in nearly every enterprise conversation, the first questions from senior leadership are about spend. Can you show CapEx versus OpEx? Can you help substantiate an R&D credit?" He argues that the organizations that cannot answer usually cannot see quality drift either, because both gaps come from "measuring activity instead of analyzing what the code actually shows." He has described Flux's code-based approach in more detail on The Lantern podcast. Aaron Beals, Flux's CTO, speaks to the engineering side. "AI made these problems move faster than the old measurement systems can keep up with," he said, and the ticket describing the work and the codebase showing the result have drifted far enough apart to create serious blind spots. He ties two problems to one gap: Finance in the quarterly review "asking what the AI spend actually bought," and an engineer paged over a dependency change nobody reviewed. His proposed remedy, "Analyzing the codebase itself closes that gap for both questions without slowing anything down," is a vendor's claim, and I would test it in a pilot before believing it. Whatever you conclude about any product, both executives point to the same requirement. The evidence has to come from what was built, not only from what was recorded about it. A Measurement Framework I find it useful to ask four questions about every AI-assisted change, and to keep each question's metrics in its own view. What shipped? Merged changes and lead time. This is the layer most dashboards already have, and it should never appear alone. Did it hold? Change failure rate, time to restore service, revert rate, and rework within thirty days of merge. DORA already defines the first two, so you can adopt them without inventing anything. What did it cost to trust? Review rounds per change, comments per change, and time from first review to approval. These are proxies, and the honest label for them is "review effort," not "review quality." What produced it, and what did it cost? Whether the change was AI-assisted, and how the spend maps to teams and repositories. This layer is the least mature in most organizations, and Qodo's 90% versus 45% gap suggests it is where confidence most outruns evidence. Flux's briefing uses three labels for the resulting blind spots: unproven spend, velocity theater, and hidden work. They map cleanly onto this framework. Unproven spend is a gap in the fourth question. Velocity theater is the first question answered without the second. Hidden work is the third question left unmeasured. Implementing It In 30 Days Baseline first. Pull the previous two quarters of lead time, change failure rate, time to restore, and revert rate for each repository before you change any process. An ROI claim without a before picture is an estimate.Mark provenance. Agree on a convention, such as a commit trailer or a pull request label, for AI-assisted changes. Expect it to be incomplete and partly self-reported at first, and cross-check it against any usage data your tools expose.Split the dashboards. Put throughput on one view and stability and rework on another. Do not average them into a single velocity score.Compare like with like. Within the same team and repository, compare AI-assisted and unassisted changes on rework and escaped defects. Comparing across teams mostly measures the teams.Map spend to work. Attribute licenses and usage to teams and repositories, and hand Finance that mapping. Questions such as CapEx versus OpEx treatment and R&D credit eligibility belong to your finance and tax advisors, and code-level evidence is an input to their judgment, not a substitute for it. Where This Approach Can Fail Any metric becomes a target once people know it is watched, so pair each measure with its counterweight, and review the set regularly. Provenance tagging depends on honest, consistent use and on what your tools can report. The DORA findings and the surveys are correlational, and cohort comparisons inside one company are not randomized experiments. Finally, all of this evidence describes tooling from 2025 and early 2026, and agentic workflows are changing quickly. Treat the framework as something to recalibrate, not to install once. Disclosure Sonar and Qodo sell verification and code review products and produced the surveys cited above. Flux supplied the executive commentary and sells code-first engineering analytics. The independent sources, DORA, Stack Overflow, and METR, point in the same direction, but each has its own limits, noted above. AI removed the bottleneck at the keyboard and moved it to trust. Trust is built from evidence about what the code does, and a ticket cannot supply that. Sources Google Cloud, Announcing the 2025 DORA Report:https://cloud.google.com/blog/products/ai-machine-learning/announcing-the-2025-dora-reportStack Overflow, 2025 Developer Survey for Leaders:https://stackoverflow.co/internal/resources/2025-stack-overflow-developer-survey-for-leaders/ai-adoption/ ADTmag coverage with fieldwork dates:https://adtmag.com/blogs/watersworks/2026/01/stack-overflow-survey.aspxMETR, Early 2025 developer productivity study:https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/ METR, February 2026 update:https://metr.org/blog/2026-02-24-uplift-update/Sonar, State of Code Developer Survey report (January 8, 2026):https://www.sonarsource.com/blog/state-of-code-developer-survey-report-the-current-reality-of-ai-coding/The Register, Devs doubt AI-written code, but don't always check it:https://www.theregister.com/software/2026/01/09/devs-doubt-ai-written-code-but-dont-always-check-it/4932910Qodo, 2026 State of AI Code Quality Report (September 23, 2026):https://www.qodo.ai/blog/state-of-ai-code-quality-report-2026/ GlobeNewswire release with survey dates:https://www.globenewswire.com/news-release/2026/09/23/3367496/0/en/qodo-s-2026-state-of-ai-code-quality-report-reveals-growing-verification-challenge-as-agentic-development-scales.htmlFlux, About page:https://www.askflux.ai/about/MGMT Boston, Ted Julian on The Lantern:https://mgmtboston.com/lantern/ted-julian-flux
Treat network documentation as versioned infrastructure, not optional project paperwork. Software teams have learned to fear technical debt because a shortcut taken today becomes a constraint tomorrow. Physical network infrastructure develops the same problem, but the consequences are harder to reverse. A stale topology diagram can send an engineer to the wrong room, hide a shared failure point, or turn a controlled change into an outage. This is documentation debt. It accumulates whenever the installed system changes but its record does not. In a long-lived facility, that gap can survive multiple upgrades, contractors, ownership transitions, and equipment generations. The Network Has Two States, and Both Must Match Every production network has a physical state and a documented state. The physical state includes switches, ports, fiber strands, copper pairs, racks, pathways, power sources, and endpoints. The documented state includes the topology, identifiers, relationships, and change history that engineers use to reason about those assets. If the two states diverge, the document becomes a misleading model rather than an operational tool. NIST defines configuration management as maintaining system integrity through controlled initialization, change, and monitoring across the lifecycle. Its security-focused configuration management guidance explicitly includes documentation in configuration control. The practical lesson is simple: a drawing is not done because it was accurate at handover. It is done only while it continues to describe the installed system. A Diagram Is Useful, But a Data Model Is Safer Traditional drawings are excellent for orientation. They show rooms, pathway routes, rack elevations, and logical groupings in a form that humans can scan quickly. They become fragile when they are the only source of truth. A structured inventory makes relationships testable. Each asset can carry a stable identifier, location, role, upstream connection, pathway, media type, owner, and last-verified date. The record can then generate views for different audiences instead of forcing one drawing to answer every question. Plain Text asset_id: SW-TR-042 role: access-switch location: facility: hub-a room: telecom-03 rack: R07 uplinks: - port: te1/1 peer: SW-CORE-002 pathway: FP-03-17 medium: single-mode-fiber power: source_a: PDU-R07-A source_b: PDU-R07-B last_verified: 2026-07-18 The format is less important than the discipline. An identifier must remain stable, required fields must be enforced, and physical labels must match the digital record exactly. Free-text notes can supplement the model, but they should not carry relationships that software could validate. Documentation Should Fail the Build The biggest improvement is to stop treating documentation as a final administrative step. Make it part of the change package and reject incomplete records before work reaches the field. Suppose every planned connection is stored as structured data. A small validation script can catch missing pathway IDs, duplicate asset names, and stale verification dates before a reviewer studies the drawing. Python from datetime import date, timedelta REQUIRED = {"asset_id", "role", "location", "uplinks", "last_verified"} def validate_asset(asset): errors = [] missing = REQUIRED - asset.keys() if missing: errors.append(f"missing fields: {sorted(missing)}") for uplink in asset.get("uplinks", []): if not uplink.get("peer") or not uplink.get("pathway"): errors.append("every uplink needs a peer and pathway") verified = date.fromisoformat(asset["last_verified"]) if verified < date.today() - timedelta(days=365): errors.append("physical verification is older than one year") return errors This does not prove that the cable is installed correctly. It proves that the change contains the minimum information needed to inspect, test, and maintain it. That is the same role a compiler plays for syntax: it eliminates avoidable ambiguity before execution. Constructability Reviews Are Architecture Reviews Many maintainability failures begin during design, not installation. A logical connection may be correct while the proposed pathway is inaccessible, overfilled, exposed to a shared hazard, or impossible to service without interrupting another system. A constructability review traces the real route before construction starts. Reviewers should follow the connection from building entry to distribution frame, patch panel, switch port, field outlet, and endpoint. They should also ask whether technicians can identify, reach, test, and replace each segment safely after the facility is live. The review must include failure relationships. Two uplinks are not redundant if both cross the same room, tray, conduit, or power domain. A clean logical diagram can conceal that physical dependency, which is why logical and pathway records must be reviewed together. Labeling Is an Interface Contract Labels are often dismissed as field details. In practice, they are the interface between the physical network and its source of truth. If an engineer cannot move from a rack label to a record and back again without interpretation, the interface is broken. Good identifiers describe identity, not mutable properties. Avoid names that encode a temporary department, device model, or current port purpose. Use stable IDs, then store changeable attributes in the record. The same rule applies to cable and pathway identifiers. A label should point to one record, and that record should expose both endpoints, the route, media, test result, and change history. NIST's Cybersecurity Framework 2.0 calls for maintained inventories of systems, software, and services, reinforcing that asset visibility must remain current, not merely exist at commissioning. Detect Drift Before the Next Emergency Documentation debt grows quietly because normal operations reward speed. A technician moves a patch, restores service, and plans to update the record later. After enough "later" changes, the database becomes a historical guess. Drift detection makes accuracy measurable. Scheduled checks can compare discovered neighbor data, switch-port descriptions, address assignments, and monitoring inventory with the approved model. Physical pathways still require field verification, but automated comparison can identify where inspection is most valuable. Shell # Export the approved topology and the discovered state. topology export --format json > approved.json discovery snapshot --format json > observed.json # Block silent drift and produce a reviewable report. topology-diff approved.json observed.json \ --require-owner \ --require-pathway \ --fail-on-untracked-asset This should create a review queue, not an automatic rewrite. Discovery can see a neighbor without understanding why the connection exists or how its cable is routed. The human decision remains essential, but software can make discrepancies visible before a high-pressure incident exposes them. The Change Record Must Survive the Project Long-lived infrastructure outlasts the team that installed it. The source of truth must therefore survive personnel changes, contract boundaries, tool migrations, and vendor turnover. A proprietary drawing stored in one person's folder is not a durable operating model. Store records in exportable formats, version them, and back them up separately from the systems they describe. CISA's ransomware guidance recommends maintaining inventories of logical and physical assets, recording interdependencies, and keeping protected offline copies of critical documentation. That asset-management guidance matters during recovery because the network record may be needed when normal management systems are unavailable. Ownership also needs to be explicit. Every field change should identify who approved it, who installed it, who verified the final state, and which record changed. A ticket number alone is not enough if the ticket disappears when a project platform is retired. Documentation Debt Is Operational Risk The cost of poor documentation is rarely the hour spent correcting a drawing. It is the uncertainty added to every future change. Engineers compensate with extra site visits, broader maintenance windows, duplicated tracing work, and cautious assumptions about paths they cannot trust. Treat topology and pathway data like production code. Give it a schema, owners, reviews, tests, version history, and a release condition. Pair digital records with durable physical labels and routine field verification. Infrastructure can remain in service for decades. Its documentation should be engineered for the same lifespan.
Agile
Career Development
Methodologies
Team Management
Kill the Worker, Keep the Research: Build a Recoverable LangGraph Agent on Temporal
October 5, 2026
by Akhil Madineni
CORE
Agentic Test Creation: From Plain-Language Requirements to End-to-End Test Cases
October 2, 2026
by John Vester
CORE
AI on Top of a Dysfunctional System
October 2, 2026
by Stefan Wolpers
CORE