DZone
Thanks for visiting DZone today,
Edit Profile
  • Manage Email Subscriptions
  • How to Post to DZone
  • Article Submission Guidelines
Sign Out View Profile
  • Post an Article
  • Manage My Drafts
Newsletter
Log In / Join
Refcards Trend Reports
Events Video Library
Refcards
Trend Reports

Events

View Events Video Library

DZone Spotlight

Friday, October 2 View All Articles »
A New Chapter for DZone Newsletters

A New Chapter for DZone Newsletters

By Dominique Roller
Hello DZone community! We’re refreshing our newsletters to help you follow the topics that matter most to your work, explore new ideas, and stay connected with the developer community. Our Zone newsletters are becoming six focused newsletters, with new names and related topics that were developed with input from some of our fantastic community members, and we’re excited to bring them to your inbox: Beyond the Rows: Big data and databasesMind the Model: AI for developers and engineersShip & Scale: Cloud, DevOps, performance, and AgileThe Attack Surface: SecurityDistributed by Design: Microservices, integration, and IoTCode & Craft: Java and web development Each topical newsletter will arrive twice a month, with editions scheduled on Tuesdays and Thursdays. Meet DZone Digest We’re also bringing DZone Daily and DZone Weekly together into DZone Digest, arriving every Wednesday. It will be your weekly roundup of articles and insights from across DZone. What Else Is Changing? All newsletters got a refreshed design, making room for the content you care about and more opportunities to discover upcoming events. We’re also opening newsletter subscriptions to everyone, including developers who aren’t DZone members. Our goal is to make DZone’s newsletters more useful, with clearer topic choices and a regular cadence that helps you keep learning. We’d love to hear from you: Which newsletter are you most interested in, and what topics would you like us to cover? Share your thoughts in the comments. You can subscribe to them here. (If you’re subscribed to any of our Zone newsletters, you’ll begin receiving the updated newsletter covering your topics. If you’re subscribed to DZone Daily or DZone Weekly, you’ll now receive DZone Digest.) More
Part 3: End-to-End Tracing and Observability Across Goose, agentgateway, and Quarkus

Part 3: End-to-End Tracing and Observability Across Goose, agentgateway, and Quarkus

By Daniel Oh DZone Core CORE
Enterprise context — Acme FinServ. SOC 2 CC7 (system monitoring) requires that Acme can detect and investigate anomalous activity. When an agent-driven workflow touches customer data at 2 AM, "we have logs somewhere" is not an answer an auditor accepts. The distributed trace built in this part is the forensic evidence trail: a single trace ID that ties the Goose prompt to every agentgateway policy decision and every Quarkus tool call, so a post-incident review can reconstruct exactly which agent did what, in what order, and how long each governed hop took. The Core Problem In Part 1, we built a Quarkus MCP tool server. In Part 2, we secured it with agentgateway's JWT authentication, RBAC, and ExtMCP guardrails. The architecture works — but when something goes wrong in production, you're flying blind. Agentic workflows are fundamentally different from traditional request-response APIs. A single user prompt like "Debug customer CUST-4091" triggers a multi-round-trip loop: Goose calls tools/list to discover available toolsThe LLM selects getCustomerStatus and Goose sends tools/callThe LLM reads the response, sees primaryRegion: US-EAST-1, and chains a second tools/call to getZoneHealthLogsThe LLM correlates both results and generates a diagnostic summary Each of these hops crosses process boundaries: Goose → agentgateway → Quarkus. Without distributed tracing, you see four isolated HTTP requests in your access logs. You cannot tell they belong to the same agentic workflow. When step 3 takes 12 seconds instead of 200ms, you have no waterfall to pinpoint whether the latency came from agentgateway policy evaluation, Quarkus bean validation, or a slow downstream call. This creates black holes in telemetry dashboards — the exact gap that autonomous agents exploit to degrade silently. The Solution: W3C Trace Context Across All Three Layers The fix is standard distributed tracing, applied to the MCP transport layer: agentgateway exports spans for every proxied MCP request and propagates traceparent headers to the backend.Quarkus with quarkus-opentelemetry picks up the incoming traceparent, creates child spans for tool execution and bean validation, and exports them to the same Jaeger instance.Jaeger correlates both sides into a single trace waterfall — one view from agent prompt to tool result. Prerequisites Everything from Parts 1 and 2, plus: Podman – for running Jaeger (podman compose) Verify Podman is available: Shell podman --version Step 1: Launching the Observability Backend We use Jaeger v2 as both the OTLP collector and the trace UI. A single container accepts traces from agentgateway on port 4317 (OTLP gRPC) and from Quarkus on port 4318 (OTLP HTTP), and serves the query UI on port 16686. Shell cd part3-observability podman compose up -d This starts Jaeger v2 with OTLP collection enabled by default. Verify it's running: Shell curl -sf http://localhost:16686/ > /dev/null && echo "Jaeger UI is ready" Open http://localhost:16686 — you'll see an empty Jaeger UI. We'll populate it with MCP traces in the following steps. Production Alternative: Grafana Tempo For production deployments, replace Jaeger with Grafana Tempo backed by object storage (S3/GCS). The OTLP endpoint stays the same — only the compose.yml changes. Grafana provides richer dashboards, alerting, and long-term trace retention. Step 2: Enabling OpenTelemetry in Quarkus Add the quarkus-opentelemetry extension to Part 1's pom.xml: Properties files <dependency> <groupId>io.quarkus</groupId> <artifactId>quarkus-opentelemetry</artifactId> </dependency> Configure the exporter in application.properties: Properties files # OpenTelemetry quarkus.otel.service.name=customer-tools quarkus.otel.exporter.otlp.traces.endpoint=http://localhost:4318 quarkus.otel.exporter.otlp.traces.protocol=http/protobuf quarkus.otel.traces.sampler=always_on quarkus.otel.traces.suppress-non-application-uris=false PropertyPurposeservice.nameIdentifies this service in Jaeger's service dropdowntraces.endpointOTLP HTTP receiver — Jaeger's port 4318 (base URL only; Quarkus appends /v1/traces)traces.protocolhttp/protobuf — Quarkus uses its Vert.x-based HTTP exportertraces.sampleralways_on — sample every span (reduce in production)suppress-non-application-urisfalse — include MCP endpoint spans (they'd be filtered otherwise) When no OTLP collector is running (Parts 1 and 2 without Jaeger), Quarkus logs a connection warning, but the MCP server works normally. When the collector IS running (Part 3), traces flow automatically. Zero code changes to the MCP tools. Rebuild Part 1: Shell cd part1-quarkus-mcp mvn package -DskipTests What Quarkus Auto-Instruments With quarkus-opentelemetry on the classpath and the SDK enabled, Quarkus automatically creates spans for: LayerSpan NameWhat It CapturesHTTP serverPOST /mcpInbound MCP request with method, status, latencyCDI beansCustomerServiceTools.getCustomerStatusTool execution time within the MCP handlerBean ValidationHibernateValidatorParameter validation before tool logic runsREST clientOutbound HTTP callsAny downstream API calls (future extensions) No @WithSpan annotations needed. The Quarkus OpenTelemetry extension instruments the reactive pipeline automatically. Step 3: Configuring W3C Trace Context in agentgateway agentgateway supports native OpenTelemetry trace export. Add a tracing block to the gateway configuration: Properties files config: adminAddr: localhost:15000 tracing: otlpEndpoint: http://localhost:4317 otlpProtocol: grpc randomSampling: 1.0 FieldPurposeotlpEndpointOTLP receiver — Jaeger's port 4317otlpProtocolgrpc for OTLP/gRPC (also supports http)randomSamplingSample 100% of traces (reduce to 0.01–0.1 in production) How Trace Propagation Works When agentgateway receives an MCP request: Creates a root span for the proxy operation (e.g., agentgateway.mcp.proxy)Injects a traceparent header into the forwarded request to Quarkus:traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01Quarkus reads the traceparent, creates a child span under the same trace ID, and records tool executionBoth spans export to Jaeger via OTLP, where they appear as a single correlated trace This is standard W3C Trace Context propagation — the same mechanism used across all OpenTelemetry-instrumented services. Configuration Files Part 3 provides two agentgateway configurations: ConfigUse Caseconfig-traced.yamlTracing only — proxy + OTLP export, no security layersconfig-traced-guardrails.yamlTracing + ExtMCP guardrails — observe the guardrail evaluation spans too Step 4: Running the Interactive Demo Start all services with the one-command script: Shell cd part3-observability ./start-all.sh The script starts Jaeger, Quarkus (with OTel enabled), and agentgateway (with trace export), then launches the demo SPA on :8890. Open the MCP Observability Console at http://localhost:8890/index.html and walk through the three demo steps: Initialize – Establishes an MCP session through agentgateway. The architecture diagram animates the trace propagation: root span creation in agentgateway, traceparent injection, child span in Quarkus, and OTLP export to Jaeger.List Tools – Discovers all 5 tools through the traced proxy. The trace waterfall panel shows the agentgateway proxy span and the Quarkus HTTP span side by side with timing.Multi-Tool Workflow – Simulates Goose's multi-turn reasoning: getCustomerStatus (finds region US-EAST-1) → getZoneHealthLogs (checks zone health) → getSLACompliance (correlates SLA metrics). Each step generates a full trace with waterfall visualization. The stat tiles track traces generated, spans collected, and Jaeger status. Click Open Jaeger to view the real trace waterfalls in the Jaeger UI at http://localhost:16686. Step 5: Generating Traces via CLI To generate additional traces manually, simulate a multi-turn agentic workflow: Shell # Step 1: Initialize MCP session export MCP_SESSION_ID=$(curl -s -D - http://localhost:3000/mcp \ -H "Content-Type: application/json" \ -H "Accept: application/json, text/event-stream" \ -H "MCP-Protocol-Version: 2025-03-26" \ -d '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-03-26","capabilities":{},"clientInfo":{"name":"curl","version":"1.0"}}' \ | grep -i "mcp-session-id:" | sed 's/.*: //' | tr -d '\r') # Step 2: Discover tools curl -s http://localhost:3000/mcp \ -H "Content-Type: application/json" \ -H "Accept: application/json, text/event-stream" \ -H "MCP-Protocol-Version: 2025-03-26" \ -H "mcp-session-id: $MCP_SESSION_ID" \ -d '{"jsonrpc":"2.0","id":2,"method":"tools/list","params":{}' \ | grep '^data: ' | sed 's/^data: //' | jq . # Step 3: Agent calls getCustomerStatus (first tool invocation) curl -s http://localhost:3000/mcp \ -H "Content-Type: application/json" \ -H "Accept: application/json, text/event-stream" \ -H "MCP-Protocol-Version: 2025-03-26" \ -H "mcp-session-id: $MCP_SESSION_ID" \ -d '{"jsonrpc":"2.0","id":3,"method":"tools/call","params":{"name":"getCustomerStatus","arguments":{"customerId":"CUST-4091"}}' \ | grep '^data: ' | sed 's/^data: //' | jq . # Step 4: Agent chains getZoneHealthLogs based on the region from step 3 curl -s http://localhost:3000/mcp \ -H "Content-Type: application/json" \ -H "Accept: application/json, text/event-stream" \ -H "MCP-Protocol-Version: 2025-03-26" \ -H "mcp-session-id: $MCP_SESSION_ID" \ -d '{"jsonrpc":"2.0","id":4,"method":"tools/call","params":{"name":"getZoneHealthLogs","arguments":{"zoneId":"US-EAST-1"}}' \ | grep '^data: ' | sed 's/^data: //' | jq . # Step 5: Agent fetches SLA compliance for correlation curl -s http://localhost:3000/mcp \ -H "Content-Type: application/json" \ -H "Accept: application/json, text/event-stream" \ -H "MCP-Protocol-Version: 2025-03-26" \ -H "mcp-session-id: $MCP_SESSION_ID" \ -d '{"jsonrpc":"2.0","id":5,"method":"tools/call","params":{"name":"getSLACompliance","arguments":{"serviceId":"api-gateway"}}' \ | grep '^data: ' | sed 's/^data: //' | jq . Each of these requests generates a trace that flows through agentgateway into Quarkus and lands in Jaeger. Step 6: Visualizing the Trace Waterfall in Jaeger Open http://localhost:16686 in your browser. Finding Traces In the Service dropdown, select customer-tools (Quarkus) or agentgatewayClick Find TracesClick on any trace to open the waterfall view Reading the Waterfall A typical tools/call trace shows the following span hierarchy: Shell agentgateway.mcp.proxy [12ms] └─ POST /mcp [8ms] ← Quarkus HTTP server └─ CustomerServiceTools.getCustomerStatus [2ms] ← CDI tool execution SpanServiceWhat It Tells Youagentgateway.mcp.proxyagentgatewayTotal proxy overhead including policy evaluationPOST /mcpcustomer-toolsQuarkus HTTP handling time for the MCP requestgetCustomerStatuscustomer-toolsPure tool execution time (business logic) What to Look For Proxy overhead: The gap between the agentgateway span and the Quarkus span shows network + policy evaluation time. If this grows, check guardrail server latency.Validation time: Bean Validation spans appear before tool execution. Regex-heavy patterns like ^CUST-[0-9]{4,8}$ are fast, but complex validators on large payloads can add latency.Multi-turn correlation: When Goose chains multiple tool calls (e.g., getCustomerStatus → getZoneHealthLogs), each appears as a separate trace. The mcp-session-id tag lets you filter all traces belonging to one agent session.Error traces: Failed validations (invalid customer ID format) or guardrail rejections (blocked poison payloads) produce error spans with exception details. Connecting Goose for Real Traces Launch Goose pointed at agentgateway and prompt a multi-tool workflow: Shell goose session "Debug customer CUST-4091 — check their account status, then pull health logs for their region and SLA compliance for api-gateway." This generates a burst of correlated traces in Jaeger showing Goose's multi-turn tool orchestration from the proxy layer down to individual tool execution spans. What We Achieved Starting from the secured architecture in Part 2, we added full observability without changing any MCP tool code: LayerWhat We AddedConfig ChangeQuarkusquarkus-opentelemetry dependencypom.xml + application.propertiesagentgatewaytracing block in config YAMLconfig-traced.yamlObservability backendJaeger all-in-one via Podman Composecompose.yml The entire stack runs locally with a single ./start-all.sh command and produces end-to-end trace waterfalls in Jaeger. Production Considerations ConcernLocal (this tutorial)ProductionTrace backendJaeger all-in-one (in-memory)Grafana Tempo + object storageSampling rate100% (default: 1.0)1-10% or adaptive samplingTrace retentionContainer lifetimeDays/weeks in durable storageAlertingManual Jaeger inspectionGrafana alerting on span latency/error rateMetricsTraces onlyAdd Prometheus + quarkus-micrometer for RED metrics Coming Up in Part 4 With tracing in place, you can now see every MCP tool call flowing through the system. In Part 4, we will move beyond single-agent tool calls to multi-agent orchestration — using the Agent-to-Agent (A2A) protocol to coordinate autonomous agents that can delegate work, enforce governance via AGENTS.md, and call back into our MCP tool services. More
Docker Sandboxes Beyond the Laptop: Running AI Agents in the Cloud
Docker Sandboxes Beyond the Laptop: Running AI Agents in the Cloud
By Naga Santhosh Reddy Vootukuri DZone Core CORE
How a 30B Model Runs on Your Laptop
How a 30B Model Runs on Your Laptop
By Akash Lomas

Refcard #291

Code Review Core Practices

By Vidyasagar (Sarath Chandra) Machupalli FBCS DZone Core CORE
Code Review Core Practices

Refcard #267

Getting Started With DevSecOps

By Akanksha Pathak DZone Core CORE
Getting Started With DevSecOps

More Articles

Embabel vs LangGraph4j: Two Agentic Philosophies for Investment and Risk Analysis in BFSI
Embabel vs LangGraph4j: Two Agentic Philosophies for Investment and Risk Analysis in BFSI

Quick Summary Both Embabel and LangGraph4j let a Java developer build multi-step AI agents without leaving the JVM.Embabel hands the framework a goal and a bag of typed actions, and lets a planner decide the order on its own.LangGraph4j asks the developer to draw the exact graph of nodes and edges by hand.We will understand both philosophies through a real Embabel agent, a small Kolkata street-crossing example, a comparison table, and finally a bigger question — is Java catching up with Python in enterprise AI work? Where the Story Starts If you are a Java developer today, you are watching two worlds collide. On one side are large language models, which grew up almost entirely in Python. On the other side is enterprise Java, which has spent twenty-five years learning to build systems that banks, insurance companies, and hospitals can actually trust. Two frameworks are now trying to bring these two worlds together on the JVM: Embabel and LangGraph4j. Both help you build an "agent" — a piece of software that uses an LLM to complete a task in several steps, rather than in one single prompt. But the way they think about "steps" is completely different. That difference is what this article is about. Meeting Embabel Through a Real Piece of Code The best way to understand Embabel is to look at actual code, not a slide. Here is a small agent that gives retirement planning advice, written the Embabel way. Java package com.example.demo; import com.embabel.agent.api.annotation.AchievesGoal; import com.embabel.agent.api.annotation.Action; import com.embabel.agent.api.annotation.Agent; import com.embabel.agent.api.common.OperationContext; import com.embabel.agent.domain.io.UserInput; import java.util.Arrays; import java.util.List; @Agent(name = "RetirementPlannerAgent", description = "This agent provides retirement planning advice.") public class RetirementPlannerAgent { record RetirementUserInput(int presentAge, int targetRetirementAge, double annualIncome) { } record RetirementPlanAdvice(String advice) { } record RetirementPlanAdvices(List<RetirementPlanAdvice> advices) { } enum RiskToleranceLevel { LOW, MEDIUM, HIGH } @Action(description = "Identify the present age, target retirement age, and annual income of the user from the user message.") public RetirementUserInput identifyRetirementUserInput(UserInput userInput, OperationContext context) { String content = userInput.getContent(); return context.ai().withDefaultLlm() .creating(RetirementUserInput.class) .fromPrompt(""" Identify the present age, target retirement age, and annual income of the user from the following message and return them as a JSON object. User message: %s """.formatted(content)); } @Action(description = "Identify the risk tolerance level of the user based on the present age, target retirement age, and annual income.") public RiskToleranceLevel identifyRiskToleranceLevel(RetirementUserInput retirementUserInput, OperationContext context) { return context.ai().withDefaultLlm() .creating(RiskToleranceLevel.class) .fromPrompt(""" Identify the risk tolerance level of the user based on the following information and return it as a JSON object. Permitted values for risk tolerance level are: %s Present age: %d Target retirement age: %d Maximum years to retirement: %d Annual income: %.2f """.formatted(Arrays.toString(RiskToleranceLevel.values()), retirementUserInput.presentAge(), retirementUserInput.targetRetirementAge(), (retirementUserInput.targetRetirementAge() - retirementUserInput.presentAge()), retirementUserInput.annualIncome())); } @Action(description = "Provide retirement plan advice based on the user's risk tolerance level.") @AchievesGoal(description = "Provide retirement plan advice based on the user's risk tolerance level.") public RetirementPlanAdvices provideRetirementPlanAdvice(RiskToleranceLevel riskToleranceLevel, OperationContext context) { String systemPrompt = """ You are a retirement planning advisor. Based on the user's risk tolerance level, provide a list of retirement plan advices. """; return context.ai().withDefaultLlm() .creating(RetirementPlanAdvices.class) .fromPrompt(""" %s User's risk tolerance level: %s """.formatted(systemPrompt, riskToleranceLevel.name())); } } Now look closely at what is missing from this code. There is no method called runAgent() that calls identifyRetirementUserInput(), then identifyRiskToleranceLevel(), then provideRetirementPlanAdvice(), in that order. Nowhere did the developer type out the sequence. Instead, each @Action simply states two things: What type it needs as input (its precondition).What type it produces as output (its effect). provideRetirementPlanAdvice needs a RiskToleranceLevel. identifyRiskToleranceLevel happens to produce a RiskToleranceLevel from a RetirementUserInput. And identifyRetirementUserInput produces that RetirementUserInput from the raw UserInput. Embabel's planner looks at all this at runtime and works out, on its own, that this is the only order in which the goal (@AchievesGoal) can be reached. This is Embabel's whole philosophy in one sentence: give the framework a goal and a set of typed building blocks, and let it plan. The Same Investment Advisory Workflow, Wired by Hand in LangGraph4j Now let us build the exact same three-step advisory flow — read the user's numbers, work out risk tolerance, give advice — the LangGraph4j way. Here, the developer draws the graph explicitly, and the shared state is a plain key-value map (an AgentState) rather than Soham's strongly typed records. Java package com.example.demo; import org.bsc.langgraph4j.CompiledGraph; import org.bsc.langgraph4j.StateGraph; import org.bsc.langgraph4j.action.NodeAction; import org.bsc.langgraph4j.state.AgentState; import java.util.List; import java.util.Map; import java.util.Optional; import static org.bsc.langgraph4j.StateGraph.END; import static org.bsc.langgraph4j.StateGraph.START; import static org.bsc.langgraph4j.action.AsyncNodeAction.node_async; /** * @author Soham Sengupta * @since 2026-09-13 * @description The retirement/investment advisory workflow from the * Embabel example above, this time wired explicitly as a LangGraph4j * graph. Every step, and the order between the steps, is declared * here by the developer - there is no planner discovering it. */ public class RetirementPlannerGraph { // Step 1: pull the present age, target retirement age, and income out of free text. static class IdentifyUserInputNode implements NodeAction<AgentState> { @Override public Map<String, Object> apply(AgentState state) throws Exception { String userMessage = state.<String>value("userMessage").orElseThrow(); // In a real system: call the LLM here (say, via langchain4j) and parse // presentAge / targetRetirementAge / annualIncome out of userMessage. return Map.of( "presentAge", 32, "targetRetirementAge", 60, "annualIncome", 1200000.0); } } // Step 2: classify how much investment risk this user can reasonably take. static class IdentifyRiskToleranceNode implements NodeAction<AgentState> { @Override public Map<String, Object> apply(AgentState state) throws Exception { int presentAge = state.<Integer>value("presentAge").orElseThrow(); int targetRetirementAge = state.<Integer>value("targetRetirementAge").orElseThrow(); // In a real system: call the LLM here with these values and ask it to // return LOW, MEDIUM, or HIGH as the risk tolerance level. String riskTolerance = (targetRetirementAge - presentAge) > 20 ? "HIGH" : "MEDIUM"; return Map.of("riskTolerance", riskTolerance); } } // Step 3: turn the risk tolerance into a concrete list of investment advice. static class ProvideAdviceNode implements NodeAction<AgentState> { @Override public Map<String, Object> apply(AgentState state) throws Exception { String riskTolerance = state.<String>value("riskTolerance").orElseThrow(); // In a real system: call the LLM here to draft actual advice - suitable // Indian investment instruments for this riskTolerance level, and so on. List<String> advice = List.of("Suggested investment mix for a " + riskTolerance + " risk profile."); return Map.of("advice", advice); } } public static void main(String[] args) throws Exception { StateGraph<AgentState> graph = new StateGraph<>(AgentState::new) .addNode("identifyUserInput", node_async(new IdentifyUserInputNode())) .addNode("identifyRiskTolerance", node_async(new IdentifyRiskToleranceNode())) .addNode("provideAdvice", node_async(new ProvideAdviceNode())) .addEdge(START, "identifyUserInput") .addEdge("identifyUserInput", "identifyRiskTolerance") .addEdge("identifyRiskTolerance", "provideAdvice") .addEdge("provideAdvice", END); CompiledGraph<AgentState> workflow = graph.compile(); Optional<AgentState> result = workflow.invoke( Map.of("userMessage", "I am 32, want to retire at 60, and earn 12 lakh a year.")); result.ifPresent(state -> System.out.println(state.data())); } } Two things stand out next to the Embabel version. First, there is a main method here that explicitly lists identifyUserInput -> identifyRiskTolerance -> provideAdvice as edges — in the Embabel version, that sequence was never written down anywhere; it was worked out by the planner. Second, the shared state (AgentState) is just a bag of string keys and values, read back out with state.value("presentAge"), instead of Soham's own strongly typed RetirementUserInput and RiskToleranceLevel. For this particular workflow, which happens to be a strict straight line with no branching, LangGraph4j's graph is refreshingly easy to read top to bottom. The real difference shows up once branching enters the picture, which is exactly where our next example — crossing a Kolkata road — comes in. A Bit of History: From Servlet to Spring, Now From Spring AI to Embabel To understand why Embabel is built this way, it helps to know who built it. Embabel comes from Rod Johnson — the same person who created the Spring Framework more than two decades ago. Back in the early 2000s, enterprise Java was drowning in heavy J2EE application servers and Enterprise Java Beans. Rod was solving real problems in the finance industry at the time, found the existing tools too heavy, and wrote a book and a framework that simplified things a great deal. That framework became Spring, and it changed how an entire generation of Java developers worked. In 2025, Rod did something similar again, this time for AI agents. When he introduced Embabel to the Java community, he framed it using a comparison that Java developers will find very familiar: Spring AI is to Embabel roughly what the plain old Servlet API once was to Spring MVC. Spring AI gives you the low-level plumbing — talking to a model, building a prompt, calling a tool. Embabel sits one level above that, giving you the actual application framework — goals, actions, planning, and a proper domain model — the same kind of jump in abstraction that Spring itself brought to raw Servlets and EJBs, twenty years back. Embabel (pronounced "Em-BAY-bel") is written mainly in Kotlin, but as you can see from the retirement planner code above, it feels completely natural to use from plain Java. It is also built to sit closely with Spring, which is exactly why an existing Spring shop can pick it up without much friction. GOAP: The Planning Engine Hiding Inside Embabel The planning idea inside Embabel is not new — it is borrowed from video games, and it is called GOAP, short for Goal-Oriented Action Planning. GOAP was built to make game characters (think of soldiers in an old shooter game) decide, on their own, a believable sequence of actions to reach a goal, instead of following a fixed script. GOAP needs three things: A state of the world as it stands right now.A set of actions, each with a precondition (what must be true to run it) and an effect (what becomes true after it runs).A goal, which is simply a desired state. A search algorithm (usually the well-known A* algorithm) then works out the cheapest chain of actions that gets you from where you are to where you want to be. Embabel uses exactly this idea, but instead of asking the LLM to "think step by step" about the plan (which is often unreliable) or asking the developer to hard-code the entire flow (which is rigid), it asks a deterministic planner to search over your own typed Java or Kotlin methods. The precondition of an action is simply the input type it needs. The effect is simply the output type it produces. Your domain classes — RetirementUserInput, RiskToleranceLevel, RetirementPlanAdvices — literally become the "world state" the planner reasons about. No LLM guesswork is involved in deciding the order; the LLM is only used inside each action, for the part it is actually good at — understanding and generating language. Understanding the GOAP + OOAD Loop, With a Kolkata Traffic Signal All this can feel a bit abstract, so let us make it concrete with something every Kolkata resident understands very well — crossing a busy road. Picture Soham standing at a signal with his four-year-old son, Kit, holding his hand. It is a typical Kolkata crossing — buses, yellow taxis, and autos, and the signal, while present, is not always fully obeyed. Sometimes there is a traffic constable standing in the middle of the road, waving vehicles through by hand, overriding the signal completely. Step one — model the world as an object (this is the OOAD part). In Object-Oriented Analysis and Design, we are trained to represent a real-world situation as a class with clearly named fields. Here, the "world" Soham is observing can be written as one simple record: record RoadSituation(boolean signalGreen, boolean trafficClear, boolean handHeld, boolean policeOnDuty) { } Step two — define the goal. The goal is not "the signal is green." The goal is "Soham and Kit have reached the other side, safely." Reaching a green signal is only useful if it actually leads there. Step three — define the actions, each with its precondition and effect (this is the GOAP part). Notice that none of these actions know about each other. Each one only knows what it needs and what it produces: holdChildHand – produces handHeld = true. This is usually the very first thing a responsible parent does, well before even looking at the signal.waitForSafeSignal – needs the current RoadSituation, and produces signalGreen = true once either the light turns green or a traffic constable is present and waving pedestrians across.confirmTrafficClear – because a green signal in Kolkata does not always mean an auto will not sneak through, this action looks both ways and produces trafficClear = true.crossTheRoad (the @AchievesGoal action) – only runs once handHeld, trafficClear, and (signalGreen or policeOnDuty) are all true. Step four — let the planner loop. This is the actual "GOAP + OOAD loop": the planner looks at the current RoadSituation object, picks whichever action's precondition is already satisfied and whose effect moves the world closer to the goal, executes it, updates the RoadSituation, and checks again if the goal is reached. It keeps looping — plan, act, update state, re-check — until crossTheRoad finally fires. The beautiful part is what happens on a day when the signal itself is not working — a fairly common event in Kolkata, especially during a power cut or during Puja season when the police fully take charge of a crossing. If tomorrow you add one more action, say waitForPoliceWave, which also produces a "safe to proceed" fact, the planner will simply discover this new path on its own the next time it runs. Nobody needs to redraw anything, because nobody drew anything explicit in the first place. The Same Scenario as Embabel Code Here is a simplified sketch of the above scenario, written in the same style as the retirement planner. Treat it as a teaching example rather than a compiled, production-ready class: Java /** * @author Soham Sengupta * @since 2026-09-13 * @description A small agent that plans how Soham and his four-year-old * son Kit can safely cross a busy Kolkata road. Written purely to show * how Embabel's GOAP-style planner reasons over typed domain objects * (OOAD) to reach a goal, without the developer wiring the order by hand. */ @Agent(name = "StreetCrossingAgent", description = "Plans a safe road crossing for a parent and a young child.") public class StreetCrossingAgent { record RoadSituation(boolean signalGreen, boolean trafficClear, boolean handHeld, boolean policeOnDuty) { } record CrossingPlan(String narrative) { } @Action(description = "Hold Kit's hand before anything else - the first rule of the road.") public RoadSituation holdChildHand(UserInput userInput) { return new RoadSituation(false, false, true, false); } @Action(description = "Wait till the signal turns green, or till a traffic constable waves pedestrians across.") public RoadSituation waitForSafeSignal(RoadSituation situation, OperationContext context) { boolean safeToMove = situation.signalGreen() || situation.policeOnDuty(); return new RoadSituation(safeToMove, situation.trafficClear(), situation.handHeld(), situation.policeOnDuty()); } @Action(description = "Look right, then left, then right again - the signal alone is not a guarantee in Kolkata traffic.") public RoadSituation confirmTrafficClear(RoadSituation situation) { return new RoadSituation(situation.signalGreen(), true, situation.handHeld(), situation.policeOnDuty()); } @Action(description = "Cross only when hand is held, the way is clear, and the signal or constable allows it.") @AchievesGoal(description = "Soham and Kit have reached the other side of the road safely.") public CrossingPlan crossTheRoad(RoadSituation situation) { return new CrossingPlan( "Soham held Kit's hand tight, waited for the green man, checked both sides once more, then crossed together."); } } Now compare this with how the same scenario would look in LangGraph4j's philosophy — as an explicit graph you draw yourself, branches and all: Java package com.example.demo; import org.bsc.langgraph4j.CompiledGraph; import org.bsc.langgraph4j.StateGraph; import org.bsc.langgraph4j.action.AsyncEdgeAction; import org.bsc.langgraph4j.state.AgentState; import java.util.Map; import java.util.Optional; import static org.bsc.langgraph4j.StateGraph.END; import static org.bsc.langgraph4j.StateGraph.START; import static org.bsc.langgraph4j.action.AsyncNodeAction.node_async; public class StreetCrossingGraph { // Stand-ins for a real signal sensor and a quick look both ways. private static boolean checkSignal() { return true; } private static boolean lookBothWays() { return true; } public static void main(String[] args) throws Exception { StateGraph<AgentState> graph = new StateGraph<>(AgentState::new) .addNode("holdHand", node_async(state -> Map.of("handHeld", true))) .addNode("waitForSignal", node_async(state -> { boolean policeOnDuty = state.<Boolean>value("policeOnDuty").orElse(false); return Map.of("signalGreen", checkSignal() || policeOnDuty); })) .addNode("checkTraffic", node_async(state -> Map.of("trafficClear", lookBothWays()))) .addNode("cross", node_async(state -> Map.of("crossed", true))) .addEdge(START, "holdHand") .addEdge("holdHand", "waitForSignal") .addConditionalEdges("waitForSignal", AsyncEdgeAction.edge_async(state -> state.<Boolean>value("signalGreen").orElse(false) ? "proceed" : "retry"), Map.of("proceed", "checkTraffic", "retry", "waitForSignal")) .addConditionalEdges("checkTraffic", AsyncEdgeAction.edge_async(state -> state.<Boolean>value("trafficClear").orElse(false) ? "proceed" : "retry"), Map.of("proceed", "cross", "retry", "checkTraffic")) .addEdge("cross", END); CompiledGraph<AgentState> workflow = graph.compile(); Optional<AgentState> result = workflow.invoke(Map.of("policeOnDuty", false)); result.ifPresent(state -> System.out.println(state.data())); } } Notice the two addConditionalEdges calls — this is how LangGraph4j handles a branch: an edge action returns a label ("proceed" or "retry"), and a small map resolves that label to the actual next node. This is exactly the graph the developer must draw by hand, action by action and branch by branch, for a scenario that Embabel's planner worked out on its own. As a flowchart, that graph looks like this: Both pieces of code reach the same goal. But in Embabel, nobody drew this flowchart — the planner found it. In LangGraph4j, this flowchart is the code. If a new real-world case turns up tomorrow, Embabel's planner can absorb it automatically as long as the new action's types fit; LangGraph4j needs a human to open the graph and add a new node or edge. Comparing the Two Philosophies Aspect Embabel LangGraph4j Core idea Give a goal and typed actions; a GOAP/A* planner works out the order Developer explicitly wires nodes and edges into a graph Mental model "What do I want, and what building blocks do I have?" "What are my steps, and how do they branch?" Control flow Discovered at runtime by the planner Declared upfront by the developer Role of the LLM Used only inside actions, never for deciding sequence Can be used inside nodes; sequence is still fixed by the graph Adapting to a new case Often automatic, if a new action's types fit the gap Needs a human to add a new node or edge Tracing "why this order" Needs the planner's own logging/tooling to see the chosen path Very direct — the graph is already the flowchart Roots Kotlin-first, Java-friendly, close to Spring A faithful Java port of Python's LangGraph, works with Langchain4j and Spring AI Maturity (as of late 2026) Young, pre-1.0, moving fast Older and more widely adopted, with a large existing community When to Reach for Which Reach for Embabel when: Your goal is clear, but the exact path to it can honestly vary depending on the situation.You want a deterministic, non-LLM planner deciding the order, not the LLM guessing it.You are already deep in the Spring ecosystem and like strongly typed domain models.New cases keep appearing over time, and you would rather add one new action than redraw a graph. Reach for LangGraph4j when: You already know the exact stages of your workflow — say, a well-understood pipeline of retrieve, rerank, generate, and validate.You want the flow to be visible as an actual graph, easy to explain to a non-technical stakeholder.Your team is porting an existing Python LangGraph pipeline and wants the Java version to mirror it closely.You value a larger, more mature community with more examples to learn from, at least for now. Neither approach is "better" in an absolute sense. Embabel bets on planning; LangGraph4j bets on explicitness. Pick the one that matches how well you actually know your workflow in advance. To Conclude: Java Still Has a Say in Enterprise AI Python remains, without question, the home of AI research — the notebooks, the training loops, the enormous ecosystem of machine learning libraries were built there first, and will likely stay there. Nobody sensible is arguing Java should train the next large language model. But training a model is only one part of the story. The other part — the much bigger part, in terms of sheer lines of code running in the real world — is taking an already-trained model and safely wiring it into systems that already exist: a bank's core banking platform, an insurance company's policy engine, a hospital's records system. The overwhelming majority of that existing code, in most large enterprises, is written in Java and Spring, not Python. That is Java's home ground, built up over more than two decades. This is exactly the ground both Embabel and LangGraph4j are fighting on. Java also tends to run this kind of orchestration work faster than Python at execution time, which matters once you are calling these agents thousands of times a day inside a live enterprise system. And when your applications are already written in Java, keeping the AI layer in Java too — rather than routing every call out to a separate Python service — often turns out to be the simpler, safer choice. So, the real contest in enterprise AI is perhaps not "who trains the smarter model" — Python wins that one comfortably. It is "who can be trusted to make that model's decisions reliably inside a bank's core system, an insurance engine, or a hospital record system." That is precisely the kind of trust Java has spent two decades earning. With Rod Johnson effectively writing a sequel to his own Spring story, and with LangGraph4j bringing a proven Python pattern faithfully onto the JVM, Java is not sitting out this wave of AI. It has simply chosen to fight the battle it already knows how to win. This piece focused on Embabel's goal-and-planner philosophy. Embabel also has other ideas worth a separate deep-dive later — like its approach to agentic search and enterprise memory. A hands-on, step-by-step guide to setting up Embabel from scratch will follow as a companion piece. Here's the link to the source code: https://github.com/trainerpb/embabel-hello-world/tree/feature/revision.

By Soham Sengupta
Gossips on Cryptography: Part 4
Gossips on Cryptography: Part 4

In this blog, we will continue our discussion from the previous parts. If you have not read them, please read them first. Parts 1 & 2 – Caesar Cipher, Vigenere Cipher, Symmetric Encryption, AES, Convergent Encryption, IVPart 3 – Hashing, Salting, Rainbow Table Attacks, Asymmetric Encryption, RSA In Part 3, we teased a few topics for Part 4 — Envelope Encryption, PKI, and more. Today we gossip about exactly those! Let's go. First, A Quick Revisit: Convergent Encryption We discussed Convergent Encryption back in Parts 1 & 2, but let's revisit it here because it connects beautifully to Envelope Encryption. Remember? Convergent Encryption means — if you encrypt the same plaintext with the same key and the same IV, you will always get the same ciphertext. So Why Is That Useful? Imagine you work at a big company. 500 employees all upload the same file — let's say the company's HR policy PDF. If you use regular encryption (different ciphertext every time), your storage system stores 500 different encrypted copies. That's 500x storage wasted! With Convergent Encryption, since the same file + same key = same ciphertext, the storage system realizes — "Hey, I already have this encrypted file!" — and stores only ONE copy. All 500 employees point to the same encrypted blob. This is called deduplication. This is exactly how Dropbox, Google Drive, and AWS S3 save enormous amounts of storage at their scale. Another Use Case — Searching Over Encrypted Data Here is another very powerful use case of Convergent Encryption that most people don't think about — searching. Imagine you have a database where all the data is encrypted. A user wants to search for records where the email is "[email protected]." With regular encryption, every time "[email protected]" is encrypted, it produces a **different** ciphertext (because of a random IV). So to search, you would have to: Decrypt every single record in the databaseCompare the plaintextReturn the matches That is insanely expensive! Imagine doing this on a database with 100 million records. Your server will cry. Now with Convergent Encryption — "[email protected]" always produces the **same** ciphertext. So to search, you just: Encrypt the search term "[email protected]" onceLook for that ciphertext in the database — just like a normal indexed search!Return the matches No decryption needed at all! The data stays encrypted at rest, and you can still do fast, exact-match searches on it. This is called searchable encryption, and it is used in scenarios like: Encrypted databases where you still need to support queriesHealthcare systems — searching patient records without ever exposing raw dataEmail systems — searching your encrypted inbox without the server ever seeing your emails in plain text Pretty powerful, right? Same property (deterministic output) — two completely different superpowers (deduplication + searchable encryption). But wait — there's a catch. If two people can produce the same ciphertext, can someone guess your file? Yes, this is called a confirmation attack. Someone could hash a known file, compare it with stored hashes, and confirm whether you uploaded that file. So Convergent Encryption is great for performance and deduplication but is used carefully in highly sensitive scenarios. Now Let's Talk About Envelope Encryption Okay, so now we know — encryption needs keys. And those keys need to be stored somewhere safely. But here's the problem — who encrypts the key itself? If your key is lying around in plain text, a hacker who gets access to your server gets everything. So the answer is — we encrypt the key too! This is the core idea of Envelope Encryption. The Two Keys in Envelope Encryption DEK — Data Encryption Key. This is the key that directly encrypts your actual data. Think of it as the key to your diary.KEK — Key Encryption Key. This is the master key that encrypts the DEK. Think of it as the key to your locker — inside which you keep your diary key. So the flow looks like this: Your Data → encrypted with DEK → Encrypted Data DEK → encrypted with KEK → Encrypted DEK You store both — the Encrypted Data and the Encrypted DEK — together. The KEK lives safely inside a highly secure system (like AWS KMS or Vault). Real-Life Example — The Bank Locker Imagine you have an important document (your data). You put it in a box and lock it with a small key (DEK). Now you don't want to carry this small key everywhere — so you put the small key inside your bank locker (encrypt DEK with KEK). The bank locker key (KEK) stays with the bank in a highly secure vault. To read your document: Go to the bank → get your small key out (decrypt DEK using KEK)Use the small key to open the box (decrypt data using DEK) Simple! And very secure. Why Not Just Encrypt Data Directly With KEK? Two very practical reasons: 1. Performance The KEK usually lives inside a secure hardware vault or cloud service (like AWS KMS). If you send your entire 10GB file to KMS every time you want to encrypt or decrypt — that's painfully slow and expensive. Instead, you only send the tiny DEK (a few bytes) to KMS. The heavy lifting of encrypting actual data is done locally with the DEK. 2. Key Rotation Say after 6 months you want to change your encryption key (key rotation is a security best practice). Without envelope encryption — you'd have to decrypt ALL your data and re-encrypt it with a new key. Imagine doing that for terabytes of data! With Envelope Encryption — you only re-encrypt the DEK with the new KEK. Your actual data stays untouched. Much faster, much cheaper. Where Is Envelope Encryption Used? Literally everywhere in the cloud world: AWS S3 – When you enable server-side encryption on a bucketAWS RDS – When you enable encryption on a databaseGCP Cloud Storage – Envelope encryption is the defaultAzure Key Vault – Same pattern Every time you see that little "encryption enabled" checkbox on a cloud service — envelope encryption is what's happening under the hood. Now Let's Talk About PKI PKI stands for Public Key Infrastructure. In Part 3, we discussed Asymmetric Encryption — where you have a Public Key and a Private Key. Sounds great in theory. But here's a real problem. The Trust Problem Imagine Rahul wants to send an encrypted message to Priya. Priya shares her public key with Rahul. Rahul encrypts the message with Priya's public key and sends it. But wait — how does Rahul know that the public key he received is actually Priya's? What if a hacker intercepted the communication and swapped Priya's public key with their own? Rahul encrypts with the hacker's public key → hacker decrypts → reads the message. This is called a Man-in-the-Middle (MITM) Attack. We need someone that both Rahul and Priya trust, who can say — "Yes, this public key truly belongs to Priya." That trusted someone is called a certificate authority (CA). PKI — The Complete Picture PKI is a system made up of several components that together solve the trust problem. Let's go one by one. 1. Certificate Authority (CA) A CA is a trusted organization whose job is to verify identities and issue digital certificates. Think of them like the government passport office — they verify who you are and give you an official identity document (passport). Well-known CAs in the real world: DigiCert, Let's Encrypt, GlobalSign, Comodo. Your browser/OS comes pre-loaded with a list of trusted CAs. That's how your browser automatically trusts websites — because their certificates were signed by a CA your browser already trusts. 2. Digital Certificate A Digital Certificate is like a government-issued ID card for websites (or people or servers). It contains: The owner's name (e.g., google.com)The owner's Public KeyThe CA's name (who issued it)Expiry dateA digital signature from the CA When you open https://google.com — your browser checks Google's certificate. It sees the CA that signed it. It checks if that CA is in its trusted list. If yes — green light, connection is secure! That's the lock you see. 3. Digital Signature We talked about hashing in Part 3 — that you cannot reverse a hash. Digital Signatures use this + asymmetric encryption together in a clever way. When a CA wants to sign a certificate, it: Takes the certificate content and hashes itEncrypts that hash with its own Private Key That encrypted hash = Digital Signature. Anyone can verify the signature using the CA's Public Key (which is publicly available). If the decrypted hash matches the actual certificate content → the certificate is genuine and untampered! Think of it like a wax seal on an envelope. Anyone can see the seal, but only the king's ring (private key) could have made it. 4. Certificate Chain (Chain of Trust) In the real world, CAs have a hierarchy: Root CA → Intermediate CA → Your Website Certificate The Root CA is the ultimate trusted authority. It signs Intermediate CAs. Intermediate CAs sign individual website certificates. This chain is called the Chain of Trust. Why this hierarchy? Security! Root CA private keys are kept in ultra-secure, air-gapped hardware. They are almost never used directly. Intermediate CAs do the day-to-day certificate signing. If an Intermediate CA is ever compromised, it can be revoked without affecting the Root CA. Where Is PKI Used in Real Life? PKI is literally everywhere. You just don't see it because it works silently in the background. 1. HTTPS Websites (SSL/TLS) Every https:// website uses PKI. When you open your bank's website — a PKI handshake happens in milliseconds: Your browser asks the bank's server for its certificateThe bank's server sends its Digital CertificateBrowser verifies the certificate using the CA's public keyIf valid → browser and server agree on a secret key (using asymmetric encryption)All further communication uses that secret key with AES (fast symmetric encryption) That's TLS in a nutshell. And PKI is the backbone of all of it. 2. Email Signing (S/MIME) When your company sends you a digitally signed email — PKI is involved. The sender signs the email with their private key. You verify it with their public key from their certificate. You can be sure the email is genuinely from them and was not tampered with in transit. 3. Code Signing When you download an app or a software update — how does your phone/OS know it's legit and not malware? The developer signs the app with their private key. Your phone verifies it using the developer's certificate. This is why on Android you see the "Install from unknown sources" warning — no valid certificate found! 4. Government and Banking Your Aadhaar card, digital signatures on GST filings, net banking OTPs — all of these use PKI infrastructure behind the scenes. India's government runs its own CA called CCA (Controller of Certifying Authorities) under the IT Act. 5. VPNs and Internal Company Networks When you connect to your company's VPN, PKI certificates are used to verify that you are connecting to the genuine company server and not an imposter. Terms We Have Learned So Far (All 4 Parts) Cryptography, Algorithm, Plain Text, Key, Cipher TextSymmetric Encryption, Convergent Encryption, Initialization Vector (IV), Searchable EncryptionHashing, Hash/Digest, Avalanche EffectSalt, Rainbow Table Attack, Confirmation AttackAsymmetric Encryption, Public Key, Private Key, RSADEK (Data Encryption Key)KEK (Key Encryption Key)Envelope EncryptionKey Rotation, DeduplicationPKI (Public Key Infrastructure)Certificate Authority (CA)Digital CertificateDigital SignatureChain of TrustTLS/SSL That's a solid vocabulary now! Coming in Part 5... (Part 5 is in progress — stay tuned!) In the next part, we will gossip about: SSL/TLS Deep Dive – The full step-by-step TLS handshake explained simplymTLS (Mutual TLS) – How microservices talk to each other securelyHSM (Hardware Security Module) – The physical vault where the most sensitive keys liveZero-Knowledge Proofs – Proving you know something without revealing what it isAnd more... Stay tuned for Part 5! If you liked this blog, do give it a like and share it with someone who you think should learn this. Let's spread the knowledge! Read the previous parts here: Part 1 & 2 Part 3

By Sahil Aggarwal
Build Software Faster With Three Simple Principles
Build Software Faster With Three Simple Principles

Development in small companies and startups often slows down at the boundaries between people and teams. A developer waits for a product decision. Frontend work stalls because the API response is unclear. QA discovers that two teams interpreted the same requirement differently. Three practical habits can reduce these delays: document the feature, discuss its technical risks, and agree on interface contracts before dependent implementations diverge. Keep each step proportional to the size and uncertainty of the task. Start With the Biggest Uncertainty Before distributing work, identify what the team needs to learn first. Unclear user experience? Start with a UI prototype. Frontend developers can use mocks to explore the flow while the team clarifies requirements.Uncertain business logic, performance, or integration? Start with a focused backend investigation or technical prototype.Unclear user problem or business value? Start with product discovery, involving users, product stakeholders, and technical specialists as needed. Having a UI does not automatically make frontend work the first priority. Choose the starting point that resolves the most important uncertainty, then agree on enough shared detail to let work proceed in parallel. 1. Maintain Feature Documentation and Define Responsibilities A concise Product Requirements Document (PRD) gives the team a shared explanation of what to build, why it matters, and how to evaluate it. For a small feature, a short page may be enough. QuestionWhat to recordWhat are we building?The intended behavior, MVP scope, and explicit exclusions.Who is it for?The users and the problem they face.Why is it needed?The expected benefit to users and the business.How will we evaluate it?Acceptance criteria, success metrics, a baseline or comparison group, and a measurement period. Name the person responsible for product decisions and identify the people needed to implement and validate the feature. These are responsibilities; a small team may combine several of them in one role. ResponsibilityTypical ownerSet goals, decide scope, and resolve product questionsProduct ownerAssess feasibility, design interfaces, and implement the featureDevelopers and technical leadDefine test scenarios and verify behaviorQA and developersDesign the user experienceDesigner, where neededDefine instrumentation and evaluate impactAnalyst or another explicitly assigned team member Example: A Post Recommendation System The following is an illustrative PRD. Its numerical targets are examples to agree on for a particular product, rather than measured results or universal benchmarks. SectionExampleProblem and hypothesisUsers may struggle to find relevant posts in a chronological feed. We expect recommendations based on their interests to improve content discovery.UsersReaders discovering posts; creators whose eligible posts can be recommended.MVP behaviorReturn up to 20 eligible posts, ranked by popularity within topics inferred from likes and subscriptions. Recompute lists every 24 hours. Exclude deleted posts and posts the requesting user cannot access when serving the response. Use a general popularity list when there is insufficient interest history.Experiment supportAssign eligible users to stable control and treatment groups and record recommendation impressions and clicks. Include this in the MVP so the first version can be evaluated.Out of scopeML models, similarity between users, and advertising recommendations. Decide whether to add these after evaluating the MVP.PerformanceIllustrative target: server-side p95 response time below 200 ms for serving precomputed lists at 1,000 requests per second on a representative test dataset. Measure the batch recomputation job separately and require it to finish within the 24-hour refresh window.Acceptance criteriaTests verify ranking on a fixed dataset, the popularity fallback, access filtering, and feedback recording. Likes and subscription changes affect the next scheduled recomputation. Load tests meet the stated latency target.Success measurementIllustrative primary target: a 10% relative increase in recommendation CTR versus the control group. Define CTR as clicks divided by recorded impressions. Plan an initial two-week experiment; estimate sample size before launch and report uncertainty if the result is inconclusive. Monitor seven-day retention and API error rate as guardrails, with acceptable thresholds agreed before launch.RisksWeak relevance, overexposure of already popular posts, and load spikes. Inspect recommendation diversity, provide a fallback, and test capacity before rollout. Technical notes can accompany the PRD, with implementation decisions reviewed during technical refinement. For example, a Go service could use PostgreSQL for source data and Redis for precomputed lists. A draft API might expose GET /recommendations for the authenticated user and POST /feedback for interactions. The team still needs to define schemas, error responses, authorization behavior, and pagination before implementation. Keep release acceptance separate from business success. A correctly implemented feature can fail to improve the chosen metric. That result should inform the next product decision. What if the Product Has No Documentation? Start with the workflow you are changing and the business rules most likely to be misunderstood. Link the relevant code, record open questions, and expand the documentation as the team learns. Documenting the entire system does not need to become a prerequisite for the next useful change. After release, compare outcomes with the original hypothesis. A weak result is a reason to investigate the feature, measurement, and assumptions. Some work provides value through reliability, lower operating costs, or reduced risk, and its evaluation should reflect that purpose. 2. Conduct Technical Refinement Technical refinement, sometimes called technical grooming, connects product expectations with implementation decisions. Its output should be a workable approach and a clear record of remaining questions. Ask a developer or technical lead to review the PRD and identify: Existing constraints and integration dependencies.Performance, security, and operational risks relevant to the change.Decisions that require input from product or another team.Unknowns that need a short investigation or prototype. Bring the relevant people together to resolve those questions. If the implementation cost changes the original assumptions, revisit the scope with the product owner. Keep the Process Proportional Agree on a timebox based on scope and uncertainty. For example, a modest feature might receive two days of initial technical analysis, with a named owner and deadline for each product question. A complex migration may require several investigations. These are planning choices to review as new information appears. Record decisions, tradeoffs, unresolved questions, and owners. There is usually no need to transcribe every discussion. Finish with a small implementation plan: what can run in parallel, what must happen first, and what evidence will show that a risky assumption holds. This process can reduce avoidable rework. It does not guarantee that every issue will be discovered in advance, so leave room to revise the plan during development. 3. Develop Using the Specification-First Principle For work that crosses an API boundary, agree on the contract early. An OpenAPI description can capture HTTP operations, request and response schemas, and other interface details. The contract should also be supported by examples and documented behavior where a schema alone is insufficient. Backend and frontend engineers review the contract together, including empty states, errors, and compatibility expectations.Frontend developers build against mocks that reflect the agreed contract.Backend developers implement the API and verify that its behavior matches the contract.QA and developers prepare API contract checks and derive end-to-end scenarios from the PRD, user flows, and business rules. API contract checks and end-to-end tests serve different purposes. Matching a response schema does not establish that a complete user journey works correctly. How This Reduces Waiting Consider an illustrative recommendation feature. If frontend developers invent a response shape while waiting for the backend, they may need to rewrite rendering and error handling during integration. Agreeing on the response schema, empty-list behavior, and refresh semantics first lets both sides work against the same assumptions. Mocks still need to match the implementation. Run compatibility checks in CI and update the specification, mocks, and tests together when the contract changes. Integrate early enough to expose mismatches before release. Specification-first development also leaves room for exploratory code. A short prototype may be necessary to discover whether a proposed interface is feasible. The aim is to agree on the contract before teams invest heavily in dependent implementations. For more on the approach, see Boost Efficiency With the Specification-First Principle. Keep Documentation Useful as the Project Evolves Documentation helps future team members understand behavior and the reasons behind earlier decisions. DORA's research on documentation quality links high-quality internal documentation with organizational performance and finds that it strengthens the impact of technical practices. Keep a small set of maintained resources close to the work: The PRD and the results of the feature's evaluation.Architecture decisions, API contracts, and relevant data models.Deployment and recovery instructions.Links to changes and the decisions behind them, in GitLab, a README, or another shared system.Test scenarios and acceptance checklists. Assign ownership and update these resources when behavior changes. Outdated documentation can mislead the next person just as missing documentation can leave them guessing. Conclusion These three habits address common sources of delay: Concise feature documentation gives the team a shared goal, scope, and definition of success.Technical refinement surfaces constraints and assigns owners to unresolved questions.Agreed API contracts support parallel development and reduce integration rework. Try them on one feature. Track time spent waiting for decisions, integration rework, and time from an agreed scope to release, alongside quality and product outcomes. Use what you learn to adjust the process. The practical goal is a team that can make decisions, build, and validate changes with less avoidable waiting.

By Ilia Ivankin
How Go Maps Work: From Buckets to Swiss Tables
How Go Maps Work: From Buckets to Swiss Tables

Hey Mates! “How does a map work in Go?” is one of my favorite interview questions. It sounds simple, but it opens up a conversation about hashing, collisions, memory layout, and why two implementations with the same average O(1) lookup complexity can behave quite differently. With Go 1.24, that conversation got more interesting: the map implementation switched to Swiss Tables. Let’s walk through how the old implementation worked, what changed, and how the new design finds your keys. TL;DR Go 1.24, released in February 2025, completely replaced the internal implementation of the built-in `map` with a design based on Google's Swiss Tables. The syntax and the behavior guaranteed by the Go specification did not change, so existing programs required no migration. In microbenchmarks, some map operations became up to 60% faster. Across the Go team's full-application benchmark suite, geometric mean CPU time improved by about 1.5%. Datadog reported roughly 70% less memory for one exceptionally large map—an impressive case study, not a universal promise. Why Replace map at All? map is one of the most frequently used data structures in Go. It appears in caches, configuration, indexes, and every `map[string]any` we would rather not discuss. The old implementation served Go well for more than a decade. Hash-table research did not stop, however. At CppCon 2017, Google engineers Sam Benzaquen, Alkis Evlogimenos, Matt Kulukundis, and Roman Perepelitsa presented a new cache-friendly design. It became known as Swiss Tables, after the Google Zürich office where the team worked. In 2018, Google released an implementation in the C++ Abseil library. The design then spread across ecosystems: C++: absl::flat_hash_map in Abseil;Rust: the standard HashMap is built on hashbrown, a Swiss Tables implementation, since Rust 1.36;Go: first through third-party packages such as dolthub/swiss and cockroachdb/swiss, then in the built-in map starting with Go 1.24. The route into Go was collaborative. Community members built early prototypes. Peter Mattis of CockroachDB combined those ideas with solutions for Go-specific requirements in cockroachdb/swiss. The Go 1.24 runtime implementation is heavily based on that work. Before we get to the “Swiss” part, let us quickly review hash tables. If collisions and load factors are already familiar, skip to section 3. Hash Tables From First Principles Imagine a theater coat check where coats are retrieved by surname. You say “Smith,” and the attendant applies a simple rule to decide which section to search first, section 17, perhaps. They do not scan the entire room; they go directly to one small area and inspect a few tags. The key is the surname.The value is the coat.The hashfunction turns a key into the starting section. The same key always produces the same result within a particular map.A slot stores one key/value pair. Lookup is O(1) on average because the hash takes us to a small part of the table instead of forcing us to scan every entry. This is an average-case property, not an unconditional guarantee. The unavoidable problem is a collision: there are finitely many locations, so different keys eventually choose the same starting point. Two classic strategies handle this: Chaining. Section 17 holds a list of key/value pairs. Lookup walks that small list. Hans Peter Luhn of IBM described this approach in 1953.Open addressing. Slot 17 is occupied, so try another slot, then another, until a suitable one is found. The order of locations is called the probe sequence. Open addressing was used in 1954 and formally published in 1957. Both ideas are about 70 years old. The interesting part is how modern implementations make them friendly to modern CPUs. The Old Implementation: Go 1.23 and Earlier The old Go map was a hybrid. It used fixed-size buckets, while excess collisions were handled with chains of overflow buckets. In spirit, it was closer to chaining. A map had an hmap header pointing to an array of buckets. Each `bmap` bucket contained exactly 8 slots: Plain Text flowchart LR subgraph HMAP["hmap header"] direction TB C["count: number of entries"] B["B: log2 bucket count"] P["buckets: array pointer"] end P --> ARR["bucket array: 2^B buckets"] ARR --> BKT subgraph BKT["bmap bucket: 8 slots"] direction TB TH["tophash: 8 filter bytes"] K["8 keys"] V["8 values"] OV["overflow pointer"] end OV --> OB1["overflow bucket"] OB1 --> OB2["another overflow bucket"] Three details matter: tophash: Eight bytes at the start of the bucket, one per slot. Each byte contains the top bits of that key's hash. Comparing a byte is cheaper than comparing a full string key, so it acts as a fast filter.Keys and values are stored separately: Eight keys followed by eight values. This reduces alignment padding.Overflow buckets: When the primary bucket cannot hold another entry, the runtime allocates another bucket and links it into a chain. To find grape, the runtime roughly did this: Calculate `hash("grape")`.Use the low `B` bits to select a bucket.Check the bucket's eight `tophash` bytes one at a time.When a byte matches, compare the full key. If the key matches, return the value.If the bucket has no match, follow its pointer to the overflow bucket and repeat.If the chain ends, the key is absent. The old map grew when average occupancy exceeded 6.5 entries per 8-slot bucket, a load factor of 81.25%, or when too many overflow buckets accumulated. The number of primary buckets doubled. Crucially, growth was incremental. Go did not move the entire old array at once. Each write evacuated a little more data. One unlucky insertion therefore did not have to copy a gigabyte-sized map in a single pause. Pointer chasing. Every overflow hop reads another part of the heap and increases the risk of a CPU-cache miss. A long overflow chain can therefore be expensive.Serial metadata checks. The runtime inspected `tophash` byte by byte and slot by slot. Eight slots could mean eight loop iterations and several branches.Memory overhead. Overflow buckets and their pointers cost memory. Raising the load factor much above 81% made overflow chains more common and lookup slower. The garbage collector could also have more pointers to scan. The goal was clear: fewer pointers, better locality, and less serial work. Swiss Tables provide exactly that. Meet Swiss Tables A Swiss Table is open addressing adapted to modern CPUs. Three decisions drive the design: Split the hash into an address and a short fingerprint.Pack the metadata for a group of slots into one machine word.Compare the fingerprints of 8 slots in parallel, then compare full keys only for the candidates. A 64-bit hash is divided into two unequal pieces: Plain Text 64-bit key hash ┌─────────────────────────────────────────────┬─────────────┐ │ h1: upper 57 bits │ h2: 7 bits │ │ chooses where probing begins │ fingerprint │ └─────────────────────────────────────────────┴─────────────┘ h1 selects the initial group.h2 is a 7-bit fingerprint used as a cheap filter before a full-key comparison. In the coat-check analogy, h1 chooses a section, and h2 is a short mark on each tag that quickly rules out unrelated coats. Slots are arranged in groups of 8. Each group has a 64-bit control word, one byte per slot: Plain Text control word: 64 bits = 8 bytes ┌────┬────┬────┬────┬────┬────┬────┬────┐ │ ∅ │ 15 │ 27 │ 5C │ ∅ │ 3A │ 71 │ † │ └────┴────┴────┴────┴────┴────┴────┴────┘ │ │ │ │ └─ occupied, h2 = 0x15 └─ deleted tombstone └─ empty Each byte describes its slot: Plain Text | Byte value | Meaning | | `0b1000_0000` (`0x80`) | **empty** slot | | `0b1111_1110` (`0xFE`) | **deleted** slot, or tombstone | | `0b0xxx_xxxx` | occupied; the low seven bits contain h2 | Occupied slots always have a zero high bit, while empty and deleted slots have a one. That encoding enables efficient parallel tests. Suppose we are looking for h2 = 0x27. Conceptually, the operation looks like this: Plain Text probe word: 27 │ 27 │ 27 │ 27 │ 27 │ 27 │ 27 │ 27 == │ == │ == │ == │ == │ == │ == │ == control word: ∅ │ 15 │ 27 │ 5C │ ∅ │ 3A │ 71 │ † result: 0 │ 0 │ 1 │ 0 │ 0 │ 0 │ 0 │ 0 ↑ candidate slot 2 On amd64, Go recognizes this operation as an intrinsic and uses SIMD instructions. Other architectures have a portable SWAR implementation — SIMD Within A Register — that processes all eight bytes using arithmetic on one 64-bit word. The following simplified code is optional. The important result is a bit mask of slots whose h2 may match: Go func matchH2(ctrl uint64, h2 uint8) uint64 { broadcast := uint64(h2) * 0x0101010101010101 x := ctrl ^ broadcast return (x - 0x0101010101010101) &^ x & 0x8080808080808080 } There are two reasons a candidate is not yet a result: h2 has only seven bits, so an occupied slot has a 1/128 chance of sharing the same fingerprint by accident;The portable bit trick can produce a rare extra candidate because of subtraction borrow. Neither affects correctness. Go always performs a full-key comparison before returning a value. If the initial group has no matching key, open addressing continues at group granularity. Go uses a quadratic probe sequence. Three rules are enough to understand lookup: Matching h2 bytes produce candidate slots whose full keys are checked;If the group contains an empty slot, stop — the key is absent;If the group has no match and no empty slot, visit the next group in the probe sequence. One probe handles the metadata for eight slots, so densely populated tables can still be searched efficiently. Because group metadata is cheaper to inspect, Swiss Tables can remain more densely populated. Go allows an average load up to 7/8 = 87.5%, compared with 81.25% in the old implementation. More useful entries in the same backing storage usually means less memory per key. Go's Extension: A Directory of Tables Go could not simply port Abseil's implementation. Two language and runtime requirements needed additional design work. A conventional Swiss Table grows all at once: allocate a larger array and move every entry. For a gigabyte-sized map, the insertion that triggers growth would suffer a noticeable pause. Go is widely used for latency-sensitive services, and its old maps already bounded growth work per insertion. A large Go map is a directory of independent Swiss Tables, not one unbounded table. Each table covers part of the hash space and has a maximum capacity of 1024 slots, or 128 groups: Plain Text flowchart TD M["map[K]V"] --> DIR["directory: array of table pointers"] DIR --> T0["table 0: up to 1024 slots"] DIR --> T1["table 1: up to 1024 slots"] DIR --> TN["table N"] T1 --> G0["group 0: control word + 8 slots"] T1 --> G1["group 1: control word + 8 slots"] T1 --> GK["up to 128 groups"] A variable number of upper hash bits selects the table. This technique is a form of extendible hashing. Growth is local. When a table below the limit fills, only that table grows. Once a table reaches the limit, it splits into two. The amount of growth work caused by one insertion is therefore bounded by one table of at most 1024 slots: roughly at most 896 live entries at the 7/8 load threshold (not the entire map). Small maps get a special fast path. A map that starts small and never exceeds 8 entries lives directly in a single group with no table directory. A map that has already grown, or was created with a large `hint`, does not shrink back into this representation after deletions. Unlike many hash tables, Go explicitly permits modifying a map during iteration: an entry deleted before it is reached must not be produced;an entry updated before it is reached must produce its latest value;a newly inserted entry may or may not be produced. Growth reshuffles storage, so an iterator cannot simply walk the current array. Go's iterator retains the old table to determine traversal order. Before returning an entry, it consults the current table to confirm that the key still exists and to obtain the latest value. According to the Go team, iteration is the most complex part of the implementation. Deletion and Tombstones With open addressing, deleting an entry cannot always turn its slot into an ordinary empty slot. Lookup stops at an empty slot. Creating one in the middle of another key's probe sequence could therefore make a live key unreachable. A tombstone, encoded as deleted (`0xFE`), means: “an entry used to be here; continue probing.” Insertions may reuse tombstones. Go avoids a tombstone when the group already contains another empty slot. Any lookup reaching that group would stop there anyway, so turning the deleted slot into empty cannot break a probe sequence. Accumulated tombstones disappear during a later grow or split. Live entries are copied into new groups and deleted slots are not. Go 1.24 does not implement a separate same-size grow for this purpose. And Performance Numbers According to the Go team and early production reports: Microbenchmarks: some map operations are up to 60% faster than in Go 1.23. Results vary widely, and a few edge cases regress.Full-application benchmarks: the Go team's suite showed about 1.5% geometric-mean CPU-time improvement for the whole application. A specific program may behave differently.Memory: Datadog measured roughly 70% less map memory for one very large map with about 3.5 million entries in a high-traffic environment. This favorable case benefited from denser storage, no overflow buckets, and growth that did not retain two enormous bucket arrays at once. Savings were much smaller in another environment. Large maps often benefit more because overflow chains and cache misses hurt the old design more strongly. The exact result depends on key and value sizes and on the mix of reads, writes, deletes, hits, and misses. Measure your own workload. Takeaways The Go Swiss Tables story is a good example of changing a foundational component carefully: Start with a proven design already used by Abseil and Rust's `hashbrown`.Adapt it to Go's invariants: bounded growth latency and modification during iteration.Preserve the public API and specified behavior. Hash tables are seven decades old, yet a better match between data layout and modern hardware can still remove CPU time and memory across a large ecosystem. “Solved” problems often have room for another good engineering pass. References [Faster Go maps with Swiss Tables] — the primary Go team article[Go 1.24 Release Notes] — runtime changes and `GOEXPERIMENT=noswissmap`[Go 1.24 `internal/runtime/maps` source] — original source[Go 1.26 runtime source] — the release in which the old map implementation was removed[Abseil Swiss Tables Design Notes] — original sourceMatt Kulukundis at CppCon 2017Datadog's workload-specific memory case study[`hashbrown`] — the Swiss Tables implementation behind Rust's standard `HashMap`

By Ilia Ivankin
Are Passphrases Still Secure in the Age of AI?
Are Passphrases Still Secure in the Age of AI?

The short answer up front: Yes, but only if the passphrase is genuinely random. Passphrases emerged as a direct response to a well-known problem with classic password policies: requirements like "at least one special character, one number, one uppercase letter" tend to produce short but complex-looking passwords ("Tr0ub4dor&3") that are hard for humans to remember but comparatively easy for machines to guess because users keep falling back on the same patterns (capital letter at the start, number and special character at the end). The passphrase flips this principle: instead of relying on character variety packed into a few positions, it relies on length through multiple words strung together. The math of passphrases makes passphrases very safe because the number of possible combinations grows exponentially with each additional word. You can find passphrases in many parts of your life: Master passwords for password managers: here it's the one passphrase you still have to remember, while every other credential is generated and stored for you.Disk and container encryption (e.g., VeraCrypt, LUKS, BitLocker), where an attacker could attack the encrypted drive offline with no rate limiting.Private SSH/PGP keys, to add an extra layer of protection to the key itself in case the key file is stolen.Cryptocurrency wallets in the form of seed phrases (e.g., following the BIP-39 standard), which directly represent the private key.Wi-Fi access (WPA2/WPA3), where long passphrases make offline attacks on the handshake harder.Corporate and cloud logins, often combined with a second factor (MFA), replacing classic password policies. Let's discuss for a minute what a passphrase is. A passphrase is a string of several, usually randomly chosen, words used to protect access to accounts, encrypted files, or systems. A passphrase doesn't rely on complexity through special characters — unlike passwords — but on length. The security gain comes from the sheer number of possible combinations: a random word from a list of 7,776 entries (as used in Diceware, more on that shortly) provides about 12.9 bits of entropy. Four such words combined yield roughly 51.6 bits — more than a typical eight-character password with mixed case, numbers, and special characters achieves, while being easier to remember. The main types include: Diceware Passphrases With Diceware, words are selected using real dice from a fixed word list (e.g., the EFF Long Wordlist with 7,776 entries). Every combination of numbers from five dice rolls (1–6) points to exactly one word. The process can be done entirely offline and produces demonstrably random, and therefore cryptographically solid, results. XKCD-Style ("correct horse battery staple") Popularized by the well-known XKCD comic: four to six random, independent words strung together, often separated by hyphens or spaces. The trick: no grammatical structure, which makes it harder to guess than a meaningful sentence. Sentence-Based Passphrases A complete but unusual sentence, e.g., "MyCatJugglesFiveBananasOnMondays!" Advantage: easy to remember if it's personally meaningful. Disadvantage: real sentences follow linguistic patterns and are therefore easier to guess than random word lists — attackers increasingly use language models for such attacks. Weaknesses of Passphrases Passphrases solved well-known problems with classic passwords in many areas, but they were not the silver bullet. They had weaknesses in the past, and they will continue to have weaknesses in the future. In the age of AI, one of those weaknesses deserves particular attention, because it really puts the "long and random" principle to the test. False randomness in AI-generated passphrases: If you have an AI generate a passphrase, the result often looks complex but follows predictable patterns and, in practice, achieves noticeably less entropy than genuine random selection. Let's discuss this problem in more detail. The core point is that an LLM doesn't generate text through genuine randomness, but through statistical probability. Every word is chosen because it seems most plausible in context. That's exactly the opposite of what a secure passphrase needs. Even when explicitly asked to produce "random words," a model tends to favor: Common, everyday words over rare onesCertain "interesting-sounding" combinations that are overrepresented in the training corpusThe same words across repeated requests, because the model tends toward similar outputs for similar prompts To a human, the result looks complex and random ("fluorescence-marmalade-quantumleap"), but it's actually drawn from a much smaller effective "word pool" than a genuine random selection would be. The true number of equally likely options is smaller than the apparent complexity suggests. An attacker who knows that a passphrase was generated by a specific AI doesn't need to search the entire character space — they can instead build a targeted dictionary from that model's preferred patterns and smaller effective word pool, and run a guessing attack against it. Similar Effects Have Already Been Observed for AI-Generated Passwords The problem described above isn't merely an academic exercise — early data on this topic already exists, though so far it has been studied analogously for AI-generated passwords. In February 2026, security firm Irregular published a report defining this exact concept as a viable attack method against AI-generated passwords, and Kaspersky's cracking tests confirmed its practical effectiveness. Specifically: once an attacker identifies which LLM generated a target credential, an exhaustive brute-force attack against the full 94^16 character space is no longer necessary. Using a model-specific attack dictionary, candidates can be ranked by their empirical generation frequency and searched via a probabilistically optimized attack against a key space that's smaller by several orders of magnitude. A separate analysis published on VPN Central in February 2026 examined the same type of attack and found that, combined with brute-force techniques, credentials could be recovered within minutes. Conclusion For passwords, several independent studies have documented with concrete figures that AI-generated passwords carry less entropy. For passphrases specifically, the current body of data is considerably thinner. However, a plausible extension of the same underlying mechanism — LLMs optimize for probability rather than randomness, which applies to words just as much as to characters - is a legitimate inference, and it suggests that AI-generated passphrases likely carry noticeably less entropy than genuinely random ones. The study "LLM-Generated Passphrases That Are Secure and Easy to Remember" (ACL Anthology, NAACL 2025 Findings) provides indirect data on this question. It's worth noting, though, that the study is primarily concerned with defining the conditions under which an LLM can be made to produce usable passphrases in the first place. In short: numeric, evidence-backed data exists for passwords; for passphrases, the same mechanism — statistical bias rather than genuine randomness — is plausible and architecturally well-grounded, but still thinly supported empirically.

By Constantin Kwiatkowski
Designing a Role-Aware Runtime for AI Avatar Systems
Designing a Role-Aware Runtime for AI Avatar Systems

AI avatar projects often begin with the most visible question: What should the character look like? That question matters, but it comes much later in the architecture than many teams expect. A responsive AI avatar is not simply an animated face connected to a large language model. It is a real-time distributed system that must listen, understand, retrieve context, make decisions, authorize tools, generate speech, synchronize animation, support interruptions, and record enough telemetry to explain failures. When all these responsibilities are placed inside one application or one oversized system prompt, every new character becomes a separate implementation. A cartoon guide, a robotic assistant, an enterprise representative, and an entertainment host may end up with different prompts, APIs, analytics pipelines, safety rules, and release processes. The visible experiences are different, but the underlying engineering problems are largely the same. A more maintainable architecture separates the system into three layers: A reusable conversational runtimeA versioned role profileAn embodiment adapter The runtime manages the conversation. The role profile defines behavior. The embodiment adapter translates an approved response into speech, facial animation, gestures, screen actions, or physical-device commands. This separation allows teams to introduce new characters without rebuilding the security, observability, retrieval, and session-management layers each time. Start With Three Explicit Layers 1. Conversational Runtime The runtime owns the operational parts of the system: Session and turn stateAutomatic Speech RecognitionRetrieval and groundingModel orchestrationTool executionText-to-Speech generationStreaming and interruptionPolicy enforcementLogging, tracing, and metricsHuman escalation These capabilities should remain stable regardless of whether the user is interacting with a stylized character, a digital employee, or a physical robot. 2. Role Profile The role profile defines what makes one avatar different from another. It can contain: Identity and purposeTone and vocabularyPermitted knowledge sourcesRestricted topicsTool permissionsEscalation rulesGesture limitsVoice configurationAudience suitabilityResponse-length preferencesLanguage support A role profile should be stored as versioned configuration rather than buried inside a single prompt. 3. Embodiment Adapter The embodiment adapter converts a validated response plan into output for a specific interface. For example, one adapter might generate speech, visemes, and facial expressions for a browser-based character. Another might coordinate an on-screen host with media cues. A robotic adapter might translate approved intentions into a restricted device-command format. The adapter should not become a second reasoning system. Its job is to translate approved decisions into embodiment-specific output. Figure 1: A role-aware AI avatar architecture that separates the conversational runtime, role configuration, control layer, telemetry, and embodiment-specific output. Use a Stable Event Contract A reusable runtime needs a stable event model. Passing unstructured strings between services makes cancellation, retries, replay tests, and end-to-end tracing unnecessarily difficult. Each event should carry at least: A session identifierA turn identifierA sequence numberAn event timestampA typed event nameThe minimum payload required by the next service A simplified TypeScript contract could look like this: TypeScript type AvatarEvent = | { type: "user.audio.partial"; sessionId: string; turnId: string; sequence: number; text: string; final: false; timestamp: number; } | { type: "user.utterance"; sessionId: string; turnId: string; sequence: number; text: string; final: true; timestamp: number; } | { type: "response.plan"; sessionId: string; turnId: string; text: string; toolIntents: ToolIntent[]; sourceIds: string[]; timestamp: number; } | { type: "speech.chunk"; sessionId: string; turnId: string; audioReference: string; animationMarks: AnimationMark[]; timestamp: number; } | { type: "tool.result"; sessionId: string; turnId: string; toolName: string; status: "ok" | "error" | "denied"; timestamp: number; } | { type: "safety.event"; sessionId: string; turnId: string; ruleId: string; action: "block" | "rewrite" | "handoff"; timestamp: number; }; A shared session and turn identifier should follow the request through speech recognition, retrieval, model inference, tool execution, speech synthesis, and animation. Without that correlation key, a team may know that an individual service was slow but still be unable to explain why the user waited several seconds before hearing a response. A typed contract also makes it easier to replace one speech service, model provider, renderer, or transport without changing every downstream integration. Make Interruption a First-Class Control Path A conversational avatar must be designed for interruption before the team selects a model or animation system. In a sequential pipeline, the system performs each operation one after another: Wait for the user to stop speaking Finalize the transcriptRetrieve supporting contextGenerate the complete responseGenerate the complete audioStart playbackStart animation This architecture may work in a controlled demo, but it feels slow during real interaction. A streaming pipeline begins useful work as soon as enough evidence is available. Partial transcription can begin retrieval. The model can stream response segments. Text-to-Speech can synthesize approved segments incrementally. Animation cues can be produced alongside the audio. Web Real-Time Communication, commonly known as WebRTC, defines browser APIs for exchanging real-time media and application data between browsers or compatible devices, making it a practical transport option for interactive voice interfaces. [1] The transport alone does not solve responsiveness. Every stage must share the same cancellation model. When the user interrupts: Current audio playback should stopPending animation cues should be cancelledObsolete model generation should terminateUnnecessary retrieval should stopUnsafe or stale tool actions should not executeThe new user turn should become authoritative One approach is to associate an AbortController with every active turn: TypeScript class TurnCoordinator { private activeTurn?: AbortController; beginTurn(): AbortSignal { this.activeTurn?.abort("Superseded by a new user turn"); this.activeTurn = new AbortController(); return this.activeTurn.signal; } cancelActiveTurn(reason = "Session cancelled"): void { this.activeTurn?.abort(reason); this.activeTurn = undefined; } } Every asynchronous operation launched during the turn should receive the resulting signal. A cancellation implementation is incomplete when it stops only the visible audio. The underlying model request, retrieval job, animation queue, and pending tool operation must also respond to cancellation. Figure 2: An overlapping streaming pipeline from user audio through transcription, retrieval, generation, speech synthesis, and synchronized avatar output. Use Role Adapters Instead of Separate Applications The same runtime can support very different experiences when role-specific requirements are explicit. Stylized and Cartoon Avatar Characters For cartoon avatar characters, expression may communicate as much meaning as speech. A stylized-character adapter might: Convert emphasis into larger gesturesHold expressions for longer durationsUse fewer subtle facial movementsApply a gesture frequency limitMap emotional states to a smaller animation vocabularyEnforce audience-appropriate languageRestrict the character to an approved story or learning corpus The conversational runtime should not need to know how large a smile should be or how long an eyebrow movement should remain visible. It should emit an abstract expression instruction such as encouraging, surprised, or concerned. The embodiment adapter then maps that instruction to the available animation system. Robotic Embodiments A phrase such as “real avatar robot” usually describes a conversational character connected to a physical robotic body. Architecturally, this introduces a hard safety boundary. The model should never write directly to: Motor controllersNavigation systemsDoors or access controlsIndustrial equipmentDevice firmwareUnrestricted operating-system commands Instead, the model should propose a typed intent. A deterministic command broker should validate the intent before it reaches the certified device controller. For example: JSON { "intent": "MOVE_TO_LOCATION", "arguments": { "locationId": "reception-zone-a", "maximumSpeed": 0.4 }, "requestedBy": "session-1842" } The command broker can then check whether:The command is supportedThe requested location is permittedThe robot is currently availableA person or obstacle is in the pathThe speed is within the approved limitHuman approval is requiredAn emergency-stop condition exists The avatar may explain what the robot intends to do, but it should not claim that an action has happened until the control system confirms it. Enterprise Assistants The phrase “AI avatar business solutions” covers several possible applications, but the underlying enterprise requirements are more specific: Tenant isolationRole-based access controlApproved knowledge sourcesData-retention policiesAudit logsHuman escalationIdentity verificationTool authorizationRegional deployment controlsClear failure handling Retrieval authorization and tool authorization should be treated as separate decisions. A user may be allowed to read an internal policy document without being allowed to execute the action described inside it. Access to information is not automatically permission to change a customer record, issue a refund, modify an account, or trigger an external workflow. The response plan should therefore keep retrieved evidence and proposed actions separate. Entertainment Hosts The future of AI in entertainment will depend on more than visually impressive characters. Interactive systems must preserve narrative state while remaining interruptible, safe, and operationally predictable. An entertainment host may need to: Follow a show rundownReact to live audience inputRemain inside licensed character boundariesCoordinate with lighting, audio, or video cuesRecover when a media asset failsAvoid repeating a segmentRespect age and content restrictionsHand control back to a human operator This is closer to event-driven orchestration than a conventional chatbot loop. A show-state service can remain authoritative for what happens next, while the model generates language within the permitted scene and character constraints. Put a Deterministic Gateway in Front of Tools Prompt instructions alone are not a sufficient authorization system. Security guidance for large language model applications identifies risks such as prompt injection, insecure output handling, sensitive-information disclosure, insecure plugin or tool design, and excessive agency. These risks support keeping model output behind deterministic validation and authorization controls. [2] The model should propose a tool call. It should not decide whether that call is permitted. A gateway can evaluate the request using normal application-security controls: TypeScript async function authorizeToolIntent( intent: ToolIntent, context: SessionContext ): Promise<AuthorizationResult> { assertKnownTool(intent.name); validateToolSchema(intent.name, intent.arguments); requireScope( context.identity, requiredScopeFor(intent.name) ); enforceTenantBoundary( context.tenantId, intent.arguments ); enforceRateLimit( context.sessionId, intent.name ); enforceCostBudget( context.tenantId, intent.name ); if (isHighImpactOperation(intent)) { return { decision: "approval_required", reason: "Human approval is required for this operation" }; } return { decision: "allow", idempotencyKey: createStableIdempotencyKey( context, intent ) }; } State-changing tools should support idempotency. Model-driven systems can repeat actions because: A user rephrases the same requestAn orchestrator retries after a timeoutA network response arrives lateThe model proposes the same call twiceThe user interrupts during execution An idempotency key allows the execution service to return the existing result instead of repeating the operation. The tool gateway should also record: Requested actionAuthenticated identityTenantPolicy decisionApproval statusFinal execution resultRelated session and turnDuration and error state Figure 3: The model operates inside deterministic policy, tool-authorization, safety, execution, and observability boundaries. Observe the User Turn, Not Only Individual Services Traditional service metrics are still necessary, but avatar systems also need interaction-specific telemetry. OpenTelemetry provides a vendor-neutral framework for generating, collecting, and exporting traces, metrics, and logs. Its signal model can be used to correlate activity across the distributed components involved in one conversational turn. [3] Useful avatar-specific signals include: Time to First Audio Measure the time between the end of the user’s utterance and the first audible response. Break the trace into: Stylized and cartoon avatar charactersSpeech endpointingFinal transcriptionRetrievalFirst model tokenFirst approved response segmentFirst synthesized audioPlayback start Barge-In Cancellation Track whether a user interruption successfully cancelled: Audio playbackAnimationModel generationRetrievalPending tool actions A visible interruption that leaves an expensive model or tool request running is not a complete cancellation. Grounding Coverage Record which response statements were supported by approved sources. The response plan can attach source identifiers before the speech and animation layers receive the content. Tool Authorization Outcomes Record whether each proposed tool action was: AllowedDeniedRewrittenSent for approvalCancelledExecuted successfullyFailed Audio and Animation Drift Measure the difference between expected and displayed speech marks, visemes, gestures, and facial movements. Review median values as well as tail latency. A system can perform well during most interactions while still producing noticeable failures at the 95th or 99th percentile. Human-Handoff Reasons Use structured reasons such as: User requestLow confidenceMissing sourcePolicy restrictionAuthentication failureTool failureSafety boundaryUnsupported task These signals help teams distinguish model-quality problems from infrastructure, retrieval, authorization, or rendering failures. Test the Persona as Software Persona quality should be tested as a versioned software behavior rather than evaluated only through manual demonstrations. A useful test suite should include: Golden-Turn Tests Use fixed inputs with expected: Policy decisionsTool decisionsSource usageHandoff behaviorResponse constraintsAnimation categories Avoid asserting an exact sentence unless exact wording is a requirement. Instead, assert the properties the response must satisfy. Property Tests Examples include: The system never retrieves another tenant’s private sourceA tool cannot execute without the required scopeA robotic command cannot bypass the command brokerA restricted topic always triggers the defined policyA user interruption invalidates the previous turnAn entertainment role cannot access an unlicensed character profile Interruption Tests Trigger interruption during: Speech recognitionRetrievalModel generationSpeech synthesisAudio playbackAnimation playbackTool authorizationTool execution Verify which operations are cancellable and which must complete safely. Failure and Chaos Tests Introduce: Slow speech recognitionUnavailable retrievalDuplicate tool resultsExpired authenticationDelayed media packetsRenderer disconnectsPartial model outputText-to-Speech failureMissing animation marks The expected result should be a controlled fallback, not a broken or misleading character performance. Human Review Automated evaluation should be supplemented with human review for: Factual accuracyToneClarityAudience suitabilityEscalation qualityCharacter consistencyGesture appropriatenessMisleading capability claims Risk-management frameworks such as the NIST AI Risk Management Framework can provide a broader structure for incorporating trustworthiness considerations into the design, development, deployment, and evaluation of AI systems. F A Practical Migration Path Teams with an existing avatar application do not need to replace the entire stack at once. A gradual migration can follow these steps. Step 1: Trace the Existing Turn Record the path from user input to visible and audible output. Identify: Missing correlation identifiersUnmeasured latencyUncancelled operationsDirect model-to-tool connectionsRole-specific logic inside shared services Step 2: Introduce an Event Envelope Add shared session and turn identifiers across speech, retrieval, model, tool, and rendering services. This creates the foundation for tracing, replay, and cancellation. Step 3: Move Tools Behind a Gateway Introduce: Typed schemasIdentity checksRole and tenant scopesApproval statesIdempotencyAudit events Start with state-changing or high-impact tools. Step 4: Extract the Role Profile Move identity, knowledge boundaries, policy, tool permissions, voice, and embodiment constraints into versioned configuration. Keep operational security outside the profile. Step 5: Add a Second Embodiment Implement one additional character or device through a new adapter. This is the strongest architecture test. When the core runtime requires role-specific branching throughout the codebase, the separation is incomplete. Conclusion A production AI avatar should be treated as a distributed system with a face, not as a face attached to a prompt. The reusable asset is not one character model. It is the runtime that manages streaming sessions, retrieval, policy, tool authorization, safety, observability, testing, and interruption. Role profiles can then define identity and behavior, while embodiment adapters translate approved responses into the output required by a stylized character, robotic body, enterprise assistant, or entertainment host. This architecture does not remove the difficulty of building interactive characters. It puts the difficult parts in explicit boundaries where they can be measured, authorized, tested, and improved. That is what makes one conversational core capable of supporting multiple embodiments without multiplying the most sensitive and failure-prone parts of the system.

By Himanshu Gautam
Your Cloud Diagram Is Already Out of Date: An Operating Model for Continuous Security Architecture
Your Cloud Diagram Is Already Out of Date: An Operating Model for Continuous Security Architecture

The Diagram-to-Deployment Gap Cloud security architecture often begins with a strong design, account boundaries are defined, identity federation and vending patterns are selected, centralized security services are planned, and diagrams show how telemetry, governance, and incident response should work together. The design may be reviewed by experienced architects and approved by risk stakeholders. Yet the most difficult part starts after deployment. Cloud environments are not static. New accounts are created, workloads are modernized, emergency changes are made, teams adopt new services, and temporary exceptions accumulate. Over time, the deployed environment can diverge from the approved design even when no single team intends to weaken security. A diagram captures architectural intent while the running cloud environment represents operational reality. A mature security program must continuously reconcile the two. The central challenge is therefore not only whether an organization can design a secure cloud architecture. It is whether the architecture can continuously determine that its assumptions still hold. In practice, applying the operating model in this article reduced the time to detect architecture drift from weeks to hours - because divergence from intent is checked continuously against explicit invariants rather than discovered during periodic reviews or audits. I have watched this play out on several architectures I have helped design over the years. The architecture was approved, the launch was clean, and a couple of quarters later the environment no longer matched the diagram we had signed off on. No single team broke it. It drifted organically through dozens of individually reasonable decisions, each of which looked fine on its own review. This article presents a five-stage Continuous Security Architecture Loop — Define, Prevent, Observe, Validate, and Improve — for turning architecture from a one-time deliverable into an operating system for cloud assurance. Why Traditional Security Architecture Becomes Static Traditional architecture processes are strongest at design time. Teams conduct threat modeling, select controls, review trust boundaries, approve exceptions, and publish reference patterns. Once a workload is approved, responsibility often shifts to platform engineering, application teams, security operations, compliance, and audit. The handoff creates a structural gap where architecture defines the intended state, while operations manages the actual state. This gap becomes particularly visible in multi-stakeholder environments. A central team may define a standard account structure, but application teams control many day-to-day decisions inside each account. A security service may be enabled at launch but disconnected later. A centralized log destination may exist while selected workloads stop delivering critical events. A temporary administrative role may become a permanent access path. Each change may look like a configuration issue, yet the combined effect can break the original trust model. Configuration Drift vs. Architecture Drift Configuration drift and architecture drift are different problems. Configuration drift means a resource no longer matches an expected setting. Architecture drift means a security property of the system is no longer true. Logging may still be enabled while workload administrators can now alter the evidence. Encryption may still be present while key access violates the intended separation of duties. The resources look compliant, but the architecture has quietly lost an assumption it depended on. None of this means engineering teams are careless. Drift is what you get when architectural decisions are never wired to enforceable boundaries, current evidence, and operational feedback. Defining Continuous Security Architecture Continuous security architecture is an operating model that translates security principles into testable requirements, applies those requirements through preventive and detective mechanisms, validates deployed environments against architectural intent, and feeds operational findings back into future designs. It is related to continuous compliance, policy as code, infrastructure as code, cloud security posture management, and security monitoring, but it is not equivalent to any one of them. Continuous compliance asks whether a defined requirement is satisfied; continuous security architecture asks a broader question: does the implemented environment continue to preserve the trust assumptions, control objectives, and security outcomes of the approved architecture? This distinction matters because an architecture can satisfy many individual configuration checks and still fail as a system. A useful model must therefore connect four elements: the principle being protected, the architecture requirement derived from that principle, the controls that implement the requirement, and the evidence used to determine whether the outcome remains true. How continuous security architecture relates to adjacent practices: practicecore questionrelationship to this loop Policy as code Is this rule encoded and enforced? A mechanism used inside Prevent and Validate - not the model itself. Continuous compliance Is a defined requirement satisfied? A subset: answers per-control conformance, not whether the system’s trust assumptions still hold. CSPM Are resources misconfigured against a benchmark? Feeds the Observe stage; scores configurations, not architectural invariants. Continuous security architecture Do the architecture’s trust assumptions still hold in production? The superset - connects principle, requirement, control, and evidence across all five stages. The Continuous Security Architecture Loop The Continuous Security Architecture Loop contains five stages. The stages are not a maturity sequence that an organization completes once. They form a recurring operating cycle. Each stage produces information needed by the next, and the final stage feeds learning back into the beginning. Figure 1. The Continuous Security Architecture Loop connects architectural intent with prevention, evidence, validation, and operational learning. Running example throughout this section: We follow one invariant, “Security logs cannot be modified by workload administrators” (INV-LOG-001), through all five stages, so “invariant” stops being an abstract term and becomes a single property each stage acts on. 1. Define: Convert Principles Into Testable Requirements Security principles are often written as apply least privilege, centralize visibility, minimize blast radius, protect administrative access, or encrypt sensitive data. These are directionally correct, but they do not specify what evidence proves implementation, so different teams interpret the same principle differently. The Define stage converts broad principles into testable architecture requirements. Consider centralizing security visibility. A testable requirement might state that security-relevant activity from every account must be delivered to a centrally governed logging environment. That statement decomposes into measurable conditions: required audit sources are enabled, destinations are centrally controlled, workload teams cannot delete retained evidence, delivery failures generate alerts, newly created accounts are automatically enrolled, and retention supports investigation needs. You are not trying to turn every architecture document into a long checklist. Rather, you are identifying the properties that materially support the security model. Each requirement should describe the intended outcome, identify the evidence needed to validate it, specify an owner, and state whether the implementation is preventive, detective, responsive, or compensating. The strongest requirements are technology-aware without being tool-bound: state the security property first, then map it to the chosen cloud services, so the design survives a change of implementation technology. Express the requirement as a structured, adoptable artifact rather than prose: YAML # invariant: logs-immutable-by-workload id: INV-LOG-001 principle: Centralize security visibility requirement: > Security-relevant events from every account are delivered to a centrally governed logging destination that workload administrators cannot alter or delete. outcome: Log evidence remains complete and tamper-resistant for investigation. control_type: preventive + detective owners: control: platform-engineering evidence: platform-engineering risk: security-architecture evidence_sources: - cloudtrail: organization trail delivery status - config: S3 bucket policy + KMS key policy on the log destination - scp: effective policy on workload OUs validation_frequency: near-real-time # log-delivery failure is high-consequence exception_policy: allowed: false # no standing exceptions to this invariant residual_risk_target: none Define — running example: The principle centralize security visibility becomes the spec above (INV-LOG-001). The security property is stated first, and then the AWS services are mapped underneath it. Figure 2. A security principle becomes continuously testable only when it is connected to an architecture requirement, control implementation, and current validation evidence. 2. Prevent: Enforce High-Confidence Architectural Boundaries Some architectural decisions should not depend on detection after a violation occurs. Disabling central logging, moving data into an unapproved region, disconnecting an account from governance, or creating an unmanaged administrative path can undermine the security model immediately. Preventive controls make selected decisions non-optional. In a multi-account environment, platform teams can use organizational policies, account foundations, identity boundaries, deployment controls, and protected service configurations to prevent workload accounts from changing central audit destinations, leaving the organization, disabling designated security services, or creating resources in restricted locations. The goal is to protect the boundaries that preserve the architecture. A single service control policy (SCP) makes the boundary concrete-the action is simply unavailable to workload accounts: JSON { "Version": "2012-10-17", "Statement": [ { "Sid": "ProtectCentralLogDestination", "Effect": "Deny", "Action": [ "s3:DeleteBucket", "s3:PutBucketPolicy", "s3:PutEncryptionConfiguration", "s3:PutLifecycleConfiguration" ], "Resource": "arn:aws:s3:::org-central-security-logs*", "Condition": { "StringNotEquals": { "aws:PrincipalArn": "arn:aws:iam::*:role/PlatformLoggingAdmin" } } }, { "Sid": "PreventLeavingOrgAndDisablingAudit", "Effect": "Deny", "Action": [ "organizations:LeaveOrganization", "cloudtrail:StopLogging", "cloudtrail:DeleteTrail" ], "Resource": "*" } ] } The design challenge is selectivity. Excessive preventive control makes the platform brittle, blocks legitimate engineering work, and creates pressure that leads to bypasses. Not every deviation has the same consequence. A useful rule: prevent actions that would materially break the security architecture, and detect-and-remediate lower-risk deviations where flexibility is necessary. Only a small number of actions, like those above, warrant a hard Deny. Before enforcing a preventive control, evaluate service behavior, failure modes, exception needs, and recovery paths. A guardrail that cannot be safely changed during an incident introduces a different form of risk. Prevent — running example: The SCP above, applied to all workload OUs, denies StopLogging, DeleteTrail, and mutating actions on the central bucket for every principal except the platform logging role. INV-LOG-001’s boundary is now unavailable, not merely monitored. 3. Observe: Collect Evidence About the Deployed Architecture Architecture cannot be validated using resource configuration alone. The evidence layer may need configuration state, identity activity, network flows, deployment events, security findings, data-access records, control exceptions, and account-lifecycle events. Observation turns a running environment into evidence that can be compared with architectural intent. Consider an approved privileged-access model requiring federation, strong authentication, time-limited role assumption, and centralized activity logging. A configuration scan will happily confirm that the administrative role exists. What it will not show you is the engineer who skips that path entirely - using a long-lived key, an alternate role, or direct access from an unmanaged identity. That only shows up in behavioral evidence. So the questions worth asking are practical ones: does the property still hold, what proves it, how fresh is that proof, and who gets paged when it goes missing? Treat absence of telemetry as a finding in its own right — a gap in the logs is a gap in your ability to say anything true about that account. Centralization should not eliminate local ownership. Workload teams still need visibility into their findings and operational context, while the organization needs a protected evidence plane that cannot be selectively altered by the systems being observed. Observe — running example: For INV-LOG-001, the platform collects trail delivery status, bucket and key policy state, and every S3 mutation event against the log destination. Absent delivery telemetry is itself recorded as a finding. 4. Validate: Compare Deployed Reality With Architectural Intent Security programs often report findings as isolated events: one missing log source, one excessive role, one unmanaged connection, or one failed service enrollment. Continuous architecture validation asks whether those findings indicate that a larger architectural property is no longer true. Validation can be organized around architectural invariants which represent properties that must remain true regardless of application changes. Examples include: security logs cannot be modified by workload administrators; production identities originate only from approved sources; all production accounts inherit baseline governance; privileged access is time-bound and attributable; and external connectivity passes through approved control points. An invariant is machine-checkable. A query for INV-LOG-001, where a zero-row result is the passing state: MS SQL -- Did any non-platform principal touch the central log destination? SELECT eventtime, useridentity.arn AS principal, eventname, requestparameters FROM security_events WHERE eventsource = 's3.amazonaws.com' AND eventname IN ('PutBucketPolicy','DeleteBucket', 'PutEncryptionConfiguration','PutLifecycleConfiguration') AND element_at(requestparameters, 'bucketName') LIKE 'org-central-security-logs%' AND useridentity.arn NOT LIKE '%role/PlatformLoggingAdmin' AND eventtime > date_add('day', -1, now()); A returned row is an architecture-conformance failure rather than a lone configuration finding. Even when the SCP already blocked the action, the attempt itself is a signal the Improve stage should see. Each invariant should be connected to a validation record containing the expected state, evidence sources, current state, owner, validation frequency, exception status, residual risk, and remediation target. The record is not simply an audit artifact; it provides a shared language for architecture, platform, operations, and workload teams. A single failing record communicates the thesis better than a page of prose: fieldexample value Invariant Security logs cannot be modified by workload administrators (INV-LOG-001) Expected state Central log bucket + KMS key policies deny workload principals; org trail delivering Evidence sources Config rule s3-log-bucket-policy; org CloudTrail delivery status; effective SCP Current state FAIL — Account 1234 delivering, but bucket policy drifted to allow WorkloadAdmin Owner Platform engineering (control) / Security architecture (risk) Validation frequency Near-real-time Residual risk High until remediated — evidence integrity not guaranteed Remediation target 4 hours Validation frequency should reflect risk and rate of change. Public exposure, privileged access, and log-delivery failures may require near-real-time evaluation. Account ownership may be checked daily. Exception reviews may occur monthly or quarterly. Reference architectures should also be reviewed when new services, threats, incidents, or business models invalidate prior assumptions. Validate — running example: A daily and event-driven check finds Account 1234’s bucket policy has drifted to grant a workload role write access. It is recorded as an architecture-conformance failure against INV-LOG-001 - with owner, residual risk, and a 4-hour remediation target - not as a standalone S3 finding. 5. Improve: Feed Operational Learning Back Into Architecture Security incidents and recurring findings are frequently remediated at the workload level. The immediate problem is fixed, but the reusable architecture remains unchanged, allowing the same design weakness to appear in other environments. The Improve stage treats operational events as feedback about the architecture itself. Incidents, near misses, recurring findings, control bypasses, exception patterns, deployment failures, new threat intelligence, and cloud-service changes should all influence future design decisions. Suppose an incident reveals that a workload role could redirect security logs. Correcting the role is necessary but not sufficient. The organization should ask whether the reference architecture clearly separates log ownership, whether organizational controls should protect the destination, whether other accounts share the condition, whether new validation logic is required, and whether the account-vending process should be updated. A practical rule: every significant incident should produce both a workload-level corrective action and an architecture-level learning decision—a revised reference pattern, a new preventive control, an additional validation test, clearer ownership, an improved deployment template, or an explicit acceptance of residual risk. I remember a stretch where three newly vended accounts failed the same detection-enrollment step inside a single month. Fixing them one at a time felt like progress until the pattern became impossible to ignore. It revealed that the account vending pipeline was the defect, not the accounts. Once we corrected the pipeline, the failure class stopped appearing. Improve — running example: Root cause of the Account 1234 drift: the account-vending template applied the bucket policy once but did not protect it against later edits. The fix is not just repairing Account 1234 - it is (a) adding the mutating actions to the SCP, (b) adding a Config rule to catch policy drift, and (c) updating the vending template so future accounts start protected. The invariant, not the incident, drives the change. From Control Deployment to Control Effectiveness A control being deployed does not prove that it is effective. Logging may be enabled while important events are excluded. Encryption may be enabled while key access is broader than intended. Threat detection may be active while findings have no response owner. Backup policies may exist while restoration has never been tested. Organizational guardrails may be present while alternative paths bypass the intended restriction. Continuous assurance therefore needs more than a binary deployed/not-deployed status. A practical evaluation can examine five dimensions: dimensionquestion it answers Coverage How much of the intended environment is protected. Correctness Whether the implementation matches the requirement. Resilience Whether the control can be bypassed, altered, or disabled. Responsiveness Whether failure produces timely action. Outcome Whether the control measurably reduces the intended risk. These dimensions should not be collapsed into a universal score without context. A high coverage percentage can conceal a critical gap, while a small number of exceptions may carry disproportionate risk. The architecture team should define what effective means for each important control objective and how that effectiveness will be demonstrated. The specific dimension I have seen fail most quietly is Resilience. A control can be present, correct, and even alerting, and a workload role can still disable it or route around it. Coverage and Correctness are the easy numbers to put on a slide, while Resilience is the one that decides whether those numbers mean anything. Central Governance With Distributed Ownership No central team can manage every workload configuration, and no workload team can set enterprise-wide requirements on its own. The shared responsibility that works in practice is that architecture owns the principles, reference patterns, invariants, and the hard exception calls. Platform turns those into the account foundations, guardrails, enrollment workflows, and evidence collection everyone else inherits. Workload teams own what is specific to their application — the controls, the context behind a finding, and their own exceptions. Security operations watches for threats and control failures and feeds what it learns back into the architecture. Risk and compliance tie the evidence to obligations and to whatever risk has been formally accepted. The model must distinguish among control ownership, evidence ownership, and risk ownership. The platform team may operate centralized logging, while the workload owner remains accountable for producing the application events needed for investigation. Security operations may own an alerting process, while architecture owns the invariant that the process is meant to protect. Ambiguity at these boundaries is a common cause of unaddressed findings. Figure 3. Continuous assurance requires explicit coordination among architecture, platform, security operations, and workload teams. End-to-End Scenario: Creating a New Production Account Consider the creation of a new production account. The organization defines several requirements: the account must join the production governance hierarchy, use approved identity federation, deliver security logs to a protected destination, enroll in centralized detection, and restrict deployment to approved regions. During Prevent, organizational policies enforce these boundaries with concrete mechanisms. SCPs on the production OU deny organizations:LeaveOrganization, deny mutating actions on the central log bucket, and deny non-approved regions via an aws:RequestedRegion condition. Account Factory (or an equivalent vending pipeline) provisions the account directly into the production OU so the guardrails apply from creation, not after. During Observe, the security platform collects the literal events: CloudTrail CreateAccount and the account’s move into the OU, the AWS Config recorder status, GuardDuty/Security Hub enrollment state, identity-federation configuration, and organization-trail delivery status. Validation compares the account with the production architecture invariants. Suppose the account is delivering logs but has not enrolled in centralized detection. The check returns a failing record: YAML invariant: all-prod-accounts-inherit-baseline-governance (INV-GOV-002) account: 1234 expected_state: GuardDuty + Security Hub enrolled via delegated admin current_state: FAIL - Security Hub not enrolled (Config recorder ON, trail OK) owner: platform-engineering (control) / security-architecture (risk) residual_risk: medium - threat findings not aggregated for this account remediation_by: 24h This is recorded as an architecture-conformance issue rather than an isolated configuration finding. The record names the account owner, the missing evidence, the remediation target, and any approved exception. The Improve stage examines patterns across accounts. If several new accounts fail at the same enrollment step, the organization does not continue fixing them individually. It treats the repeated failure as evidence that the account-vending or enrollment architecture is incomplete, and the reusable process is corrected so future accounts begin in the expected state. This example illustrates the essential shift where continuous architecture is not a larger collection of controls; it is a system that connects design intent, platform implementation, operational evidence, validation, and learning. Common Anti-Patterns Architecture by diagram. There is a beautiful target-state picture on the wiki, and no way to answer the only question that matters: does the running environment still look like it?Guardrail accumulation. Controls pile up over years. Nobody removes them, nobody re-checks whether they still fire, and eventually the platform is so encrusted that engineers route around it - which is its own risk.Dashboard assurance. The board is green, so leadership feels safe - even though the dashboard only measures the handful of configurations someone remembered to wire up.Permanent “temporary” exceptions. The exception was granted for two weeks in 2023. It has no expiry, no owner, no compensating control, and no evidence the original risk still exists. It is now load-bearing.Finding-by-finding remediation. Teams close tickets faster than the architecture produces them, treating each symptom as new while the weakness that generates them stays untouched.Tool-defined architecture. The security model quietly shrinks to whatever the chosen product happens to detect. The tool should serve an architecture you defined independently - not the other way around. Adopting the Loop Incrementally Continuous security architecture does not require building all five stages at once. A workable sequence can be: Start with 3-5 invariants, not a catalog. Pick the properties whose failure would most damage the trust model, such as log integrity, production identity origin, privileged-access time-bounding, network egress control, baseline governance inheritance. Write each as a spec like INV-LOG-001.Validate before you observe everything. You do not need a complete evidence lake to begin. For each invariant, identify the single signal that proves it and check that. Breadth of telemetry comes later.Prevent only the highest-consequence actions first. A small set of well-chosen Deny guardrails, like leaving the org, disabling logging, or altering the log destination, protects more than a large, brittle policy set. Add preventive controls where a violation is irreversible; detect-and-remediate everywhere else.Wire the feedback loop early, even if manual. A monthly review that turns recurring findings into architecture changes delivers most of the Improve-stage value before any automation exists.Automate by consequence and rate of change. Move the highest-risk, fastest-changing invariants to near-real-time checks first. Slower-moving properties can stay on a daily or weekly cadence. A team can reach a useful state with five invariants, a handful of SCPs, one validation query per invariant, and a recurring review, and then expand coverage as the model proves its value. This is also how the weeks-to-hours improvement in drift detection is realized in practice: near-real-time validation of a few high-consequence invariants, not full automation on day one. When I have taken teams through this, the first few invariants we picked mattered far more than the tooling around them. Choose the properties whose failure would genuinely hurt, prove those, and resist the pull to boil the ocean on day one. Conclusion: Architecture as an Operating System A reference architecture captures intended security design. Continuous security architecture is how you find out whether that intent still holds as systems, teams, threats, and cloud services change. The loop gives you a practical model: Define testable requirements, Prevent the actions that would break critical boundaries, Observe the evidence you need to understand the environment, Validate reality against intent, and Improve the design from what operations teaches you. The value of a security architecture is not determined by the quality of its diagram. It is determined by how reliably the organization preserves its security assumptions in production and how quickly it learns when those assumptions no longer hold. Applied to a multi-account environment, this model cut architecture-drift detection from weeks to hours, turning drift from a condition found in periodic reviews into one that is continuously observed.

By Avik Mukherjee
The Telemetry Tax: Architecting Zero-Allocation Event Observability at 15B+ Daily Event Scale
The Telemetry Tax: Architecting Zero-Allocation Event Observability at 15B+ Daily Event Scale

Throughout my career, observability and monitoring systems have played a very important role in system stability, availability, and scalability when architecting, whether a system handles millions or billions of events per day. Having architected core infrastructure across communication platforms, enterprise compliance pipelines, and high-throughput transaction systems, I have repeatedly seen how background instrumentation can silently become a bottleneck under heavy production load. Telemetry and observability come with their own challenges. Instrumenting telemetry without degrading live business application performance, increasing request latency, or inflating cloud costs is often deprioritized initially but becomes very important as the system scales. Furthermore, it is a much harder problem to solve than processing domain events themselves in production. When an event-driven platform scales past 15 billion events per day, peaking between 200,000 and 400,000 TPS (transactions per second), the infrastructure costs and performance overhead of logging and tracing libraries create a severe operational bottleneck known as the Telemetry Tax. This article explains why addressing this telemetry tax is important and also provides you with brief insight into how you can start handling it with architectural changes at the initial level. The Multiplier Effect In a very scalable, high-throughput microservices architecture, business transactions rarely execute as isolated operations. An incoming event arriving at an API gateway typically passes through five to ten downstream services, caching layers, database instances, queues, message brokers (e.g., Kafka), etc. If every microservice hop generates three telemetry events in addition to the core business logic events, assuming two metric updates and a single log line, the resulting metric data scales exponentially: Regular workload (Core business Logic events): 15,000,000,000 events per dayMetrics and logs: 100,000,000,000 signals per dayStorage estimation: 15+ TBs of uncompressed JSON payloads daily Without specialized memory management and proper architecture, this massive volume of background instrumentation introduces critical failures in live production systems. Example: Garbage collection: Short-lived heap allocation for JSON log stringsThread Lock Contention: Sync metric exporters hold mutex locks while completing the request to send telemetry metrics, adding 10-20ms of latencyCascading downstream failures: Failure in downstream telemetry collector or failure in processing existing logs and metrics causes issues or backlog in upstream transaction workers, which may result in queue back-pressure and API timeouts. Engineering a Zero-Allocation Telemetry Pipeline To eliminate garbage collection pauses and thread lock delays, high-throughput microservices must decouple metric and log generation from core business logic, utilizing a middleware concept with pre-allocated memory pools and asynchronous buffers. Under heavy execution loads, standard string concatenations from logging and JSON marshaling for telemetry signals escape to the heap during Go's escape analysis. When millions of goroutines allocate short-lived objects on the heap per second, Go's runtime is forced into frequent concurrent sweep phases, consuming 15-20% of available CPU purely on garbage collection. By utilizing sync.Pool, we pre-allocate backing byte arrays that survive across the request lifecycle. When a worker goroutine finishes formatting a payload, the buffer pointer is reset (buf[:0]) and returned to the pool without invoking runtime.newobject . This keeps memory allocation in the critical execution path at zero bytes. Go package telemetry import ( "context" "sync" "go.opentelemetry.io/otel" "go.opentelemetry.io/otel/attribute" "go.opentelemetry.io/otel/trace" ) var tracer = otel.Tracer("zero-alloc-telemetry") // BufferPool pre-allocates memory var bufferPool = sync.Pool{ New: func() interface{} { b := make([]byte, 0, 1024) return &b }, } type RingBufferPipeline struct { telemetryChannel chan *trace.Span } func NewPipeline(bufferSize int) *RingBufferPipeline { return &RingBufferPipeline{ telemetryChannel: make(chan trace.Span, bufferSize), } } func (p *RingBufferPipeline) InstrumentEvent(ctx context.Context, eventID string) { // Retrieve pre-allocated byte slice from pool bufPtr := bufferPool.Get().(*[]byte) buf := (*bufPtr)[:0] defer func() { *bufPtr = buf bufferPool.Put(bufPtr) }() // Non-blocking trace span initialization ctx, span := tracer.Start(ctx, "ExecuteTransaction", trace.WithAttributes(attribute.String("event.id", eventID))) defer span.End() buf = append(buf, []byte("event_processed:")...) buf = append(buf, eventID...) // Async non-blocking dispatch select { case p.telemetryChannel <- span: default: // Add metric here to protect API response SLOs } } How to Control Ingestion Costs With Tail-based Adaptive Sampling Collecting 100% of telemetry traces across 100+ billion daily events leads to unsustainable tool ingestion costs (Regardless of the tool you use, e.g., Datadog, OpenTelemetry, Splunk). Standard Head-based sampling (deciding whether to keep a trace at the start of the request) drops fatal error traces while keeping millions of redundant HTTP 200 success traces. By deploying OpenTelemetry collectors configured with Tail-based adaptive sampling, trace spans are held in a 500 ms sliding memory buffer prior to routing: HTTP 200 / Successful Executions: Sampled at 0.1% to maintain baseline latency metricsHTTP 5xx / Latency exceptions (>200ms) / Server Errors: Retained at 100% for debugging, or any other analysis Production Performance Benchmark (Approx.) Refactoring telemetry infrastructure from sync logging to zero-allocation ring buffers and tail-based adaptive sampling helps with substantial performance improvements. (Note: The numbers in the table below are rough estimations based on high-scale modeling and past operational experience. Actual metrics may vary depending on your system and other aspects of architecture) Performance MetricSync TelemtryZero-allocation Adaptive TelemetryIngestion Throughput35k events/sec400k events/secp99 API response latency200ms40msAverage Monthly Cost$35-45k+$15-20kCPU overhead15-20% total CPU time spent on GC2-3% total CPU time spent on GC These estimations illustrate that the telemetry tax is not an inevitable cost of scale, but a consequence of applying synchronous, allocation-heavy architectural patterns to hyper-scale architectures. By refactoring memory management at the application level and applying adaptive filtering at the collector layer, engineering teams can achieve deep operational visibility while protecting system performance and cloud infrastructure budgets. Key Takeaways for System Architects Decouple the path: Never allow telemetry exporters to execute synchronously on the core business logic pathPre-allocate memory: Use memory buffer pools to eliminate garbage collection pauses during high-throughput event processingSample at the tail: Evaluate trace retention based on the execution outcome rather than making static decisions at the request ingress.

By Brindal Patel
The Silent Container Death: A TCP Dial That Never Times Out
The Silent Container Death: A TCP Dial That Never Times Out

A pod goes into CrashLoopBackOff. You pull the logs expecting a stack trace, a panic, an error string - anything that points you somewhere. Instead, you get one line: Plain Text Loading config... And then nothing. No error. No exit message. The container is just gone, and a few seconds later it’s back, prints the exact same line, and disappears again. Magic. This is the story of chasing that silence to its root cause. TCP connection that was never going to succeed, and never going to fail either. At least not on any timescale a Kubernetes health check was willing to wait for. The Setup We were migrating backend services from a legacy message queue to Kafka. The new consumers ran side by side with the old ones in a “shadow mode.” That let us compare behavior before the real cutover. Part of that work meant pointing a dev environment at a managed Kafka cluster. (Think AWS MSK — the specifics don’t matter here.) We also updated the broker endpoint in config. In shadow mode, we send to both old and new queues, but only one of them processes the message. The other queue infrastructure just logs what it receives. The change looked trivial: swap one connection string for another, restart the pods, watch them come up. Instead, every pod that touched Kafka went straight into a crash loop. The only clue was that single “Loading config” line. Repeated forever. Why “No Error” Is the Error The instinct when a service crashes is to look for what it logged right before dying. Here that instinct is a trap. The absence of any further log output isn’t a hint — it’s the symptom itself. Two things had to be true simultaneously for this to happen: Something blocked the process long enough that Kubernetes’ health checks gave up on it and sent SIGKILL.Whatever the process wanted to log about being blocked never made it out of its internal buffers before the kill. That second point matters more than it looks. Go’s standard logger writes to os.Stdout. How a container runtime attaches to that stream determines whether output appears immediately or sits in a buffer. Buffering is common under load, or when the write target isn’t a real TTY. Consider a process blocked inside a library call, say dialing a broker. It never gets back to the point in its code where it would flush or print the next line. SIGKILL doesn’t give a process the chance to clean up. Whatever was sitting in a buffer is gone. From the outside, a service that’s actually deep in a hung network call looks identical to one that exited silently. Both just print “Loading config” and stop. The lesson here generalizes past Kafka. If a container’s logs stop dead with no error and no clean shutdown message, assume a hang-then-kill. Not a fast crash. Until proven otherwise. Reaching for the Network Layer Once “look at the application logs” stopped being useful, the next step was to get underneath the application entirely. Shelling into a node and watching the raw traffic (tcpdump) tells you what actually happened at the OS level. So does tracing the process’s syscalls with strace. Neither depends on whether the application ever got to log anything about it. What that showed: a TCP handshake that started and never finished. A SYN packet went out toward the broker; no SYN-ACK ever came back, and critically, no RST came back either. That distinction is the whole story. Connection refused is fast and loud. The remote host, or a firewall in front of it, actively sends back an RST packet. Your client’s connect() call fails almost immediately.Connection blackholed is slow and silent. Packets go out, and nothing comes back. The OS has no way to know if the remote end is down, unreachable, or just very far away. So it retransmits the SYN a few times with exponential backoff, then gives up. The kernel’s default TCP connect timeout can be well over a minute. In this case, the broker endpoint we’d configured was a private, VPC-internal address. It was reachable from some parts of the network, but not from the specific node group these pods landed on. No security group or routing rule was actively rejecting the connection – the packets were simply going nowhere. That’s the worst kind of network failure to debug from inside an application. Everything about it looks like the process is just slow, right up until it isn’t. Where the Health Check Made Things Worse None of this would have been quite so opaque if the failure had surfaced immediately. But the service’s startup path connected to Kafka before reporting itself healthy. On top of that, the Kubernetes startup probe carried a generous timeout, meant to avoid flapping on slow boots. That combination left the platform with no opinion about what was wrong. It just saw a container that hadn’t become healthy in time. So it did the only thing it can do here: kill it and try again. The pod restart count climbed. The backoff delay between restarts grew too - Kubernetes doubles it after repeated failures, up to roughly five minutes. Every fix we tried afterward seemed to take forever to take effect. That’s because we were still watching a container that hadn’t actually restarted yet. It was just waiting out its backoff window. Deleting the pod outright forced an immediate restart. That turned out to be the fastest way to test each hypothesis, rather than waiting for the backoff timer. The Fix, and the More Useful Part The actual fix was almost anticlimactic: switch to the broker’s public endpoint. In a real production setup, you’d instead fix the VPC routing or peering. That makes the private endpoint reachable from every node group that needs it. Once the TCP path was real, the connection succeeded instantly, and the crash loop stopped. The useful part isn’t the fix. It’s the checklist that could have saved us time: Handy Checklist Treat “one log line then silence” as a hang, not a crash. A clean crash logs an error. A silent one usually means something upstream killed the process mid-blocking-call.Go to the network layer early, not last. tcpdump or strace on the affected node will show you a stuck SYN in seconds. That’s far faster than adding print statements and waiting through several crash-loop cycles.Know the difference between “refused” and “blackholed” in your bones. An RST means someone answered and said no. Check credentials, ports, and application-level config. Silence means the packet never arrived. Check routing, VPC peering, security groups, and whether you’re using the right endpoint for the network you’re actually in.Set explicit, short connect timeouts in your client libraries. Don’t let a startup path inherit the OS’s default TCP connect timeout. The OS optimizes that default for general robustness, not for failing fast during a health check window.Make sure your logger flushes before anything that can block indefinitely. If a call to an external system can hang, log “attempting to connect to X” first. Then make sure that line is actually out the door, synchronously if necessary, before making the call. It costs you nothing when the call succeeds and saves you hours when it doesn’t.When you’re mid-debug, delete the pod instead of waiting out the backoff. Kubernetes’ exponential backoff on repeated CrashLoopBackOff restarts is helpful in production and actively annoying when you’re iterating on a fix. None of this is exotic — it’s TCP fundamentals and container basics that everyone technically knows. What makes it worth writing down is how convincingly a blackholed connection disguises itself as an application bug. Right up until you stop looking at the application and start looking at the wire.

By Alexander Fo
Six Degrees of Ayrton Senna: Learn Neo4j by Connecting 75 Years of Formula 1
Six Degrees of Ayrton Senna: Learn Neo4j by Connecting 75 Years of Formula 1

One of my favorite things about F1 racing is the data behind it. F1 cars are the most complex and advanced in any racing series. They collect huge amounts of telemetry data. The tracks also gather data during events, and race engineers analyze it week after week. They study everything from weather and tire temperatures to corner exit speeds. Data drives the sport forward in a major way. While learning about graph databases and Neo4j, I realized it was the perfect tool for answering a question I was curious about. We've all seen or heard of "Six Degrees of Kevin Bacon," where nearly any actor can be traced back to the Footloose star. I wondered: could this work for F1 drivers? By comparison, it's a much smaller dataset than famous actors. Photos via Wikimedia Commons, licensed under CC BY‑SA 4.0. Is Max Verstappen connected to Juan Manuel Fangio? Could a driver who retired in 1958, decades before Max was born, connect to him through a chain of teammates? And if so, how many links does it take? This post is how I answered that, and it doubles as a gentle introduction to Neo4j and graph databases. By the end, you'll have built a real graph of every F1 driver since 1950 on your own machine, and you'll run a query that answers my Verstappen-to-Fangio question in a single line. No prior graph experience needed. Let's get into it. Why This Is a Graph Problem Before we start: My question isn't really about drivers. It's about the connections between drivers. If all I wanted was a list of drivers, or each driver's win count, or how many races happened at Monza, a plain old relational table handles that beautifully. Even a spreadsheet can do it. Spreadsheets are wonderful at facts about things. Where they start to sweat is questions about relationships and chains of relationships. Specifically, long relationship chains. Think about what "is Verstappen connected to Fangio?" actually requires. You don't know in advance whether the answer is three hops or nine. So in SQL you'd be writing a recursive common table expression that joins a results table to itself, over and over, to a depth you can't predict, while trying not to drown in duplicate paths. I tried to do this very thing and locked up the application trying. It's possible to do queries like this, but they rarely run fast, if they run at all. Relational databases weren't designed for things like this. A graph database flips the whole thing around. Instead of storing drivers in one table and hoping to reconstruct their connections later with joins, it stores the connections themselves as useful entities. That's the one idea underneath everything else in this post: In a graph database, the relationships are first class data. They're not something you compute at query time. They're something you store, traverse, and count directly. That single design choice is what turns this tough question into a one-liner. Let me show you the model before we build it. The Property Graph Model, in Four Pieces Neo4j uses the Labeled Property Graph model. It sounds fancy; it's just four building blocks. I'll introduce each one using our F1 data. Nodes are the things in your domain. The entities. For us, that's drivers and teams. In a diagram, you draw them as circles. Ayrton Senna is a node. McLaren is a node. Labels are the type of a node, written with a colon: :Driver, :Constructor. (Constructor is just F1's official word for "team".) Labels are how Neo4j knows a Senna node is a driver and a McLaren node is a team. By convention, they're written in PascalCase. Relationships are the connections between nodes, and this is where graphs shine. Every relationship has a type in SCREAMING_SNAKE_CASE, a direction, and a start and end node. Senna DROVE_FOR McLaren is a relationship. Crucially, that connection is stored in the database. Neo4j keeps a pointer from one node to the next. This is why hopping across relationships stays fast even when your graph gets huge. The cost of following one relationship remains small, whether your database has a thousand nodes or a billion. Properties are key-value pairs you can hang on either a node or a relationship. A :Driver node has forename: 'Ayrton', surname: 'Senna', nationality: 'Brazilian'. Here's something that might be surprising if you come from a table world: a relationship can carry properties too. Our DROVE_FOR relationship will carry season: 1988, because which season someone drove for a team is a fact about the connection, not about the driver or the team on their own. That last point is worth thinking about, because it was an "aha" moment for relational-to-graph thinking. Senna drove for McLaren, but when he did, it doesn't belong to Senna and doesn't belong to McLaren. It belongs to the link between them. Put it on the relationship, and a whole category of modeling headaches evaporates. Here's our entire starting model: Two circles, one labeled arrow between them. If you want to make that concrete right now, open arrows.neo4jlabs.com (a free browser diagramming tool) and draw it. Click to make a node, give it a label and some properties, drag from its edge to a second node to create the relationship. It's helpful to sketch your schema there before writing any code. Design for the Question You Want to Ask Before building anything, I did the single most useful thing you can do when modeling a graph: I wrote down the question I actually wanted to answer first, and let it drive every decision afterward. My question: "How are two drivers connected through shared teammates?" This tells me exactly what my graph needs. It needs drivers. It needs some notion of two drivers being teammates. Everything else is optional scaffolding. But notice the dataset doesn't hand me "teammate" directly. It gives me who drove for which team in which season. Two drivers are teammates when they drove for the same team in the same season. So my plan has a nice shape to it: Load Driver and Constructor nodes.Connect them with DROVE_FOR relationships (one per driver, per team, per season).Derive a brand-new TEAMMATE_OF relationship between any two drivers who share a team and season.Walk the TEAMMATE_OF web to answer my question. In step three, we are creating relationships that weren't in the raw data, by reasoning about the graph you already have. This is one of my favorite things about working in Neo4j, and you'll see why shortly. Let's build. What You'll Need Neo4j Desktop – free, from neo4j.com/download. I'm on the current version (Desktop 2.x) on a Mac; Windows and Linux are the same journey.The dataset – the "Formula 1 World Championship (1950–2024)" dataset by Rohan Rao on Kaggle. Free Kaggle account, one download, clean CSVs.An afternoon. Realistically, a couple of hours, most of it spent going "oh that's cool" at the results.No prior Cypher. I'll explain every query as we go. Cypher is Neo4j's query language, and it's genuinely readable. If you already know SQL, you'll be nodding along within minutes. A quick note on the data: For about a decade, the go-to source for F1 data was the Ergast API. It shut down at the end of 2024. The Kaggle dataset we're using preserves Ergast's exact structure as downloadable CSVs, and if you later want live current-season data, the community-run Jolpica-F1 API (api.jolpi.ca/ergast/f1/) serves the same schema. Build against the CSVs today; top up from Jolpica whenever you like. Nothing in this tutorial changes. Once Neo4j Desktop is installed, under local instances, click create instance to create a new local instance. Give it a name and a password you'll remember, and start it. Then open the Query tool (in Desktop 2.x this is where you run Cypher — it's the modern replacement for what older tutorials call "Neo4j Browser"). That's our workbench. We need four files from the Kaggle download: drivers.csvconstructors.csvraces.csvresults.csv We can stage these files for import by placing them in our imports folder. Select your local instance, and look for the ... button. Select Open then Instance folder: This is the folder for your Neo4j instance. Next, select the import folder. This is where you want to place the files. macOS: Plain Text /Users/[your username]/neo4j-community-2026.x/import Windows: Plain Text C:\Neo4j\import Linux: Plain Text /var/lib/neo4j/import Now that the files are in the import folder, we can access them with Cypher later. Step 1: Constraints First Before loading a single row, I created uniqueness constraints. A constraint guarantees you'll never accidentally create two copies of the same driver. Neo4j automatically builds an index behind each one, which makes all the lookups during import dramatically faster. In Neo4j Desktop, you can run queries by selecting Query from the left-hand panel and entering your queries in the window in the upper right. Here's the query to create the constraints: Cypher CREATE CONSTRAINT driver_id IF NOT EXISTS FOR (d:Driver) REQUIRE d.driverId IS UNIQUE; CREATE CONSTRAINT constructor_id IF NOT EXISTS FOR (c:Constructor) REQUIRE c.constructorId IS UNIQUE; CREATE CONSTRAINT race_id IF NOT EXISTS FOR (r:Race) REQUIRE r.raceId IS UNIQUE; Run SHOW CONSTRAINTS to confirm all three landed. That's your data-integrity seatbelt fastened. Step 2: Load the Nodes Now we bring in the entities. LOAD CSV reads a file row by row; MERGE is Cypher's "create this if it doesn't already exist, otherwise match the existing one" command. Drivers: Cypher LOAD CSV WITH HEADERS FROM 'https://raw.githubusercontent.com/JeremyMorgan/Six-Degrees-Senna/main/import/drivers.csv' AS row MERGE (d:Driver {driverId: toInteger(row.driverId)}) SET d.forename = row.forename, d.surname = row.surname, d.fullName = row.forename + ' ' + row.surname, d.nationality = row.nationality, d.dob = CASE WHEN row.dob <> '\\N' THEN date(row.dob) END; That CASE WHEN row.dob <> '\\N' is guarding against a quirk you'll hit constantly with this dataset: missing values are stored as the literal text \N. If you don't filter them out, you'll end up with drivers whose birthday is the string "backslash-N", which is exactly as useful as it sounds. Consider that your first real-world data-cleaning lesson, delivered by Formula 1. Constructors: Cypher LOAD CSV WITH HEADERS FROM 'https://raw.githubusercontent.com/JeremyMorgan/Six-Degrees-Senna/main/import/constructors.csv' AS row MERGE (c:Constructor {constructorId: toInteger(row.constructorId)}) SET c.name = row.name, c.nationality = row.nationality; Races (we mostly need these to know which season a result belongs to, but full Race nodes cost nothing and set you up for future projects): Cypher LOAD CSV WITH HEADERS FROM 'https://raw.githubusercontent.com/JeremyMorgan/Six-Degrees-Senna/main/import/races.csv' AS row MERGE (r:Race {raceId: toInteger(row.raceId)}) SET r.year = toInteger(row.year), r.round = toInteger(row.round), r.name = row.name, r.date = date(row.date); Quick check: you should see something in the ballpark of 861 drivers, 212 constructors, and 1,100-plus races: Cypher MATCH (n) RETURN labels(n)[0] AS label, count(*) ORDER BY label; If those numbers look right, you've just loaded three-quarters of a century of motorsport into a graph. We haven't done anything clever yet, but we're about to. Step 3: Connect Drivers to Teams The results.csv file has one row per driver, per race — around 26,000 rows. I don't want 26,000 relationships cluttering my graph. I want one clean fact per driver, per team, per season: this person drove for this team that year. So I aggregate as I load. Cypher :auto LOAD CSV WITH HEADERS FROM 'https://raw.githubusercontent.com/JeremyMorgan/Six-Degrees-Senna/main/import/results.csv' AS row CALL (row) { MATCH (d:Driver {driverId: toInteger(row.driverId)}) MATCH (c:Constructor {constructorId: toInteger(row.constructorId)}) MATCH (r:Race {raceId: toInteger(row.raceId)}) MERGE (d)-[s:DROVE_FOR {season: r.year, constructorId: c.constructorId}]->(c) ON CREATE SET s.entries = 1, s.wins = CASE WHEN row.positionOrder = '1' THEN 1 ELSE 0 END ON MATCH SET s.entries = s.entries + 1, s.wins = s.wins + CASE WHEN row.positionOrder = '1' THEN 1 ELSE 0 END } IN TRANSACTIONS OF 2000 ROWS; There's a lot of learning packed into that one statement, so let's unpack it: MERGE (d)-[:DROVE_FOR {season, constructorId}]->(c) is the key move. The first time we see Senna-at-McLaren-in-1988, this creates the relationship. Every subsequent race that season just finds the existing one and bumps its counters. Twenty-six thousand rows collapse into roughly 3,600 clean driver-season-team facts. This is the graph-modeling principle in action: store relationships at the granularity you plan to query.ON CREATE / ON MATCH let you do one thing when the relationship is brand new and a different thing when it already exists — here, initialize the counters versus increment them.:auto and IN TRANSACTIONS OF 2000 ROWS tell Neo4j to commit the import in batches rather than one giant transaction, which keeps memory happy on a big file. Version note: CALL (row) { … } is the modern syntax (Neo4j 5.23 and up) for passing a variable into a subquery. If your Neo4j is older, you'll get a syntax error on that line — use the legacy form CALL { WITH row … } instead. And don't copy an abbreviated snippet with ... in the middle into the Query editor; Cypher will try to parse the dots. Use the full block above. Let's make sure it worked by looking at a career I know: Cypher MATCH (d:Driver {surname:'Senna', forename:'Ayrton'})-[s:DROVE_FOR]->(c:Constructor) RETURN c.name AS team, s.season AS season, s.wins AS wins ORDER BY season; Toleman in '84, Lotus '85 to '87, McLaren '88 to '93, Williams in '94. If that's what you see, your graph is alive and correct. Step 4: Derive the Teammate Network Everything so far was set up. This is the payoff of graph thinking. Nowhere in the data does it say "Senna and Prost were teammates." But we can derive that: two drivers are teammates if they each have a DROVE_FOR relationship to the same constructor with the same season. And in Cypher, describing that pattern is close to describing it in English: Cypher MATCH (d1:Driver)-[r1:DROVE_FOR]->(c:Constructor)<-[r2:DROVE_FOR]-(d2:Driver) WHERE r1.season = r2.season AND d1.driverId < d2.driverId MERGE (d1)-[t:TEAMMATE_OF {season: r1.season, team: c.name}]->(d2); Read that MATCH line like a picture: driver one points to a constructor, and driver two points to the same constructor from the other side. The WHERE says "same season." And then we MERGE a shiny new TEAMMATE_OF relationship between them. We just created around 10,000 relationships that didn't exist in the source data, purely by reasoning about the shape of the graph. Two small things worth understanding: d1.driverId < d2.driverId stops us creating each pairing twice (Senna→Prost and Prost→Senna). By only linking the lower ID to the higher one, each pair gets a single relationship. When we query it, we'll just ignore direction — because "teammate" goes both ways, and Cypher happily traverses a relationship in either direction when you leave the arrowhead off.A deliberately imperfect definition, and why I'm keeping it. "Same team, same season" isn't exactly "raced side by side." Midseason driver swaps mean, for example, that Senna and David Coulthard both count as 1994 Williams drivers. Coulthard was Senna's replacement, and they never actually raced as teammates. I could tighten this up by deriving teammate links per-race instead of per-season. I'm keeping the looser version on purpose, for two reasons. First, it makes the network richer and more connected across eras, which is the whole point. Second, every graph model is an argument about what a relationship means. There's no universally correct answer; there's only the definition that serves your question. Naming that trade-off out loud is important. Step 5: Answer the Question Here it is. The reason I built the whole thing. Two drivers separated by half a century, and one line of Cypher to connect them: Cypher MATCH (max:Driver {surname:'Verstappen', forename:'Max'}), (fangio:Driver {surname:'Fangio'}) MATCH p = shortestPath((max)-[:TEAMMATE_OF*]-(fangio)) RETURN p; That * after TEAMMATE_OF is the star of the show. It means "follow this relationship any number of times" — a variable-length path. shortestPath then finds the tightest chain of teammate links between the two drivers. This is the exact query that would've been a page of recursive SQL. In Cypher, it fits on a napkin. When you run it, the Query tool draws the answer as a chain of driver nodes, each link labeled with the team and season that connects them. This is how close Max really is to Fangio. Once that lands, you'll want to push further. Here are the three queries I couldn't stop running. Every driver's "Senna number" — like the Bacon number, but for F1. How many teammate-hops is each driver from Ayrton Senna? Cypher MATCH (senna:Driver {surname:'Senna', forename:'Ayrton'}) MATCH (d:Driver) WHERE d <> senna MATCH p = shortestPath((senna)-[:TEAMMATE_OF*..25]-(d)) RETURN length(p) AS sennaNumber, count(d) AS drivers ORDER BY sennaNumber; The shape of those results is the real insight: almost the entire history of the sport sits within a handful of hops of Senna. That's a "small-world network," demonstrated with race cars. The most connected drivers in history — the human bridges holding the whole web together: Cypher MATCH (d:Driver)-[:TEAMMATE_OF]-(other:Driver) RETURN d.fullName AS driver, count(DISTINCT other) AS teammates ORDER BY teammates DESC LIMIT 10; Watch who tops this list: long-career journeymen and team-hoppers, not necessarily the champions. Connectedness rewards longevity and movement, not podiums. I find that interesting. The people stitching F1's social fabric together are often not the ones holding the trophies. Is it really all one network? Are there isolated islands of drivers? Cypher MATCH (d:Driver) WHERE NOT (d)-[:TEAMMATE_OF]-() RETURN count(d) AS unconnectedDrivers; A small handful of true loners from F1's chaotic early days, and then one enormous connected web containing basically everyone else. Seventy-five years, one family. What You Just Learned (It Wasn't Really About F1) The property graph model – nodes, labels, relationships, and properties, including the quietly powerful idea that a relationship can carry properties of its own.Designing for questions, not entities – writing the question first and letting it shape the model.LOAD CSV, MERGE, and constraints – the everyday mechanics of getting real data into Neo4j cleanly.Deriving new relationships – creating structure that wasn't in your source data by reasoning about the graph you already have.Variable-length paths and shortestPath – the thing graphs do effortlessly and relational databases do through gritted teeth. And here's the part that matters beyond motorsport: swap the dataset and every one of these skills transfers directly. The teammate network is structurally identical to a fraud ring, a supply chain, a social graph, an org chart, or the knowledge graph behind an AI application. "Who is connected to whom, and how?" is one of the most valuable questions in software, and you now know how to ask it. Download the code here. Where to Go Next If this clicked for you, the best next move is to get the fundamentals properly, in order. That's exactly what GraphAcademy is for. It's Neo4j's free, hands-on learning platform. For future articles, I'm thinking: turn this same graph into a fair fight, deriving "who-beat-whom" links between teammates and running an algorithm called PageRank to settle the greatest-of-all-time argument without ever touching the points table.

By Jeremy Morgan

Culture and Methodologies

Agile

Career Development

Methodologies

Team Management

Build Software Faster With Three Simple Principles

October 1, 2026 by Ilia Ivankin

Can Your Team Name the Work It Already Runs With AI?

September 25, 2026 by Stefan Wolpers DZone Core CORE

One Agent, Two Runtimes: Defining State Ownership Between Temporal and LangGraph

September 25, 2026 by Akhil Madineni DZone Core CORE

Data Engineering

AI/ML

Big Data

Databases

IoT

Docker Sandboxes Beyond the Laptop: Running AI Agents in the Cloud

October 2, 2026 by Naga Santhosh Reddy Vootukuri DZone Core CORE

How a 30B Model Runs on Your Laptop

October 1, 2026 by Akash Lomas

Embabel vs LangGraph4j: Two Agentic Philosophies for Investment and Risk Analysis in BFSI

October 1, 2026 by Soham Sengupta

Software Design and Architecture

Cloud Architecture

Integration

Microservices

Performance

Docker Sandboxes Beyond the Laptop: Running AI Agents in the Cloud

October 2, 2026 by Naga Santhosh Reddy Vootukuri DZone Core CORE

Gossips on Cryptography: Part 4

October 1, 2026 by Sahil Aggarwal

Build Software Faster With Three Simple Principles

October 1, 2026 by Ilia Ivankin

Coding

Frameworks

Java

JavaScript

Languages

Tools

Docker Sandboxes Beyond the Laptop: Running AI Agents in the Cloud

October 2, 2026 by Naga Santhosh Reddy Vootukuri DZone Core CORE

Embabel vs LangGraph4j: Two Agentic Philosophies for Investment and Risk Analysis in BFSI

October 1, 2026 by Soham Sengupta

How Go Maps Work: From Buckets to Swiss Tables

October 1, 2026 by Ilia Ivankin

Testing, Deployment, and Maintenance

Deployment

DevOps and CI/CD

Maintenance

Monitoring and Observability

Docker Sandboxes Beyond the Laptop: Running AI Agents in the Cloud

October 2, 2026 by Naga Santhosh Reddy Vootukuri DZone Core CORE

The Telemetry Tax: Architecting Zero-Allocation Event Observability at 15B+ Daily Event Scale

September 30, 2026 by Brindal Patel

The Silent Container Death: A TCP Dial That Never Times Out

September 30, 2026 by Alexander Fo

Popular

AI/ML

Java

JavaScript

Open Source

How a 30B Model Runs on Your Laptop

October 1, 2026 by Akash Lomas

Embabel vs LangGraph4j: Two Agentic Philosophies for Investment and Risk Analysis in BFSI

October 1, 2026 by Soham Sengupta

Are Passphrases Still Secure in the Age of AI?

October 1, 2026 by Constantin Kwiatkowski

  • RSS
  • X
  • Facebook

ABOUT US

  • About DZone
  • Support and feedback
  • Community research

ADVERTISE

  • Advertise with DZone

CONTRIBUTE ON DZONE

  • Article Submission Guidelines
  • Become a Contributor
  • Core Program
  • Visit the Writers' Zone

LEGAL

  • Terms of Service
  • Privacy Policy

CONTACT US

  • 3343 Perimeter Hill Drive
  • Suite 215
  • Nashville, TN 37211
  • [email protected]

Let's be friends:

  • RSS
  • X
  • Facebook
×