Building Time-Series Applications With Java and InfluxDB
Agentic Test Creation: From Plain-Language Requirements to End-to-End Test Cases
Code Review Core Practices
Getting Started With DevSecOps
In a recent DZone article, Running Sentiment Analysis Inside Neo4j With a Java Plugin, we explored several approaches to running sentiment analysis inside the Neo4j database engine. One of those approaches — embedding a Wasm runtime inside a Java UDF — was described like this: Theoretically, we could embed a Wasm runtime such as wasmtime inside a Java UDF and execute the VADER Wasm module from within Neo4j, getting Wasm's sandbox guarantees inside Neo4j's plugin model. It's technically feasible but no published working example appears to exist and the complexity cost is high relative to the alternatives. An interesting idea to watch, but not practical today. This article builds that working example. We'll show how to embed a real VADER sentiment analyzer compiled to Wasm inside a Neo4j Java UDF, returning a full polarity score map callable directly from Cypher. We'll cover the tools and inspection techniques needed to understand what the Wasm compiler generates and why the Java calling convention looks the way it does. The full source code is available on GitHub. What We're Building We're embedding a wasmtime Wasm runtime inside a Neo4j Java UDF using wasmtime-java, a community JNI binding for the Wasmtime runtime. It's not an official Bytecode Alliance product, but it ships prebuilt native libraries for all major platforms and is sufficient for this proof-of-concept. A Rust function compiled to WebAssembly rides inside the plugin JAR alongside the Java code. When Cypher calls the UDF, Java initializes the Wasm runtime, loads the binary, and invokes the Rust function — all inside the Neo4j JVM process with no external API calls and no network round-trips. Note: This article was tested specifically against wasmtime-java 0.19.0. The API used here is version-specific; newer releases or alternative JVM Wasm runtimes may expose different interfaces and calling conventions. Prerequisites You'll need the following installed if you wish to follow along. We're using Apple Silicon (ARM64) as our development platform, so we'll note where the setup differs from other platforms. Java We're using OpenJDK 21 (tested with 21.0.12.1). Install it using your platform's package manager or download it directly from adoptium.net. On macOS via Homebrew: Shell brew install openjdk@21 On Ubuntu/Debian: Shell sudo apt install openjdk-21-jdk On Windows, download and run the installer from Adoptium. Confirm your Java version: Shell java -version You should see a Java 21 runtime. If you're on Apple Silicon, also confirm you're running a native ARM64 JVM with: Shell uname -m You should see arm64. Not running under ARM64 will likely break the wasmtime-java JNI library loading. Maven We're using Maven 3.9.6. On Apple Silicon, be cautious about installing Maven via Homebrew as, at the time of writing, the Homebrew Maven formula pulls in OpenJDK 26 as a dependency, which conflicts with a Java 21 installation. If your package manager installs an incompatible JDK alongside Maven, verify the runtime with mvn -version and configure JAVA_HOME as necessary. Installing Maven manually is the safest approach: Shell cd ~ curl -O https://archive.apache.org/dist/maven/maven-3/3.9.6/binaries/apache-maven-3.9.6-bin.tar.gz tar xzf apache-maven-3.9.6-bin.tar.gz Then add Maven to your PATH and make it persist across terminal sessions: Shell echo 'export PATH="$HOME/apache-maven-3.9.6/bin:$PATH"' >> ~/.zshrc source ~/.zshrc On Linux, add the same line to ~/.bashrc instead.On Windows, download the zip from maven.apache.org and add the bin folder to your system PATH via System Properties. Confirm Maven is using Java 21: Shell mvn -version You should see Java version: 21 in the output. Rust We're using Rust 1.96.0. Install via rustup if not already present: Shell curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh On Windows, download and run rustup-init.exe from rustup.rs. To pin to the specific Rust version we tested with: Shell rustup toolchain install 1.96.0 rustup default 1.96.0 Then add the WASI target: Shell rustup target add wasm32-wasip1 This target works identically across macOS, Linux, and Windows. WABT The WebAssembly Binary Toolkit gives us wasm-objdump for inspecting Wasm binaries. We tested with version 1.0.41. On macOS: Shell brew install wabt On Ubuntu/Debian: Shell sudo apt install wabt On Windows, download the latest release from github.com/WebAssembly/wabt/releases. wit-bindgen This is the interface types generator for WebAssembly. There are two distinct version numbers to be aware of: the wit-bindgen-cli command-line tool and the wit-bindgen Rust crate used as a dependency inside the Wasm module. These can differ. We tested with CLI version 0.59.0 and Rust crate version 0.40.0 (specified in Cargo.toml). The generated binary identifies the crate version through the export name cabi_realloc_wit_bindgen_0_40_0. Install the pinned CLI version via Cargo on all platforms: Shell cargo install wit-bindgen-cli --version 0.59.0 Confirm it's installed: Shell wit-bindgen --version Neo4j Desktop We're using Neo4j Desktop with a local database instance. Download from Neo4j for Desktop. The pom.xml in this article is pinned to Neo4j 2026.07.0 — update the neo4j.version property to match your own Desktop installation. wasmtime-java Platform Support The wasmtime-java library ships prebuilt JNI native libraries for: macOS aarch64macOS x86_64Linux aarch64Linux x86_64Windows x86_64 No additional setup is needed, as Maven pulls the correct native library for your platform automatically. Version Summary For reference, here are all the component versions used in this article: ComponentVersionOpenJDK21.0.12.1Maven3.9.6Rust1.96.0WABT1.0.41wit-bindgen CLI0.59.0wit-bindgen crate0.40.0vader_sentiment crate0.1.1wasmtime-java0.19.0Neo4j2026.07.0 Getting the Code Clone the repository before following along. All source files are provided so you don't need to create them manually. Shell cd ~ git clone --filter=blob:none --sparse https://github.com/VeryFatBoy/neo4j.git cd neo4j git sparse-checkout set wasm-udf mv wasm-udf ../wasm-udf cd ../wasm-udf Project Structure Before creating any files, here's the final layout we're building toward. There are two separate projects: A Rust crate that compiles to Wasm.A Maven project that hosts the Neo4j UDF. First, the Rust crate: Plain Text sentimentable/ ├── Cargo.toml ├── src/ │ └── lib.rs └── wit/ └── sentimentable.wit Second, the Maven project: Plain Text neo4j-wasm-udf/ ├── pom.xml └── src/ └── main/ ├── java/ │ └── com/example/ │ ├── WasmUDF.java │ └── SentimentUDF.java └── resources/ ├── add.wat ├── add.wasm └── sentimentable.wasm The Wasm binaries in resources/ are bundled into the plugin JAR at build time. The Rust crate and Maven project are kept separate, and the Wasm binary is the handoff point between them. The project layout is also shown in Figure 1. Figure 1. Two-Project Layout How the Wasm Plumbing Works Before diving into the code, it's worth understanding the three layers that make this possible. Core Wasm and WASI WebAssembly defines a portable binary format and a stack-based execution model. On its own, it only understands numbers, such as integers and floats. When a Wasm module needs system capabilities, like memory allocation or I/O, it uses WASI (WebAssembly System Interface), a standardized set of system calls that a host runtime implements. Our Rust code targets wasm32-wasip1, which means it compiles to Wasm with WASI preview 1 system calls. The wasmtime runtime implements those calls on the host side. wasmtime-java This library wraps the wasmtime Wasm runtime in a JNI binding, making it callable from Java. It ships prebuilt native libraries for all major platforms, so adding it as a Maven dependency is all that's needed — no separate wasmtime installation required. The Java API lets us load a Wasm binary, set up a WASI context, and call exported functions directly. wit-bindgen and the String ABI Core WebAssembly functions operate on Wasm value types such as integers and floats. WIT (WebAssembly Interface Types) and the Component Model provide higher-level interface types such as strings, tuples, and records; wit-bindgen generates the lowering and lifting code needed to represent those types at the Wasm boundary. For strings, it uses a pointer-and-length convention: the caller allocates memory inside the Wasm module using a generated cabi_realloc function, writes the string bytes there and passes the memory address and byte length as two integers. The Rust code reads the string from that address. For return values, the lowering strategy depends on the type, which we'll see when we inspect the generated binary. With those three pieces in place, the calling chain looks like this: Plain Text Cypher query -> Neo4j routes to @UserFunction -> Java initializes wasmtime engine + WASI context -> Java allocates string in Wasm memory -> Java calls exported Wasm function -> Rust executes VADER scoring -> Java reads result from Wasm memory -> Java returns Map<String, Double> to Neo4j -> Neo4j returns result to Cypher Graphically, the calling chain is also shown in Figure 2. Figure 2. Calling Chain We built up to this through two simpler stepping-stone cases: Case 1: A trivial integer addition to prove the chain works.Case 2: A single compound score to introduce string passing and WASI. The full walkthrough of both, including the WasmUDF.java implementation is in a technical report on the GitHub repo. The Maven Project XML <project xmlns="http://maven.apache.org/POM/4.0.0" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://maven.apache.org/POM/4.0.0 http://maven.apache.org/xsd/maven-4.0.0.xsd"> <modelVersion>4.0.0</modelVersion> <groupId>com.example</groupId> <artifactId>neo4j-wasm-udf</artifactId> <version>1.0-SNAPSHOT</version> <packaging>jar</packaging> <properties> <maven.compiler.source>21</maven.compiler.source> <maven.compiler.target>21</maven.compiler.target> <project.build.sourceEncoding>UTF-8</project.build.sourceEncoding> <neo4j.version>2026.07.0</neo4j.version> </properties> <dependencies> <dependency> <groupId>org.neo4j</groupId> <artifactId>neo4j</artifactId> <version>${neo4j.version}</version> <scope>provided</scope> </dependency> <dependency> <groupId>io.github.kawamuray.wasmtime</groupId> <artifactId>wasmtime-java</artifactId> <version>0.19.0</version> </dependency> </dependencies> <build> <plugins> <plugin> <artifactId>maven-compiler-plugin</artifactId> <configuration> <source>21</source> <target>21</target> </configuration> </plugin> <plugin> <groupId>org.apache.maven.plugins</groupId> <artifactId>maven-shade-plugin</artifactId> <version>3.5.1</version> <executions> <execution> <phase>package</phase> <goals><goal>shade</goal></goals> <configuration> <artifactSet> <excludes> <exclude>org.neo4j:*</exclude> </excludes> </artifactSet> <shadedArtifactAttached>false</shadedArtifactAttached> </configuration> </execution> </executions> </plugin> </plugins> </build> </project> Two things worth noting here: org.neo4j:neo4j is declared as provided scope — Neo4j is already present in the database JVM at runtime, so we exclude it from the bundled JAR.We use maven-shade-plugin rather than maven-jar-plugin to produce a fat JAR that bundles wasmtime-java and its native libraries alongside our code. Update the neo4j.version property to match your own Neo4j Desktop installation. Case 3: Full Polarity Map VADER produces four scores: compound, positive, negative, and neutral. In this case, we update the Rust function to return all four and the Java UDF to return them as a Map<String, Double> — matching the return shape of the Java VADER UDF from the previous article. The sentimentable.wit File We change the return type from a single f32 to a tuple of four f32 values: Plain Text package local:sentimentable; world sentimentable { export sentimentable: func(input: string) -> tuple<f32, f32, f32, f32>; } We use a tuple rather than a named record. Both would work, but a tuple is simpler on the Java side — we read four consecutive f32 values from memory at known offsets without needing to decode field names. The lib.rs File Rust wit_bindgen::generate!({ world: "sentimentable", }); struct Component; impl Guest for Component { fn sentimentable(input: String) -> (f32, f32, f32, f32) { lazy_static::lazy_static! { static ref ANALYZER: vader_sentiment::SentimentIntensityAnalyzer<'static> = vader_sentiment::SentimentIntensityAnalyzer::new(); } let scores = ANALYZER.polarity_scores(input.as_str()); ( *scores.get("compound").unwrap_or(&0.0) as f32, *scores.get("pos").unwrap_or(&0.0) as f32, *scores.get("neg").unwrap_or(&0.0) as f32, *scores.get("neu").unwrap_or(&0.0) as f32, ) } } export!(Component); Build: Shell cd ~/wasm-udf/sentimentable cargo build --target wasm32-wasip1 --release Inspecting the Binary Let's check the exports first: Shell wasm-objdump -x target/wasm32-wasip1/release/sentimentable.wasm | grep "^Export" -A 6 Four exports should be present: memory, sentimentable, cabi_realloc, and cabi_realloc_wit_bindgen_0_40_0. Now let's find the type signature of the sentimentable function. Find the sig index: Shell wasm-objdump -x target/wasm32-wasip1/release/sentimentable.wasm | grep "func\[9\]" | head -1 Then look it up: Shell wasm-objdump -x target/wasm32-wasip1/release/sentimentable.wasm | grep "type\[9\]" You should see: Plain Text - type[9] (i32, i32) -> i32 The signature is (i32, i32) -> i32 . This is the key difference between returning a single scalar and returning a tuple: wit-bindgen uses a direct f32 return for a single value, but switches to an indirect result pointer when returning a tuple. What's written at that pointer is four f32 values (16 bytes) at consecutive 4-byte offsets. The Java side reads all four. This illustrates an important distinction between the WIT interface definition and the generated core Wasm ABI. The WIT signature and the Wasm-level signature are different layers: wit-bindgen lowers WIT types to a core Wasm ABI, and the lowering strategy depends on the return type. A single scalar such as f32 is returned directly as a Wasm value. A tuple is returned indirectly through linear memory, with the caller receiving a pointer to where the values were written. The Java calling code must match the generated ABI rather than the WIT definition, which is why inspecting the binary with wasm-objdump before writing the Java wrapper is essential. Figure 3 shows the memory layout. Figure 3. Memory Layout. Figure 4 compares Cases 2 and 3. Figure 4. Case 2 vs. Case 3 ABI Comparison The SentimentUDF.java file Java package com.example; import io.github.kawamuray.wasmtime.Engine; import io.github.kawamuray.wasmtime.Func; import io.github.kawamuray.wasmtime.Linker; import io.github.kawamuray.wasmtime.Memory; import io.github.kawamuray.wasmtime.Module; import io.github.kawamuray.wasmtime.Store; import io.github.kawamuray.wasmtime.WasmFunctions; import io.github.kawamuray.wasmtime.WasmValType; import io.github.kawamuray.wasmtime.wasi.WasiCtx; import io.github.kawamuray.wasmtime.wasi.WasiCtxBuilder; import org.neo4j.procedure.Description; import org.neo4j.procedure.Name; import org.neo4j.procedure.UserFunction; import java.io.InputStream; import java.nio.ByteBuffer; import java.nio.ByteOrder; import java.nio.charset.StandardCharsets; import java.util.HashMap; import java.util.Map; public class SentimentUDF { @UserFunction("com.example.wasm.sentiment") @Description("Scores text using VADER sentiment analysis compiled to Wasm. Returns compound, positive, negative, neutral.") public Map<String, Double> sentiment(@Name("text") String text) throws Exception { if (text == null || text.isBlank()) { return Map.of("compound", 0.0, "positive", 0.0, "negative", 0.0, "neutral", 1.0); } byte[] wasmBytes; try (InputStream is = SentimentUDF.class.getResourceAsStream("/sentimentable.wasm")) { if (is == null) throw new RuntimeException("sentimentable.wasm not found in resources"); wasmBytes = is.readAllBytes(); } WasiCtx wasi = new WasiCtxBuilder().inheritStdout().inheritStderr().build(); try (Store<Void> store = Store.withoutData(wasi); Engine engine = store.engine(); Module module = Module.fromBinary(engine, wasmBytes); Linker linker = new Linker(engine)) { WasiCtx.addToLinker(linker); linker.module(store, "", module); Memory memory = linker.get(store, "", "memory").get().memory(); Func reallocFn = linker.get(store, "", "cabi_realloc").get().func(); WasmFunctions.Function4<Integer, Integer, Integer, Integer, Integer> realloc = WasmFunctions.func(store, reallocFn, WasmValType.I32, WasmValType.I32, WasmValType.I32, WasmValType.I32, WasmValType.I32); byte[] inputBytes = text.getBytes(StandardCharsets.UTF_8); int len = inputBytes.length; int strPtr = realloc.call(0, 0, 1, len); ByteBuffer buf = memory.buffer(store); buf.position(strPtr); buf.put(inputBytes); Func sentimentFn = linker.get(store, "", "sentimentable").get().func(); WasmFunctions.Function2<Integer, Integer, Integer> scoreFn = WasmFunctions.func(store, sentimentFn, WasmValType.I32, WasmValType.I32, WasmValType.I32); int resultPtr = scoreFn.call(strPtr, len); // read four f32 values at 4-byte offsets: compound, pos, neg, neu ByteBuffer resultBuf = memory.buffer(store); resultBuf.order(ByteOrder.LITTLE_ENDIAN); float compound = resultBuf.getFloat(resultPtr); float positive = resultBuf.getFloat(resultPtr + 4); float negative = resultBuf.getFloat(resultPtr + 8); float neutral = resultBuf.getFloat(resultPtr + 12); Map<String, Double> result = new HashMap<>(); result.put("compound", (double) compound); result.put("positive", (double) positive); result.put("negative", (double) negative); result.put("neutral", (double) neutral); return result; } } } The return type is (Map<String, Double> ), the null guard returning a neutral map and the four getFloat() reads at consecutive 4-byte offsets from the result pointer. Build and Deploy Copy the Wasm binary, build and deploy: Shell cp ~/wasm-udf/sentimentable/target/wasm32-wasip1/release/sentimentable.wasm \ ~/wasm-udf/neo4j-wasm-udf/src/main/resources/ cd ~/wasm-udf/neo4j-wasm-udf mvn -q clean package cp target/neo4j-wasm-udf-1.0-SNAPSHOT.jar \ ~/Library/Application\ Support/neo4j-desktop/Application/Data/dbmss/<your-dbms-id>/plugins/ Stop Neo4j, restart it, and run the verification queries. Positive sentence: Cypher RETURN com.example.wasm.sentiment('The movie was great') AS scores; Result: JSON { "compound": 0.624893307685852, "positive": 0.577464759349823, "negative": 0.0, "neutral": 0.4225352108478546 } Capitalization test: Cypher RETURN com.example.wasm.sentiment('The movie was GREAT!') AS scores; Result: JSON { "compound": 0.7290259003639221, "positive": 0.6307692527770996, "negative": 0.0, "neutral": 0.3692307770252228 } Empty string guard: Cypher RETURN com.example.wasm.sentiment('') AS scores; Result: JSON { "compound": 0.0, "positive": 0.0, "negative": 0.0, "neutral": 1.0 } All three cases behave correctly. Summary We set out to build the working example that our previous article said didn't exist. Here's what we showed. We embedded a real VADER sentiment analyzer compiled to Wasm inside a Neo4j Java UDF, returning a full polarity map — compound, positive, negative and neutral — matching the return shape of the Java VADER UDF from the previous article. The wit-bindgen tuple ABI writes four f32 values to consecutive memory addresses; the Java side reads them back with a LITTLE_ENDIAN ByteBuffer. All four scores are correct, capitalization sensitivity works, and the empty string guard returns a sensible neutral map. The result is a workable integration pattern rather than a universal replacement for a native Java implementation. With the per-call initialization used in this proof of concept, the approach is best suited to low-frequency workloads where Wasm isolation and portability justify the additional complexity. For high-throughput workloads, the natural next step is to benchmark and reuse the Wasmtime engine and compiled module while keeping execution state appropriately isolated between calls. In the next article, we'll look at running TypeSafe AI's Jev inside Neo4j for calibrated sentiment decisions. Stay tuned! The full source code is available on GitHub.
Time is now a critical dimension in enterprise data. Modern applications generate a constant stream of events, such as payments, sensor readings, infrastructure metrics, user activity, logistics movements, market data, and application telemetry. The value exists not only in what happened, but also when it occurred, what came before it, and how the data changes over time. While traditional relational and NoSQL databases can store this information, increasing data volume and frequency make it essential to efficiently query recent states, historical windows, trends, and ordered events. Time-series databases fulfill these needs by treating time as a primary dimension. They make it easy to retrieve the latest measurements, examine historical intervals, aggregate data over time, and handle high-frequency writes. For enterprise applications, time-series databases now support not only monitoring and IoT, but also observability, financial systems, industrial platforms, logistics, energy, and any architecture where tracking data evolution is as important as the data itself. Why Time Series Databases Matter At first glance, time-series data appears similar to data stored in relational or NoSQL databases. SQL tables can include timestamp columns, document databases can store dated events, and key-value stores can use time-based keys. The distinction appears when time becomes the primary access pattern rather than a secondary attribute. If typical queries include “what is the latest value?”, “what happened during this interval?”, or “how did this metric change over time?”, a time-series model is often a better fit. Relational databases can support these queries, but they frequently require complex indexes, partitions, retention policies, and aggregation logic as data volume increases. Document and key-value databases typically require the application to manage temporal structure. Time-series databases are designed for append-heavy workloads, ordered data, time-window queries, downsampling, retention, and time-based aggregation. They do not replace SQL or general-purpose NoSQL databases, but serve as a specialized solution when tracking data changes over time is essential. A useful way to think about the difference is this: Plain Text Relational database: "What is the current state of this entity?" Document database: "What does this aggregate/document look like?" Key-Value database: "What value is associated with this key?" Time Series database: "How has this value changed over time, and what is happening now?" Consider a temperature sensor. In a relational model, we could create a table like: SQL CREATE TABLE sensor_reading ( timestamp TIMESTAMP, sensor VARCHAR(100), temperature DOUBLE ); This approach is functional, but typical workloads often require queries such as: Plain Text Give me the latest value for sensor-01. Give me the last 100 measurements. Give me the average temperature every five minutes. Show me the period where the temperature exceeded 30°C. Compare today's measurements with yesterday's. Delete or compress measurements older than six months. These are not isolated queries; they represent the typical data lifecycle. Time-series databases are specifically designed to support this pattern. Beyond IoT IoT is a clear example, as sensors naturally generate timestamped measurements. However, this model is common across many enterprise systems. Observability and infrastructure monitoring are classic use cases. Measurements such as CPU usage, memory consumption, request latency, error rates, queue depth, database connections, and network throughput are sampled repeatedly over time. Rather than focusing on a single point, teams typically seek trends, spikes, rolling averages, anomalies, and interval comparisons. Financial systems also generate highly temporal data. Market prices, exchange rates, trades, account activity, risk measurements, and portfolio valuations are all chronological streams. Trading applications may require the latest price, recent ticks, or price ranges within specific windows. Banking systems regularly retain transactional data in relational databases while sending temporal metrics and activity streams to time-series stores for analysis. Logistics and transportation also benefit from this approach. Vehicle position, speed, fuel consumption, delivery status, warehouse temperature, and route telemetry are all continuously changing metrics. Relevant queries include: Plain Text Where was the vehicle during the last hour? How has fuel consumption changed during this route? When did the delivery temperature leave the acceptable range? What was the most recent location reported by each vehicle? Energy and utilities are also inherently time-oriented. Smart meters, electricity consumption, voltage, water usage, solar generation, battery charge, and grid load all generate continuous measurements. While billing may remain in a relational system, raw measurements and their aggregation suit time-series databases well. Application and business metrics further demonstrate that time-series data goes beyond infrastructure. Enterprise applications may record: orders per minutepayments approved per minutefailed logins per houractive userscheckout latencyinventory changesfraud scoresmessages processedAPI calls by customer These are business signals, not simply technical telemetry. Persisting them over time enables organizations to understand both system and business processes as they evolve. Event-driven and distributed architectures present another important use case. Microservices continuously emit events and measurements. While event stores and message brokers are effective for transporting and preserving events, they are not optimized for queries such as: Plain Text What was the average latency for this service during the last 30 minutes? What is the latest measurement from every region? How has throughput changed since the deployment? Which customers experienced the highest error rate over the last hour? A time-series database can complement Kafka, Pulsar, or other event backbones by providing a queryable temporal view of these streams. This distinction is important: selecting a time-series database does not require replacing relational databases, document stores, or event brokers. Modern enterprise architectures are increasingly polyglot. Relational databases may remain the system of record, document databases may manage flexible aggregates, key-value databases may provide fast lookups, and time-series databases may handle continuously changing measurements. The architectural advantage lies in choosing a data model that corresponds to the application's key questions. Java and Time Series Traditionally, using time-series databases in Java required working with database-specific drivers and APIs. Developers needed to learn unique programming models, configuration styles, query APIs, and object-mapping strategies for each database. This increased implementation complexity and made maintenance challenging, especially when multiple time-series technologies were used. Eclipse JNoSQL 1.1.18 addresses this complexity by introducing Time Series as a supported database type. This release adds support for InfluxDB 3, Apache IoTDB, and QuestDB, along with a new Jakarta NoSQL Template specialization: TimeSeriesTemplate. Java @Inject private TimeSeriesTemplate template; The mapping model remains consistent. Time-series entities continue to use Jakarta NoSQL annotations: Java @Entity public class SensorReading { @Id private Instant timestamp; @Column private String sensor; @Column private double value; // constructors, getters, and setters } As with other Jakarta NoSQL database types, the primary difference lies in the data model. In time-series models, the identifier usually represents the temporal dimension. In this example, @Id is an Instant marking when the measurement occurred. The programming model remains familiar: Java @Inject private TimeSeriesTemplate template; SensorReading reading = new SensorReading( Instant.now(), "sensor-1", 21.5); template.insert(reading); Optional<SensorReading> result = template.find(SensorReading.class, reading.getTimestamp()); List<SensorReading> readings = template.select(SensorReading.class) .where("sensor").eq("sensor-1") .result(); The same model applies to Jakarta Data repositories: Java @Repository public interface SensorReadingRepository extends BasicRepository<SensorReading, Instant> { List<SensorReading> findBySensor(String sensor); } Then the repository can be injected normally: Java @Inject private SensorReadingRepository repository; The key improvement is not Java’s ability to connect to time-series databases, as native drivers have always enabled this. The real value is that developers can now use a consistent Jakarta NoSQL and Jakarta Data programming model across different time-series databases, reducing database-specific code within applications. Conclusion Time-series databases matter because many modern systems are no longer only about storing the current state of data, but about understanding how that data changes over time. From observability and finance to logistics, energy, IoT, and business metrics, a time-oriented model can simplify both the architecture and the queries when temporal behavior is central to the problem.
TL;DR: Polished Artifacts, Unchanged Decisions, Or Ten Backlog Anti-Patterns AI Makes Worse Add AI to a Product Backlog process that already struggles, and everything seems to improve within an afternoon. The problem is that polishing artifacts with AI doesn’t fix the root cause: the basis for the team’s decisions doesn’t change; AI only removes the visible discomfort that used to signal something was broken, along with some of the pressure to fix it. AI applied to a dysfunctional system – here, the Product Backlog process – makes the dysfunction look like progress. This is the first article of a new series that explains the uselessness of bolting AI onto a dysfunctional system. It explains why generative AI worsens ten Product Backlog anti-patterns, how polished AI artifacts can hide missing evidence and authority, and how teams can check whether their process is fit for AI. Thesis: Adding AI onto a dysfunctional system, for example, the Product Backlog, makes ten typical backlog anti-patterns worse, because polished AI artifacts hide missing customer evidence, decision authority, and feedback; the article explains these mechanisms, their costs, and how teams test whether their process is ready for AI. Disclaimer: I read Charniak/McDermott’s book on “Artificial Intelligence” decades ago; of course, I use AI for research, translations, proofreading, challenging story arcs and article structures, and summarization. It is a production tool, not a substitute for thinking. A Scenario You May Recognize Consider the following scenario: A Product Owner under pressure feeds stakeholders' requirement documents into an LLM, adds context such as the product roadmap or access to Jira. Soon, the Product Backlog holds 40 new items, each with a user story, five acceptance criteria, and an "estimated" size. Next Monday, Sprint Planning runs smoothly, and the stakeholders are delighted: finally, the team seems "ready." Now ask the four questions nobody asked: Which customer problem does this work address?Whose evidence says it matters?Which alternative did the team consider and reject?And who could have stopped it? If the answers are "whatever the stakeholders requested," "none," "none," and "nobody," the process didn't improve; it just stopped showing its flaws. (You may recall that Agile is really good at exposing problems, issues, and flaws.) What Must Work Before You Accelerate the Process Product Backlog management runs in a loop. What the team knows and assumes informs its product choices, and refinement exposes the uncertainty in those choices. Experiments, research, and delivered Increments then put those choices in front of customers and produce new evidence, which changes the next choices. That is why the Product Backlog is emergent, and every item comes with an implicit "based on what we know now." Work may legitimately enter the Product Backlog to investigate an open question; what matters is that people distinguish what they know from what they still need to learn and pick a suitable next step. If a team fails to feed valuable work into this loop, however, Scrum will be just as effective at building things no one needs: garbage in, garbage out. (Keep in mind that Scrum can be an effective delivery framework, but it completely lacks the product discovery part.) Refinement is a continuous activity within that loop, and it produces a shared understanding sufficient to make the next investment decision, including the decision not to build something at all. Selecting work for a Sprint happens later, in Sprint Planning, once the team has decided on a Sprint Goal: what is most valuable to create? A dysfunctional Product Backlog process usually announces itself: thin items, awkward Sprint Planning sessions, arguments about scope in the middle of a Sprint. That discomfort is useful, as it is the evidence a Scrum Master brings to the Retrospective and a Product Owner brings to stakeholders. AI can remove it without removing its cause. The Scrum Guide 2020 warning then applies literally: "Transparency enables inspection. Inspection without transparency is misleading and wasteful." Google's announcement of the 2025 DORA report, based on survey responses from nearly 5,000 technology professionals, describes the general pattern: "AI doesn't fix a team; it amplifies what's already there." The same announcement reports positive associations between AI adoption and both delivery throughput and product performance, alongside a continued negative association with delivery stability (Announcing the 2025 DORA Report). The findings cited do not establish the ten mechanisms described below; applying the amplification argument to them is my interpretation. AI can help repair a dysfunctional process if you deliberately use it to expose gaps: list the assumptions behind an item, find contradictions between a stakeholder request and the Product Goal (or the equivalent planning objective), or organize customer evidence the team already has. However, that only works while everyone stays aware of what the tool is doing. Polished artifacts fool people quickly into believing that AI is doing the real job, and teams need to be very careful not to fall into that trap. The useful question is whether an AI output makes evidence, assumptions, and alternatives easier to examine, or easier to skip. The dangerous intervention is the second: generating more apparently implementation-ready work before addressing the gaps. (AI, given the right context, is really good at papering over flaws and polishing its output.) What "Functional" Means Compared to a Dysfunctional System A functional Product Backlog process does not need to be flawless, fully documented, or unchanging. It needs to detect and correct poor decisions, which depends on three things: Direction: A customer problem or Product Goal makes a piece of work worth considering in the first place.Authority: Someone can challenge, change, or reject proposed work, and the organization respects that decision.Feedback: When delivery shows an assumption was wrong, later decisions change. The ten anti-patterns resulting from putting AI on top of a dysfunctional system below fail on one or more of these, and they fall into three groups: AI on Top of a Dysfunctional System, Group 1: Choosing Work Without Evidence 1. Prioritization by Proxy Stakeholders decide what goes into the Product Backlog, and the Product Owner passes their decisions on. With an LLM, a stakeholder now arrives with a ten-page requirements document, including personas and success metrics, produced in an afternoon. The polish does not create the power imbalance; it makes the existing one harder to confront, because rejecting a document that looks finished feels like obstruction and puts your career at risk. A better refinement session will not restore product management here, since the problem lies in who holds decision authority. The team becomes an internal development agency, only faster, and the question of whether something else would serve customers better never gets asked, because the polished document seems to have answered it already. 2. The Oracle Product Owner The Product Owner involves neither stakeholders nor subject matter experts, and the Product Backlog contains no research tasks, such as building prototypes or spikes. The AI version consults a model instead: synthetic personas and a generated market analysis. Source-backed desk research is legitimate work; the failure occurs when generated hypotheses acquire the status of observed customer demand. Product discovery gets skipped while looking complete, and the team then builds the wrong thing with considerable confidence, which may well be the most expensive way of building it, even if agentic coding makes building simpler. Without relevant customer evidence, generated personas and market narratives remain hypotheses, no matter how convincing they sound. 3. The Copy and Paste Product Owner The Product Owner breaks stakeholder requirement documents into smaller chunks. A model does this in seconds and returns neatly sliced items. Smaller items, however, do not make the underlying solution appropriate, nor does decomposing a document turn it into incremental learning. The Product Backlog commits prematurely to one solution, and nobody on the team has to understand the problem well enough to propose a cheaper or better-fitting one. The Product Backlog thereby fills with the stakeholder's solution, cut into Sprint-sized portions, and the later slices quietly turn into assumed commitments, which discourages the team from learning anything from the first slice before building the next one. 4. 100% in Advance The organization asks the team to rebuild a legacy application one-for-one, so the team prepares the complete Product Backlog upfront. With AI, a model reads the old codebase and produces an apparently complete specification with hundreds of items. The danger lies in treating that extracted description of the existing system as the approved specification for its replacement, including every workaround users have complained about for years. A rebuild is a rare opportunity to improve processes and usability; treating a generated inventory as fixed scope quietly closes that opportunity. I worked with a Scrum Team on such a rebuild for a large utility, and the team delivered two weeks early and under budget while streamlining processes and improving usability; all gains came from questioning inherited requirements. An extracted specification can even make existing behavior easier to question. AI on Top of a Dysfunctional System, Group 2: Turning Assumptions Into Apparent Readiness 5. Refinement Without the Team The Product Owner refines with the "lead engineer" and a designer, while the other Developers prefer creating code by herding agents. Asynchronous comments and preparation by a subset of the team are not dysfunctional as such; the problem arises when consequential disagreement never gets resolved. Estimation shows it best. The number was never the point; comparing estimates reveals whether Developers have the same work in mind before they start, as I argued in Estimates Are Useful, Just Ditch the Numbers. Accept an AI-suggested size without independent examination, and no one ever surfaces the split in opinion; the different ideas still exist, but nobody learns about them until the Sprint is underway. 6. Cosmetic Readiness All items look thoroughly detailed and estimated, yet the team has not applied INVEST in substance. AI makes every item look ready: correct template, complete fields, a checklist passed. Proper formatting, however, cannot prove what readiness depends on: credible evidence, understood uncertainty, workable dependencies, and shared understanding. A model can propose alternatives and analyze potential value; it cannot assert customer value or make an organization willing to negotiate. Sprint Planning then selects work that looks transparent but isn't, and the Developers discover missing understanding mid-Sprint, after implementation has started, and correcting it may require rework. 7. Acceptance Criteria Inflation The Product Owner covers every conceivable edge case without negotiating with the Developers. Ask a model for acceptance criteria, and you will typically receive a long list of plausible conditions. My heuristic is that three to five acceptance criteria usually suffice, and needing more often indicates the item needs splitting. The count, as such, is not the problem; the problem is criteria that expand scope without support or add constraints nobody examined, which the team's agents then build without evidence that real users need them. A separate risk arises when criteria start prescribing the solution instead of describing the required behavior: then the Product Owner drifts from the why and the what into the how, which belongs to the Developers, and the Developers end up negotiating with a document instead of a colleague. This is a particular challenge nowadays, when many organizations expect their product managers or product owners to be product builders, vibe-coding the prototype themselves. 8. The Last-Minute Product Backlog The team does not invest in Product Backlog management and rushes to fill the Product Backlog right before Sprint Planning. Before AI, the scramble was visible: thin items, a random assortment of stuff to fill the Sprint and appear busy, consequently resulting in a "Sprint Goal" nobody could explain. Now, an hour of prompting produces items that look as if the team had refined them for weeks. The preparation gap becomes much harder to see, although rework, confusion, and missed goals may still expose it later, by which point the cost has already been incurred. AI on Top of a Dysfunctional System, Group 3: Expanding Work Beyond the Capacity to Inspect and Sustain It 9. The Infinite Product Backlog Oversized Product Backlogs, Product Backlogs used as storage for ideas, and items nobody has touched for months predate AI. The economics change, though: generating candidates becomes nearly free, while assessing, ordering, and maintaining them still consumes the same human attention. If nobody moves anything to a separate list of permanently or temporarily discarded ideas, what I call the "Anti-Product Backlog," the most valuable items get buried in noise, and the ordering stops informing any decision. At that point, the Product Owner administers a database, and the team, facing hundreds of plausible items, stops reading the Product Backlog altogether. (Of course, you can instruct an agent to go through that list and flag all entries the agent believes offer no value. But this is a cleanup operation, not a proper process.) 10. What Technical Debt? The team ships feature after feature, and the Product Backlog reserves no capacity for bugs, refactoring, or platform maintenance. AI-assisted development increases the volume of change, and DORA's findings on delivery stability point to the risk when control systems, such as automated testing and fast feedback loops, are weak. This is, above all, a product decision: if the organization rewards feature output and never negotiates quality expectations or capacity trade-offs, AI will produce more features faster, and customers will experience the result as instability, not to mention that they are unlikely to pay for the flood of new features they never asked for. The Cost of Apparent Efficiency "We created the Product Backlog faster" is an incomplete business case, but typical for bolting AI onto a dysfunctional system. Consider an illustration, not a measured result: saving six hours of preparation has little value if unsupported scope then consumes three Sprints of delivery capacity, plus the review effort for AI output nobody inspected closely, the rework once users react, and the maintenance of features nobody needed. Meanwhile, the more valuable problem waits, and your competitors move ahead. When reporting after an AI rollout focuses on preparation time, the number of refined items, time spent on creating shared understanding, and the duration of Sprint Planning, none of these costs show up. The dashboard looks like a success while the expensive part of the path, building and maintaining the wrong things, sits in a different budget line and appears months later. Therefore, evaluate the whole path, from identifying a customer problem to learning whether the delivered change helped, including the effort it takes to inspect what the model produced. The costs also compound: every unexamined decision that "worked" teaches the organization that examination is optional. You can apply the A3 Delegation System to the problem to get a better understanding of where AI may support the process, where you should avoid using AI, what "good quality of AI outputs" looks like, and how to bring transparency to your team’s use of AI. Conclusion: Can Your Process Reject the Wrong Work? Before you accelerate product work by employing AI on top of a dysfunctional system, establish that your organization can question the work's justification, reject it, and learn whether it helped. Inspect a recent, important Product Backlog item and ask yourself which of the following criteria are true for your team: Grounding work in a problem: The team can name the evidence supporting the item and distinguish it from assumptions.Making choices: Someone can explain why the item took precedence and name a plausible request that was rejected or deferred.Exposing uncertainty: The people involved can name open questions and explain how they affect the next step.Exercising authority: The accountable person can change or stop the work when its justification fails.Learning from results: Feedback from delivered work has changed a later product decision. A first round producing convincing answers is rarely enough, so follow up and ask for recent examples: Which proposed item did the team stop or materially change because its evidence was weak?Which discovery or delivery result changed the next product decision?When did someone last exercise the authority the team claims to have? If nobody can point to such examples, generating more apparently ready work will only make the gap harder to see, although AI may still help you expose it. Key Questions This Article on the Effects of Employing AI on Top of a Dysfunctional System Answers How Does AI Make Product Backlog Anti-Patterns Worse? AI makes a dysfunctional Product Backlog process look functional, a typical first impression of bolting AI onto a dysfunctional system. It produces polished items, acceptance criteria, and estimates within minutes, while the basis for the decisions stays the same: missing customer evidence, unclear decision authority, and no working feedback loop. The visible discomfort that used to expose the problem, such as thin items or chaotic Sprint Planning sessions, disappears, and with it some of the pressure to fix the process. Should a Scrum Team Use AI to Write User Stories and Acceptance Criteria? Only if the team's Product Backlog process can already detect and correct poor decisions. AI generates user stories and long lists of acceptance criteria quickly, yet formatting cannot prove credible evidence, understood uncertainty, or shared understanding. Generated criteria also tend to expand scope without support; as a heuristic, three to five acceptance criteria usually suffice, and needing more often signals that the item needs splitting. How Can a Team Tell Whether Its Product Backlog Process Is Ready for AI? A process is ready when the organization can question the justification of work, reject it, and learn whether it helped. Test it with recent behavior instead of convincing answers: Which proposed item did the team stop or materially change because its evidence was weak? Which delivery result changed the next product decision? And when did someone last exercise the authority to stop work? Can AI Help Fix a Dysfunctional Scrum Process? Yes, when the team deliberately uses AI to expose gaps and stays aware of what the tool is doing. Useful applications include listing the assumptions behind a Product Backlog item, finding contradictions between a stakeholder request and the Product Goal, and organizing customer evidence the team already has. The dangerous use is generating more apparently implementation-ready work before those gaps are addressed. Why Does Creating a Product Backlog Faster With AI Not Save Money? Hours saved in preparation lose their value when unsupported scope consumes weeks of delivery capacity. The full cost includes reviewing AI output, reworking after users react, maintaining features nobody needed, and delaying work on more valuable problems. Reports that focus on preparation time, refined item counts, or Sprint Planning duration miss these costs because they appear months later in different budget lines.
Software organizations require rules to ensure predictability, knowledge sharing, and sound technical decisions. However, excessive rules, processes, and standards can obstruct progress if they eclipse desired outcomes. This longstanding tension in software engineering is now more pronounced as AI accelerates code generation, solution exploration, and implementation of change. The main challenge is not the number of rules, but how they fit with context, autonomy, and responsibility. Some organizations have minimal constraints, others tightly control decisions, and a few adapt principles to specific situations. Noticing these models helps engineering leaders assess their organization, understand associated risks, and move toward more effective technical governance. Software Architecture Is Also an Organizational Problem Software architecture is often defined by technical elements such as components, APIs, databases, infrastructure, quality attributes, and architectural styles. While these matter, they exist within a wider context. People design and maintain software, influenced by business constraints, communication structures, incentives, policies, and authority levels. Architecture is not only a technical artifact; it also reflects the environment in which teams make technical decisions. This broader perspective reflects the socio-technical nature of software engineering and is consistent with standards such as ISO/IEC/IEEE 42010, which recognize that architectural concerns encompass organizational, economic, regulatory, and community factors. Conway's Law shows that organizational communication structures often mirror system design. If organizational structure shapes software, then the distribution of authority, management of disagreement, response to failure, and governance of technical decisions likewise influence architecture. Therefore, engineering culture and governance are architectural concerns, not just management issues. Culture Shapes How Engineering Decisions Are Made If architecture is formed by its organizational context, organizational culture also directly influences engineers’ daily decisions. Sociologist Ron Westrum offers a system for understanding how organizations respond to problems and opportunities, particularly in high-risk environments. His model focuses on information flow, including cooperation, safe communication of bad news, responsibility management, failure investigation, and acceptance of new ideas. Westrum identified three main cultural patterns: pathological (power-oriented), bureaucratic (rule-oriented), and generative (performance-oriented) organizations. This distinction matters in software engineering, where architectural decisions depend on information moving across teams, technical boundaries, and leadership levels. Organizations that conceal problems, discourage disagreement, or fragment responsibility make different technical choices than those that discuss risks openly, encourage collaboration, and treat failures as chances to learn. DORA adopted Westrum's model in its research and consistently links generative, high-trust cultures to stronger software delivery and organizational performance. Culture isn't just the environment around engineering work; it shapes how teams interpret technical information, who can challenge decisions, how rules are enforced, and how architecture evolves. PathologicalBureaucraticGenerativePower orientedRule orientedPerformance orientedLow cooperationModest cooperationHigh cooperationMessengers “shot”Messengers neglectedMessengers trainedResponsibilities shirkedNarrow responsibilitiesRisks are sharedBridging discouragedBridging toleratedBridging encouragedFailure leads to scapegoatingFailure leads to justiceFailure leads to inquiryNovelty crushedNovelty leads to problemsNovelty implemented Mode 1: Wild West Engineering Wild West Engineering refers to environments with weak governance, unclear ownership, and few shared engineering principles. Teams may appear highly autonomous, but such autonomy lacks architectural direction, reliable feedback, and consistent accountability. Technical decisions are frequently reactive and localized, with tools and approaches selected for immediate needs or personal preference rather than the wider system context. The problem is not simply a lack of rules. Experienced teams can operate effectively with minimal rules if they share strong principles and understand the impact of their decisions. Wild West Engineering occurs when autonomy exists without sufficient shared context, accountability, or coordination. Over time, architectural knowledge becomes siloed, similar problems are solved inconsistently, and changes become riskier as no one fully understands the system. Westrum’s pathological culture is a relevant comparison: cooperation is limited, responsibilities are avoided, bad news is suppressed, and failures result in blame rather than inquiry. In software organizations, this often creates a hero culture where a few engineers become indispensable for managing disorder. AI can worsen this problem; when teams generate code, add dependencies, or make architectural decisions without shared guardrails, inconsistency and technical debt can spread quickly. Indicators Typical signs involve unclear ownership, multiple solutions for similar problems, undocumented architectural decisions, reliance on tribal knowledge, fear of changing legacy components, recurring blame after incidents, and dependence on a few key engineers. Another important sign is resignation: engineers stop proposing improvements because changing the system seems riskier than leaving existing problems unresolved. How to Move Forward Imposing heavy bureaucracy is not the solution to chaos. The first step is to establish minimum shared guardrails: clarify ownership, document key architectural decisions, define core engineering principles, set baseline expectations for security and operability, and create safer ways to discuss failures and technical debt. The goal is to support autonomy with enough structure to ensure local decisions conform to a coherent system. Mode 2: Bureaucratic Engineering Bureaucratic Engineering represents the opposite extreme, where governance is pervasive. Standards, approval processes, architectural reviews, mandatory technologies, and detailed procedures dictate how teams operate. While this structure creates predictability and consistency, problems arise when adherence to process outweighs knowing its purpose. Teams may focus on compliance rather than evaluating whether rules remain relevant. Westrum’s bureaucratic culture reflects similar traits: limited cooperation, narrow responsibilities, minimal cross-boundary collaboration, and a tendency to view novelty as problematic. In engineering, this occurs when architectural decisions are centralized, outdated standards persist, and exceptions are hard to secure. While these mechanisms may guarantee a baseline of quality, they may also limit performance. The core issue is not the presence of rules, but the erosion of judgment. Governance becomes rigid when organizations treat principles as permanent mandates and compliance as a substitute for engineering quality. AI can increase this tension, as organizations may respond to new risks with broad restrictions, approved-tool lists, or complex authorization processes. Although these measures may reduce risk, they can also hinder teams from pursuing valuable opportunities, making the organization safer but less nimble and flexible. Indicators Typical signs include rules missing explicit justification, standards defended only by “we have always done it this way,” centralized approval for minor technical decisions, difficulty obtaining exceptions, and architectural decisions that ignore local context. Status quo bias helps explain this persistence. Samuelson and Zeckhauser found that people often prefer existing or default options. In engineering organizations, this reinforces architectural inertia: once a technology, methodology, or rule becomes standard, replacing it requires far more justification than maintaining it. As a result, rules persist not because they are optimal, but because they are easier to retain. Organizational silence is another key warning sign. Morrison and Milliken define this as a collective tendency to withhold concerns when employees believe speaking up is unwise or ineffective. In engineering, this occurs when developers privately disagree with decisions but remain silent in meetings, believing that contesting established processes will not lead to change. Persistent silence should not be mistaken for agreement; it may signal that the organization discourages dissent. Another warning sign is when success is measured mainly by process compliance rather than software outcomes. Teams may meet all requirements yet become slower, less innovative, and disconnected from the initial intent of these controls. How to Move Forward To move beyond Bureaucratic Engineering, organizations should redefine governance rather than eliminate it. Leaders must frequently review the purpose of rules, distinguish primary constraints from preferences, and replace unnecessary approvals with clear principles. Teams need defined boundaries, authority to make decisions within them, and a clear process for exceptions when justified by context. A practical test is to ask whether a rule is still justified by its intended outcome. If no one can explain the risk it addresses, the value it protects, or evidence of its need, the organization may be maintaining the process for its own sake. The aim is to shift from prescriptive control to outcomes-based control, with accountability and feedback. Mode 3: Context-Driven Engineering Context-Driven Engineering integrates governance with local judgment. Rules are tools for managing risk and ensuring consistency, not ends in themselves. Teams work within defined principles, constraints, and responsibilities, while retaining the autonomy to adapt decisions to their particular technical and commercial context. The core assumption is that the same rule may yield different outcomes in different environments. This approach corresponds with Westrum’s generative culture, defined by high cooperation, shared risk, cross-boundary collaboration, inquiry following failures, and openness to innovation. In engineering, teams are expected to follow standards, understand their purpose, and recognize when exceptions are warranted. Architectural decisions are treated as explicit trade-offs, not automatic pattern applications. Teams evaluate technologies such as microservices, hexagonal architecture, event-driven systems, or specific cloud platforms based on the problem, constraints, and desired outcomes, rather than assuming they are universally correct. Applying a rule in the wrong context can lead to poor decisions. The same applies to architectural patterns: they address recurring problems in specific contexts, and applying them without considering context can create anti-patterns. For example, a lifeboat is essential in the ocean but irrelevant in a desert. Similarly, providing water is life-saving for someone dehydrated, but ineffective for someone drowning. A tool or rule's effectiveness depends on the context in which it is used. This environment requires greater maturity, not reduced governance. Autonomy must be balanced with accountability, feedback, and a readiness to revisit decisions as new information arises. AI fits well with this model, as its use can be governed based on context and risk: low-risk tasks may permit broad experimentation, while sensitive data, production changes, or high-impact decisions demand stricter controls. The objective is responsible judgment within clear boundaries, not unrestricted freedom. Indicators Key indicators include clear engineering principles, documented decision rationales, explicit ownership, and constructive disagreement. Teams should be able to explain both the rule and its rationale. Exceptions are allowed but must be justified. Failures prompt learning rather than blame, and architecture teams periodically review standards instead of letting them become outdated. Another strong indicator is that teams can contest established practices without equating disagreement with disloyalty or process violations. A context-driven organization distinguishes between non-negotiable constraints and situation-dependent choices. Security, legal requirements, data protection, and regulatory obligations remain strict, while implementation details such as architectural style, framework selection, or deployment strategy are evaluated locally. This distinction ensures autonomy does not devolve into unstructured or uncontrolled practices. How to Sustain This Model The main risk is regression. Insufficient discipline can lead to disorder, whereas excessive concern for consistency can create bureaucracy. Upkeeping this model requires ongoing review of principles, active feedback loops, documented architectural decisions, clear exception processes, and leaders willing to reexamine their assumptions. A practical test is whether the organization can answer four questions: What problem are we solving? Which constraints are truly non-negotiable? What trade-off are we accepting? What evidence would prompt us to revisit the decision? When teams can answer these questions consistently, governance functions as a learning system rather than a control mechanism. Conclusion Three Engineering Governance models inspired by Ron Westrum Engineering organizations improve not by simply adding or removing rules, but by ensuring teams understand their purpose, have the required context, and are trusted to exercise judgment within clear boundaries. Wild West Engineering lacks the structure for long-term growth, while Bureaucratic Engineering frequently leads to rigidity and limits learning. Context-Driven Engineering seeks balance through prioritizing principles over prescriptions, fostering autonomy with accountability, and encouraging decisions based on trade-offs rather than routine. As AI accelerates code generation and change, maintaining this harmony is increasingly important. The most adaptable organizations will foster responsible, context-aware decision-making.
Hello DZone community! We’re refreshing our newsletters to help you follow the topics that matter most to your work, explore new ideas, and stay connected with the developer community. Our Zone newsletters are becoming six focused newsletters, with new names and related topics that were developed with input from some of our fantastic community members, and we’re excited to bring them to your inbox: Beyond the Rows: Big data and databasesMind the Model: AI for developers and engineersShip & Scale: Cloud, DevOps, performance, and AgileThe Attack Surface: SecurityDistributed by Design: Microservices, integration, and IoTCode & Craft: Java and web development Each topical newsletter will arrive twice a month, with editions scheduled on Tuesdays and Thursdays. Meet DZone Digest We’re also bringing DZone Daily and DZone Weekly together into DZone Digest, arriving every Wednesday. It will be your weekly roundup of articles and insights from across DZone. What Else Is Changing? All newsletters got a refreshed design, making room for the content you care about and more opportunities to discover upcoming events. We’re also opening newsletter subscriptions to everyone, including developers who aren’t DZone members. Our goal is to make DZone’s newsletters more useful, with clearer topic choices and a regular cadence that helps you keep learning. We’d love to hear from you: Which newsletter are you most interested in, and what topics would you like us to cover? Share your thoughts in the comments. You can subscribe to them here. (If you’re subscribed to any of our Zone newsletters, you’ll begin receiving the updated newsletter covering your topics. If you’re subscribed to DZone Daily or DZone Weekly, you’ll now receive DZone Digest.)
Browser agents become useful when they can do more than reach the right page and trigger the right control. A click is only an attempted action; it is not proof that a business operation completed. A form can submit while validation fails, a checkout button can respond while the API returns an error, and a success-looking route can render stale state. Playwright MCP is well suited to closing that gap because it exposes browser interaction through structured accessibility snapshots and adds explicit testing tools for checking visible elements, text, lists, and form values. The result is an agent loop that can treat verification as a first-class phase rather than as an optimistic interpretation of the previous action. An Action Is Not a Result The central design rule is simple: every state-changing action should have a postcondition. In an ordinary browser automation script, success is often inferred from the absence of an exception. That standard is too weak for an autonomous agent. Playwright MCP interaction tools such as browser_click operate on element references taken from accessibility snapshots, and most actions return an updated snapshot after triggered browser work settles. That makes the post-action page state immediately available, but the controller still has to decide what evidence counts as success. A useful contract separates intent, action, and evidence. For a task such as submitting an order, the action is a click on the final submission control. The evidence might be a visible confirmation heading plus an order identifier. The agent should not report completion until those conditions are independently checked. With the testing capability enabled, Playwright MCP provides browser_verify_element_visible, browser_verify_text_visible, browser_verify_list_visible, and browser_verify_value. Verification calls return Done on success and an error on failure, which creates a clean boundary between a browser action and a verified outcome. A minimal controller can therefore make the verification step explicit: TypeScript await client.callTool({ name: "browser_click", arguments: { target: submitRef } }); await client.callTool({ name: "browser_wait_for", arguments: { text: "Order confirmed" } }); const proof = await client.callTool({ name: "browser_verify_element_visible", arguments: { role: "heading", accessibleName: "Order confirmed" } }); if (proof.isError) { throw new Error("Order submission could not be verified"); } The important property is not the specific wrapper around callTool; it is the control flow. Action and verification are different operations, and a failed verifier changes the task status from “completed” to “unconfirmed.” browser_wait_for is appropriate when a specific asynchronous transition must finish, while Playwright MCP already waits for triggered navigation and network activity after most actions. Fixed sleeps should remain a last resort because the server supports waiting for text to appear or disappear directly. Verification Should Match the Business Outcome A reliable verifier checks the state that matters to the task rather than a convenient visual change. A button becoming disabled proves only that the button changed. A toast saying “Saved” is stronger, but still may not prove persistence if the application updates optimistically. Browser-level evidence becomes stronger when multiple independent signals agree: semantic UI state, the resulting page structure, and relevant network activity. Playwright MCP exposes each of these forms of evidence through snapshots, verification tools, and network inspection. Playwright MCP exposes network inspection through browser_network_requests and browser_network_request, allowing an agent to locate a relevant request and inspect its details. Console messages are also available through the core browser_console_messages tool. Those channels are useful when a task appears successful in the DOM while a background request fails or the page emits an uncaught error. Network evidence should still be tied to application semantics; an HTTP response by itself does not establish that the intended record contains the correct data. For a profile update, the strongest browser-side check may be a round trip: submit the change, wait for the completion signal, navigate away or reload, then verify the field value from the newly rendered state. browser_verify_value supports textboxes, checkboxes, radios, comboboxes, and sliders, so persistence checks can stay semantic instead of scraping raw HTML. TypeScript await client.callTool({ name: "browser_verify_value", arguments: { type: "textbox", element: "Display name", target: displayNameRef, value: expectedName } }); This pattern is especially important for agentic workflows because planning logic can be probabilistic while verification can remain deterministic. The model may choose among several valid ways to reach a form, but the acceptance criterion can still be exact: a heading exists, a field equals an expected value, a list contains required entries, or confirmation text is visible. Playwright MCP’s verification tools are designed around those concrete browser states. Accessibility Snapshots Make the Feedback Loop Precise Playwright MCP uses accessibility snapshots as the primary representation for agent interaction. A snapshot contains roles, accessible names, text, and element references used by subsequent tool calls. This is materially different from relying on screenshots as the main control surface. Screenshots remain valuable for visual diagnostics, but the MCP documentation explicitly directs actions toward snapshots rather than screenshot coordinates. That distinction improves verification quality. A semantic check such as “heading named Order confirmed is visible” is less ambiguous than a model deciding whether a collection of pixels resembles a success page. It also aligns the agent’s evidence with the same roles and names used by Playwright locators. When only part of a large page matters, browser_find can search the accessibility snapshot and return matching nodes with local context, reducing the need to repeatedly consume the full tree. Screenshots still have a place when the requirement is inherently visual, such as confirming layout, clipping, or rendering. For correctness of transactional browser work, however, semantic evidence should dominate. Tracing can then provide failure forensics rather than primary success criteria. With the devtools capability, Playwright MCP can record traces containing DOM snapshots, screenshots, network activity, console logs, and timing, making an unverified or failed run reproducible after the fact. Successful Agent Runs Can Become Regression Tests Verification becomes more valuable when it survives beyond a single agent session. Playwright MCP’s testing capability records matching expect(...) code for verification tools, and action responses can include generated Playwright code. The documentation explicitly shows an exploratory sequence being assembled into a conventional Playwright test. That creates a productive path from autonomous exploration to deterministic regression coverage. A verified flow can therefore graduate into a compact test instead of remaining hidden inside an agent transcript: TypeScript test("submits an order", async ({ page }) => { await page.getByRole("button", { name: "Place order" }).click(); await expect( page.getByRole("heading", { name: "Order confirmed" }) ).toBeVisible(); await expect( page.getByText(expectedOrderNumber) ).toBeVisible(); }); The same principle should shape production configuration. Only required capabilities should be exposed, isolated sessions should be preferred for repeatable runs, and origin restrictions can reduce accidental navigation. Playwright MCP supports capability selection, isolated profiles, allowed and blocked origins, secrets redaction, and configurable timeouts. Its documentation also warns that origin controls and secret handling are convenience defenses rather than security boundaries, so client-level permissions remain necessary when an agent can perform consequential actions. Conclusion A browser agent becomes dependable only when completion means more than “the click happened.” Playwright MCP provides the pieces required for a verification-centered design: structured accessibility snapshots for precise targeting, explicit verification tools for semantic postconditions, waiting primitives for asynchronous transitions, network and console evidence for deeper diagnosis, and traces for failed-run analysis. The strongest implementation treats every consequential action as a hypothesis that must be proven by observable browser state. That shift turns browser automation from a sequence of hopeful interactions into a controlled execution loop whose results can be checked, explained, and eventually converted into durable Playwright regression tests.
Enterprise context — Acme FinServ. SOC 2 CC7 (system monitoring) requires that Acme can detect and investigate anomalous activity. When an agent-driven workflow touches customer data at 2 AM, "we have logs somewhere" is not an answer an auditor accepts. The distributed trace built in this part is the forensic evidence trail: a single trace ID that ties the Goose prompt to every agentgateway policy decision and every Quarkus tool call, so a post-incident review can reconstruct exactly which agent did what, in what order, and how long each governed hop took. The Core Problem In Part 1, we built a Quarkus MCP tool server. In Part 2, we secured it with agentgateway's JWT authentication, RBAC, and ExtMCP guardrails. The architecture works — but when something goes wrong in production, you're flying blind. Agentic workflows are fundamentally different from traditional request-response APIs. A single user prompt like "Debug customer CUST-4091" triggers a multi-round-trip loop: Goose calls tools/list to discover available toolsThe LLM selects getCustomerStatus and Goose sends tools/callThe LLM reads the response, sees primaryRegion: US-EAST-1, and chains a second tools/call to getZoneHealthLogsThe LLM correlates both results and generates a diagnostic summary Each of these hops crosses process boundaries: Goose → agentgateway → Quarkus. Without distributed tracing, you see four isolated HTTP requests in your access logs. You cannot tell they belong to the same agentic workflow. When step 3 takes 12 seconds instead of 200ms, you have no waterfall to pinpoint whether the latency came from agentgateway policy evaluation, Quarkus bean validation, or a slow downstream call. This creates black holes in telemetry dashboards — the exact gap that autonomous agents exploit to degrade silently. The Solution: W3C Trace Context Across All Three Layers The fix is standard distributed tracing, applied to the MCP transport layer: agentgateway exports spans for every proxied MCP request and propagates traceparent headers to the backend.Quarkus with quarkus-opentelemetry picks up the incoming traceparent, creates child spans for tool execution and bean validation, and exports them to the same Jaeger instance.Jaeger correlates both sides into a single trace waterfall — one view from agent prompt to tool result. Prerequisites Everything from Parts 1 and 2, plus: Podman – for running Jaeger (podman compose) Verify Podman is available: Shell podman --version Step 1: Launching the Observability Backend We use Jaeger v2 as both the OTLP collector and the trace UI. A single container accepts traces from agentgateway on port 4317 (OTLP gRPC) and from Quarkus on port 4318 (OTLP HTTP), and serves the query UI on port 16686. Shell cd part3-observability podman compose up -d This starts Jaeger v2 with OTLP collection enabled by default. Verify it's running: Shell curl -sf http://localhost:16686/ > /dev/null && echo "Jaeger UI is ready" Open http://localhost:16686 — you'll see an empty Jaeger UI. We'll populate it with MCP traces in the following steps. Production Alternative: Grafana Tempo For production deployments, replace Jaeger with Grafana Tempo backed by object storage (S3/GCS). The OTLP endpoint stays the same — only the compose.yml changes. Grafana provides richer dashboards, alerting, and long-term trace retention. Step 2: Enabling OpenTelemetry in Quarkus Add the quarkus-opentelemetry extension to Part 1's pom.xml: Properties files <dependency> <groupId>io.quarkus</groupId> <artifactId>quarkus-opentelemetry</artifactId> </dependency> Configure the exporter in application.properties: Properties files # OpenTelemetry quarkus.otel.service.name=customer-tools quarkus.otel.exporter.otlp.traces.endpoint=http://localhost:4318 quarkus.otel.exporter.otlp.traces.protocol=http/protobuf quarkus.otel.traces.sampler=always_on quarkus.otel.traces.suppress-non-application-uris=false PropertyPurposeservice.nameIdentifies this service in Jaeger's service dropdowntraces.endpointOTLP HTTP receiver — Jaeger's port 4318 (base URL only; Quarkus appends /v1/traces)traces.protocolhttp/protobuf — Quarkus uses its Vert.x-based HTTP exportertraces.sampleralways_on — sample every span (reduce in production)suppress-non-application-urisfalse — include MCP endpoint spans (they'd be filtered otherwise) When no OTLP collector is running (Parts 1 and 2 without Jaeger), Quarkus logs a connection warning, but the MCP server works normally. When the collector IS running (Part 3), traces flow automatically. Zero code changes to the MCP tools. Rebuild Part 1: Shell cd part1-quarkus-mcp mvn package -DskipTests What Quarkus Auto-Instruments With quarkus-opentelemetry on the classpath and the SDK enabled, Quarkus automatically creates spans for: LayerSpan NameWhat It CapturesHTTP serverPOST /mcpInbound MCP request with method, status, latencyCDI beansCustomerServiceTools.getCustomerStatusTool execution time within the MCP handlerBean ValidationHibernateValidatorParameter validation before tool logic runsREST clientOutbound HTTP callsAny downstream API calls (future extensions) No @WithSpan annotations needed. The Quarkus OpenTelemetry extension instruments the reactive pipeline automatically. Step 3: Configuring W3C Trace Context in agentgateway agentgateway supports native OpenTelemetry trace export. Add a tracing block to the gateway configuration: Properties files config: adminAddr: localhost:15000 tracing: otlpEndpoint: http://localhost:4317 otlpProtocol: grpc randomSampling: 1.0 FieldPurposeotlpEndpointOTLP receiver — Jaeger's port 4317otlpProtocolgrpc for OTLP/gRPC (also supports http)randomSamplingSample 100% of traces (reduce to 0.01–0.1 in production) How Trace Propagation Works When agentgateway receives an MCP request: Creates a root span for the proxy operation (e.g., agentgateway.mcp.proxy)Injects a traceparent header into the forwarded request to Quarkus:traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01Quarkus reads the traceparent, creates a child span under the same trace ID, and records tool executionBoth spans export to Jaeger via OTLP, where they appear as a single correlated trace This is standard W3C Trace Context propagation — the same mechanism used across all OpenTelemetry-instrumented services. Configuration Files Part 3 provides two agentgateway configurations: ConfigUse Caseconfig-traced.yamlTracing only — proxy + OTLP export, no security layersconfig-traced-guardrails.yamlTracing + ExtMCP guardrails — observe the guardrail evaluation spans too Step 4: Running the Interactive Demo Start all services with the one-command script: Shell cd part3-observability ./start-all.sh The script starts Jaeger, Quarkus (with OTel enabled), and agentgateway (with trace export), then launches the demo SPA on :8890. Open the MCP Observability Console at http://localhost:8890/index.html and walk through the three demo steps: Initialize – Establishes an MCP session through agentgateway. The architecture diagram animates the trace propagation: root span creation in agentgateway, traceparent injection, child span in Quarkus, and OTLP export to Jaeger.List Tools – Discovers all 5 tools through the traced proxy. The trace waterfall panel shows the agentgateway proxy span and the Quarkus HTTP span side by side with timing.Multi-Tool Workflow – Simulates Goose's multi-turn reasoning: getCustomerStatus (finds region US-EAST-1) → getZoneHealthLogs (checks zone health) → getSLACompliance (correlates SLA metrics). Each step generates a full trace with waterfall visualization. The stat tiles track traces generated, spans collected, and Jaeger status. Click Open Jaeger to view the real trace waterfalls in the Jaeger UI at http://localhost:16686. Step 5: Generating Traces via CLI To generate additional traces manually, simulate a multi-turn agentic workflow: Shell # Step 1: Initialize MCP session export MCP_SESSION_ID=$(curl -s -D - http://localhost:3000/mcp \ -H "Content-Type: application/json" \ -H "Accept: application/json, text/event-stream" \ -H "MCP-Protocol-Version: 2025-03-26" \ -d '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-03-26","capabilities":{},"clientInfo":{"name":"curl","version":"1.0"}}' \ | grep -i "mcp-session-id:" | sed 's/.*: //' | tr -d '\r') # Step 2: Discover tools curl -s http://localhost:3000/mcp \ -H "Content-Type: application/json" \ -H "Accept: application/json, text/event-stream" \ -H "MCP-Protocol-Version: 2025-03-26" \ -H "mcp-session-id: $MCP_SESSION_ID" \ -d '{"jsonrpc":"2.0","id":2,"method":"tools/list","params":{}' \ | grep '^data: ' | sed 's/^data: //' | jq . # Step 3: Agent calls getCustomerStatus (first tool invocation) curl -s http://localhost:3000/mcp \ -H "Content-Type: application/json" \ -H "Accept: application/json, text/event-stream" \ -H "MCP-Protocol-Version: 2025-03-26" \ -H "mcp-session-id: $MCP_SESSION_ID" \ -d '{"jsonrpc":"2.0","id":3,"method":"tools/call","params":{"name":"getCustomerStatus","arguments":{"customerId":"CUST-4091"}}' \ | grep '^data: ' | sed 's/^data: //' | jq . # Step 4: Agent chains getZoneHealthLogs based on the region from step 3 curl -s http://localhost:3000/mcp \ -H "Content-Type: application/json" \ -H "Accept: application/json, text/event-stream" \ -H "MCP-Protocol-Version: 2025-03-26" \ -H "mcp-session-id: $MCP_SESSION_ID" \ -d '{"jsonrpc":"2.0","id":4,"method":"tools/call","params":{"name":"getZoneHealthLogs","arguments":{"zoneId":"US-EAST-1"}}' \ | grep '^data: ' | sed 's/^data: //' | jq . # Step 5: Agent fetches SLA compliance for correlation curl -s http://localhost:3000/mcp \ -H "Content-Type: application/json" \ -H "Accept: application/json, text/event-stream" \ -H "MCP-Protocol-Version: 2025-03-26" \ -H "mcp-session-id: $MCP_SESSION_ID" \ -d '{"jsonrpc":"2.0","id":5,"method":"tools/call","params":{"name":"getSLACompliance","arguments":{"serviceId":"api-gateway"}}' \ | grep '^data: ' | sed 's/^data: //' | jq . Each of these requests generates a trace that flows through agentgateway into Quarkus and lands in Jaeger. Step 6: Visualizing the Trace Waterfall in Jaeger Open http://localhost:16686 in your browser. Finding Traces In the Service dropdown, select customer-tools (Quarkus) or agentgatewayClick Find TracesClick on any trace to open the waterfall view Reading the Waterfall A typical tools/call trace shows the following span hierarchy: Shell agentgateway.mcp.proxy [12ms] └─ POST /mcp [8ms] ← Quarkus HTTP server └─ CustomerServiceTools.getCustomerStatus [2ms] ← CDI tool execution SpanServiceWhat It Tells Youagentgateway.mcp.proxyagentgatewayTotal proxy overhead including policy evaluationPOST /mcpcustomer-toolsQuarkus HTTP handling time for the MCP requestgetCustomerStatuscustomer-toolsPure tool execution time (business logic) What to Look For Proxy overhead: The gap between the agentgateway span and the Quarkus span shows network + policy evaluation time. If this grows, check guardrail server latency.Validation time: Bean Validation spans appear before tool execution. Regex-heavy patterns like ^CUST-[0-9]{4,8}$ are fast, but complex validators on large payloads can add latency.Multi-turn correlation: When Goose chains multiple tool calls (e.g., getCustomerStatus → getZoneHealthLogs), each appears as a separate trace. The mcp-session-id tag lets you filter all traces belonging to one agent session.Error traces: Failed validations (invalid customer ID format) or guardrail rejections (blocked poison payloads) produce error spans with exception details. Connecting Goose for Real Traces Launch Goose pointed at agentgateway and prompt a multi-tool workflow: Shell goose session "Debug customer CUST-4091 — check their account status, then pull health logs for their region and SLA compliance for api-gateway." This generates a burst of correlated traces in Jaeger showing Goose's multi-turn tool orchestration from the proxy layer down to individual tool execution spans. What We Achieved Starting from the secured architecture in Part 2, we added full observability without changing any MCP tool code: LayerWhat We AddedConfig ChangeQuarkusquarkus-opentelemetry dependencypom.xml + application.propertiesagentgatewaytracing block in config YAMLconfig-traced.yamlObservability backendJaeger all-in-one via Podman Composecompose.yml The entire stack runs locally with a single ./start-all.sh command and produces end-to-end trace waterfalls in Jaeger. Production Considerations ConcernLocal (this tutorial)ProductionTrace backendJaeger all-in-one (in-memory)Grafana Tempo + object storageSampling rate100% (default: 1.0)1-10% or adaptive samplingTrace retentionContainer lifetimeDays/weeks in durable storageAlertingManual Jaeger inspectionGrafana alerting on span latency/error rateMetricsTraces onlyAdd Prometheus + quarkus-micrometer for RED metrics Coming Up in Part 4 With tracing in place, you can now see every MCP tool call flowing through the system. In Part 4, we will move beyond single-agent tool calls to multi-agent orchestration — using the Agent-to-Agent (A2A) protocol to coordinate autonomous agents that can delegate work, enforce governance via AGENTS.md, and call back into our MCP tool services.
In my previous article, I walked through running coding agents inside Docker Sandboxes on a local machine. We installed the sbx CLI, started with a small project, and covered the commands needed to run, stop, and remove a sandbox. This time, I want to take that same workflow off the laptop. Docker added cloud sandboxes in version 0.42.0. You can now use sbx --cloud to run an agent on Docker-managed infrastructure instead of using your machine for the sandbox’s compute. The command is simple to use. The part that is worth understanding is how you get your code into that environment, work with the agent, and bring the changes back locally. That is what we will do here. Nothing complicated; we will start with a small Python project, one coding task, and a cloud sandbox. We will remove the sandbox when we are done with the work. What Changes With a Cloud Sandbox? The sbx CLI still runs in your terminal. With --cloud, supported commands target Docker’s cloud service rather than your local sandbox environment. For example: PowerShell sbx ls Lists your local sandboxes. PowerShell sbx --cloud ls Lists your cloud sandboxes. That distinction matters throughout this walkthrough. If you forget --cloud, you are not asking about the same environment. Cloud sandboxes also have separate credentials and network policies. Do not assume that an agent login or network policy you configured locally is already available in the cloud. For this example, we will copy individual files explicitly. That keeps it easy to see what we send to the sandbox and what we bring back. Before You Start You will need: An updated sbx CLI with cloud support, introduced in version 0.42.0.A Docker account with an active Docker Agentic Platform plan for cloud compute.Authentication for the coding agent you want to use. This walkthrough uses Claude.Python 3 available in the sandbox image for the example. Note: The free sbx CLI does not mean cloud compute is free. Docker bills cloud compute based on usage, and your model provider bills inference separately. Check your account’s pricing before starting. Also, use a small sample project first. Running remotely means sending code off your machine. For company repositories, make sure that is allowed before uploading anything. The host-side commands below use PowerShell. Paths inside the cloud sandbox use Linux-style paths. Step 1: Sign In and Configure the Agent First, check your installed version: PowerShell sbx version If you are still using an older version from the previous walkthrough, update it before continuing. Sign in to Docker: PowerShell sbx login For Claude, Docker documents a cloud OAuth flow: PowerShell sbx --cloud secret set anthropic --oauth Complete the provider sign-in with an account that has the required access. Notice the --cloud flag here, too. These credentials are stored for cloud use, separately from your local sandbox credentials. There is no reason to put a token in our Python files or paste it into an agent prompt. Step 2: Create a Small Project Let us give the agent something specific to fix. Create a project folder: PowerShell New-Item -ItemType Directory -Path .\cloud-sandbox-demo Set-Location .\cloud-sandbox-demo Inside it, create a file named slug.py: Python def make_slug(text): return text.lower().replace(" ", "-") This converts "Docker Sandboxes" into "docker-sandboxes". It works for that input, but it does not handle whitespace very well. Leading spaces become leading hyphens. Repeated spaces become repeated hyphens. Tabs are not handled at all. Now create test_slug.py: Python import unittest from slug import make_slug class SlugTests(unittest.TestCase): def test_two_words(self): self.assertEqual(make_slug("Docker Sandboxes"),"docker-sandboxes") if __name__ == "__main__": unittest.main() We have one passing case and a clear improvement to make. The point is not that this function needs cloud compute. It is small enough that we can focus on the sandbox workflow without spending half the article explaining an application. Step 3: Start a Cloud Sandbox Run the following command: PowerShell sbx --cloud run --detached --name cloud-demo --ttl 1h claude This creates a cloud sandbox and starts the agent without attaching your terminal to it. The flags in the above command are for doing useful things: --cloud selects the cloud environment.--detached returns control to your terminal.--name cloud-demo gives the sandbox a recognizable name.--ttl 1h requests a one-hour lifetime. Important: The documented default action when the TTL expires is deletion. Treat this as a disposable environment, and copy your work out before the deadline. The command prints a sandbox ID. You can also find it with: PowerShell sbx --cloud ls Copy that ID into a PowerShell variable: PowerShell $sandbox = "PASTE_YOUR_SANDBOX_ID_HERE" Use the real ID returned by Docker, not the placeholder above. Keep using this terminal for the remaining commands. One detail to remember is a detached cloud run creates a new sandbox. It is not the command to run repeatedly when you want to reconnect to the same one. Step 4: Copy the Project Into the Sandbox Create a directory for our example: PowerShell sbx --cloud exec $sandbox mkdir -p /workspace/demo The mkdir command runs inside the Linux sandbox, not on Windows. Now copy the two files: PowerShell sbx --cloud cp .\slug.py "${sandbox}:/workspace/demo/slug.py" sbx --cloud cp .\test_slug.py "${sandbox}:/workspace/demo/test_slug.py" The ${sandbox} syntax is intentional. In PowerShell, it separates the variable name from the colon used in Docker’s SANDBOX:PATH format. This is also why I am copying individual files rather than uploading the entire folder. We do not need a virtual environment, local configuration, or an accidentally included .env file for this task. Run the existing test inside the sandbox: PowerShell sbx --cloud exec --workdir /workspace/demo $sandbox python3 -m unittest discover -v If your selected image does not include Python 3, add it inside the sandbox before continuing. The existing test only covers two words separated by one space. Passing it does not mean the whitespace handling is correct yet. Step 5: Give the Agent a Narrow Task Attach to the running cloud sandbox: PowerShell sbx --cloud attach $sandbox Now give Claude a concrete task: Plain Text Work on the Python project in /workspace/demo. Update make_slug so that: - The output remains lowercase. - Leading and trailing whitespace is removed. - Consecutive whitespace becomes a single hyphen. - Spaces, tabs, and newlines are handled consistently. - Empty input returns an empty string. Add unit tests for these cases using unittest. Keep the existing test. Do not add third-party dependencies or modify files outside this project. Run the tests and summarize which files you changed. This is much more useful than asking the agent to “improve the project.” We have told it what the function should do, which edge cases matter, and how much freedom it has. There is no reason for it to introduce a framework or reorganize the project. The prompt is task guidance, though — not a security policy. File access, network access, and credentials still need the appropriate sandbox controls. Once the agent finishes, use Ctrl + backslash to detach and return to your local terminal. Detaching does not stop the cloud sandbox. Step 6: Run the Tests and Bring the Changes Back Run the test command again from your terminal: PowerShell sbx --cloud exec --workdir /workspace/demo $sandbox python3 -m unittest discover -v This executes inside the cloud sandbox. It is not running against your original local files. For this task, a straightforward implementation could look like: Python def make_slug(text): return "-".join(text.lower().split()) Calling split() without a separator handles consecutive whitespace and removes leading and trailing whitespace. Joining those words with a hyphen gives us the requested behavior. The agent may arrive at a different implementation. Read it rather than assuming that passing tests makes every change worth keeping. Create a separate folder for the returned files: PowerShell New-Item -ItemType Directory -Path .\review Copy the modified files into it: PowerShell sbx --cloud cp "${sandbox}:/workspace/demo/slug.py" .\review\slug.py sbx --cloud cp "${sandbox}:/workspace/demo/test_slug.py" .\review\test_slug.py Your original files are still untouched. If you have Git installed, compare the versions: PowerShell git diff --no-index -- .\slug.py .\review\slug.py git diff --no-index -- .\test_slug.py .\review\test_slug.py You can also compare them in your editor. Look at the tests as closely as the implementation. Did the agent actually add cases for tabs and newlines? Did it keep the original test? Did it add anything unrelated? For a real repository, I would bring the changes into a working branch and use the normal review process. The sandbox changes where the agent works. It does not replace code review. What About Web Applications? Our Python example does not start a server. If you use a web project instead, cloud sandboxes can expose an application through a public HTTPS URL. For an application already listening on sandbox port 3000: PowerShell sbx --cloud ports $sandbox --publish 3000 sbx --cloud ports $sandbox Use the URL returned by Docker. This is different from publishing a local port such as localhost:3000. In cloud mode, the command accepts the sandbox port, and Docker assigns the public URL. Note: Publicly reachable is not the same as private. Do not expose an unauthenticated admin page, secrets, or sensitive test data. Remove the exposure when you no longer need it: PowerShell sbx --cloud ports $sandbox --unpublish 3000 Step 7: Clean Up the Cloud Sandbox Before cleanup, make sure the files you want to keep are on your machine. If you want to pause rather than delete, Docker documents cloud stop as preserving the sandbox’s memory and disk state: PowerShell sbx --cloud stop $sandbox Do not assume that preserved resources have no cost. Check your plan’s billing terms. For this small exercise, we have already copied the results out, so we can remove the sandbox: PowerShell sbx --cloud rm $sandbox Confirm the removal when prompted, then list your cloud sandboxes: PowerShell sbx --cloud ls There is an important difference from my earlier article: sbx --cloud rm --all is intentionally disabled. Cloud cleanup requires explicit sandbox identifiers. That is a useful safeguard. A cloud credential may have access to more than the one environment you were experimenting with. A Few Things That Can Slow You Down If the agent cannot authenticate, check its cloud credentials. A successful local session does not prove that cloud authentication is configured. If it cannot reach a service, check the cloud network policy. Do not immediately open access to everything just to make an error disappear. If your local files have not changed, remember the workflow we used: we copied files into the cloud and copied the results back. Those copies are not a live synchronization mechanism. And if you are coming back to a running sandbox, use attach. Repeating the detached creation command gives you another sandbox, not another connection to the original one. Conclusion What I like about this addition is that it keeps the workflow familiar. We are still using sbx, still giving the agent a specific project, and still deciding what work to keep. The difference is where that work happens. Start small. Send only the files the agent needs, give it one clear task, and bring the results back into your normal development process. Once that feels comfortable, move on to a larger repository or a task that actually benefits from remote compute. And copy the changes back before the sandbox expires. A useful fix is not very useful if the only copy disappears with the environment.
Quick Summary Both Embabel and LangGraph4j let a Java developer build multi-step AI agents without leaving the JVM.Embabel hands the framework a goal and a bag of typed actions, and lets a planner decide the order on its own.LangGraph4j asks the developer to draw the exact graph of nodes and edges by hand.We will understand both philosophies through a real Embabel agent, a small Kolkata street-crossing example, a comparison table, and finally a bigger question — is Java catching up with Python in enterprise AI work? Where the Story Starts If you are a Java developer today, you are watching two worlds collide. On one side are large language models, which grew up almost entirely in Python. On the other side is enterprise Java, which has spent twenty-five years learning to build systems that banks, insurance companies, and hospitals can actually trust. Two frameworks are now trying to bring these two worlds together on the JVM: Embabel and LangGraph4j. Both help you build an "agent" — a piece of software that uses an LLM to complete a task in several steps, rather than in one single prompt. But the way they think about "steps" is completely different. That difference is what this article is about. Meeting Embabel Through a Real Piece of Code The best way to understand Embabel is to look at actual code, not a slide. Here is a small agent that gives retirement planning advice, written the Embabel way. Java package com.example.demo; import com.embabel.agent.api.annotation.AchievesGoal; import com.embabel.agent.api.annotation.Action; import com.embabel.agent.api.annotation.Agent; import com.embabel.agent.api.common.OperationContext; import com.embabel.agent.domain.io.UserInput; import java.util.Arrays; import java.util.List; @Agent(name = "RetirementPlannerAgent", description = "This agent provides retirement planning advice.") public class RetirementPlannerAgent { record RetirementUserInput(int presentAge, int targetRetirementAge, double annualIncome) { } record RetirementPlanAdvice(String advice) { } record RetirementPlanAdvices(List<RetirementPlanAdvice> advices) { } enum RiskToleranceLevel { LOW, MEDIUM, HIGH } @Action(description = "Identify the present age, target retirement age, and annual income of the user from the user message.") public RetirementUserInput identifyRetirementUserInput(UserInput userInput, OperationContext context) { String content = userInput.getContent(); return context.ai().withDefaultLlm() .creating(RetirementUserInput.class) .fromPrompt(""" Identify the present age, target retirement age, and annual income of the user from the following message and return them as a JSON object. User message: %s """.formatted(content)); } @Action(description = "Identify the risk tolerance level of the user based on the present age, target retirement age, and annual income.") public RiskToleranceLevel identifyRiskToleranceLevel(RetirementUserInput retirementUserInput, OperationContext context) { return context.ai().withDefaultLlm() .creating(RiskToleranceLevel.class) .fromPrompt(""" Identify the risk tolerance level of the user based on the following information and return it as a JSON object. Permitted values for risk tolerance level are: %s Present age: %d Target retirement age: %d Maximum years to retirement: %d Annual income: %.2f """.formatted(Arrays.toString(RiskToleranceLevel.values()), retirementUserInput.presentAge(), retirementUserInput.targetRetirementAge(), (retirementUserInput.targetRetirementAge() - retirementUserInput.presentAge()), retirementUserInput.annualIncome())); } @Action(description = "Provide retirement plan advice based on the user's risk tolerance level.") @AchievesGoal(description = "Provide retirement plan advice based on the user's risk tolerance level.") public RetirementPlanAdvices provideRetirementPlanAdvice(RiskToleranceLevel riskToleranceLevel, OperationContext context) { String systemPrompt = """ You are a retirement planning advisor. Based on the user's risk tolerance level, provide a list of retirement plan advices. """; return context.ai().withDefaultLlm() .creating(RetirementPlanAdvices.class) .fromPrompt(""" %s User's risk tolerance level: %s """.formatted(systemPrompt, riskToleranceLevel.name())); } } Now look closely at what is missing from this code. There is no method called runAgent() that calls identifyRetirementUserInput(), then identifyRiskToleranceLevel(), then provideRetirementPlanAdvice(), in that order. Nowhere did the developer type out the sequence. Instead, each @Action simply states two things: What type it needs as input (its precondition).What type it produces as output (its effect). provideRetirementPlanAdvice needs a RiskToleranceLevel. identifyRiskToleranceLevel happens to produce a RiskToleranceLevel from a RetirementUserInput. And identifyRetirementUserInput produces that RetirementUserInput from the raw UserInput. Embabel's planner looks at all this at runtime and works out, on its own, that this is the only order in which the goal (@AchievesGoal) can be reached. This is Embabel's whole philosophy in one sentence: give the framework a goal and a set of typed building blocks, and let it plan. The Same Investment Advisory Workflow, Wired by Hand in LangGraph4j Now let us build the exact same three-step advisory flow — read the user's numbers, work out risk tolerance, give advice — the LangGraph4j way. Here, the developer draws the graph explicitly, and the shared state is a plain key-value map (an AgentState) rather than Soham's strongly typed records. Java package com.example.demo; import org.bsc.langgraph4j.CompiledGraph; import org.bsc.langgraph4j.StateGraph; import org.bsc.langgraph4j.action.NodeAction; import org.bsc.langgraph4j.state.AgentState; import java.util.List; import java.util.Map; import java.util.Optional; import static org.bsc.langgraph4j.StateGraph.END; import static org.bsc.langgraph4j.StateGraph.START; import static org.bsc.langgraph4j.action.AsyncNodeAction.node_async; /** * @author Soham Sengupta * @since 2026-09-13 * @description The retirement/investment advisory workflow from the * Embabel example above, this time wired explicitly as a LangGraph4j * graph. Every step, and the order between the steps, is declared * here by the developer - there is no planner discovering it. */ public class RetirementPlannerGraph { // Step 1: pull the present age, target retirement age, and income out of free text. static class IdentifyUserInputNode implements NodeAction<AgentState> { @Override public Map<String, Object> apply(AgentState state) throws Exception { String userMessage = state.<String>value("userMessage").orElseThrow(); // In a real system: call the LLM here (say, via langchain4j) and parse // presentAge / targetRetirementAge / annualIncome out of userMessage. return Map.of( "presentAge", 32, "targetRetirementAge", 60, "annualIncome", 1200000.0); } } // Step 2: classify how much investment risk this user can reasonably take. static class IdentifyRiskToleranceNode implements NodeAction<AgentState> { @Override public Map<String, Object> apply(AgentState state) throws Exception { int presentAge = state.<Integer>value("presentAge").orElseThrow(); int targetRetirementAge = state.<Integer>value("targetRetirementAge").orElseThrow(); // In a real system: call the LLM here with these values and ask it to // return LOW, MEDIUM, or HIGH as the risk tolerance level. String riskTolerance = (targetRetirementAge - presentAge) > 20 ? "HIGH" : "MEDIUM"; return Map.of("riskTolerance", riskTolerance); } } // Step 3: turn the risk tolerance into a concrete list of investment advice. static class ProvideAdviceNode implements NodeAction<AgentState> { @Override public Map<String, Object> apply(AgentState state) throws Exception { String riskTolerance = state.<String>value("riskTolerance").orElseThrow(); // In a real system: call the LLM here to draft actual advice - suitable // Indian investment instruments for this riskTolerance level, and so on. List<String> advice = List.of("Suggested investment mix for a " + riskTolerance + " risk profile."); return Map.of("advice", advice); } } public static void main(String[] args) throws Exception { StateGraph<AgentState> graph = new StateGraph<>(AgentState::new) .addNode("identifyUserInput", node_async(new IdentifyUserInputNode())) .addNode("identifyRiskTolerance", node_async(new IdentifyRiskToleranceNode())) .addNode("provideAdvice", node_async(new ProvideAdviceNode())) .addEdge(START, "identifyUserInput") .addEdge("identifyUserInput", "identifyRiskTolerance") .addEdge("identifyRiskTolerance", "provideAdvice") .addEdge("provideAdvice", END); CompiledGraph<AgentState> workflow = graph.compile(); Optional<AgentState> result = workflow.invoke( Map.of("userMessage", "I am 32, want to retire at 60, and earn 12 lakh a year.")); result.ifPresent(state -> System.out.println(state.data())); } } Two things stand out next to the Embabel version. First, there is a main method here that explicitly lists identifyUserInput -> identifyRiskTolerance -> provideAdvice as edges — in the Embabel version, that sequence was never written down anywhere; it was worked out by the planner. Second, the shared state (AgentState) is just a bag of string keys and values, read back out with state.value("presentAge"), instead of Soham's own strongly typed RetirementUserInput and RiskToleranceLevel. For this particular workflow, which happens to be a strict straight line with no branching, LangGraph4j's graph is refreshingly easy to read top to bottom. The real difference shows up once branching enters the picture, which is exactly where our next example — crossing a Kolkata road — comes in. A Bit of History: From Servlet to Spring, Now From Spring AI to Embabel To understand why Embabel is built this way, it helps to know who built it. Embabel comes from Rod Johnson — the same person who created the Spring Framework more than two decades ago. Back in the early 2000s, enterprise Java was drowning in heavy J2EE application servers and Enterprise Java Beans. Rod was solving real problems in the finance industry at the time, found the existing tools too heavy, and wrote a book and a framework that simplified things a great deal. That framework became Spring, and it changed how an entire generation of Java developers worked. In 2025, Rod did something similar again, this time for AI agents. When he introduced Embabel to the Java community, he framed it using a comparison that Java developers will find very familiar: Spring AI is to Embabel roughly what the plain old Servlet API once was to Spring MVC. Spring AI gives you the low-level plumbing — talking to a model, building a prompt, calling a tool. Embabel sits one level above that, giving you the actual application framework — goals, actions, planning, and a proper domain model — the same kind of jump in abstraction that Spring itself brought to raw Servlets and EJBs, twenty years back. Embabel (pronounced "Em-BAY-bel") is written mainly in Kotlin, but as you can see from the retirement planner code above, it feels completely natural to use from plain Java. It is also built to sit closely with Spring, which is exactly why an existing Spring shop can pick it up without much friction. GOAP: The Planning Engine Hiding Inside Embabel The planning idea inside Embabel is not new — it is borrowed from video games, and it is called GOAP, short for Goal-Oriented Action Planning. GOAP was built to make game characters (think of soldiers in an old shooter game) decide, on their own, a believable sequence of actions to reach a goal, instead of following a fixed script. GOAP needs three things: A state of the world as it stands right now.A set of actions, each with a precondition (what must be true to run it) and an effect (what becomes true after it runs).A goal, which is simply a desired state. A search algorithm (usually the well-known A* algorithm) then works out the cheapest chain of actions that gets you from where you are to where you want to be. Embabel uses exactly this idea, but instead of asking the LLM to "think step by step" about the plan (which is often unreliable) or asking the developer to hard-code the entire flow (which is rigid), it asks a deterministic planner to search over your own typed Java or Kotlin methods. The precondition of an action is simply the input type it needs. The effect is simply the output type it produces. Your domain classes — RetirementUserInput, RiskToleranceLevel, RetirementPlanAdvices — literally become the "world state" the planner reasons about. No LLM guesswork is involved in deciding the order; the LLM is only used inside each action, for the part it is actually good at — understanding and generating language. Understanding the GOAP + OOAD Loop, With a Kolkata Traffic Signal All this can feel a bit abstract, so let us make it concrete with something every Kolkata resident understands very well — crossing a busy road. Picture Soham standing at a signal with his four-year-old son, Kit, holding his hand. It is a typical Kolkata crossing — buses, yellow taxis, and autos, and the signal, while present, is not always fully obeyed. Sometimes there is a traffic constable standing in the middle of the road, waving vehicles through by hand, overriding the signal completely. Step one — model the world as an object (this is the OOAD part). In Object-Oriented Analysis and Design, we are trained to represent a real-world situation as a class with clearly named fields. Here, the "world" Soham is observing can be written as one simple record: record RoadSituation(boolean signalGreen, boolean trafficClear, boolean handHeld, boolean policeOnDuty) { } Step two — define the goal. The goal is not "the signal is green." The goal is "Soham and Kit have reached the other side, safely." Reaching a green signal is only useful if it actually leads there. Step three — define the actions, each with its precondition and effect (this is the GOAP part). Notice that none of these actions know about each other. Each one only knows what it needs and what it produces: holdChildHand – produces handHeld = true. This is usually the very first thing a responsible parent does, well before even looking at the signal.waitForSafeSignal – needs the current RoadSituation, and produces signalGreen = true once either the light turns green or a traffic constable is present and waving pedestrians across.confirmTrafficClear – because a green signal in Kolkata does not always mean an auto will not sneak through, this action looks both ways and produces trafficClear = true.crossTheRoad (the @AchievesGoal action) – only runs once handHeld, trafficClear, and (signalGreen or policeOnDuty) are all true. Step four — let the planner loop. This is the actual "GOAP + OOAD loop": the planner looks at the current RoadSituation object, picks whichever action's precondition is already satisfied and whose effect moves the world closer to the goal, executes it, updates the RoadSituation, and checks again if the goal is reached. It keeps looping — plan, act, update state, re-check — until crossTheRoad finally fires. The beautiful part is what happens on a day when the signal itself is not working — a fairly common event in Kolkata, especially during a power cut or during Puja season when the police fully take charge of a crossing. If tomorrow you add one more action, say waitForPoliceWave, which also produces a "safe to proceed" fact, the planner will simply discover this new path on its own the next time it runs. Nobody needs to redraw anything, because nobody drew anything explicit in the first place. The Same Scenario as Embabel Code Here is a simplified sketch of the above scenario, written in the same style as the retirement planner. Treat it as a teaching example rather than a compiled, production-ready class: Java /** * @author Soham Sengupta * @since 2026-09-13 * @description A small agent that plans how Soham and his four-year-old * son Kit can safely cross a busy Kolkata road. Written purely to show * how Embabel's GOAP-style planner reasons over typed domain objects * (OOAD) to reach a goal, without the developer wiring the order by hand. */ @Agent(name = "StreetCrossingAgent", description = "Plans a safe road crossing for a parent and a young child.") public class StreetCrossingAgent { record RoadSituation(boolean signalGreen, boolean trafficClear, boolean handHeld, boolean policeOnDuty) { } record CrossingPlan(String narrative) { } @Action(description = "Hold Kit's hand before anything else - the first rule of the road.") public RoadSituation holdChildHand(UserInput userInput) { return new RoadSituation(false, false, true, false); } @Action(description = "Wait till the signal turns green, or till a traffic constable waves pedestrians across.") public RoadSituation waitForSafeSignal(RoadSituation situation, OperationContext context) { boolean safeToMove = situation.signalGreen() || situation.policeOnDuty(); return new RoadSituation(safeToMove, situation.trafficClear(), situation.handHeld(), situation.policeOnDuty()); } @Action(description = "Look right, then left, then right again - the signal alone is not a guarantee in Kolkata traffic.") public RoadSituation confirmTrafficClear(RoadSituation situation) { return new RoadSituation(situation.signalGreen(), true, situation.handHeld(), situation.policeOnDuty()); } @Action(description = "Cross only when hand is held, the way is clear, and the signal or constable allows it.") @AchievesGoal(description = "Soham and Kit have reached the other side of the road safely.") public CrossingPlan crossTheRoad(RoadSituation situation) { return new CrossingPlan( "Soham held Kit's hand tight, waited for the green man, checked both sides once more, then crossed together."); } } Now compare this with how the same scenario would look in LangGraph4j's philosophy — as an explicit graph you draw yourself, branches and all: Java package com.example.demo; import org.bsc.langgraph4j.CompiledGraph; import org.bsc.langgraph4j.StateGraph; import org.bsc.langgraph4j.action.AsyncEdgeAction; import org.bsc.langgraph4j.state.AgentState; import java.util.Map; import java.util.Optional; import static org.bsc.langgraph4j.StateGraph.END; import static org.bsc.langgraph4j.StateGraph.START; import static org.bsc.langgraph4j.action.AsyncNodeAction.node_async; public class StreetCrossingGraph { // Stand-ins for a real signal sensor and a quick look both ways. private static boolean checkSignal() { return true; } private static boolean lookBothWays() { return true; } public static void main(String[] args) throws Exception { StateGraph<AgentState> graph = new StateGraph<>(AgentState::new) .addNode("holdHand", node_async(state -> Map.of("handHeld", true))) .addNode("waitForSignal", node_async(state -> { boolean policeOnDuty = state.<Boolean>value("policeOnDuty").orElse(false); return Map.of("signalGreen", checkSignal() || policeOnDuty); })) .addNode("checkTraffic", node_async(state -> Map.of("trafficClear", lookBothWays()))) .addNode("cross", node_async(state -> Map.of("crossed", true))) .addEdge(START, "holdHand") .addEdge("holdHand", "waitForSignal") .addConditionalEdges("waitForSignal", AsyncEdgeAction.edge_async(state -> state.<Boolean>value("signalGreen").orElse(false) ? "proceed" : "retry"), Map.of("proceed", "checkTraffic", "retry", "waitForSignal")) .addConditionalEdges("checkTraffic", AsyncEdgeAction.edge_async(state -> state.<Boolean>value("trafficClear").orElse(false) ? "proceed" : "retry"), Map.of("proceed", "cross", "retry", "checkTraffic")) .addEdge("cross", END); CompiledGraph<AgentState> workflow = graph.compile(); Optional<AgentState> result = workflow.invoke(Map.of("policeOnDuty", false)); result.ifPresent(state -> System.out.println(state.data())); } } Notice the two addConditionalEdges calls — this is how LangGraph4j handles a branch: an edge action returns a label ("proceed" or "retry"), and a small map resolves that label to the actual next node. This is exactly the graph the developer must draw by hand, action by action and branch by branch, for a scenario that Embabel's planner worked out on its own. As a flowchart, that graph looks like this: Both pieces of code reach the same goal. But in Embabel, nobody drew this flowchart — the planner found it. In LangGraph4j, this flowchart is the code. If a new real-world case turns up tomorrow, Embabel's planner can absorb it automatically as long as the new action's types fit; LangGraph4j needs a human to open the graph and add a new node or edge. Comparing the Two Philosophies Aspect Embabel LangGraph4j Core idea Give a goal and typed actions; a GOAP/A* planner works out the order Developer explicitly wires nodes and edges into a graph Mental model "What do I want, and what building blocks do I have?" "What are my steps, and how do they branch?" Control flow Discovered at runtime by the planner Declared upfront by the developer Role of the LLM Used only inside actions, never for deciding sequence Can be used inside nodes; sequence is still fixed by the graph Adapting to a new case Often automatic, if a new action's types fit the gap Needs a human to add a new node or edge Tracing "why this order" Needs the planner's own logging/tooling to see the chosen path Very direct — the graph is already the flowchart Roots Kotlin-first, Java-friendly, close to Spring A faithful Java port of Python's LangGraph, works with Langchain4j and Spring AI Maturity (as of late 2026) Young, pre-1.0, moving fast Older and more widely adopted, with a large existing community When to Reach for Which Reach for Embabel when: Your goal is clear, but the exact path to it can honestly vary depending on the situation.You want a deterministic, non-LLM planner deciding the order, not the LLM guessing it.You are already deep in the Spring ecosystem and like strongly typed domain models.New cases keep appearing over time, and you would rather add one new action than redraw a graph. Reach for LangGraph4j when: You already know the exact stages of your workflow — say, a well-understood pipeline of retrieve, rerank, generate, and validate.You want the flow to be visible as an actual graph, easy to explain to a non-technical stakeholder.Your team is porting an existing Python LangGraph pipeline and wants the Java version to mirror it closely.You value a larger, more mature community with more examples to learn from, at least for now. Neither approach is "better" in an absolute sense. Embabel bets on planning; LangGraph4j bets on explicitness. Pick the one that matches how well you actually know your workflow in advance. To Conclude: Java Still Has a Say in Enterprise AI Python remains, without question, the home of AI research — the notebooks, the training loops, the enormous ecosystem of machine learning libraries were built there first, and will likely stay there. Nobody sensible is arguing Java should train the next large language model. But training a model is only one part of the story. The other part — the much bigger part, in terms of sheer lines of code running in the real world — is taking an already-trained model and safely wiring it into systems that already exist: a bank's core banking platform, an insurance company's policy engine, a hospital's records system. The overwhelming majority of that existing code, in most large enterprises, is written in Java and Spring, not Python. That is Java's home ground, built up over more than two decades. This is exactly the ground both Embabel and LangGraph4j are fighting on. Java also tends to run this kind of orchestration work faster than Python at execution time, which matters once you are calling these agents thousands of times a day inside a live enterprise system. And when your applications are already written in Java, keeping the AI layer in Java too — rather than routing every call out to a separate Python service — often turns out to be the simpler, safer choice. So, the real contest in enterprise AI is perhaps not "who trains the smarter model" — Python wins that one comfortably. It is "who can be trusted to make that model's decisions reliably inside a bank's core system, an insurance engine, or a hospital record system." That is precisely the kind of trust Java has spent two decades earning. With Rod Johnson effectively writing a sequel to his own Spring story, and with LangGraph4j bringing a proven Python pattern faithfully onto the JVM, Java is not sitting out this wave of AI. It has simply chosen to fight the battle it already knows how to win. This piece focused on Embabel's goal-and-planner philosophy. Embabel also has other ideas worth a separate deep-dive later — like its approach to agentic search and enterprise memory. A hands-on, step-by-step guide to setting up Embabel from scratch will follow as a companion piece. Here's the link to the source code: https://github.com/trainerpb/embabel-hello-world/tree/feature/revision.
In this blog, we will continue our discussion from the previous parts. If you have not read them, please read them first. Parts 1 & 2 – Caesar Cipher, Vigenere Cipher, Symmetric Encryption, AES, Convergent Encryption, IVPart 3 – Hashing, Salting, Rainbow Table Attacks, Asymmetric Encryption, RSA In Part 3, we teased a few topics for Part 4 — Envelope Encryption, PKI, and more. Today we gossip about exactly those! Let's go. First, A Quick Revisit: Convergent Encryption We discussed Convergent Encryption back in Parts 1 & 2, but let's revisit it here because it connects beautifully to Envelope Encryption. Remember? Convergent Encryption means — if you encrypt the same plaintext with the same key and the same IV, you will always get the same ciphertext. So Why Is That Useful? Imagine you work at a big company. 500 employees all upload the same file — let's say the company's HR policy PDF. If you use regular encryption (different ciphertext every time), your storage system stores 500 different encrypted copies. That's 500x storage wasted! With Convergent Encryption, since the same file + same key = same ciphertext, the storage system realizes — "Hey, I already have this encrypted file!" — and stores only ONE copy. All 500 employees point to the same encrypted blob. This is called deduplication. This is exactly how Dropbox, Google Drive, and AWS S3 save enormous amounts of storage at their scale. Another Use Case — Searching Over Encrypted Data Here is another very powerful use case of Convergent Encryption that most people don't think about — searching. Imagine you have a database where all the data is encrypted. A user wants to search for records where the email is "[email protected]." With regular encryption, every time "[email protected]" is encrypted, it produces a **different** ciphertext (because of a random IV). So to search, you would have to: Decrypt every single record in the databaseCompare the plaintextReturn the matches That is insanely expensive! Imagine doing this on a database with 100 million records. Your server will cry. Now with Convergent Encryption — "[email protected]" always produces the **same** ciphertext. So to search, you just: Encrypt the search term "[email protected]" onceLook for that ciphertext in the database — just like a normal indexed search!Return the matches No decryption needed at all! The data stays encrypted at rest, and you can still do fast, exact-match searches on it. This is called searchable encryption, and it is used in scenarios like: Encrypted databases where you still need to support queriesHealthcare systems — searching patient records without ever exposing raw dataEmail systems — searching your encrypted inbox without the server ever seeing your emails in plain text Pretty powerful, right? Same property (deterministic output) — two completely different superpowers (deduplication + searchable encryption). But wait — there's a catch. If two people can produce the same ciphertext, can someone guess your file? Yes, this is called a confirmation attack. Someone could hash a known file, compare it with stored hashes, and confirm whether you uploaded that file. So Convergent Encryption is great for performance and deduplication but is used carefully in highly sensitive scenarios. Now Let's Talk About Envelope Encryption Okay, so now we know — encryption needs keys. And those keys need to be stored somewhere safely. But here's the problem — who encrypts the key itself? If your key is lying around in plain text, a hacker who gets access to your server gets everything. So the answer is — we encrypt the key too! This is the core idea of Envelope Encryption. The Two Keys in Envelope Encryption DEK — Data Encryption Key. This is the key that directly encrypts your actual data. Think of it as the key to your diary.KEK — Key Encryption Key. This is the master key that encrypts the DEK. Think of it as the key to your locker — inside which you keep your diary key. So the flow looks like this: Your Data → encrypted with DEK → Encrypted Data DEK → encrypted with KEK → Encrypted DEK You store both — the Encrypted Data and the Encrypted DEK — together. The KEK lives safely inside a highly secure system (like AWS KMS or Vault). Real-Life Example — The Bank Locker Imagine you have an important document (your data). You put it in a box and lock it with a small key (DEK). Now you don't want to carry this small key everywhere — so you put the small key inside your bank locker (encrypt DEK with KEK). The bank locker key (KEK) stays with the bank in a highly secure vault. To read your document: Go to the bank → get your small key out (decrypt DEK using KEK)Use the small key to open the box (decrypt data using DEK) Simple! And very secure. Why Not Just Encrypt Data Directly With KEK? Two very practical reasons: 1. Performance The KEK usually lives inside a secure hardware vault or cloud service (like AWS KMS). If you send your entire 10GB file to KMS every time you want to encrypt or decrypt — that's painfully slow and expensive. Instead, you only send the tiny DEK (a few bytes) to KMS. The heavy lifting of encrypting actual data is done locally with the DEK. 2. Key Rotation Say after 6 months you want to change your encryption key (key rotation is a security best practice). Without envelope encryption — you'd have to decrypt ALL your data and re-encrypt it with a new key. Imagine doing that for terabytes of data! With Envelope Encryption — you only re-encrypt the DEK with the new KEK. Your actual data stays untouched. Much faster, much cheaper. Where Is Envelope Encryption Used? Literally everywhere in the cloud world: AWS S3 – When you enable server-side encryption on a bucketAWS RDS – When you enable encryption on a databaseGCP Cloud Storage – Envelope encryption is the defaultAzure Key Vault – Same pattern Every time you see that little "encryption enabled" checkbox on a cloud service — envelope encryption is what's happening under the hood. Now Let's Talk About PKI PKI stands for Public Key Infrastructure. In Part 3, we discussed Asymmetric Encryption — where you have a Public Key and a Private Key. Sounds great in theory. But here's a real problem. The Trust Problem Imagine Rahul wants to send an encrypted message to Priya. Priya shares her public key with Rahul. Rahul encrypts the message with Priya's public key and sends it. But wait — how does Rahul know that the public key he received is actually Priya's? What if a hacker intercepted the communication and swapped Priya's public key with their own? Rahul encrypts with the hacker's public key → hacker decrypts → reads the message. This is called a Man-in-the-Middle (MITM) Attack. We need someone that both Rahul and Priya trust, who can say — "Yes, this public key truly belongs to Priya." That trusted someone is called a certificate authority (CA). PKI — The Complete Picture PKI is a system made up of several components that together solve the trust problem. Let's go one by one. 1. Certificate Authority (CA) A CA is a trusted organization whose job is to verify identities and issue digital certificates. Think of them like the government passport office — they verify who you are and give you an official identity document (passport). Well-known CAs in the real world: DigiCert, Let's Encrypt, GlobalSign, Comodo. Your browser/OS comes pre-loaded with a list of trusted CAs. That's how your browser automatically trusts websites — because their certificates were signed by a CA your browser already trusts. 2. Digital Certificate A Digital Certificate is like a government-issued ID card for websites (or people or servers). It contains: The owner's name (e.g., google.com)The owner's Public KeyThe CA's name (who issued it)Expiry dateA digital signature from the CA When you open https://google.com — your browser checks Google's certificate. It sees the CA that signed it. It checks if that CA is in its trusted list. If yes — green light, connection is secure! That's the lock you see. 3. Digital Signature We talked about hashing in Part 3 — that you cannot reverse a hash. Digital Signatures use this + asymmetric encryption together in a clever way. When a CA wants to sign a certificate, it: Takes the certificate content and hashes itEncrypts that hash with its own Private Key That encrypted hash = Digital Signature. Anyone can verify the signature using the CA's Public Key (which is publicly available). If the decrypted hash matches the actual certificate content → the certificate is genuine and untampered! Think of it like a wax seal on an envelope. Anyone can see the seal, but only the king's ring (private key) could have made it. 4. Certificate Chain (Chain of Trust) In the real world, CAs have a hierarchy: Root CA → Intermediate CA → Your Website Certificate The Root CA is the ultimate trusted authority. It signs Intermediate CAs. Intermediate CAs sign individual website certificates. This chain is called the Chain of Trust. Why this hierarchy? Security! Root CA private keys are kept in ultra-secure, air-gapped hardware. They are almost never used directly. Intermediate CAs do the day-to-day certificate signing. If an Intermediate CA is ever compromised, it can be revoked without affecting the Root CA. Where Is PKI Used in Real Life? PKI is literally everywhere. You just don't see it because it works silently in the background. 1. HTTPS Websites (SSL/TLS) Every https:// website uses PKI. When you open your bank's website — a PKI handshake happens in milliseconds: Your browser asks the bank's server for its certificateThe bank's server sends its Digital CertificateBrowser verifies the certificate using the CA's public keyIf valid → browser and server agree on a secret key (using asymmetric encryption)All further communication uses that secret key with AES (fast symmetric encryption) That's TLS in a nutshell. And PKI is the backbone of all of it. 2. Email Signing (S/MIME) When your company sends you a digitally signed email — PKI is involved. The sender signs the email with their private key. You verify it with their public key from their certificate. You can be sure the email is genuinely from them and was not tampered with in transit. 3. Code Signing When you download an app or a software update — how does your phone/OS know it's legit and not malware? The developer signs the app with their private key. Your phone verifies it using the developer's certificate. This is why on Android you see the "Install from unknown sources" warning — no valid certificate found! 4. Government and Banking Your Aadhaar card, digital signatures on GST filings, net banking OTPs — all of these use PKI infrastructure behind the scenes. India's government runs its own CA called CCA (Controller of Certifying Authorities) under the IT Act. 5. VPNs and Internal Company Networks When you connect to your company's VPN, PKI certificates are used to verify that you are connecting to the genuine company server and not an imposter. Terms We Have Learned So Far (All 4 Parts) Cryptography, Algorithm, Plain Text, Key, Cipher TextSymmetric Encryption, Convergent Encryption, Initialization Vector (IV), Searchable EncryptionHashing, Hash/Digest, Avalanche EffectSalt, Rainbow Table Attack, Confirmation AttackAsymmetric Encryption, Public Key, Private Key, RSADEK (Data Encryption Key)KEK (Key Encryption Key)Envelope EncryptionKey Rotation, DeduplicationPKI (Public Key Infrastructure)Certificate Authority (CA)Digital CertificateDigital SignatureChain of TrustTLS/SSL That's a solid vocabulary now! Coming in Part 5... (Part 5 is in progress — stay tuned!) In the next part, we will gossip about: SSL/TLS Deep Dive – The full step-by-step TLS handshake explained simplymTLS (Mutual TLS) – How microservices talk to each other securelyHSM (Hardware Security Module) – The physical vault where the most sensitive keys liveZero-Knowledge Proofs – Proving you know something without revealing what it isAnd more... Stay tuned for Part 5! If you liked this blog, do give it a like and share it with someone who you think should learn this. Let's spread the knowledge! Read the previous parts here: Part 1 & 2 Part 3
Agile
Career Development
Methodologies
Team Management
Kill the Worker, Keep the Research: Build a Recoverable LangGraph Agent on Temporal
October 5, 2026
by Akhil Madineni
CORE
Agentic Test Creation: From Plain-Language Requirements to End-to-End Test Cases
October 2, 2026
by John Vester
CORE
AI on Top of a Dysfunctional System
October 2, 2026
by Stefan Wolpers
CORE
AI/ML
Big Data
Databases
IoT
Building Enterprise File-Heavy AI Workflows: From Secure Uploads to Governed Document Intelligence
October 5, 2026
by Dr Gopala Krishna Behara
CORE
Building Time-Series Applications With Java and InfluxDB
October 5, 2026
by Otavio Santana
CORE
Kill the Worker, Keep the Research: Build a Recoverable LangGraph Agent on Temporal
October 5, 2026
by Akhil Madineni
CORE
Cloud Architecture
Integration
Microservices
Performance
Building Enterprise File-Heavy AI Workflows: From Secure Uploads to Governed Document Intelligence
October 5, 2026
by Dr Gopala Krishna Behara
CORE
Why Time Series Databases Matter for Modern Enterprise Applications
October 2, 2026
by Otavio Santana
CORE
Part 3: End-to-End Tracing and Observability Across Goose, agentgateway, and Quarkus
October 2, 2026
by Daniel Oh
CORE
Frameworks
Java
JavaScript
Languages
Tools
YAML vs XML vs JSON: History, Trade-offs, and Where Each Wins in the Age of Agentic AI
October 5, 2026
by Kai Wähner
CORE
Building Time-Series Applications With Java and InfluxDB
October 5, 2026
by Otavio Santana
CORE
Wasm Inside Neo4j: Building the Example That Didn't Exist
October 2, 2026
by Akmal Chaudhri
CORE
Deployment
DevOps and CI/CD
Maintenance
Monitoring and Observability
Agentic Test Creation: From Plain-Language Requirements to End-to-End Test Cases
October 2, 2026
by John Vester
CORE
AI on Top of a Dysfunctional System
October 2, 2026
by Stefan Wolpers
CORE
AI/ML
Java
JavaScript
Open Source
Building Enterprise File-Heavy AI Workflows: From Secure Uploads to Governed Document Intelligence
October 5, 2026
by Dr Gopala Krishna Behara
CORE
Building Time-Series Applications With Java and InfluxDB
October 5, 2026
by Otavio Santana
CORE
Kill the Worker, Keep the Research: Build a Recoverable LangGraph Agent on Temporal
October 5, 2026
by Akhil Madineni
CORE