NVIDIA has released Nemotron 3 Ultra, a 550-billion-parameter Mixture-of-Experts model running 55 billion active parameters on a hybrid Transformer-Mamba architecture. With support for a 1-million-token context window and up to 5.9 times higher inference throughput against comparable open frontier weights, NVIDIA is explicitly positioning the model for long-running agentic workloads, including automated register-transfer level (RTL) coding and multi-turn hardware design orchestration.
Simultaneously, the research community is moving aggressively beyond single-prompt Verilog generation. Frameworks like Dr. RTL (accepted at ICCAD 2026) and HORIZON treat RTL design not as a text completion exercise, but as a closed-loop iterative process. In these frameworks, sub-agents write HDL, run downstream EDA tools (synthesis, linting, and formal equivalence checkers), parse the resulting log files, and rewrite the Verilog to eliminate slack violations or syntax errors.
For front-end RTL engineers, FPGA architects, and design leads managing tight tape-out schedules, this shift from single-shot text generation to agentic tool-use loops represents a real architectural change. It also creates a sharp division between designs that can be automated reliably and logic that breaks silently under physical synthesis.
Understanding where agentic models succeed and where they fail requires separating clean algorithmic logic from the physical constraints of silicon implementation.
The Mechanical Shift to Agentic Hardware Loops
Single-shot LLM code generation has a low ceiling in chip design. An engineer who prompts a general-purpose model for an AXI crossbar or an arithmetic block usually receives code that looks plausible at first glance. Once passed through Verilator, Synopsys SpyGlass, or Vivado, that code routinely reveals syntax bugs, undeclared nets, non-blocking assignment mix-ups, or unhandled protocol edge cases.
Agentic workflows attempt to resolve this by mimicking the human engineer's terminal loop:
- An orchestration agent reads an architectural specification or register map.
- A generation sub-agent produces an initial Verilog or SystemVerilog implementation.
- A tool execution harness runs synthesis (for example, via Yosys, Cadence Genus, or Synopsys Design Compiler) and lint checks.
- An analysis agent parses the timing report (
timing_word.jsonor standard static timing analysis logs) and identifies the top critical paths or violated setup times. - An optimizer sub-agent rewrites specific RTL blocks to reduce logic depth or fix protocol errors.
- A formal agent runs Sequential Equivalence Checking (SEC) to confirm that functional behavior remains identical across iterations.
+----------------------------------------------------------------------+
| Agentic RTL Closed Loop |
| |
| +--------------------+ RTL Source +--------------------+ |
| | Generative Agent | -------------------> | Lint & Synthesis | |
| | (Nemotron 3 Ultra) | | (Yosys / Synopsys) | |
| +--------------------+ +--------------------+ |
| ^ | |
| | Error Diagnostics / Timing Slacks | Log Files |
| | v |
| +--------------------+ Formal Pass/Fail +--------------------+ |
| | Critical Path | <------------------- | Equivalence Check | |
| | Optimizer Agent | | (SEC / Formality) | |
| +--------------------+ +--------------------+ |
+----------------------------------------------------------------------+
The architectural choices in Nemotron 3 Ultra, specifically the hybrid Mamba-Transformer layers, target the primary computational bottleneck of this multi-turn loop. Mamba state-space layers allow linear-time sequence processing over long context windows. When an agent must ingest fifty pages of synthesis log outputs, detailed timing path reports, and multi-file module hierarchies across twenty iterations, standard quadratic attention explodes in compute cost. High inference throughput (measured at 1.6x to 5.9x faster than dense or standard MoE counterparts) makes twenty-turn synthesis loops economically feasible for engineering teams.
Yet, faster iteration over bad assumptions merely generates invalid silicon faster. The real test is where the underlying reasoning breaks down.
Where Agentic RTL Succeeds: Standard Control and Bus Logic
Iterative agentic models handle structured, self-contained, single-clock designs with high competence. When the problem space consists of translating deterministic protocol rules into finite state machines (FSMs), the combination of fast generation and compiler feedback converges rapidly.
1. AXI4-Lite and APB Peripheral Slaves
Generating register read/write decoders for memory-mapped peripherals is mechanical. Given a clear address map and register field definitions, Nemotron 3 Ultra and multi-agent harnesses can generate complete AXI4-Lite slave interfaces. When the linter flags missing default assignments for unmapped address spaces or improper handshaking on bvalid/bready, the feedback loop fixes the assignment within two to three turns.
2. Standard FIFO and Stream Buffers (Synchronous)
Single-clock circular buffers, shift registers, and skid buffers with standard ready/valid handshakes converge cleanly. The agent can write the circular pointer logic, generate empty/full flags, and handle backpressure without human intervention. Verification harnesses testing these blocks against simple directed testbenches usually pass with minimal iterations.
3. Protocol Framing and Packet Header Parsing
State machines that parse Ethernet, UART, or custom packet headers based on byte-counter thresholds are well suited to autoregressive models. The structures follow standard patterns that appear thousands of times in open-source IP repositories. Lint and simulator feedback can easily catch off-by-one counter errors or unassigned output flags.
The First Wall: Clock Domain Crossing (CDC)
Where agentic generation begins to fail consistently is at the boundary between asynchronous domains. Large language models treat Verilog as text with syntax rules. They do not maintain an internal physical graph of clock tree networks, phase relationships, or silicon metastability physics.
When asked to build a multi-clock FIFO or pass control flags between unrelated clock domains, agentic models exhibit classic failure patterns.
INCORRECT CDC PATTERN (Common LLM Output)
Fast Clock Domain (clk_a) Slow Clock Domain (clk_b)
+-----------------------+ +-------------------------+
| Multi-bit Bus | -----------------> | Two-Stage Flop Sync |
| (e.g. gray_ptr[3:0]) | Data Skew Risk | (sync_reg1 -> sync_reg2)|
| Changes multiple bits | | Latches intermediate |
| simultaneously | | invalid bit states! |
+-----------------------+ +-------------------------+
CORRECT CDC SYNCHRONIZATION
Fast Clock Domain (clk_a) Slow Clock Domain (clk_b)
+-----------------------+ +-------------------------+
| Binary to Gray | | Two-Stage Flop Sync |
| Conversion | -----------------> | Only 1 bit transitions |
| (Single-bit changes) | Gray Bus [3:0] | per clock period |
+-----------------------+ +-------------------------+
The Multi-Bit Convergence Trap
When transferring a multi-bit pointer or status vector across domains, standard LLM agents frequently instantiate a basic two-stage flip-flop synchronizer directly on the multi-bit bus. A two-flop synchronizer works for a single-bit quasi-static control line. On a multi-bit bus, route delays and silicon process variations cause bit skew. The receiving domain latches intermediate, invalid states, causing catastrophic functional failure.
Even when prompted explicitly to use Gray-coded pointers, models regularly make subtle mistakes. They convert the pointer to Gray code, pass it through flip-flops, and then attempt to perform arithmetic (such as calculating full/empty capacity thresholds) directly on the Gray-coded vector in the destination domain without decoding it back to binary. Basic unit simulation might pass if the test vectors do not trigger simultaneous bit flips under randomized jitter, masking a severe bug that only appears on silicon.
Missing Physical Constraints (SDC)
An RTL engineer knows that writing CDC logic is only half the job. The design requires precise Synopsys Design Constraints (SDC) to prevent the static timing engine from trying to close setup and hold times across asynchronous domains while constraining data path skew:
# Required constraint for CDC pointer synchronizers
set_max_delay -from [get_cells rptr_reg*] -to [get_cells rptr_sync_reg*] 2.5 -datapath_only
set_bus_skew -from [get_cells rptr_reg*] -to [get_cells rptr_sync_reg*] 0.8
Agentic models rarely generate paired SDC constraints alongside Verilog modules. When an agent optimizes RTL solely against synthesis tool output without comprehensive CDC linters like SpyGlass CDC or Questa CDC in the loop, it remains blind to these crossing violations.
The Second Wall: Inferred Latches and Latch Loops
In synchronous design, every storage element must map directly to a clock-edge-triggered flip-flop. Unintentional transparent latches waste area, degrade timing closure, and create massive testability problems for Automated Test Pattern Generation (ATPG).
Agentic optimization loops frequently infer latches when attempting to fix complex combinational logic paths. Consider a multi-state decoder handling an SoC crossbar:
// Vulnerable agent-generated combinational block
always @(*) begin
case (current_state)
STATE_IDLE:
out_data = in_a;
STATE_ACTIVE:
if (valid_req)
out_data = in_b;
// Missing else branch: out_data retains state!
STATE_ERROR:
out_data = 32'hDEAD_BEEF;
// Missing default case
endcase
end
When a synthesis tool compiles this block, it must preserve out_data when current_state is STATE_ACTIVE but valid_req is low. The tool instantiates a transparent latch.
When the agent receives a synthesis log showing a latch warning, its typical fix is to add a default assignment at the top of the always block. While this eliminates the latch warning, it frequently introduces priority inversion or changes cycle-accurate behavior on downstream buses. In closed-loop systems without formal equivalence checking against the original golden specification, these local fixes cause semantic drift.
The Third Wall: Timing Optimization and Multi-Cycle Paths
When human engineers optimize RTL for timing closure, they analyze the physical path from launch flop to capture flop. If the critical path contains a deep adder tree or wide multiplexer, the engineer decides whether to:
- Pipeline the path by adding a register stage (introducing a 1-cycle latency penalty).
- Restructure the boolean logic (e.g., using carry-select or tree structures).
- Declare a multi-cycle path if the control logic guarantees that the data path has two or more clock cycles to settle.
Frameworks like Dr. RTL show that agents can successfully apply 47 or more hard-coded optimization strategies (such as arithmetic re-association and resource sharing). However, when the agent must make micro-architectural trade-offs, it struggles with the systemic consequences.
+-----------------------------------------------------------------------------+
| Illustrative Performance Comparison: |
| Single-Shot LLM vs. Tool-Grounded Agentic Iteration |
| |
| (Data compiled as an illustrative composite from recent benchmark studies, |
| including Dr. RTL ICCAD'26 metrics and CVDP hardware agent evaluations) |
+------------------------------------+-------------------+--------------------+
| Design Category | Baseline LLM | Agentic Loop |
| | (Single-Shot Ver) | (Synthesis + SEC) |
+------------------------------------+-------------------+--------------------+
| AXI4-Lite Register Decoder | 72% Syntax/Func | 98% Tape-Out Ready |
| Synchronous Skid Buffer | 64% Functional | 94% Tape-Out Ready |
| Multi-Clock Asynchronous FIFO | 18% Functional | 32% (CDC Failures) |
| 64-Bit FPU Pipeline (Timing-closed)| 12% Timing Met | 46% (SEC Failures) |
| SoC Arbiter with Round-Robin QoS | 45% Functional | 88% Functional |
+------------------------------------+-------------------+--------------------+
The Latency vs. Equivalence Conflict
If an agent inserts a pipeline register to break a critical path, it changes the latency of the module. When the Sequential Equivalence Checker (SEC) runs against the reference design, it immediately flags a non-equivalence error because output values appear one cycle later.
If the agent is forbidden from adding pipeline stages, it is forced to restructure combinational logic within a single clock cycle. Autoregressive models lack spatial awareness of cell placement and routing congestion. An agent may rewrite a wide multiplexer into a nested binary tree to reduce theoretical logic depth, only for the backend place-and-route tool to suffer severe routing congestion around dense local multiplexer clusters, worsening actual post-route slack.
Evaluating Agentic Tools: A Checklist for RTL Leads
Before adopting frontier models like Nemotron 3 Ultra or deploying multi-agent RTL frameworks into active project pipelines, engineering managers must evaluate tools against silicon realities rather than text benchmarks.
1. Does the Execution Harness Include Asynchronous Linting?
Standard Verilator or Icarus Verilog runs do not check for clock domain crossing errors or false-path timing violations. If your agentic loop only runs compilation and basic unit simulation, it will generate code containing hidden metastability and race conditions. Insist on integrating static CDC checkers (e.g., SpyGlass, VC CDC, or open-source equivalents) into the agent feedback loop.
2. Is Formal Sequential Equivalence Checking Enforced?
An agent that rewrites RTL to meet timing constraints must not alter functional behavior. The workflow must include automated SEC after every code transformation. If an agent restructures an arithmetic pipeline, formal tools must verify that all state outputs match the reference model across all reachable states.
3. Are SDC and UPF/CPF Constraints Part of the Output?
Modern silicon design is defined as much by its timing and power intent as by its RTL. If an agent generates a multi-voltage block or a cross-domain bridge, it must produce the corresponding SDC (Synopsys Design Constraints) and UPF (Unified Power Format) files. RTL without matching timing constraints cannot be handed off to physical backend teams.
4. What Is the Compute and Token Cost per Valid Module?
Nemotron 3 Ultra's 55B active parameters provide fast inference, but long multi-turn trajectories across 20 iterations consume substantial compute. Teams must track token expenditure per closed path. If an agent consumes 200,000 tokens trying to iteratively optimize an arithmetic carry chain that a junior engineer could resolve with a standard Synopsys DesignWare macro in ten minutes, the economics do not hold.
The Path Forward for Automated Silicon Design
Frontier reasoning models and fast inference architectures like Nemotron 3 Ultra are establishing a new baseline for hardware automation. They make standard register maps, bus decoders, protocol adapters, and basic state machines vastly faster to draft and iterate on.
However, production tape-outs do not fail because an engineer typed an AXI decoder slowly. They fail because of unexpected clock domain skew, unconstrained multi-cycle paths, unhandled reset synchronizer de-assertion, and post-layout timing violations.
Hardware design platforms must bridge this gap by grounding generation models in rigorous verification engines. Systems like Silicode are built around this reality: ensuring that generated RTL is never treated as complete until it is proven against formal assertions, timing closure requirements, and physical constraints.
For design teams building modern ASICs and complex FPGAs, the winning strategy is clear. Treat agentic models as fast, highly capable junior engineers for standard, single-domain logic. Keep CDC verification, physical timing analysis, and architectural partitioning under the strict control of verified toolchains and experienced sign-off engineers.
Sources
- NVIDIA Developer Blog: Nemotron 3 Ultra in Agentic RTL Coding
- NVIDIA Research: Nemotron 3 Ultra Technical Overview
- Dr. RTL: Autonomous Agentic RTL Optimization through Tool-Grounded Self-Improvement (arXiv:2604.14989)
- Dr. RTL ICCAD'26 GitHub Repository
- Empowering Chip Designers through LLMs (arXiv:2601.13815)
