silicode · 2026-10-08 · 11 min

Why Tokenization Breaks Verilog Delta Cycles and Concurrency

Autoregressive language models predict tokens sequentially, treating Verilog like procedural C. Here is why delta cycles fail and how to catch race conditions.

Diagram contrasting sequential token streams with parallel hardware delta cycle evaluation queues

A recent evaluation on arXiv assessing large language models on concurrent programs (arXiv:2501.14326) confirmed what RTL leads encounter daily in code reviews. Even when frontier models achieve high pass rates on isolated algorithmic puzzles, their reliability collapses when code execution depends on fine-grained concurrent scheduling. In digital design, this failure is not a superficial syntax hallucination. It is an architectural mismatch between left-to-right causal language modeling and the stratified event scheduler defined by IEEE 1364 and IEEE 1800 standards.

For an ASIC or FPGA design lead, this failure mode carries real risk. When an engineer prompts an LLM for a multi-stage pipeline, an arbitration tree, or a decoupled FIFO interface, the model routinely emits syntactically valid SystemVerilog that compiles cleanly in Verilator or Synopsys VCS. Under the hood, however, the model often arranges assignments as if instructions execute sequentially in program counter order, exactly like C or Python. The resulting RTL passes single-cycle smoke tests but exhibits race conditions, simulation-synthesis mismatches, and delta-cycle dependency inversion once integrated into a broader SoC testbench.

Understanding the exact mechanical breakdown between token probability distributions and simulator event regions is necessary to prevent defective RTL from slipping past automated lint stages.

The Mechanical Clash Between Token Chains and Stratified Event Queues

Autoregressive transformers operate on a simple mathematical premise: given a context sequence of tokens $(t_1, t_2, \dots, t_n)$, the model computes conditional probability distributions to predict $t_{n+1}$. This process is fundamentally causal, directional, and sequential in time. The internal attention mechanism establishes dependencies across the token stream, but the generation mechanics unfold strictly one step after another.

Hardware description languages model physical silicon, where millions of logic gates evaluate concurrently under continuous voltage changes. To simulate physical concurrency on sequential microprocessors, EDA engines rely on discrete-event scheduling algorithms. Under the IEEE standard, simulation time advances only when all activity for the current time step is exhausted. Within a single time step ($t = 0$), the simulator evaluates multiple delta cycles, which are infinitesimal slices of simulation time with zero physical duration.

The standard stratified event queue splits each time step into distinct execution regions:

  1. Active Region: Simulation evaluates continuous assignments (assign), evaluates inputs, and executes blocking procedural statements (=). Non-blocking assignments (<=) evaluate their right-hand-side (RHS) expressions here, but do not update the left-hand-side (LHS).
  2. Inactive Region: Statements with zero-delay controls (#0) execute, a legacy practice that creates simulator-dependent execution races.
  3. NBA (Non-Blocking Assignment Update) Region: The simulator applies the previously evaluated RHS values to the LHS target registers.
  4. Observed Region: SystemVerilog Assertions (SVA) evaluate properties against stable signal values sampled before state updates.
  5. Reactive Region: Testbench code and verification programs execute response stimulus.
  6. Postponed Region: Read-only probes, such as $strobe and $monitor, capture the final settled state of the current time step.

When a human designer writes an RTL block, their mental model evaluates all statements within an always_ff block as simultaneous transfers taking place during the NBA update phase. The language model, by contrast, assigns high probability to patterns learned from billions of lines of sequential software. It arranges logic in an imperative order where intermediate signal variables are updated immediately and read downstream within the same delta iteration.

+-------------------------------------------------------------------------+
|                       IEEE 1800 Stratified Time Step                     |
|                                                                         |
|  [ Active Region ]                                                      |
|    - Evaluate Continuous Assignments (assign out = a & b)                |
|    - Execute Blocking Statements (var = val)                            |
|    - Evaluate RHS of Non-Blocking Assignments (stage2 <= stage1)        |
|         |                                                               |
|         v                                                               |
|  [ NBA Update Region ]                                                  |
|    - Commit Scheduled LHS Values to Target Registers                    |
|         |                                                               |
|         v                                                               |
|  [ Observed / Reactive / Postponed Regions ]                            |
|    - Sample Assertions, Execute Testbench Drivers, Monitor Settled Net  |
+-------------------------------------------------------------------------+

Because the LLM generates tokens sequentially, it frequently mixes blocking assignments into sequential clock domains or non-blocking assignments into combinatorial trees. The resulting code masks temporal races behind apparent functional correctness.

Tokenization Pitfalls in Temporal and Bit-Vector Representations

Subword tokenization schemes (such as Byte-Pair Encoding used across modern foundational models) compound the concurrency problem. Standard tokenizers treat arbitrary sequences of characters as statistical subwords rather than structured hardware constructs. Research analyzing LLM performance on time-series and hardware data (PMC11339515) shows that tokenizers split signal identifiers, vector widths, and temporal timestamps across arbitrary token boundaries.

Consider an explicit bus slicing operation:

logic [63:0] packet_header;
assign payload_tag = packet_header[47:32];

Depending on the specific vocabulary, a tokenizer may split packet_header[47:32] into tokens like ["packet", "_", "header", "[", "47", ":", "32", "]"] or merge substrings arbitrarily like ["packet_header", "[4", "7:3", "2]"]. When the model must maintain structural awareness across fifty parallel bus assignments, this subword splitting degrades the attention layer's ability to track spatial and temporal relationships across distinct signals.

More critically, the concept of a delta cycle does not exist in text tokens. When an LLM produces five successive non-blocking statements in an always_ff @(posedge clk) block:

stage_a <= data_in;
stage_b <= stage_a;
stage_c <= stage_b;
stage_d <= stage_c;

The model generates stage_b <= stage_a; by predicting tokens sequentially after stage_a <= data_in;. In software execution (e.g., C), stage_b would immediately receive the new value of data_in. In IEEE Verilog simulation semantics, stage_b captures the value that stage_a held in the previous clock cycle because the assignment update is deferred to the NBA queue.

When prompted to implement slightly more complex logic, such as an accumulator with an inline valid flag, models routinely slip into sequential procedural patterns:

// Antipattern emitted by LLMs attempting sequential accumulation
always_ff @(posedge clk or negedge rst_n) begin
    if (!rst_n) begin
        acc_reg   <= '0;
        acc_valid <= 1'b0;
    end else if (in_valid) begin
        acc_reg   = acc_reg + in_data; // INCORRECT: Blocking assignment used
        acc_valid <= (acc_reg > THRESHOLD); // Captures new acc_reg in C, but old in Verilog
    end
end

In the code above, the LLM intended for acc_valid to trigger immediately when the updated acc_reg crosses THRESHOLD. If the simulator evaluates acc_reg = acc_reg + in_data as a blocking assignment in the active region, acc_reg updates immediately. But if the model alternatively writes acc_reg <= acc_reg + in_data, the comparison acc_reg > THRESHOLD samples the value of acc_reg before the addition, delaying the assertion of acc_valid by an entire clock cycle. The model writes the tokens assuming synchronous software evaluation, creating a cycle-accurate functional bug that lint checkers might miss if blocking and non-blocking assignments are mixed carelessly.

Failure Receipts: Reproducing Concurrency Degradation

To demonstrate how autoregressive models stumble over delta cycles, consider a classic benchmark: a three-stage shift register with an integrated parity check and tap feedback. This structure requires strict preservation of previous-state variables across concurrent updates.

The following illustrative composite summarizes typical structural errors observed when prompting frontier models for high-throughput pipeline control blocks across a suite of standard micro-architectural specifications.

Structural Concurrency Failure Patterns in AI-Generated RTL

Pipeline Pattern Target Architecture Dominant LLM Failure Mode Simulator / Synthesis Symptom
Two-Stage Synchronizer Clock Domain Crossing Emits blocking assignments (=) inside always_ff, collapsing shift register to single flop Synthesis infers single flip-flop; CDC violation occurs instantly in hardware
Pipelined Ring Counter 8-bit Token Ring Uses continuous assignment dependent on un-clocked intermediate variables Delta-cycle zero-delay infinite simulation oscillation loop
Multi-Bank RAM Arbiter Round-Robin Priority Tree Procedural for loop modifying shared grant vector with blocking assignments Simulator order dependence; mismatch between Verilator and Cadence Xcelium
Decoupled Skid Buffer Ready-Valid Flow Control Reads un-latched data out of order during simultaneous push/pop transitions Single-cycle data corruption on back-to-back stall assertion
Shift Register with Parity Linear Feedback Shift Reg Evaluates feedback polynomial using post-updated register state in NBA queue Parity calculation is phase-shifted by one cycle relative to output stream

Note: This block represents an illustrative composite of failure categories derived from standardized lint and simulation evaluations across multi-agent RTL generation workflows.

The Shift Register Breakdown

When a model is asked to generate a concurrent pipeline stage with registered outputs, it frequently produces code resembling this pattern:

// Defective AI Generation: Sequential Bleed in Synchronous Logic
module pipeline_hazard_example (
    input  logic        clk,
    input  logic        rst_n,
    input  logic [15:0] d_in,
    output logic [15:0] d_out,
    output logic        parity_err
);

    logic [15:0] r1, r2, r3;

    always_ff @(posedge clk or negedge rst_n) begin
        if (!rst_n) begin
            r1         <= '0;
            r2         <= '0;
            r3         <= '0;
            parity_err <= 1'b0;
        end else begin
            r1 = d_in;                  // Flaw 1: Blocking assignment in clocked block
            r2 = r1 ^ 16'hA5A5;         // Flaw 2: Evaluates instantly in Active region
            r3 <= r2;                   // Flaw 3: Non-blocking scheduled for NBA region
            parity_err <= ^r3;          // Flaw 4: Samples r3 prior to NBA update
        end
    end

    assign d_out = r3;

endmodule

In this output, synthesis tools such as Synopsys Design Compiler or Yosys will interpret r1 and r2 not as hardware pipeline registers, but as transparent combinational wires feeding directly into the r3 register stage. Two intended clock cycles of pipeline latency vanish. The static timing analyzer will flag a critical path violation because combinational XOR logic has been folded into the input path of r3, increasing data path delay ($T_{data}$) and destroying the designed clock frequency target.

The Rule Set: Enforcing Hardware Concurrency

To prevent language models from emitting procedural software antipatterns in hardware descriptions, engineering teams must constrain the model using strict architectural guardrails, explicit style schemas, and post-generation automated linting passes.

1. The Two-Process (Gaisler) FSM Constraint

For control paths and state machines, enforce a strict separation between sequential state registers and combinational next-state logic. Prompting the model to use the classic Two-Process method eliminates ambiguity regarding variable updates:

  • The Sequential Process: Contains strictly non-blocking assignments (<=), updating current state registers (state_q <= state_d) and nothing else.
  • The Combinational Process: Contains purely blocking assignments (=), computing next-state values (state_d) and control outputs based on state_q and primary inputs.

When the LLM is barred from computing intermediate logic inside the sequential block, the risk of blocking assignment contamination drops to zero.

2. Zero-Tolerance Lint Rules

Configure commercial and open-source linters to halt automated pull request pipelines when sequential-combinational mixing appears. Essential lint checks include:

  • Verilator BLKSEQ: Triggers an error whenever a blocking assignment (=) appears inside a clocked always_ff or always @(posedge clk) block.
  • Verilator COMBDLY: Triggers an error when delayed or non-blocking assignments (<=) appear inside purely combinational always_comb or always @(*) blocks.
  • Synopsys SpyGlass STARC05-2.1.5.3: Prohibits mixing blocking and non-blocking assignments within the same procedural block.
  • SpyGlass NoMixedAssign-ML: Rejects any module where a single signal is driven by both continuous assignments and procedural assignments.
+-------------------------------------------------------------------------+
|                       Prompt-to-GDSII Lint Filter                       |
|                                                                         |
|   [ LLM Output Stream ]                                                 |
|             |                                                           |
|             v                                                           |
|   [ AST Static Parse ]  --> Rejects mixed '=' and '<=' in always blocks |
|             |                                                           |
|             v                                                           |
|   [ Verilator BLKSEQ ]  --> Flags blocking assignments in always_ff     |
|             |                                                           |
|             v                                                           |
|   [ Formal Check ]      --> Verifies cycle-accurate pipeline latency    |
|             |                                                           |
|             v                                                           |
|   [ Verified Synthesizable RTL ]                                        |
+-------------------------------------------------------------------------+

3. Prompt-Level Structural Invariants

Never prompt an LLM with free-form requests like "Write a Verilog FIFO." Provide a deterministic module signature and an explicit operational invariant section that explicitly defines clock-cycle latency and signal assignment semantics.

An effective system prompt framing includes:

You are an ASIC RTL engineer writing synthesizable IEEE 1800-2017 SystemVerilog.
You must adhere to the following concurrency rules:
1. All sequential logic MUST use 'always_ff @(posedge clk or negedge rst_n)' with non-blocking assignments ('<=').
2. No blocking assignments ('=') are permitted inside sequential blocks under any circumstances.
3. Combinational logic MUST use 'always_comb' or continuous 'assign' statements with blocking assignments ('=').
4. Do not perform intermediate variable updates inside sequential blocks; allocate dedicated registers for pipeline stages.
5. Guarantee that every multi-stage register transfer preserves state across distinct delta cycles.

What This Means for Silicode

At Silicode, the objective is to eliminate the gap between plausible conversational code and tape-out-grade silicon. Generating RTL requires more than predicting the next probable token; it demands rigorous validation against physical hardware semantics.

Silicode's compilation pipelines do not rely on raw frontier models generating unverified text streams. By pairing generative reasoning with deterministic Abstract Syntax Tree (AST) validation, automated linting sweeps via Verilator, and formal equivalence checking, Silicode enforces IEEE stratified scheduling rules before any code touches a repository. This guarantees that non-blocking updates, delta-cycle progressions, and clock domain boundaries behave identically in simulation, static timing analysis, and physical synthesis.

Direct Q&A: Why Language Models Struggle with Hardware Concurrency

Question: Why do modern LLMs default to sequential software logic when generating hardware description languages?

Answer: Autoregressive language models predict tokens sequentially based on causal text training dominated by imperative software languages (C, Python, Java). They lack an intrinsic concept of the IEEE stratified event queue, where concurrent processes evaluate across zero-duration delta cycles. As a result, they frequently mix blocking and non-blocking assignments, treating hardware descriptions as step-by-step execution scripts rather than spatial gate topologies.

Actionable Takeaways for Design Leads

If your silicon team is integrating generative AI into RTL drafting workflows, implement these three verification steps immediately:

  1. Enforce AST-level lint gates: Add -Werror-BLKSEQ and -Werror-COMBDLY to your standard Verilator continuous integration runs to reject code that violates assignment semantics before it reaches verification engineers.
  2. Mandate the two-process FSM style: Restrict generative tools to pure Gaisler-style state machines where sequential state transitions and combinational output decodes reside in separate, dedicated blocks.
  3. Run multi-simulator delta checks: Execute unit tests across two distinct simulation engines (such as Verilator and an event-driven tool like Icarus Verilog, VCS, or Questa) to verify that signal propagation does not rely on simulator-specific event-queue ordering.

By treating AI-generated RTL as unverified structural proposals that require strict semantic confinement, engineering teams can capture the drafting speed of generative models without compromising functional verification or tape-out schedules.

Sources

More Silicode Insight

VerilogRTL VerificationEDADigital Design