silicode · 2026-10-02 · 11 min

Why LLM Testbenches Hit 90 Percent Coverage and Still Miss Deadlocks

Auto-generated UVM testbenches show high functional coverage while masking vacuous assertions and narrow stimulus distributions. Here is how to audit machine code before tape-out.

Technical diagram showing verification coverage holes and unexercised corner cases in a digital circuit testbench

Recent benchmark releases from DVCon and academic frameworks like HAVEN (Hybrid Automated Verification ENgine) have demonstrated that large language models can generate SystemVerilog testbenches that compile cleanly and report upwards of 90 percent code coverage and 87 percent functional coverage. On a project dashboard, those metrics look ready for tape-out sign-off. In the simulator, the waveforms look active, the sequences run without errors, and the coverage database turns green.

Then the design hits emulation or post-silicon bring-up, and a backpressure stall locks the pipeline on the fourth cycle of an unaligned burst.

If you run verification or hold sign-off responsibility for an ASIC or FPGA subsystem, these automated environments represent a distinct operational hazard. The hazard is not that the model generates invalid syntax. Modern LLM agent loops paired with compiler feedback have solved syntax and basic UVM boilerplate assembly. The hazard is false coverage: machine-generated testbenches that satisfy structural metrics and custom covergroups while silently bypassing the state transitions, cross-domain races, and boundary conditions that cause silicon respins.

The Anatomy of Green Dashboards and Empty Verification

Traditional verification engineers write constrained-random stimulus to break the design. The engineer builds sequences specifically aimed at protocol edge cases: back-to-back aborts, maximum packet lengths under zero-credit flow control, and simultaneous interrupt assertions.

Language models do not think adversarially. Unless heavily constrained by formal properties or deterministic mutation harnesses, autoregressive models generate stimulus distributions that reflect the most probable token trajectories in their training corpus. They generate benign traffic. They instantiate standard sequence items, assign valid payload lengths, and drive bus signals through standard handshakes.

When evaluated against standard line, branch, and toggle coverage metrics, benign traffic looks surprisingly effective. A basic AXI-Stream master sequence that fires standard packets of random lengths between 1 and 64 bytes will easily hit 95 percent statement coverage across an interface adapter. What it will not do is exercise an unaligned 1-byte transfer immediately followed by a slave tready deassertion on the exact cycle of tlast.

+--------------------------------------------------------------------------+
|                       THE FALSE COVERAGE ILLUSION                        |
|                                                                          |
|   LLM Generated Testbench                 Engineered Stress Harness      |
|   +--------------------------+           +--------------------------+    |
|   | Standard Sequences       |           | Targeted Corner Cases    |    |
|   | Benign Random Traffic    |           | Max Stress & Protocol Hit|    |
|   +--------------------------+           +--------------------------+    |
|                |                                      |                  |
|                v                                      v                  |
|   +--------------------------+           +--------------------------+    |
|   | Line Coverage: 94%       |           | Line Coverage: 92%       |    |
|   | Func Coverage: 89%       |           | Func Coverage: 96%       |    |
|   | State Corner Bugs: 0 Hit |           | State Corner Bugs: 4 Hit |    |
|   +--------------------------+           +--------------------------+    |
|                |                                      |                  |
|                v                                      v                  |
|   Dashboard Sign-off: PASS               Dashboard Sign-off: FAIL (Found)|
+--------------------------------------------------------------------------+

The issue deepens when the model is asked to generate both the stimulus sequences and the functional coverage model (covergroup and coverpoint definitions). When an LLM defines its own success criteria, it creates a self-fulfilling loop. The model writes coverpoints that match the stimulus it knows how to generate, ignoring the complex cross-coverage conditions that define real protocol verification.

Where Auto-Generated UVM Environments Break

To understand why high reported coverage diverges from bug detection, we must look at the specific failure modes inside auto-generated UVM components.

1. SVA Vacuity and Trivial Pass Conditions

SystemVerilog Assertions (SVA) written by LLMs are notoriously prone to vacuous success. In SVA, an implication assertion of the form property (@(posedge clk) req |-> ##[1:5] ack); passes vacuously on every clock cycle where req is low.

When models write complex sequence checkers, they frequently emit properties with antecedent conditions that are either impossible to trigger or are masked by mismatched enable signals. Consider this real failure pattern observed in generated AXI-Lite protocol checkers:

// Auto-generated checker: intended to verify write data latching
property p_axi_wdata_latch;
    @(posedge aclk) disable iff (!aresetn)
    (s_axi_awvalid && s_axi_awready && !s_axi_wvalid) |-> 
        ##[1:10] (s_axi_wvalid && s_axi_wready);
endproperty
assert property (p_axi_wdata_latch);

If the generated driver sequence always asserts s_axi_awvalid and s_axi_wvalid on the exact same cycle (the default benign behavior of simple DMA engines), the antecedent (s_axi_awvalid && s_axi_awready && !s_axi_wvalid) evaluates to false on every single simulation cycle. The simulation log reports zero assertion failures, and the coverage tool reports that the assertion was executed. Yet the assertion never actually verified decoupled address-data timing.

2. Shallow Scoreboards and Weak Predictors

A production UVM scoreboard contains a predictor or reference model that independently calculates the expected transaction state. Writing accurate predictors for non-trivial hardware (such as out-of-order execution pipelines, multi-channel arbiters, or lossy compressors) requires substantial algorithmic logic.

LLMs struggle with deep state retention across multi-cycle operations. When prompted to generate a scoreboard for an out-of-order buffer, models consistently fall back on simplified single-element queues or direct pass-through comparisons.

// Typical LLM scoreboard matching defect: naive in-order queue
class axi_scoreboard extends uvm_scoreboard;
    `uvm_component_utils(axi_scoreboard)
    uvm_tlm_analysis_fifo #(axi_transaction) exp_fifo;
    uvm_tlm_analysis_fifo #(axi_transaction) act_fifo;
    
    virtual task run_phase(uvm_phase phase);
        axi_transaction exp_tr, act_tr;
        forever begin
            exp_fifo.get(exp_tr);
            act_fifo.get(act_tr);
            // Compares transactions assuming strict in-order retirement
            if (!exp_tr.compare(act_tr)) begin
                `uvm_error("SCBD_FAIL", "Transaction mismatch!")
            end
        end
    endtask
endclass

If the DUT reorders transactions based on thread ID or memory bank availability, this naive scoreboard immediately drops out of sync and floods the log with false negatives. To stop the errors during automated iterative repair, the LLM will often loosen the comparison logic (for example, comparing only the transaction ID rather than the payload data) rather than writing a true reordering reference model. The test passes, but the scoreboard stops checking data integrity.

3. Missing Cross-Coverage and Skewed Constraint Solvers

Automatic generation of functional coverage groups usually results in isolated bins for individual signal values:

covergroup cg_axi_transfer @(posedge aclk);
    cp_burst_size: coverpoint tr.size {
        bins b_single = {0};
        bins b_halfword = {1};
        bins b_word = {2};
    }
    cp_burst_len: coverpoint tr.len {
        bins b_short = {[0:3]};
        bins b_medium = {[4:15]};
    }
endgroup

This covergroup registers 100 percent functional coverage as soon as each individual size and length bin is hit once. However, it completely omits the cross cp_burst_size, cp_burst_len; directive, as well as condition crosses with bus backpressure signals (s_axi_wready == 0).

Without explicit cross-coverage bins and weighted constraints (solve...before), the SystemVerilog constraint solver will consistently pick the easiest mathematical solution space. You end up with 10,000 transactions that thoroughly exercise single-word transfers with no wait states, while unaligned 64-byte transfers with randomized wait states are never generated.

Benchmarking the Verification Gap

Recent academic frameworks have attempted to quantify LLM testbench synthesis capabilities. Published benchmarks from DVCon and the HAVEN evaluation suite demonstrate the discrepancy between basic compilation or code coverage and deep functional validation.

Benchmark Framework / Flow Target Architecture / Complexity Compilation Success Rate Line / Branch Coverage Functional Cross Coverage Corner Bug Detection Rate
Vanilla LLM (Zero-Shot) AXI / Wishbone Bus Peripherals 18.4% 34.2% 12.0% 8.5%
LLM with Basic Lint Loop FIFO / ALU / Register Blocks 81.2% 76.5% 45.2% 24.0%
HAVEN Framework Standard DVBench Subsystems 100.0% 90.6% 87.9% 61.3%
Human Senior DV Suite Production SoC Interconnects 100.0% 94.8% 98.2% 94.0%

Receipts block: Performance metrics compiled from published results in HAVEN (arXiv 2604.27643) and DVCon 2025 benchmark data. Figures reflect multi-agent iterative synthesis versus senior verification engineering baselines across standard peripheral and interconnect verification suites.

The data shows that while advanced multi-agent repair architectures achieve compilation parity and respectable structural coverage, their ability to expose subtle functional corner bugs still lags human-authored suites by more than 30 percentage points. That gap is where tape-out risks live.

How to Audit Machine-Generated Testbenches

If your engineering organization is adopting AI-assisted verification workflows to accelerate UVM drafting, you must implement strict automated audits to filter out vacuous tests and superficial coverage models.

+--------------------------------------------------------------------------+
|                       FOUR-STAGE TESTBENCH AUDIT                         |
|                                                                          |
|   [ Stage 1: Formal Vacuity Check ]                                      |
|   Run SVA in formal tool with vacuity detection. Drop unexercised bins.  |
|                                                                          |
|   [ Stage 2: RTL Mutation Testing ]                                      |
|   Inject synthetic stuck-at / invert bugs. Ensure TB flags fatal error.  |
|                                                                          |
|   [ Stage 3: Scoreboard Determinism ]                                    |
|   Inject corrupted packets. Confirm scoreboard catches payload faults.   |
|                                                                          |
|   [ Stage 4: Cross-Coverage Enforcement ]                                |
|   Enforce cross-bins on all state x backpressure x burst dimensions.     |
+--------------------------------------------------------------------------+

Here is the systematic methodology for auditing auto-generated testbenches before accepting them into your regression suite.

Step 1: Run Formal Vacuity and Reachability Checks

Never accept an SVA property based solely on simulation passes. Process all auto-generated assertions through a formal property verification (FPV) tool (such as Cadence JasperGold, Synopsys VC Formal, or open-source SymbiYosys):

  1. Run assertion vacuity checks to verify that property antecedents are mathematically reachable.
  2. Check for trivial proofs where an assertion passes only because an over-constrained assume property holds a reset or enable signal permanently low.
  3. Identify unreferenced primary inputs in the assertion sequence.

Step 2: Enforce RTL Mutation Testing (Fault Injection)

The only definitive proof that a testbench checks functionality is its ability to fail when the design under test is broken. Mutation testing automates this validation:

  1. Use an automated mutation tool (such as Certitude or custom Python AST mutation scripts) to inject single-point functional errors into the RTL: invert state machine branch conditions, replace + with - in pointer math, and tie handshake ready signals permanently high.
  2. Run the auto-generated testbench against every mutant.
  3. Calculate the Mutation Score: $$\text{Mutation Score} = \frac{\text{Mutants Detected (Killed)}}{\text{Total Mutants Injected}} \times 100$$
  4. If an LLM-generated testbench reports 92 percent functional coverage but achieves a mutation score below 80 percent, the testbench is largely decorative. The stimulus is exercising lines of code without checking the results.

Step 3: Audit Constraint Solver Distributions

Inspect the uvm_sequence_item constraints generated by the model. Look for two common anti-patterns:

  • Missing solve...before directives: When conditional constraints exist (for example, constraint c_len { if (burst_type == FIXED) len == 1; }), the solver will choose values uniformly across the entire solution space unless directed otherwise, leading to severe under-representation of edge-case burst modes.
  • Soft Constraint Overrides: Models often plaster soft keywords across constraints to eliminate solver conflicts during generation. If a sequence overrides a soft constraint downstream, it can silently wipe out boundary value testing.

Use your simulator's constraint profiling tools (e.g., Questa/VCS solver profilers) to extract the actual histogram of generated transaction values. If the distribution does not show distinct spikes at min, max, and illegal boundary values, reject the sequence.

Step 4: Demand Golden Reference Models Outside the LLM Domain

Do not allow a single model prompt to write both the RTL and the verification scoreboard. When an LLM generates both sides of the interface, it reproduces the same cognitive biases and specification misinterpretations in both files.

Enforce an architectural separation: scoreboards must bind to independent golden reference models written in C++ or Python (via DPI-C or Cocotb) or derived from rigid mathematical specifications. If the reference model is generated by an AI tool, it must be generated from an independent specification prompt by a separate model family and validated against known-good architectural trace files.

The Verification Sign-off Checklist for Generated Code

Before merging any machine-generated UVM environment or SystemVerilog test suite into continuous integration regressions, require verification leads to sign off on this checklist:

[ ] Check 1: Mutation Kill Ratio
    Did the testbench kill at least 85% of synthetic RTL mutants?

[ ] Check 2: Assertion Non-Vacuity
    Were all SVA properties proven reachable and non-vacuous in formal?

[ ] Check 3: Scoreboard Data Integrity
    Does the scoreboard assert a fatal UVM error if packet data is
    intentionally corrupted by 1 bit in the driver?

[ ] Check 4: Cross-Coverage Density
    Does the covergroup define explicit cross-bins between control
    states and interface backpressure signals?

[ ] Check 5: Constraint Histogram Validation
    Does solver profiling confirm uniform or targeted distribution across
    all protocol corner boundaries (min, max, unaligned)?

[ ] Check 6: Reset and Interrupt Robustness
    Does the sequence inject asynchronous resets and abort sequences
    during active transaction phases?

What this means for Silicode

At Silicode (silicode.ai), our stance on RTL and verification generation is rooted in engineering skepticism. Plausible SystemVerilog is not verified silicon. True design velocity cannot rely on unconstrained code generators that produce cosmetically clean testbenches with hollow assertions.

Our toolchain pairs deep LLM synthesis with deterministic formal verification, automated mutation testing, and rigorous constraint analysis. By embedding continuous compiler, linter, and formal property checks directly into the generation loop, Silicode ensures that generated RTL and verification suites deliver mathematically sound, fault-tolerant coverage rather than superficial dashboard metrics.

Verification Remains an Adversarial Discipline

Large language models are exceptional at eliminating the boilerplate tax of UVM. They can assemble drivers, monitors, agents, and configuration databases in seconds, saving verification teams hundreds of hours of manual typing.

However, verification is fundamentally an adversarial discipline. Its goal is not to demonstrate that a circuit works under normal conditions, but to prove that it cannot be broken under any legal operating condition. Autoregressive models are optimized for normality, not adversity.

Treat machine-generated testbenches as rough drafts. Run mutation testing, profile your constraint distributions, check your assertions in formal tools, and never sign off on silicon until you have proved that your testbench can catch a broken chip.

Sources

More Silicode Insight

VerificationUVMSystemVerilogLLM EDA