\
October 8, 2026

Moores Lab AI Achieves Industry Leading Performance on the CVDP (v1.1.0) Benchmark

Fill out the form, and then you will be able to download the whitepaper.
Fill out the form
CVDP Results Post

MooresLab AI · Benchmarks

MooresLab AI Agent™ harness sets an industry-leading score on NVIDIA's CVDP

Earlier this week we explained what CVDP is. Here is how the MooresLab AI Agent™ harness performs on its verification problems, compared with Codex and Claude.

The agentic track of CVDP v1.1 contains 62 verification problems in three categories: testbench stimulus generation, testbench checker generation and assertion generation. The agent receives the prompt and the design, works inside a container with a commercial simulator, and is graded against hidden tests. Each problem is attempted once and scored pass or fail.

We evaluated three configurations:

  • Codex: OpenAI GPT-5.6
  • Claude: Anthropic Claude Opus 5.5
  • MooresLab AI Agent harness

Pass rate on CVDP v1.1 verification problems, one attempt per problem.*

Results by category

Codex GPT-5.6Claude Opus 5.5MooresLab AI Agent harness

Stimulus: testbench stimulus generation. Checkers: testbench checker generation. Assertions: SystemVerilog assertion generation. Same 56 problems as the chart above.*

Insights

Stimulus generation: 100%. The agent drives the design and confirms that its stimulus reaches every point the test plan requires.

93% overall, ahead of both models. The harness is what turns a capable model into verification work that passes.

Assertions remain the hardest category. CVDP measures assertion coverage under its own hidden stimulus, and a correct assertion that this stimulus never reaches is counted as a miss. This category is the focus of our next release.

The harness behind these results is the same one we ship to customers. Next week we open our public benchmark, where you can run your own verification environment against designs with known, hidden bugs and see which ones it catches.

* Six of the 62 problems were excluded. In each of them the hidden grader contradicts the problem's own prompt or specification, so a testbench that checks the specified behavior cannot pass. All three configurations are scored on the remaining 56 problems, and passes on the excluded six were removed from every configuration.

Sharath Giri
Founding AI Lead

Download the full white paper

This field is essential
This field is essential
This field is essential
This field is essential
Thank you! Your submission has been received!
Get pdf file
Oops! Something went wrong while submitting the form.
Yellow cube huge
Smarter Design. Lower Cost. Faster Time‑to‑Market.

Contact Us

Need help or have questions about our AI solutions?
We are always available — tell us about your challenges
This field is essential
This field is essential
This field is essential
This field is essential
Thank you!
We will contact you shortly
Okay
Oops! Something went wrong while submitting the form.