Every chip project has one artifact that everything else inherits from, and it is
the least engineered thing we own.
RTL has version control, lint, and CI. Testbenches have coverage metrics and
regression dashboards. The specification, the document that decides what the RTL
is supposed to do and what the testbench is supposed to check, is a 120-page PDF,
a Confluence page that stopped being true 4 months ago, and an email thread with the
addendum that never made it into either.
The failure mode is not that specs are wrong. Wrong is easy to catch. The failure
mode is that specs are incomplete in ways that read fluently. A paragraph says
"reset values are defined per register" and names none of them. A section
describes an arbiter without listing its states. A clock is referenced in three
places at two frequencies. None of that trips a review, because prose does not
have a compiler. It surfaces eight weeks later, when a DV engineer builds a
scoreboard against a value nobody ever wrote down, or when an escaped bug turns
out to have been an unspecified corner all along.
That gap is expensive precisely because it is upstream. Everything built on top of
an ambiguous spec has to be rebuilt when the ambiguity resolves.
Why "generate the spec with an LLM" makes this worse
The obvious move in 2026 is to point a language model at the datasheet and ask for
a hundred pages. We tried the obvious move. It produces a document that is
confident, well-organized, and quietly invented.
For verification, a fabricated reset value is strictly worse than a blank one. A
blank is a question, and a question gets asked. A fabrication is a defect that
propagates: into the register model, into the scoreboard, into the test that
passes for the wrong reason. One-shot generation optimizes for the thing that
does not matter here (fluency) and destroys the thing that does (traceability).
So SpecAgent is built around a different premise. The job is not to write the
spec faster. It is to make the spec's completeness auditable, and to make its
gaps loud instead of invisible.
What we do differently
**The output is a LiveDoc, not a document.** A LiveDoc is structured, versioned,
and sectioned: an ordered list of sections, each with a body, a status, a source
manifest, and an open-questions list. It renders as a spec for humans and parses
as a contract for machines. Publishing is a deliberate act that freezes a numbered
version, and only published versions are visible to the rest of our agent suite.
Drafts stay private. The boundary is the point.
**The agent drafts one to three sections per turn, and proposes rather than
writes.** Every authoring turn stages its work for review, section by section, and
the architect accepts or discards each one. There is no run that silently rewrites
forty pages. The engineer steers, the agent advances the document one small step,
and the run tree keeps every intermediate state as an audit trail.
**Every technical fact carries provenance.** Register addresses, signal widths,
timing numbers, clock frequencies, parameter defaults: each is cited inline back to
the source file and page, and each citation carries a literal excerpt short enough
to verify by eye and exact enough to verify byte for byte. Turn on "Show
resources" and the spec stops being an assertion and becomes an argument with
footnotes. Facts the engineer supplies in chat are marked as such, so a claim sourced from a person is never mistaken for one sourced from a datasheet.
**Gaps are a first-class state, not an authoring failure.** Sections carry
Completed, Needs clarification, or Missing data. When source material does
not support a section, the agent is required to leave it incomplete and ask a
specific, answerable question, "What is the reset value of the CTRL_STATUS
register at offset 0x10?", not "could you clarify this section?". Supporting
documents the team uploads are consulted before the question is asked, so we
never make an engineer retype what the integration guide already says. Speculative
content is prohibited by contract.
**Completeness gets a number.** Every LiveDoc is rated chapter by chapter against
a section taxonomy that defines what a complete version of each section contains.
The rating is evidence-based and deliberately unkind: mentions do not count as
specifies, and a chapter with no citable artifacts cannot score above the stub
band however well it reads. The result is a per-section completion percentage, a
weighted overall score, a letter grade, and a ranked list of what is missing, by
name. "Is the spec ready?" stops being a matter of taste.
**And the loop closes downstream.** The published LiveDoc is what VerifAgent
consumes to build the testbench and the test plan. That gives spec quality a
measurable consequence: a section that cannot be verified against is a section
that was not specified. We are extending that path so a verification agent working
against a published spec can raise an ambiguity directly on the section that
caused it, which puts the feedback where the fix belongs instead of in a bug
tracker three weeks later.
Why this shape
We build these tools because we use them. Moores Lab delivers verification work as
a service, and every hour our own engineers lose to an under-specified input is an
hour we pay for directly. That tends to focus the design.
The bet is simple. Specification is not a writing problem that happens to involve
hardware. It is an engineering artifact that has never been given the tooling
every other artifact in the flow already takes for granted: version control,
provenance, status, and a completeness metric you can argue with. Give the spec
that, and the rest of the flow gets faster on its own.




