ClaimMark-P
Wanda, a patent attorney with a magician's wand, strides along a stone causeway through towering gates marked §101, §102, §103, §112 and PTAB, while Milo follows with a torch and a stack of patent drafts, ducking a swinging blade.

Our benchmark

The ClaimMark Gauntlet

Run the claims through the attacks they actually faced.

The Gauntlet is the set of benchmarks we hold ourselves to, and invite every patent AI tool to run. It takes real applications and patents from the public record, lets a tool refine the claims, and then replays the examiner’s and the Board’s actual attacks against the result.

Our commitment

We publish every result, whatever it shows, including rows where a plain AI model does as well as ClaimMark-P.

Six gates

Every claim set passes through each gate in turn. The later the gate, the closer the attack is to what really happened to the patent.

  1. Gate I

    Prior-art search

    The Gate of Search

    XRef

    What runs
    Search from the imported matter, capped at its priority date.
    Answer key
    The examiner’s cited references, grouped by family; the X, Y, and A categories assigned in European and international search reports; the examiner’s mapping of references to claim limitations.
    Measured
    Examiner references found in the top 5, 20, and 100; graded relevance; agreement with the examiner on which limitations a reference discloses; references dated after the priority date, counted as leaks.
  2. Gate II

    Statute by statute

    The Gate of Statutes

    Detect

    What runs
    Each check the tool runs: antecedent basis, claim dependency, §112, §101, §102, §103.
    Answer key
    The examiner’s or the Board’s rejections under each statute.
    Measured
    Share of actual rejections the tool flagged; flags per hundred claims; the share of a sample of flags that attorneys judge real.
  3. Gate III

    Claim form

    The Gate of Form

    Form

    What runs
    The claims before and after refinement.
    Answer key
    None. PatentScore’s published prompts, run unchanged.
    Measured
    A form-conformance score, labeled as a model-judged rubric that says nothing about outcome.
  4. Gate IV

    The headline

    The Examiner’s Gate

    Replay: Examiner

    What runs
    A published application, imported at the point where drafting ends and hardening begins. The tool runs every review and recommendation it has; the recommendations are applied automatically.
    Answer key
    The examiner’s actual first rejection, by claim and statute, replayed with the examiner’s own references against the revised claims.
    Measured
    Rejections the revised claims would no longer face, counted only when the revised claim is at least as broad as the claim the applicant eventually got allowed.
  5. Gate V

    The hardest test

    The Board’s Gate

    Replay: Board

    What runs
    A patent later challenged at the Patent Trial and Appeal Board, imported as issued.
    Answer key
    The Board’s final written decision: the references, combinations, and reasoning it relied on, claim by claim.
    Measured
    Claims that would survive the Board’s own attack, counted only when the revised claim still covers the product accused in the public litigation record. Also: does the tool separate the claims later invalidated from the claims later upheld?
  6. Gate VI

    Every gate, twice

    The Twin Gate

    Stability and cost

    What runs
    Every case, run twice.
    Answer key
    The tool’s own second run.
    Measured
    Agreement between runs; cost and time per case.

The headline is the end of a chain

A survived attack counts only if the tool saw the weakness coming. Each step is a smaller number than the one before, and the last one is the headline.

  1. 1FlaggedThe tool flagged the claim as weak.
  2. 2AimedIt proposed a change directed at the ground the examiner or Board later used.
  3. 3SurvivedThe revised claim survives the replayed attack.
  4. 4Kept its scopeThe revised claim still covers what it needed to cover.

Two traps, and how the Gauntlet closes them

The first trap

Narrowing always wins.

Add enough limitations and any claim dodges any reference. So a survived attack counts only when scope is preserved. For patents that went to litigation, the yardstick is the product accused in the public claim chart: the revised claim must still read on it. For applications, the yardstick is the claim the applicant eventually got allowed: the revised claim must be at least as broad.

The second trap

A search miss looks like a reasoning failure.

If the tool’s search never finds the reference the examiner used, its refinement cannot be aimed at it. So every Replay result is reported twice: once with the tool’s own search, which is what a customer gets, and once with the examiner’s or the Board’s references supplied. The gap between the two is the cost of search. The second figure is the refinement’s ceiling.

A survived replay shows that one specific attack would have failed. It does not show the patent would have been upheld, since a challenger facing different claims would have searched for different art. The results pages will say so.

The rules of the Gauntlet

Written before any result exists. Rules chosen after the numbers are known are not rules.

  1. IPublic material only, randomly drawn. No client, partner, or ClaimMark matters anywhere in the study, including for developing the method.
  2. IIThe headline run is fully automatic. Every recommendation is applied as proposed, with no human choice between input and output. An attorney-in-the-loop run may be published separately, never as the headline.
  3. IIIOutputs are the tool’s literal artifacts, hashed and time-stamped before scoring.
  4. IVEvery scoring prompt, sampling rule, seed, and ruling is published.
  5. VPrompts may be tuned on the pilot draw only. The headline draw is new and scored once.
  6. VILabels say exactly what was measured: “rejections the revised claims would no longer face”, never “office actions prevented”.
  7. VIIAn allowance is not a correct answer. Only positive action counts: a rejection, reasons for allowance, a holding.
  8. VIIIAnyone can dispute a ruling, and every correction is published with its date.

Registered samples

The seed, the rule, and the hash fix each sample before any tool runs on it. Anyone can re-derive the list and check the hash.

SampleHow it was drawnSeed and hash
Examiner pilotUniform random application numbers filed 2022 to 2024, kept if published with a first action on the merits and claims on file before it, until each of the eight technology centers had its quota. 154 draws, 50 kept. 38 first actions are rejections and 12 are allowances; 9 were mailed after 2026-07-01 and form a hold-out.seed claimmark-backtest-group1-pilot-2026-09-25
SHA-256 617f4a3ab1a6c568f34b209e8b4cb8209d5c613379a089d11611c1add04ae93b
Board pilotOriginal final written decisions in inter partes and post-grant review, one per patent, shuffled from a pool of 1,295 patents decided 2023-01-01 to 2026-06-30. Kept in shuffled order if the original complaint in a district-court case carries a public claim chart for that patent, until 50 are kept. A hold-out pool of decisions from 2026-07-01 is registered separately.seed claimmark-backtest-group2-charted-2026-09-26
order SHA-256 b178b6e2fba415ae35f6c43d17c73116751247056f01fdb9a3aa053d75a4f789

Standings and the frontier

Standings

Published here after each scoring round, with every tool that has run the Gauntlet, ours included.

Each event gets its own table. We also plot the frontier: each tool’s Replay result against its cost per case, per capability area, so a reader can see which tools no other tool beats on both. Any overall ranking shows its weights on the page, and readers can change them, because any fixed weighting is an opinion.

Take the challenge

Run the Gauntlet

Any patent AI tool can enter, two ways.

Run it yourself

Use the published harness on the registered samples and submit the full output bundle with its hash, plus a trial account so we can rerun a sample of cases and confirm your results.

We run it

Where a product’s public trial terms allow, we run it ourselves, publish the bundle, and send the vendor the results before publication. The vendor’s reply is published beside its row.

There is no paid placement. Results are withdrawn only for a factual error, and the correction is dated.

Governance

The method, harness, sampling code, and scoring prompts are published on GitHub, and the samples on HuggingFace. Every change to the method is dated and explained. We are inviting an outside patent attorney and an academic research group to review the method and the rulings alongside ClaimMark’s co-founder, patent attorney Jürgen Vollrath, and we will name them here when they accept.

Method sign-off

I have reviewed the Gauntlet’s method as described on this page. Its answer key is what examiners and courts actually decided, not a model’s opinion of a draft. Its cases are drawn at random from public records and registered before any tool runs on them. Every tool, ours included, is scored the same way.

I sign off on the method. I will review the scoring rulings as results come in, and my name will not appear beside any result I have not reviewed.

Jürgen Vollrath, co-founder of ClaimMark Inc. · Registered U.S. Patent Attorney, USPTO Reg. No. 49,098

Why we built it

Why did we make the Gauntlet? Existing benchmarks were not enough!

Before we built the Gauntlet, we went looking for a benchmark that could tell an attorney whether an AI tool’s claims would hold up. We could not find one.

Every benchmark we found measured something narrower than the question attorneys care about:

  • Patsnap’s PatentBench measures only whether a search finds the references the examiner cited, on a test set nobody outside Patsnap can see.
  • ABIGAIL’s PatentBench is open source and aims at drafting and Office Actions, but so far it scores only docketing arithmetic, and its dataset could not be downloaded.
  • The academic benchmarks score similarity to the granted text, or a model’s opinion of the claims.

None of them asks the question an attorney is paid to answer: would these claims survive the examiner, the Board, or a court? Finding the right art is only the first step, and a claim that reads well can still be anticipated. A tool that cannot show what happens when its claims meet a real attack has not shown that it can be trusted.

So we built a benchmark that asks that question directly. The ClaimMark Gauntlet takes real applications and patents from the public record, lets a tool refine the claims, and replays the examiner’s and the Board’s actual attacks against the result. A win counts only when the claims survive and keep their scope. And because a benchmark is only as good as its honesty, we hold the Gauntlet to every rule we ask of others: random samples, outputs locked before scoring, and every result published.

Why benchmarks matter

Patent AI benchmarks are emerging — but what do they actually measure?

AI tools have moved into patent practice fast. Clarivate reports that AI use among IP professionals rose from 57% in 2023 to 85% in 2025. The question attorneys now ask is not whether to use these tools, but whether they can trust what the tools produce. When HGF polled patent attorneys in late 2025, accuracy and hallucinations topped the list of concerns, named by 42%, more than twice the share most worried about confidentiality.

The concern is well founded. A peer-reviewed Stanford and Yale study found that leading AI legal research tools, built specifically for lawyers, still gave incorrect or misgrounded answers between 17% and 33% of the time. And the responsibility does not move. The USPTO’s guidance reminds practitioners that their signature certifies their own reasonable inquiry, whether or not an AI tool helped, and ABA Formal Opinion 512 expects lawyers to review AI output with appropriate care.

So how can an attorney tell which tools deserve that trust? Checking every output by hand erases the time the tool saved. That is the job of benchmarks: independent, repeatable tests that show how often a tool gets it right, on cases like the attorney’s own. General-purpose AI has been measured this way for years. Benchmarks for patent AI are only now emerging, and we list every one we have found below.

This section explains which benchmarks currently exist, what those numbers measure and what they don’t.

Side by side

How the benchmarks compare

The same eight questions, asked of each.

ClaimMark GauntletPatsnap PatentBenchABIGAIL PatentBenchAcademic benchmarks (typical)
What is scoredWhether the refined claims would have survived the examiner’s or the Board’s actual attack, with their scope intact.Prior-art search only: does the search return the references the examiner cited?Docketing arithmetic (deadlines, fees, classification). Drafting and Office Action quality are planned, scored by an AI judge.Similarity to the granted text, or an AI judge’s rubric score.
Answer keyPositive action in the public record: examiner rejections, Board holdings, court rulings.Examiner citations, grouped by patent family.Fixed answers for arithmetic; a rubric opinion for quality.The granted text, or a model’s opinion.
How cases are chosenA uniform random draw. The seed and the list’s hash are published before anything runs.340 curated patent families. The set is not published.Drawn from 604 real Office Actions. The dataset was not publicly downloadable on 2026-09-26.Varies. Often older patents, or disclosures written by a model from the granted patent.
Outputs locked before scoringYes. Every output is hashed and the hash is time-stamped before scoring.Not statedNot statedRarely
Can an outsider rerun it?Yes, case by case, on a trial accountNoNot yet: the code is open, the data is notOften, when code and data are released
Training-data contaminationCases decided after the model’s training cutoff are reported separately; search is capped at the priority dateNot addressedNot addressedRarely addressed
Stability and costEvery case runs twice; cost and time per case are publishedNot reportedNot reportedRarely reported
Status on 2026-09-26Published, with samples registered before any runPublishedDocketing layer published; the rest in progressPublished papers

Patsnap and ABIGAIL columns reflect each vendor’s own benchmark pages and repository on 2026-09-26. “Academic” describes the eight papers reviewed under Academic work.

A reader’s checklist

How should you read a patent AI benchmark?

Ten questions that separate a measurement from a marketing number. We ask them of every benchmark we know of, and of our own.

None of these questions requires technical knowledge of AI. They are the questions an attorney would ask of any expert report: where did the data come from, who chose it, what was it compared against, and could someone else check it.

Our own column answers the same questions, and every answer in it can be checked against what we publish.

QuestionClaimMark GauntletPatsnap PatentBenchABIGAIL PatentBenchPatentScoreDis2PatPatRe
1. Can you get the test set?
A score on data nobody else can see is a claim, not a measurement.
Samples, every output, and their hashes are published
Not published
Code is open; the dataset returned “unauthorized” on 2026-09-26
Prompts published; built from a public patent dataset
Code and data released
Code and data released, per the authors
2. Was the sample drawn at random, by a rule published first?
A hand-picked sample can make any tool look good, even when nobody intends it to.
Uniform random draw; seed and list hash published before any run
Curated for language and technology class
?Selection rule not stated
400 patents from two technology sections, 2016 to 2017
?Split from an existing corpus; rule not verified by us
?Not verified by us
3. Is the answer key an outcome, or an opinion?
An examiner’s rejection happened. A model’s rubric score is somebody’s view.
Examiner rejections, Board holdings, court rulings
Examiner citations
Fixed answers for arithmetic; an AI judge for quality
Expert and model scores of claim form
The granted text, and an AI judge
Real Office Actions as reference; scoring method not verified by us
4. Does the score track what happens to the application?
Attorneys care whether claims issue and survive, not whether text looks like other text.
The headline is survival of the replayed attack with scope preserved
Finding the art is the first step, not the result
Not for the layer that is scored today
The authors say it does not address novelty
Similarity to one granted version
Models the examination exchange
5. Are the comparisons fair?
Beating a chat window at patent search says little about beating a search tool.
Every case also runs through a plain AI model as a baseline; other vendors are invited to run the same cases
Compared with general chatbots, not with search tools or a keyword baseline
General models only; rival vendors listed as “not submitted”
Validated against three human experts
Open and closed models, plus attorney pairwise review
Proprietary and open models
6. Were outputs locked before scoring?
Without a lock, a score can be tuned after the fact, even unintentionally.
Every output is hashed and time-stamped before scoring
?Not stated
?Not stated
?Not stated
?Not stated
?Not stated
7. Can you rerun a case yourself?
The strongest proof is a reader getting the same answer.
Any case can be rerun on a trial account
No
Not until the data is available
Prompts are public
Yes
Yes, per the authors
8. Does it report stability, uncertainty, and cost?
An AI tool can give different answers to the same input. A score without its spread hides that.
Every case runs twice; agreement between runs, cost, and time per case are reported
No repeat runs, intervals, or cost
Not reported
Averages ten runs; variance not reported
?Not verified by us
?Not verified by us
9. Is training-data contamination controlled?
A model may already have seen the patent, and the examiner’s answer, during training.
Post-cutoff hold-outs reported separately; priority-date search ceiling
Not addressed
Not addressed
Patents from 2016 and 2017
Disclosures written from the granted patent being predicted
?Not verified by us
10. Does the number stay put?
A benchmark that changes should say when, and what changed.
Every change dated; the headline sample is scored once
Moved from 81% to 85%, with different comparators, and no version note
Case counts differ between the repository and the website
Versioned paper
Versioned paper
Versioned paper
Meets the question Partly, or planned but not yet done Does not? Not stated, or not verified by us

What a good answer looks like

The single most useful test is the fifth one. A number means little until you know what it beat. A specialized search tool beating a general chatbot is expected. A drafting tool scoring well on arithmetic that general models already get right is expected. The informative comparison is against the best available alternative, run on the same cases, by rules fixed before anyone saw the results.

The second most useful is the third. An outcome that happened, such as an examiner’s rejection or a Board decision, cannot be argued with. A model’s opinion of a claim can be useful as a diagnostic, but as a score it measures agreement with the model, not quality.

Benchmark report

Patsnap’s PatentBench: a real outcome, one narrow task, a test set nobody else can see.

Patsnap’s benchmark scores prior-art search against the references examiners actually cited. That is a sound answer key. What it measures is narrower than its marketing, and nobody outside Patsnap can check it.

What it is

PublisherPatsnap, which sells the search tool it ranks first
First publishedSeptember 2025; the current figures were on the page on 2026-09-26
TaskNovelty search: given an invention, return prior art
Test set340 patent families; receiving offices about 32% US, 32% China, 18% Europe, 18% WIPO; 68% English, 32% Chinese
Answer keyThe X and Y references examiners cited, deduplicated and grouped by patent family
MeasuresX hit rate: share of cases where at least one examiner X reference is returned. X recall: share of all examiner X references found in the top 100
Compared againstGeneral-purpose AI chat products with web search
PublishedHeadline numbers and a one-case worked example. Not the test set, the queries, the prompts, or per-case results

Patsnap also publishes a design-patent freedom-to-operate benchmark (261 samples) and lists a utility-patent freedom-to-operate benchmark as in progress.

The numbers, as published

SystemX hit rateX recall, top 100
Patsnap Novelty Search agent85%37%
ChatGPT 5.671.76%30.51%
Claude Opus 4.852.37%11.68%
Perplexity Pro39.16%6.4%
Kimi-k333.53%12.09%
DeepSeek-v4-flash22.09%6.25%
Gemini 3.1 Pro14.24%2.39%

From patsnap.com/benchmark, 2026-09-26.

What 85% and 37% actually mean

85% is not the share of prior art found. It is the share of test cases in which the search returned at least one reference an examiner cited as an X reference. A search that finds one of five relevant references scores the same as one that finds all five.

37% is the share of prior art found. Of all the X references examiners cited across the test set, Patsnap’s top 100 results contained 37%. The rest, close to two in three, were not in the top 100. This is the more informative of the two figures, and it is the one Patsnap’s own tool scores on, in its own benchmark.

That is not a criticism of Patsnap’s engineering. Examiners search with classification codes, full-text databases, and years of art-unit experience, and matching them is hard. It is a reason to read the headline carefully.

Where the published figures disagree

  • Which results count. The benchmark page defines the hit rate as a correct answer “among the top 1, 3, or 5 results”. The same page, and the benchmark hub, label the 85% headline “Top 100”. Those are very different tests, and the page does not say which one produced 85%.
  • What the homepage says. Patsnap’s Eureka homepage presents the figure as “85% Prior art found in the top 100 results”. By the benchmark’s own definitions, the share of prior art found is the recall figure, 37%.
  • Which competitor is quoted. The homepage compares Patsnap with ChatGPT 5.4 at 17.18%. The benchmark page reports ChatGPT 5.6 at 71.76%, within about 13 points of Patsnap.
  • How the figures moved. In November 2025, Patsnap reported 81% and 36%, compared against ChatGPT o3 and DeepSeek-R1. The benchmark page now reports 85% and 37% against a different set of products, while Patsnap’s own novelty-search product page still shows 81% on the same day. No version history explains the change.

What it does not tell you

  • Anything about drafting. It is a search benchmark. It says nothing about the quality of a drafted application or an Office Action response.
  • How relevant each result is. A result is a hit or a miss. There is no graded relevance, and no check of which claim limitations a reference actually discloses.
  • How it compares with other search tools. The comparators are general chat products with web search. No patent search tool, and no simple keyword search over the same collection, was run as a baseline.
  • Whether the answer was in the question. The page does not say what text each query was built from. If it came from a granted patent, the examiner’s citations are printed on the patent’s face, and may be in a model’s training data.
  • How much the number moves. There are no repeat runs and no confidence intervals. With 340 cases, sampling alone puts an 85% figure within about four points either way, before any run-to-run variation.

What it gets right

The answer key is an outcome, not an opinion: references that examiners actually cited, grouped by family so that a U.S. patent and its European counterpart count once. The test set spans four offices and two languages. And Patsnap published something, with a worked example, which most of the industry has not.

Why we do not report a score on Patsnap’s PatentBench

We cannot. The test set, the queries, and the per-case results are not published, so nobody outside Patsnap can run the benchmark, including Patsnap’s own customers. If Patsnap publishes the set, we will run ClaimMark-P on it and publish every output.

Instead we do two things. First, we run the same protocol on our own sample: applications drawn at random from the public record, scored against the references the examiner cited, with the search capped at each application’s priority date. That figure appears beside Patsnap’s with a plain caveat: a different sample, a different index, and a Patsnap figure nobody can verify. It is not a head-to-head, and we do not call it one.

Second, we measure what the hit rate leaves out. XRef, the search event of the ClaimMark Gauntlet, grades relevance using the X, Y, and A categories that European and international search reports assign to every citation, and checks whether a tool’s account of which claim limitations a reference discloses matches the examiner’s own mapping in the rejection.

Patsnap may still find more of the examiner’s references than we do. Its index is larger, and for most offices our worldwide search reads abstracts and bibliographic data rather than full text. That is why search recall is one row of our results, not the headline.

We also need to compare prior art search tools

There is no shortage of tools that search for prior art. Established databases such as Derwent Innovation, Questel Orbit, PatBase, and Patsnap, free public ones such as Google Patents, Espacenet, and the USPTO’s Patent Public Search, and a new wave of AI search tools all promise to find what matters. Almost all of them stop at the same place: a ranked list of references. Working out what each reference actually does to your claims is left to the attorney.

That last step is where ClaimMark-P is different, and it is why we compare search tools on more than whether they find the right documents. Below, ClaimMark-P and Patsnap, the one vendor that publishes a search benchmark, feature by feature. Patsnap’s column is taken from its public product pages on 2026-09-26.

CapabilityClaimMark-PPatsnap Eureka Novelty Search
Size and reach of the index
Over 100 million patent documents from more than 100 patent offices through the European Patent Office, plus the USPTO, a semantic patent and literature index, and web and product sources
Over 200 million patents, worldwide, per the vendor
Non-U.S. patent offices
More than 100 offices, including Europe, WIPO, China, and Japan: abstracts and bibliographic data for all of them, full text for European and international filings
Includes China, Europe, and WIPO
Starts from the invention disclosure
Built from the invention record’s inventive center, problem, and summary
From a disclosure
Searches again from the drafted claims
Recommends a revised search when the claims change, grounded in their exact language
?Not stated
No invented references
Every reference is checked against its record; one that cannot be verified is dropped, never shown
Source-linked results, per the vendor
Explains why each reference matters
Overlap type and a written relevance note per reference
Feature-level evidence, per the vendor
Which invention elements each reference discloses, and which it misses
Listed per reference, both ways
Feature-level evidence; the missing side is not described
Which claim limitations each reference discloses, and which it misses
Listed per reference once claims exist, and mapped into the claims with the cited passage
?Not stated for novelty search
Relevance score you can audit
Computed by fixed rules from the listed elements, so the same judgment always yields the same number; the rubric version is stored with it
?Results are ranked; no score definition published
Feeds the examiner-style review of your claims
References flow into §102 and §103 review and the adversarial examiner simulation in the same matter
Search is callable inside drafting and Office Action workflows, per the vendor
Search capped at the priority date
Every search can be limited to art published before the priority date
?Not stated
Published retrieval benchmark
XRef, in the ClaimMark Gauntlet, on public samples anyone can check
Yes, but on a test set nobody else can see

Beyond the result list

Not just a search — why ClaimMark-P is revolutionary

Other tools hand you a list of references and wish you luck. ClaimMark-P puts every reference to work on your claims, the moment it is found.

  • Every reference, taken apart. Which elements of your invention it discloses, which it misses, and which of your claim limitations it reads on, listed both ways, for every reference.
  • Scores you can audit. Relevance is computed by fixed rules from those lists, so the same judgment always yields the same number, and the rubric version is stored with it.
  • Mapped straight into your claims. Each limitation links to the passage of the art that reads on it, so you see the collision at the exact words where it happens.
  • Turned into an attack before the examiner makes one. The references feed the §102 and §103 review and an adversarial examiner simulation that tries to reject your claims first, and every weakness it finds comes with a proposed fix.
  • Never out of date. Change the claims and ClaimMark-P tells you the search is behind, and runs it again from the new claim language when you ask.

Search is where most tools end. In ClaimMark-P it is where hardening your claims begins.

Benchmark report

ABIGAIL’s PatentBench: the right ambition, most of it still to come.

ABIGAIL has published the most ambitious patent AI benchmark so far, as open source, and it is the only one that sets out to test drafting and Office Action work. Today the only scores are for docketing arithmetic, and the dataset could not be downloaded.

Who publishes it

ABIGAIL is a self-funded patent AI company. Its product drafts applications, prepares Office Action responses, runs prior-art searches, and handles prosecution docketing. PatentBench is published under the Apache 2.0 license on GitHub, installable as a Python package, with a white paper signed by ABIGAIL’s founder, a registered patent attorney.

ABIGAIL uses PatentBench on its website to point out that far better-funded rivals publish no benchmarks at all. That criticism is fair as far as it goes. The same standard should apply to PatentBench itself, and that is what this page does.

What it is designed to measure

7,200 test cases across five domains, derived from 604 real USPTO Office Actions from 2019 to 2024 across nine technology centers, per ABIGAIL’s site.

DomainCasesWhat it testsStated human baseline
Administration1,500Deadlines, fees, IDS completeness99.8%
Drafting500Claim scope, specification support, terminology8.5 out of 10
Prosecution2,500Rejection analysis, arguments, amendments8.6 out of 10
Analytics1,500Examiner prediction, allowance probability75%
Prior art1,200Reference relevance, anticipation85%
Scoring layerHow it scoresWeightStatus on 2026-09-26
1. DeterministicFixed answers: deadline math, event codes, fee lookups30%Published: 298 tests
2. AI judgeA separate model scores outputs against rubrics35%In progress
3. ComparativeHead-to-head ranking between systems25%Planned
4. Human calibrationRegistered attorneys check the automated scores10%Recruiting

What is actually scored today

Layer 1, 298 testsScore
ABIGAIL v3100.0%
Claude Sonnet 499.1%
Gemini 2.5 Flash99.1%
Gemini 2.5 Pro88.7%

The published leaderboard covers the deterministic layer only: classifying Office Actions, reading timelines, computing fees and deadlines. ABIGAIL’s own product scores 100%. Two general-purpose models score 99.1%.

A test that general models already pass does not tell tools apart. It confirms the arithmetic is right, which matters for docketing, but it says nothing yet about drafting or arguing an Office Action. Those layers carry most of the weight in ABIGAIL’s own formula, and they are not yet scored.

What we could not reconcile or reproduce

GitHub READMEabigail.app/patentbench
Drafting casesA drafting-focused set of 1,200A drafting domain of 500
Office Action casesAn Office Action set of 1,800A prosecution domain of 2,500
Where the data livesA HuggingFace dataset, linked from the READMEDescribed as public data
Can the public download it?The dataset returned HTTP 401 (unauthorized) to an anonymous request on 2026-09-26

The two pages may describe different cuts of the same data, a subset in one place and a domain in the other. Neither explains how they relate, and a reader cannot check, because the data could not be downloaded. Until it can, “reproducible” describes the intention rather than the current state.

The stated human baselines raise a further question. A drafting baseline of 8.5 out of 10 implies human work was scored on the same rubric, but we did not find who scored it, how many attorneys, or on which cases.

What it gets right

  • It is open source, with a license that lets anyone run and extend it.
  • It is built from real Office Actions, not synthetic cases.
  • It says plainly which layers are finished and which are not.
  • It plans to calibrate the automated scores against registered attorneys, which every AI-judged benchmark needs.

Where an AI judge falls short

When the drafting and prosecution layers are scored, they will be scored mainly by a model reading the output against a rubric. That is useful for catching form problems. It cannot tell you whether an argument would have persuaded the examiner, because the examiner is not in the loop. The strongest answer key for a response to a rejection is what the examiner did next, and for a drafted claim, whether it survived examination. The academic work page covers why rubric scores should be a diagnostic rather than a grade.

Docketing: what ABIGAIL tests, and what we do

The scored layer of PatentBench is docketing. ABIGAIL’s product syncs deadlines from the USPTO with weekend and holiday rollover, prefills USPTO forms, and offers examiner analytics. ClaimMark-P does not do docketing. It shows the period for reply exactly as the examiner wrote it and tells the attorney to use the firm’s docket. We do not calculate due dates, because a wrong date is a malpractice problem and firms already run dedicated docketing systems with a second-check rule. ClaimMark-P does generate IDS forms from the references in a matter.

So on PatentBench we would submit only the drafting and prosecution domains, and say so.

What we will do

When the dataset can be downloaded, we will run ClaimMark-P on the drafting and prosecution cases and publish every output, whatever the score. Separately, the ClaimMark Gauntlet scores against what examiners and the Board actually did, which no AI judge can stand in for. We would welcome ABIGAIL’s participation in both.

Research review

The academic work: careful, useful, and not yet asking the question that matters.

Researchers have published a wave of patent drafting and prosecution benchmarks since 2025. Each is careful within its frame. None yet asks whether an AI tool’s claims would have survived the examiner.

Wanda reads an academic paper in a grand university library with a raised eyebrow, while Milo balances on a library ladder behind her with a teetering stack of journals.

What has been published

WorkWhere, whenWhat it testsAnswer keyOpen
Patent-CRNAACL 2025Revise rejected claims toward the version that was grantedThe granted claims; professional human evaluationDataset paper
EPDEMNLP 2025 FindingsGenerate claims, trained and tested on granted European patentsThe granted claimsNot verified by us
PatentWriterarXiv, July 2025Write a patent abstract from the first claimThe published abstract, by text-overlap metricsCode released
PatentScorearXiv, 2025Score the form of a generated claim on seven rubric dimensionsThree experts’ scores of 400 claimsPrompts published
Tree-of-ClaimsarXiv, November 2025Claim search and revision against prior artExaminer actions from the USPTO Office Action Research Dataset; 1,145 wireless patentsNot verified by us
PatRearXiv, May 2026Write the examiner’s Office Action and the applicant’s rebuttal480 real examination casesCode and data released
Dis2PatEMNLP 2026Draft a full application from an inventor-style disclosureThe granted patent, text similarity, an AI judge, and 60 attorney comparisonsCode and data released
Vibe PatentingarXiv, September 2026Test whether AI judges can grade and improve drafting agentsOne patent attorney’s independent reviewNot verified by us

What they get right

Patent-CR and PatRe move in the right direction. Both are built on real examination history, so the model has to deal with an examiner’s objections rather than an idealized task. PatRe separately tests a model given the examiner’s references and a model that must find them, which is the right way to tell a search failure from a reasoning failure. Dis2Pat is the first to take the drafting problem from the inventor’s side, and it asked patent attorneys to compare outputs directly. Vibe Patenting checked its AI judges against a patent attorney and reported, honestly, where they disagreed.

Four problems they share

  1. The granted text is treated as the one right answer. Many claim sets would have been allowable. The granted claims are the product of a negotiation with the examiner that the drafter could not have seen. Scoring similarity to them rewards predicting the end of prosecution, not drafting well at the start of it.
  2. The answer leaks into the question. When an inventor’s disclosure is written by a model from the granted patent, as in Dis2Pat, it carries that patent’s structure, emphasis and vocabulary. The same risk arises for search when the query is built from a granted patent whose citations are printed on its face.
  3. The test cases are in the training data. Patents from 2016 and 2017, as in PatentScore, and granted European texts, as in EPD, are in every large model’s training corpus. A model may be recalling rather than reasoning, and none of these papers reports a set held out past the model’s training cutoff.
  4. An opinion stands in for an outcome. Where there is no granted text to compare against, a model grades the output. Vibe Patenting found that agreement between AI judges and a patent attorney depended strongly on which quality was being measured.

Underneath all four is one missing question. None of these benchmarks asks whether the claims would have held up: before the examiner, before the Board, or in court. That is the question an attorney is paid to answer.

PatentScore, up close

PatentScore is the most complete published rubric for grading patent claims, and its authors published it in full. It deserves a close reading, because rubric scores like it are likely to become the default way vendors grade drafting.

Is the rubric published?

Yes. The prompts and the scoring rubric, with example claims at each score level, are in the paper’s appendix. Seven dimensions: claim structure, punctuation, antecedent basis, element referencing, validity and uniqueness, ambiguous scope, and semantic similarity.

Is the scoring repeatable?

Not by design. A small model scores each dimension ten times and the average is taken. The paper does not report the sampling temperature or how much the ten scores vary.

How well does it agree with experts?

A correlation of 0.819 with three experts on 400 first claims. The dimension weights were fitted on the same 400 claims used to report that figure, so it is an in-sample result. The semantic dimension ends up with a weight of 0.001.

Is it a fair absolute grade?

No, and the authors do not claim it is. It measures whether a claim is well formed. By their own account it does not address novelty. A claim can score well and still be anticipated.

How we use it. ClaimMark-P already runs its own rubric reviews, per statute, with every finding tied to a claim and a recommendation, and its antecedent-basis check is rule-based rather than a model’s opinion. We treat those as diagnostics that drive recommendations, not as a score. In the Gauntlet we run PatentScore’s published prompts, unchanged, on each claim set before and after refinement, and report the result as a form-conformance row, labeled as a model-judged rubric that says nothing about outcome.

A better method

Nine principles. Each one answers a problem above, and each one is a rule of the ClaimMark Gauntlet.

  1. 1
    Score against what happened.

    The answer key is positive action in the public record: an examiner’s rejection, the examiner’s reasons for allowance, a Board holding, a court ruling. The absence of a rejection proves nothing, so an allowed claim is never treated as a correct negative.

  2. 2
    Replay the attack, not the text.

    Apply the tool’s recommendations, then replay the examiner’s actual rejection, with the examiner’s own references, against the revised claims. Similarity to the granted claims is not the question.

  3. 3
    Make narrowing count against you.

    Any claim survives if you narrow it enough. A revised claim wins only if it survives the replayed attack and keeps its scope: it still reads on the product accused in litigation, or it is at least as broad as the claims the applicant eventually got allowed.

  4. 4
    Draw at random, by a rule published first.

    Uniform random draws from the public record, with the seed and the list’s hash published before anything runs. Prompts may be tuned on a pilot draw; the headline draw is new and scored once.

  5. 5
    Lock the outputs before scoring.

    Every output is hashed, and the hash time-stamped, before any score is computed.

  6. 6
    Control contamination.

    Cases decided after the model’s training cutoff are reported separately. Every search is capped at the application’s priority date.

  7. 7
    Report the spread.

    Run every case at least twice, and publish agreement between runs, the sample’s uncertainty, and cost and time per case.

  8. 8
    Use AI judges as instruments, not verdicts.

    Where a model makes a scoring judgment, publish its prompt, calibrate it against attorneys on a sample, and label the result as model-judged.

  9. 9
    Let anyone rerun a case.

    Publish every input and output, and let a reader rerun any case and dispute any ruling.

The white paper

We are writing up this method as a paper for peer review. The samples were registered, with their hashes, before any tool ran on them, so nobody, including us, could adjust the rules after seeing the numbers. We would welcome co-authors from the research groups whose work is reviewed here, and an outside patent attorney to review the scoring rulings.

Get involved

Help build an unbiased benchmark.

A benchmark written by one vendor will always be read with suspicion, including ours. So we publish the Gauntlet’s method, code, sampling procedure, and scoring prompts on GitHub and HuggingFace, and we invite patent attorneys, researchers, and other vendors to challenge any part of it.

  • Attorneys: tell us which outcomes matter in your practice, and review our rulings on a sample of cases.
  • Researchers: co-author the method paper, or run the Gauntlet on your own system.
  • Vendors: run the Gauntlet and submit your results with proof of work, or tell us where the method is unfair to your product. Every vendor we score gets a right of reply, published beside its row.

Sources were checked on 2026-09-26 and saved to the Internet Archive on that date. Vendors change their pages, so each figure here is quoted as of that date. Product and benchmark names are trademarks of their owners; ClaimMark is not affiliated with any of them. Corrections are welcome and will be published with the date they were made.