
Our benchmark
The ClaimMark Gauntlet
Run the claims through the attacks they actually faced.
The Gauntlet is the set of benchmarks we hold ourselves to, and invite every patent AI tool to run. It takes real applications and patents from the public record, lets a tool refine the claims, and then replays the examiner’s and the Board’s actual attacks against the result.
Our commitment
We publish every result, whatever it shows, including rows where a plain AI model does as well as ClaimMark-P.
Six gates
Every claim set passes through each gate in turn. The later the gate, the closer the attack is to what really happened to the patent.
- Gate I
Prior-art search
The Gate of Search
XRef
- What runs
- Search from the imported matter, capped at its priority date.
- Answer key
- The examiner’s cited references, grouped by family; the X, Y, and A categories assigned in European and international search reports; the examiner’s mapping of references to claim limitations.
- Measured
- Examiner references found in the top 5, 20, and 100; graded relevance; agreement with the examiner on which limitations a reference discloses; references dated after the priority date, counted as leaks.
- Gate II
Statute by statute
The Gate of Statutes
Detect
- What runs
- Each check the tool runs: antecedent basis, claim dependency, §112, §101, §102, §103.
- Answer key
- The examiner’s or the Board’s rejections under each statute.
- Measured
- Share of actual rejections the tool flagged; flags per hundred claims; the share of a sample of flags that attorneys judge real.
- Gate III
Claim form
The Gate of Form
Form
- What runs
- The claims before and after refinement.
- Answer key
- None. PatentScore’s published prompts, run unchanged.
- Measured
- A form-conformance score, labeled as a model-judged rubric that says nothing about outcome.
- Gate IV
The headline
The Examiner’s Gate
Replay: Examiner
- What runs
- A published application, imported at the point where drafting ends and hardening begins. The tool runs every review and recommendation it has; the recommendations are applied automatically.
- Answer key
- The examiner’s actual first rejection, by claim and statute, replayed with the examiner’s own references against the revised claims.
- Measured
- Rejections the revised claims would no longer face, counted only when the revised claim is at least as broad as the claim the applicant eventually got allowed.
- Gate V
The hardest test
The Board’s Gate
Replay: Board
- What runs
- A patent later challenged at the Patent Trial and Appeal Board, imported as issued.
- Answer key
- The Board’s final written decision: the references, combinations, and reasoning it relied on, claim by claim.
- Measured
- Claims that would survive the Board’s own attack, counted only when the revised claim still covers the product accused in the public litigation record. Also: does the tool separate the claims later invalidated from the claims later upheld?
- Gate VI
Every gate, twice
The Twin Gate
Stability and cost
- What runs
- Every case, run twice.
- Answer key
- The tool’s own second run.
- Measured
- Agreement between runs; cost and time per case.
The headline is the end of a chain
A survived attack counts only if the tool saw the weakness coming. Each step is a smaller number than the one before, and the last one is the headline.
- 1FlaggedThe tool flagged the claim as weak.
- 2AimedIt proposed a change directed at the ground the examiner or Board later used.
- 3SurvivedThe revised claim survives the replayed attack.
- 4Kept its scopeThe revised claim still covers what it needed to cover.
Two traps, and how the Gauntlet closes them
The first trap
Narrowing always wins.
Add enough limitations and any claim dodges any reference. So a survived attack counts only when scope is preserved. For patents that went to litigation, the yardstick is the product accused in the public claim chart: the revised claim must still read on it. For applications, the yardstick is the claim the applicant eventually got allowed: the revised claim must be at least as broad.
The second trap
A search miss looks like a reasoning failure.
If the tool’s search never finds the reference the examiner used, its refinement cannot be aimed at it. So every Replay result is reported twice: once with the tool’s own search, which is what a customer gets, and once with the examiner’s or the Board’s references supplied. The gap between the two is the cost of search. The second figure is the refinement’s ceiling.
A survived replay shows that one specific attack would have failed. It does not show the patent would have been upheld, since a challenger facing different claims would have searched for different art. The results pages will say so.
The rules of the Gauntlet
Written before any result exists. Rules chosen after the numbers are known are not rules.
- IPublic material only, randomly drawn. No client, partner, or ClaimMark matters anywhere in the study, including for developing the method.
- IIThe headline run is fully automatic. Every recommendation is applied as proposed, with no human choice between input and output. An attorney-in-the-loop run may be published separately, never as the headline.
- IIIOutputs are the tool’s literal artifacts, hashed and time-stamped before scoring.
- IVEvery scoring prompt, sampling rule, seed, and ruling is published.
- VPrompts may be tuned on the pilot draw only. The headline draw is new and scored once.
- VILabels say exactly what was measured: “rejections the revised claims would no longer face”, never “office actions prevented”.
- VIIAn allowance is not a correct answer. Only positive action counts: a rejection, reasons for allowance, a holding.
- VIIIAnyone can dispute a ruling, and every correction is published with its date.
Registered samples
The seed, the rule, and the hash fix each sample before any tool runs on it. Anyone can re-derive the list and check the hash.
| Sample | How it was drawn | Seed and hash |
|---|---|---|
| Examiner pilot | Uniform random application numbers filed 2022 to 2024, kept if published with a first action on the merits and claims on file before it, until each of the eight technology centers had its quota. 154 draws, 50 kept. 38 first actions are rejections and 12 are allowances; 9 were mailed after 2026-07-01 and form a hold-out. | seed claimmark-backtest-group1-pilot-2026-09-25 SHA-256 617f4a3ab1a6c568f34b209e8b4cb8209d5c613379a089d11611c1add04ae93b |
| Board pilot | Original final written decisions in inter partes and post-grant review, one per patent, shuffled from a pool of 1,295 patents decided 2023-01-01 to 2026-06-30. Kept in shuffled order if the original complaint in a district-court case carries a public claim chart for that patent, until 50 are kept. A hold-out pool of decisions from 2026-07-01 is registered separately. | seed claimmark-backtest-group2-charted-2026-09-26 order SHA-256 b178b6e2fba415ae35f6c43d17c73116751247056f01fdb9a3aa053d75a4f789 |
Standings and the frontier
Standings
Published here after each scoring round, with every tool that has run the Gauntlet, ours included.
Each event gets its own table. We also plot the frontier: each tool’s Replay result against its cost per case, per capability area, so a reader can see which tools no other tool beats on both. Any overall ranking shows its weights on the page, and readers can change them, because any fixed weighting is an opinion.
Take the challenge
Run the Gauntlet
Any patent AI tool can enter, two ways.
Run it yourself
Use the published harness on the registered samples and submit the full output bundle with its hash, plus a trial account so we can rerun a sample of cases and confirm your results.
We run it
Where a product’s public trial terms allow, we run it ourselves, publish the bundle, and send the vendor the results before publication. The vendor’s reply is published beside its row.
There is no paid placement. Results are withdrawn only for a factual error, and the correction is dated.
Governance
The method, harness, sampling code, and scoring prompts are published on GitHub, and the samples on HuggingFace. Every change to the method is dated and explained. We are inviting an outside patent attorney and an academic research group to review the method and the rulings alongside ClaimMark’s co-founder, patent attorney Jürgen Vollrath, and we will name them here when they accept.
Method sign-off
I have reviewed the Gauntlet’s method as described on this page. Its answer key is what examiners and courts actually decided, not a model’s opinion of a draft. Its cases are drawn at random from public records and registered before any tool runs on them. Every tool, ours included, is scored the same way.
I sign off on the method. I will review the scoring rulings as results come in, and my name will not appear beside any result I have not reviewed.
Why we built it
Why did we make the Gauntlet? Existing benchmarks were not enough!
Before we built the Gauntlet, we went looking for a benchmark that could tell an attorney whether an AI tool’s claims would hold up. We could not find one.
Every benchmark we found measured something narrower than the question attorneys care about:
- Patsnap’s PatentBench measures only whether a search finds the references the examiner cited, on a test set nobody outside Patsnap can see.
- ABIGAIL’s PatentBench is open source and aims at drafting and Office Actions, but so far it scores only docketing arithmetic, and its dataset could not be downloaded.
- The academic benchmarks score similarity to the granted text, or a model’s opinion of the claims.
None of them asks the question an attorney is paid to answer: would these claims survive the examiner, the Board, or a court? Finding the right art is only the first step, and a claim that reads well can still be anticipated. A tool that cannot show what happens when its claims meet a real attack has not shown that it can be trusted.
So we built a benchmark that asks that question directly. The ClaimMark Gauntlet takes real applications and patents from the public record, lets a tool refine the claims, and replays the examiner’s and the Board’s actual attacks against the result. A win counts only when the claims survive and keep their scope. And because a benchmark is only as good as its honesty, we hold the Gauntlet to every rule we ask of others: random samples, outputs locked before scoring, and every result published.
Why benchmarks matter
Patent AI benchmarks are emerging — but what do they actually measure?
AI tools have moved into patent practice fast. Clarivate reports that AI use among IP professionals rose from 57% in 2023 to 85% in 2025. The question attorneys now ask is not whether to use these tools, but whether they can trust what the tools produce. When HGF polled patent attorneys in late 2025, accuracy and hallucinations topped the list of concerns, named by 42%, more than twice the share most worried about confidentiality.
The concern is well founded. A peer-reviewed Stanford and Yale study found that leading AI legal research tools, built specifically for lawyers, still gave incorrect or misgrounded answers between 17% and 33% of the time. And the responsibility does not move. The USPTO’s guidance reminds practitioners that their signature certifies their own reasonable inquiry, whether or not an AI tool helped, and ABA Formal Opinion 512 expects lawyers to review AI output with appropriate care.
So how can an attorney tell which tools deserve that trust? Checking every output by hand erases the time the tool saved. That is the job of benchmarks: independent, repeatable tests that show how often a tool gets it right, on cases like the attorney’s own. General-purpose AI has been measured this way for years. Benchmarks for patent AI are only now emerging, and we list every one we have found below.
This section explains which benchmarks currently exist, what those numbers measure and what they don’t.
Side by side
How the benchmarks compare
The same eight questions, asked of each.
| ClaimMark Gauntlet | Patsnap PatentBench | ABIGAIL PatentBench | Academic benchmarks (typical) | |
|---|---|---|---|---|
| What is scored | Whether the refined claims would have survived the examiner’s or the Board’s actual attack, with their scope intact. | Prior-art search only: does the search return the references the examiner cited? | Docketing arithmetic (deadlines, fees, classification). Drafting and Office Action quality are planned, scored by an AI judge. | Similarity to the granted text, or an AI judge’s rubric score. |
| Answer key | Positive action in the public record: examiner rejections, Board holdings, court rulings. | Examiner citations, grouped by patent family. | Fixed answers for arithmetic; a rubric opinion for quality. | The granted text, or a model’s opinion. |
| How cases are chosen | A uniform random draw. The seed and the list’s hash are published before anything runs. | 340 curated patent families. The set is not published. | Drawn from 604 real Office Actions. The dataset was not publicly downloadable on 2026-09-26. | Varies. Often older patents, or disclosures written by a model from the granted patent. |
| Outputs locked before scoring | Yes. Every output is hashed and the hash is time-stamped before scoring. | Not stated | Not stated | Rarely |
| Can an outsider rerun it? | Yes, case by case, on a trial account | No | Not yet: the code is open, the data is not | Often, when code and data are released |
| Training-data contamination | Cases decided after the model’s training cutoff are reported separately; search is capped at the priority date | Not addressed | Not addressed | Rarely addressed |
| Stability and cost | Every case runs twice; cost and time per case are published | Not reported | Not reported | Rarely reported |
| Status on 2026-09-26 | Published, with samples registered before any run | Published | Docketing layer published; the rest in progress | Published papers |
Patsnap and ABIGAIL columns reflect each vendor’s own benchmark pages and repository on 2026-09-26. “Academic” describes the eight papers reviewed under Academic work.
A reader’s checklist
How should you read a patent AI benchmark?
Ten questions that separate a measurement from a marketing number. We ask them of every benchmark we know of, and of our own.
None of these questions requires technical knowledge of AI. They are the questions an attorney would ask of any expert report: where did the data come from, who chose it, what was it compared against, and could someone else check it.
Our own column answers the same questions, and every answer in it can be checked against what we publish.
| Question | ClaimMark Gauntlet | Patsnap PatentBench | ABIGAIL PatentBench | PatentScore | Dis2Pat | PatRe |
|---|---|---|---|---|---|---|
1. Can you get the test set? A score on data nobody else can see is a claim, not a measurement. | Samples, every output, and their hashes are published | Not published | Code is open; the dataset returned “unauthorized” on 2026-09-26 | Prompts published; built from a public patent dataset | Code and data released | Code and data released, per the authors |
2. Was the sample drawn at random, by a rule published first? A hand-picked sample can make any tool look good, even when nobody intends it to. | Uniform random draw; seed and list hash published before any run | Curated for language and technology class | Selection rule not stated | 400 patents from two technology sections, 2016 to 2017 | Split from an existing corpus; rule not verified by us | Not verified by us |
3. Is the answer key an outcome, or an opinion? An examiner’s rejection happened. A model’s rubric score is somebody’s view. | Examiner rejections, Board holdings, court rulings | Examiner citations | Fixed answers for arithmetic; an AI judge for quality | Expert and model scores of claim form | The granted text, and an AI judge | Real Office Actions as reference; scoring method not verified by us |
4. Does the score track what happens to the application? Attorneys care whether claims issue and survive, not whether text looks like other text. | The headline is survival of the replayed attack with scope preserved | Finding the art is the first step, not the result | Not for the layer that is scored today | The authors say it does not address novelty | Similarity to one granted version | Models the examination exchange |
5. Are the comparisons fair? Beating a chat window at patent search says little about beating a search tool. | Every case also runs through a plain AI model as a baseline; other vendors are invited to run the same cases | Compared with general chatbots, not with search tools or a keyword baseline | General models only; rival vendors listed as “not submitted” | Validated against three human experts | Open and closed models, plus attorney pairwise review | Proprietary and open models |
6. Were outputs locked before scoring? Without a lock, a score can be tuned after the fact, even unintentionally. | Every output is hashed and time-stamped before scoring | Not stated | Not stated | Not stated | Not stated | Not stated |
7. Can you rerun a case yourself? The strongest proof is a reader getting the same answer. | Any case can be rerun on a trial account | No | Not until the data is available | Prompts are public | Yes | Yes, per the authors |
8. Does it report stability, uncertainty, and cost? An AI tool can give different answers to the same input. A score without its spread hides that. | Every case runs twice; agreement between runs, cost, and time per case are reported | No repeat runs, intervals, or cost | Not reported | Averages ten runs; variance not reported | Not verified by us | Not verified by us |
9. Is training-data contamination controlled? A model may already have seen the patent, and the examiner’s answer, during training. | Post-cutoff hold-outs reported separately; priority-date search ceiling | Not addressed | Not addressed | Patents from 2016 and 2017 | Disclosures written from the granted patent being predicted | Not verified by us |
10. Does the number stay put? A benchmark that changes should say when, and what changed. | Every change dated; the headline sample is scored once | Moved from 81% to 85%, with different comparators, and no version note | Case counts differ between the repository and the website | Versioned paper | Versioned paper | Versioned paper |
What a good answer looks like
The single most useful test is the fifth one. A number means little until you know what it beat. A specialized search tool beating a general chatbot is expected. A drafting tool scoring well on arithmetic that general models already get right is expected. The informative comparison is against the best available alternative, run on the same cases, by rules fixed before anyone saw the results.
The second most useful is the third. An outcome that happened, such as an examiner’s rejection or a Board decision, cannot be argued with. A model’s opinion of a claim can be useful as a diagnostic, but as a score it measures agreement with the model, not quality.
Benchmark report
Patsnap’s PatentBench: a real outcome, one narrow task, a test set nobody else can see.
Patsnap’s benchmark scores prior-art search against the references examiners actually cited. That is a sound answer key. What it measures is narrower than its marketing, and nobody outside Patsnap can check it.
What it is
| Publisher | Patsnap, which sells the search tool it ranks first |
|---|---|
| First published | September 2025; the current figures were on the page on 2026-09-26 |
| Task | Novelty search: given an invention, return prior art |
| Test set | 340 patent families; receiving offices about 32% US, 32% China, 18% Europe, 18% WIPO; 68% English, 32% Chinese |
| Answer key | The X and Y references examiners cited, deduplicated and grouped by patent family |
| Measures | X hit rate: share of cases where at least one examiner X reference is returned. X recall: share of all examiner X references found in the top 100 |
| Compared against | General-purpose AI chat products with web search |
| Published | Headline numbers and a one-case worked example. Not the test set, the queries, the prompts, or per-case results |
Patsnap also publishes a design-patent freedom-to-operate benchmark (261 samples) and lists a utility-patent freedom-to-operate benchmark as in progress.
The numbers, as published
| System | X hit rate | X recall, top 100 |
|---|---|---|
| Patsnap Novelty Search agent | 85% | 37% |
| ChatGPT 5.6 | 71.76% | 30.51% |
| Claude Opus 4.8 | 52.37% | 11.68% |
| Perplexity Pro | 39.16% | 6.4% |
| Kimi-k3 | 33.53% | 12.09% |
| DeepSeek-v4-flash | 22.09% | 6.25% |
| Gemini 3.1 Pro | 14.24% | 2.39% |
From patsnap.com/benchmark, 2026-09-26.
What 85% and 37% actually mean
85% is not the share of prior art found. It is the share of test cases in which the search returned at least one reference an examiner cited as an X reference. A search that finds one of five relevant references scores the same as one that finds all five.
37% is the share of prior art found. Of all the X references examiners cited across the test set, Patsnap’s top 100 results contained 37%. The rest, close to two in three, were not in the top 100. This is the more informative of the two figures, and it is the one Patsnap’s own tool scores on, in its own benchmark.
That is not a criticism of Patsnap’s engineering. Examiners search with classification codes, full-text databases, and years of art-unit experience, and matching them is hard. It is a reason to read the headline carefully.
Where the published figures disagree
- Which results count. The benchmark page defines the hit rate as a correct answer “among the top 1, 3, or 5 results”. The same page, and the benchmark hub, label the 85% headline “Top 100”. Those are very different tests, and the page does not say which one produced 85%.
- What the homepage says. Patsnap’s Eureka homepage presents the figure as “85% Prior art found in the top 100 results”. By the benchmark’s own definitions, the share of prior art found is the recall figure, 37%.
- Which competitor is quoted. The homepage compares Patsnap with ChatGPT 5.4 at 17.18%. The benchmark page reports ChatGPT 5.6 at 71.76%, within about 13 points of Patsnap.
- How the figures moved. In November 2025, Patsnap reported 81% and 36%, compared against ChatGPT o3 and DeepSeek-R1. The benchmark page now reports 85% and 37% against a different set of products, while Patsnap’s own novelty-search product page still shows 81% on the same day. No version history explains the change.
What it does not tell you
- Anything about drafting. It is a search benchmark. It says nothing about the quality of a drafted application or an Office Action response.
- How relevant each result is. A result is a hit or a miss. There is no graded relevance, and no check of which claim limitations a reference actually discloses.
- How it compares with other search tools. The comparators are general chat products with web search. No patent search tool, and no simple keyword search over the same collection, was run as a baseline.
- Whether the answer was in the question. The page does not say what text each query was built from. If it came from a granted patent, the examiner’s citations are printed on the patent’s face, and may be in a model’s training data.
- How much the number moves. There are no repeat runs and no confidence intervals. With 340 cases, sampling alone puts an 85% figure within about four points either way, before any run-to-run variation.
What it gets right
The answer key is an outcome, not an opinion: references that examiners actually cited, grouped by family so that a U.S. patent and its European counterpart count once. The test set spans four offices and two languages. And Patsnap published something, with a worked example, which most of the industry has not.
Why we do not report a score on Patsnap’s PatentBench
We cannot. The test set, the queries, and the per-case results are not published, so nobody outside Patsnap can run the benchmark, including Patsnap’s own customers. If Patsnap publishes the set, we will run ClaimMark-P on it and publish every output.
Instead we do two things. First, we run the same protocol on our own sample: applications drawn at random from the public record, scored against the references the examiner cited, with the search capped at each application’s priority date. That figure appears beside Patsnap’s with a plain caveat: a different sample, a different index, and a Patsnap figure nobody can verify. It is not a head-to-head, and we do not call it one.
Second, we measure what the hit rate leaves out. XRef, the search event of the ClaimMark Gauntlet, grades relevance using the X, Y, and A categories that European and international search reports assign to every citation, and checks whether a tool’s account of which claim limitations a reference discloses matches the examiner’s own mapping in the rejection.
Patsnap may still find more of the examiner’s references than we do. Its index is larger, and for most offices our worldwide search reads abstracts and bibliographic data rather than full text. That is why search recall is one row of our results, not the headline.
We also need to compare prior art search tools
There is no shortage of tools that search for prior art. Established databases such as Derwent Innovation, Questel Orbit, PatBase, and Patsnap, free public ones such as Google Patents, Espacenet, and the USPTO’s Patent Public Search, and a new wave of AI search tools all promise to find what matters. Almost all of them stop at the same place: a ranked list of references. Working out what each reference actually does to your claims is left to the attorney.
That last step is where ClaimMark-P is different, and it is why we compare search tools on more than whether they find the right documents. Below, ClaimMark-P and Patsnap, the one vendor that publishes a search benchmark, feature by feature. Patsnap’s column is taken from its public product pages on 2026-09-26.
| Capability | ClaimMark-P | Patsnap Eureka Novelty Search |
|---|---|---|
| Size and reach of the index | Over 100 million patent documents from more than 100 patent offices through the European Patent Office, plus the USPTO, a semantic patent and literature index, and web and product sources | Over 200 million patents, worldwide, per the vendor |
| Non-U.S. patent offices | More than 100 offices, including Europe, WIPO, China, and Japan: abstracts and bibliographic data for all of them, full text for European and international filings | Includes China, Europe, and WIPO |
| Starts from the invention disclosure | Built from the invention record’s inventive center, problem, and summary | From a disclosure |
| Searches again from the drafted claims | Recommends a revised search when the claims change, grounded in their exact language | Not stated |
| No invented references | Every reference is checked against its record; one that cannot be verified is dropped, never shown | Source-linked results, per the vendor |
| Explains why each reference matters | Overlap type and a written relevance note per reference | Feature-level evidence, per the vendor |
| Which invention elements each reference discloses, and which it misses | Listed per reference, both ways | Feature-level evidence; the missing side is not described |
| Which claim limitations each reference discloses, and which it misses | Listed per reference once claims exist, and mapped into the claims with the cited passage | Not stated for novelty search |
| Relevance score you can audit | Computed by fixed rules from the listed elements, so the same judgment always yields the same number; the rubric version is stored with it | Results are ranked; no score definition published |
| Feeds the examiner-style review of your claims | References flow into §102 and §103 review and the adversarial examiner simulation in the same matter | Search is callable inside drafting and Office Action workflows, per the vendor |
| Search capped at the priority date | Every search can be limited to art published before the priority date | Not stated |
| Published retrieval benchmark | XRef, in the ClaimMark Gauntlet, on public samples anyone can check | Yes, but on a test set nobody else can see |
Beyond the result list
Not just a search — why ClaimMark-P is revolutionary
Other tools hand you a list of references and wish you luck. ClaimMark-P puts every reference to work on your claims, the moment it is found.
- Every reference, taken apart. Which elements of your invention it discloses, which it misses, and which of your claim limitations it reads on, listed both ways, for every reference.
- Scores you can audit. Relevance is computed by fixed rules from those lists, so the same judgment always yields the same number, and the rubric version is stored with it.
- Mapped straight into your claims. Each limitation links to the passage of the art that reads on it, so you see the collision at the exact words where it happens.
- Turned into an attack before the examiner makes one. The references feed the §102 and §103 review and an adversarial examiner simulation that tries to reject your claims first, and every weakness it finds comes with a proposed fix.
- Never out of date. Change the claims and ClaimMark-P tells you the search is behind, and runs it again from the new claim language when you ask.
Search is where most tools end. In ClaimMark-P it is where hardening your claims begins.
Benchmark report
ABIGAIL’s PatentBench: the right ambition, most of it still to come.
ABIGAIL has published the most ambitious patent AI benchmark so far, as open source, and it is the only one that sets out to test drafting and Office Action work. Today the only scores are for docketing arithmetic, and the dataset could not be downloaded.
Who publishes it
ABIGAIL is a self-funded patent AI company. Its product drafts applications, prepares Office Action responses, runs prior-art searches, and handles prosecution docketing. PatentBench is published under the Apache 2.0 license on GitHub, installable as a Python package, with a white paper signed by ABIGAIL’s founder, a registered patent attorney.
ABIGAIL uses PatentBench on its website to point out that far better-funded rivals publish no benchmarks at all. That criticism is fair as far as it goes. The same standard should apply to PatentBench itself, and that is what this page does.
What it is designed to measure
7,200 test cases across five domains, derived from 604 real USPTO Office Actions from 2019 to 2024 across nine technology centers, per ABIGAIL’s site.
| Domain | Cases | What it tests | Stated human baseline |
|---|---|---|---|
| Administration | 1,500 | Deadlines, fees, IDS completeness | 99.8% |
| Drafting | 500 | Claim scope, specification support, terminology | 8.5 out of 10 |
| Prosecution | 2,500 | Rejection analysis, arguments, amendments | 8.6 out of 10 |
| Analytics | 1,500 | Examiner prediction, allowance probability | 75% |
| Prior art | 1,200 | Reference relevance, anticipation | 85% |
| Scoring layer | How it scores | Weight | Status on 2026-09-26 |
|---|---|---|---|
| 1. Deterministic | Fixed answers: deadline math, event codes, fee lookups | 30% | Published: 298 tests |
| 2. AI judge | A separate model scores outputs against rubrics | 35% | In progress |
| 3. Comparative | Head-to-head ranking between systems | 25% | Planned |
| 4. Human calibration | Registered attorneys check the automated scores | 10% | Recruiting |
What is actually scored today
| Layer 1, 298 tests | Score |
|---|---|
| ABIGAIL v3 | 100.0% |
| Claude Sonnet 4 | 99.1% |
| Gemini 2.5 Flash | 99.1% |
| Gemini 2.5 Pro | 88.7% |
The published leaderboard covers the deterministic layer only: classifying Office Actions, reading timelines, computing fees and deadlines. ABIGAIL’s own product scores 100%. Two general-purpose models score 99.1%.
A test that general models already pass does not tell tools apart. It confirms the arithmetic is right, which matters for docketing, but it says nothing yet about drafting or arguing an Office Action. Those layers carry most of the weight in ABIGAIL’s own formula, and they are not yet scored.
What we could not reconcile or reproduce
| GitHub README | abigail.app/patentbench | |
|---|---|---|
| Drafting cases | A drafting-focused set of 1,200 | A drafting domain of 500 |
| Office Action cases | An Office Action set of 1,800 | A prosecution domain of 2,500 |
| Where the data lives | A HuggingFace dataset, linked from the README | Described as public data |
| Can the public download it? | The dataset returned HTTP 401 (unauthorized) to an anonymous request on 2026-09-26 | |
The two pages may describe different cuts of the same data, a subset in one place and a domain in the other. Neither explains how they relate, and a reader cannot check, because the data could not be downloaded. Until it can, “reproducible” describes the intention rather than the current state.
The stated human baselines raise a further question. A drafting baseline of 8.5 out of 10 implies human work was scored on the same rubric, but we did not find who scored it, how many attorneys, or on which cases.
What it gets right
- It is open source, with a license that lets anyone run and extend it.
- It is built from real Office Actions, not synthetic cases.
- It says plainly which layers are finished and which are not.
- It plans to calibrate the automated scores against registered attorneys, which every AI-judged benchmark needs.
Where an AI judge falls short
When the drafting and prosecution layers are scored, they will be scored mainly by a model reading the output against a rubric. That is useful for catching form problems. It cannot tell you whether an argument would have persuaded the examiner, because the examiner is not in the loop. The strongest answer key for a response to a rejection is what the examiner did next, and for a drafted claim, whether it survived examination. The academic work page covers why rubric scores should be a diagnostic rather than a grade.
Docketing: what ABIGAIL tests, and what we do
The scored layer of PatentBench is docketing. ABIGAIL’s product syncs deadlines from the USPTO with weekend and holiday rollover, prefills USPTO forms, and offers examiner analytics. ClaimMark-P does not do docketing. It shows the period for reply exactly as the examiner wrote it and tells the attorney to use the firm’s docket. We do not calculate due dates, because a wrong date is a malpractice problem and firms already run dedicated docketing systems with a second-check rule. ClaimMark-P does generate IDS forms from the references in a matter.
So on PatentBench we would submit only the drafting and prosecution domains, and say so.
What we will do
When the dataset can be downloaded, we will run ClaimMark-P on the drafting and prosecution cases and publish every output, whatever the score. Separately, the ClaimMark Gauntlet scores against what examiners and the Board actually did, which no AI judge can stand in for. We would welcome ABIGAIL’s participation in both.
Research review
The academic work: careful, useful, and not yet asking the question that matters.
Researchers have published a wave of patent drafting and prosecution benchmarks since 2025. Each is careful within its frame. None yet asks whether an AI tool’s claims would have survived the examiner.

What has been published
| Work | Where, when | What it tests | Answer key | Open |
|---|---|---|---|---|
| Patent-CR | NAACL 2025 | Revise rejected claims toward the version that was granted | The granted claims; professional human evaluation | Dataset paper |
| EPD | EMNLP 2025 Findings | Generate claims, trained and tested on granted European patents | The granted claims | Not verified by us |
| PatentWriter | arXiv, July 2025 | Write a patent abstract from the first claim | The published abstract, by text-overlap metrics | Code released |
| PatentScore | arXiv, 2025 | Score the form of a generated claim on seven rubric dimensions | Three experts’ scores of 400 claims | Prompts published |
| Tree-of-Claims | arXiv, November 2025 | Claim search and revision against prior art | Examiner actions from the USPTO Office Action Research Dataset; 1,145 wireless patents | Not verified by us |
| PatRe | arXiv, May 2026 | Write the examiner’s Office Action and the applicant’s rebuttal | 480 real examination cases | Code and data released |
| Dis2Pat | EMNLP 2026 | Draft a full application from an inventor-style disclosure | The granted patent, text similarity, an AI judge, and 60 attorney comparisons | Code and data released |
| Vibe Patenting | arXiv, September 2026 | Test whether AI judges can grade and improve drafting agents | One patent attorney’s independent review | Not verified by us |
What they get right
Patent-CR and PatRe move in the right direction. Both are built on real examination history, so the model has to deal with an examiner’s objections rather than an idealized task. PatRe separately tests a model given the examiner’s references and a model that must find them, which is the right way to tell a search failure from a reasoning failure. Dis2Pat is the first to take the drafting problem from the inventor’s side, and it asked patent attorneys to compare outputs directly. Vibe Patenting checked its AI judges against a patent attorney and reported, honestly, where they disagreed.
Four problems they share
- The granted text is treated as the one right answer. Many claim sets would have been allowable. The granted claims are the product of a negotiation with the examiner that the drafter could not have seen. Scoring similarity to them rewards predicting the end of prosecution, not drafting well at the start of it.
- The answer leaks into the question. When an inventor’s disclosure is written by a model from the granted patent, as in Dis2Pat, it carries that patent’s structure, emphasis and vocabulary. The same risk arises for search when the query is built from a granted patent whose citations are printed on its face.
- The test cases are in the training data. Patents from 2016 and 2017, as in PatentScore, and granted European texts, as in EPD, are in every large model’s training corpus. A model may be recalling rather than reasoning, and none of these papers reports a set held out past the model’s training cutoff.
- An opinion stands in for an outcome. Where there is no granted text to compare against, a model grades the output. Vibe Patenting found that agreement between AI judges and a patent attorney depended strongly on which quality was being measured.
Underneath all four is one missing question. None of these benchmarks asks whether the claims would have held up: before the examiner, before the Board, or in court. That is the question an attorney is paid to answer.
PatentScore, up close
PatentScore is the most complete published rubric for grading patent claims, and its authors published it in full. It deserves a close reading, because rubric scores like it are likely to become the default way vendors grade drafting.
Is the rubric published?
Yes. The prompts and the scoring rubric, with example claims at each score level, are in the paper’s appendix. Seven dimensions: claim structure, punctuation, antecedent basis, element referencing, validity and uniqueness, ambiguous scope, and semantic similarity.
Is the scoring repeatable?
Not by design. A small model scores each dimension ten times and the average is taken. The paper does not report the sampling temperature or how much the ten scores vary.
How well does it agree with experts?
A correlation of 0.819 with three experts on 400 first claims. The dimension weights were fitted on the same 400 claims used to report that figure, so it is an in-sample result. The semantic dimension ends up with a weight of 0.001.
Is it a fair absolute grade?
No, and the authors do not claim it is. It measures whether a claim is well formed. By their own account it does not address novelty. A claim can score well and still be anticipated.
How we use it. ClaimMark-P already runs its own rubric reviews, per statute, with every finding tied to a claim and a recommendation, and its antecedent-basis check is rule-based rather than a model’s opinion. We treat those as diagnostics that drive recommendations, not as a score. In the Gauntlet we run PatentScore’s published prompts, unchanged, on each claim set before and after refinement, and report the result as a form-conformance row, labeled as a model-judged rubric that says nothing about outcome.
A better method
Nine principles. Each one answers a problem above, and each one is a rule of the ClaimMark Gauntlet.
- 1Score against what happened.
The answer key is positive action in the public record: an examiner’s rejection, the examiner’s reasons for allowance, a Board holding, a court ruling. The absence of a rejection proves nothing, so an allowed claim is never treated as a correct negative.
- 2Replay the attack, not the text.
Apply the tool’s recommendations, then replay the examiner’s actual rejection, with the examiner’s own references, against the revised claims. Similarity to the granted claims is not the question.
- 3Make narrowing count against you.
Any claim survives if you narrow it enough. A revised claim wins only if it survives the replayed attack and keeps its scope: it still reads on the product accused in litigation, or it is at least as broad as the claims the applicant eventually got allowed.
- 4Draw at random, by a rule published first.
Uniform random draws from the public record, with the seed and the list’s hash published before anything runs. Prompts may be tuned on a pilot draw; the headline draw is new and scored once.
- 5Lock the outputs before scoring.
Every output is hashed, and the hash time-stamped, before any score is computed.
- 6Control contamination.
Cases decided after the model’s training cutoff are reported separately. Every search is capped at the application’s priority date.
- 7Report the spread.
Run every case at least twice, and publish agreement between runs, the sample’s uncertainty, and cost and time per case.
- 8Use AI judges as instruments, not verdicts.
Where a model makes a scoring judgment, publish its prompt, calibrate it against attorneys on a sample, and label the result as model-judged.
- 9Let anyone rerun a case.
Publish every input and output, and let a reader rerun any case and dispute any ruling.
The white paper
We are writing up this method as a paper for peer review. The samples were registered, with their hashes, before any tool ran on them, so nobody, including us, could adjust the rules after seeing the numbers. We would welcome co-authors from the research groups whose work is reviewed here, and an outside patent attorney to review the scoring rulings.
Get involved
Help build an unbiased benchmark.
A benchmark written by one vendor will always be read with suspicion, including ours. So we publish the Gauntlet’s method, code, sampling procedure, and scoring prompts on GitHub and HuggingFace, and we invite patent attorneys, researchers, and other vendors to challenge any part of it.
- Attorneys: tell us which outcomes matter in your practice, and review our rulings on a sample of cases.
- Researchers: co-author the method paper, or run the Gauntlet on your own system.
- Vendors: run the Gauntlet and submit your results with proof of work, or tell us where the method is unfair to your product. Every vendor we score gets a right of reply, published beside its row.