- AI model
- Claude Sonnet 5
- Model runs
- 8each recorded individually
- AI work
- 250,56944,291 read · 139,814 written
- Elapsed time
- 5m 14s
- Cost
- $1.61all 8 runs combined
Office action reasonableness review · U.S. App. 15/352,952
One final rejection. Nine cited references. We reviewed it three ways on the same day — once through OneTwelve, and twice through Cowork, Anthropic’s agentic assistant, handed the very same files. One of the Cowork runs used the identical AI model OneTwelve did. It cost three times as much, and did fifty times the work to get there.
All three compare OneTwelve against Cowork running the same AI model — Claude Sonnet 5 — on the same case, the same day. That Cowork run was performed twice; the figures quoted are the average of the two.
And cost is the smaller half of the story. On a fraction of the tokens, OneTwelve returned patent-centric work product the flagship models did not produce on their own, at any price — a reasoned verdict on every rejection, grounded in the cited art, filed with the case — from a case file it assembled itself, off the serial number in the open document. The Cowork runs, including the one on the largest model available, hedged: they described each rejection without committing to whether it holds, and without the citations that would let you check the call. The judgement is always the attorney’s — the question is whether you start from a stated, checkable position or from a summary you have to work up yourself.
Application 15/352,952, final rejection mailed 15 December 2020: seven rejections across the claim set, eight patent references and one non-patent reference. Every run had the office action, the specification, the drawings and all eight reference documents in front of it. Every run produced a written reasonableness review.
Every figure on this page assumes the case file is already sitting in a folder. Getting it there is work, and it is work only one of these two does for you.
To run the comparison at all, we had to do Cowork’s gathering by hand first — and then hand OneTwelve the same folder, so both sides started from the same place. None of that manual retrieval is counted in any number on this page. In an ordinary week it lands on whoever is drafting: a docket lookup, eleven downloads, and a naming convention to keep straight, before a single word of analysis gets written.
Put whatever that is worth in your practice against it — a quarter of an hour an office action is a conservative guess, and it is a quarter of an hour that gets written off, or billed to a client for moving PDFs around. Neither is what anyone is paying a patent attorney for. Retrieval that happens on its own does not just remove a hassle; it hands the time back to work worth billing.
Same review, same evidence, same day. Longer bars are worse in every row.
Bars are proportional within each row. Hover or tab to a bar for the detail. “Units” are how AI providers bill — roughly three quarters of a word each.
The cheapest run and the most expensive run used the same AI model, so raw model intelligence is not what separates them. What separates them is how much of the case file each one had to plough through before it could say anything useful.
2.4× more of the case file read by the smaller model than the larger one — for the same job, in the same assistant.
Run inside Cowork, Claude Sonnet 5 read 13 million units. Claude Opus 5 read 5.4 million. The per-unit price of the smaller model is a fraction of the larger one, and it still finished up costing the same — $5.30 against $5.28.
The likeliest explanation is discrimination. Deciding which four of forty documents bear on a particular rejection is itself a hard judgement, and the larger model is simply better at it: it makes fewer, better-aimed passes. The smaller model compensates by reading more of everything. OneTwelve never asks the model to make that call. The pipeline already knows which documents a rejection turns on — which is exactly why a mid-sized model is enough, and why the same mid-sized model costs a third as much here as it does on its own.
Office actions, prior-art patents and file-wrapper documents arrive as scans — pictures of paper. OneTwelve reads them with its own recognition models, trained on millions of patent documents, prior-art references and prosecution filings. The text comes out clean and labelled: this is claim 1, this is the rejection, this is the passage the examiner leaned on. A general assistant gets none of that. It has to puzzle out the raw scans itself, every time it opens them — and you pay for every attempt.
OneTwelve knows which four documents a particular rejection turns on, and hands the model exactly those, once. A general assistant has no way to know, so it keeps the whole file in view and re-reads it turn after turn to stay oriented. That re-reading is the bill.
The reference copies in the file are degraded … I can verify the substance of the passages the Examiner relies on … but I cannot verify most of the pin cites … I am relying on recoverable text, which I paraphrase or reconstruct rather than pretend to quote verbatim.
We handed the same scanned references straight to a top-tier model as a separate check. It opened its answer by warning that it could not check the examiner’s citations against the art it had been given. That is the failure OneTwelve’s recognition models exist to prevent — the passages arrive clean, with the column and line numbers the examiner cited still attached, so a citation can actually be verified.
Cost is the easy half of the comparison. The half that matters when a client asks is what actually arrives at the end — and that is where the gap widens rather than closes. Fewer tokens bought the better answer, not a cheaper one: what came back is shaped like prosecution work, because the system was built for prosecution rather than pointed at it.
A verdict per rejection, not an essay. All seven rejections came back with a position — sustainable or not, and why — in a form your review and docketing process can act on.
Grounded in the actual references. Every assertion traces to text in the cited art. On this run, nothing unsupported made it into the memo and nothing had to be handed back for a person to resolve.
A memo, filed with the case. The output lands as a formatted reasonableness memo stored with the application — not a chat transcript somebody has to rewrite.
Measured, not projected. Each of the eight model runs is recorded on its own — what it read, what it wrote, what it cost — and then totalled: $1.61 for this review. Every figure on this page was read straight out of that record. Nothing here is modelled, sampled, or extrapolated from a demo.
The same difference, multiplied by a docket. Compared against the Cowork run on the same AI model OneTwelve used. These are AI processing costs for reasonableness reviews only, not a OneTwelve price list.
| Reviews per month | OneTwelve | Cowork, same model | You keep |
|---|---|---|---|
| 25 | $40 | $133 | $93 |
| 100 | $161 | $530 | $369 |
| 250 | $404 | $1,325 | $921 |
| 500 | $807 | $2,650 | $1,843 |
A case study is only worth the provenance behind it — including what it does not prove.