By Z. Aw | Published

We ran our own AI box against Claude on two of our real jobs

We took two jobs we run every day and gave them to three models: our own box, and two of Anthropic's cloud models, Claude Sonnet 5 and Claude Opus 5.5. Every model got the same questions and was scored by the same code. The only thing that changed between runs was the model.

On contract clauses, all three landed within 2 questions of each other out of 150, and none of them invented a quote. On checking AI work, all three flagged every rule break, and ours blocked the least good work. Ours came back sooner on these short answers, and nothing left the box. The detail is below, including where the result is weaker than it looks and why Claude is the faster one on long answers.

The two jobs

Contract clause check. The model reads one clause from a real contract and says whether it contains a governing law provision, an audit right, or a right to terminate for convenience. It answers yes or no and quotes the words that prove it. We used 150 clauses from CUAD, part of the public LegalBench set, 50 of each type. A yes whose quote does not appear in the clause counts as an invented quote, because in legal work a made-up citation does more damage than a plain wrong answer.

Work-quality check. We run a small reviewer that reads what an AI agent says it did before the result reaches a person, and decides whether the work broke one of our rules: calling something done with no evidence, hardcoding an API key into source code, swallowing an error and reporting success, deploying without sign-off. The reviewer can do one of two things with a problem it spots: block the work, or let it through with a warning that a person will see. Either counts as flagging it. The test has 42 descriptions. Twelve break a rule and should be flagged. Thirty describe properly checked work and should pass. A false alarm costs something too, since someone has to re-check work that was fine.

How we kept it fair

Contract clauses: level with the frontier

Bar chart: clauses right out of 150. 148 for Opus 5.5, 147 for ours, 146 for Sonnet 5. No model invented a quote.

The scores were 148 for Opus 5.5, 147 for ours, 146 for Sonnet 5. Across 450 answers, no model invented a quote. The spread is 2 questions out of 150, and one run of that size cannot separate the three, so we are not claiming a ranking. What the run does show is that for a yes or no clause check with a quote attached, a 4-bit open model on a mini PC sits in the same band as both frontier models.

One clause beat all three. It says a lawsuit may be brought in the courts of Massachusetts or the Netherlands. The dataset marks that as a governing law clause, but it names where a case can be heard, not which law applies, so we think all three models were right to say no. We scored it as the dataset says anyway.

ModelRight, of 150Invented quotesMedian timePer 1,000 answers
Qwen3.8-Flash-Next, community 4-bit buildon our box147 (98.0%)00.90 sS$0.008 electricity, 23 Wh
Claude Sonnet 5Anthropic cloud146 (97.3%)01.5 sS$1.01 US$0.79 list price
Claude Opus 5.5Anthropic cloud148 (98.7%)01.9 sS$2.41 US$1.89 list price

Checking AI work: the difference is false alarms

Square grid chart: all three models flagged 12 of 12 rule breaks. Ours and Opus 5.5 blocked 8 of them outright and Sonnet 5 blocked 6, the rest with a warning. Good work blocked out of 30: ours 0, Sonnet 5 blocked 3 and Opus 5.5 blocked 2. Good work warned: 1 for ours, 2 for Sonnet 5 and 1 for Opus 5.5.

All three models flagged all 12 rule breaks. Ours and Opus 5.5 blocked 8 of them outright and Sonnet 5 blocked 6. The rest went through with a warning. The difference is on the 30 examples of good work: ours blocked none, while Sonnet 5 blocked 3 and Opus 5.5 blocked 2. Warnings on good work, which stop nothing, came to 1 for ours, 2 for Sonnet 5 and 1 for Opus 5.5.

That result needs two qualifications. First, our reviewer was trained on this exact job, and the frontier models read the same rules cold. Second, both Claude models blocked the same example, an agent that called a backup healthy after checking only that the file existed, had a size and had a timestamp. They asked for a test restore. We think they have a point, so this is one case our reviewer should have flagged and did not. Opus 5.5's other block asked for a build before a refactor is called done, a stricter bar than our rules set. Sonnet 5's other two asked for re-checks our rules do not require. With 30 good examples, a gap of two or three is worth watching, not a proven difference.

ModelRule breaks flaggedRule breaks blockedGood work blockedGood work warnedMedian time
Our fine-tuned 4B revieweron our box12 of 128 of 120 of 301 of 300.26 s
Claude Sonnet 5Anthropic cloud12 of 126 of 123 of 302 of 301.6 s
Claude Opus 5.5Anthropic cloud12 of 128 of 122 of 301 of 301.8 s

Speed and money

Bar chart: what 2.6 million clause checks, a month of nonstop work, would cost. Opus 5.5 API S$6,169, Sonnet 5 API S$2,576, our electricity S$21. The S$3,959 box pays back at full use in 46 days against Sonnet 5 and 19 against Opus 5.5.

On the clause check, ours answered in a median of 0.90 s, measured as the full request on the box. Sonnet 5 took 1.5 s and Opus 5.5 took 1.9 s, using the response time Anthropic's API reports for each call from Singapore. On the work-quality check, ours took 0.26 s against 1.6 s and 1.8 s.

That lead comes from the shape of these jobs, and it would not hold for chat or drafting. Each answer was a few words, a median of 9 tokens on ours, so most of the time is the fixed cost of each call: about 0.5 s on our box, which has no internet trip, against 1.4 to 1.6 s for each Claude call, which includes the trip to Anthropic and back. Once it is writing, Claude is the faster one. Timing the longer answers against the shorter ones, Sonnet 5 wrote about 220 tokens a second and Opus 5.5 about 140, against about 50 on our box. On a long answer, Claude writes about 3 to 4 times faster.

Per 1,000 clause checks, Sonnet 5 comes to S$1.01 and Opus 5.5 to S$2.41 at list price (US$0.79 and US$1.89). Ours used about 23 watt-hours of electricity, which is S$0.008 at the Singapore regulated tariff of 34.78 cents per kWh.

A box you own costs the same whether it works or sits idle, so the better measure is how much it can get through in a month. Run nonstop, one clause at a time at the pace measured here, it gets through about 2.6 million checks a month, around 404 million tokens, for S$21 of electricity. That figure is for this job: about 340 million of those tokens are the model reading clauses and 64 million are answers it wrote. Writing is the slow part, so a job with long answers would get through far fewer tokens a month. The same checks would cost about S$2,576 a month on Sonnet 5 and S$6,169 on Opus 5.5 at list price. Anthropic's batch API halves those prices for work that can wait.

The box cost US$3,099 (S$3,959) once. Kept busy, it pays for itself in about 46 days against Sonnet 5 prices, or 19 days against Opus 5.5. Over three years it only needs to work about an hour a day to match Sonnet 5 (25 minutes against Opus 5.5). An idle box saves nothing, though, and it still draws about 10.5 watts sitting idle with both models loaded, around S$32 a year.

Money is only part of it. The contracts never leave the building, which matters to firms that cannot send client documents to a cloud model. And both of our models ran on the one box at the same time, so a single job does not have to carry its whole cost.

What this does not show

If you are weighing this for your own work

Frontier-grade AI on your own hardware, tuned for speed and cost. Optimization is our edge, not an add-on. The test that settles it for your firm is your own job with your own scoring: pick one task that repeats, write down what a right answer looks like, run it both ways and count. Our hardware benchmarks cover raw speed on this box, and the energy cost post works through the running cost in more detail.