Legal AI Accuracy

  • Executive finding: In the reported Harvey LAB-AA results, the leading model satisfies 94.6% of individual rubric criteria but fully passes only 26.7% of legal tasks; those two numbers measure very different things.
  • Benchmark scope: Harvey LAB-AA evaluates agents on 120 private legal tasks across 24 practice areas; in the launch round of 28 models, 13 fully passed zero tasks and the launch leader fully passed 14.2%.
  • Reliability and cost: On AA-Omniscience, Claude Fable 5 with fallback reports 65.35% accuracy, an 89.97% attempt rate, and a 63.64% conditional hallucination rate, while reported per-task costs across the field spanned roughly 950x.
  • Enterprise decision: Custom AI is not the model alone; it combines the right model with retrieval, rules, validators, permissions, workflow controls, human review, observability, and safe fallback behavior.
  • Consulting support: We help organizations evaluate, architect, secure, and operate AI systems through our AI consulting services.
  • Enterprise training: We provide hands-on training programs for AI evaluation, agents, cloud platforms, governance, and production delivery.
  • Top clients: We help Fortune 500 enterprises, large corporations, mid-size firms, SMBs, and startups with AI automation development, enterprise search, agent systems, consulting, and hands-on training; our clients include Microsoft, Google, Broadcom, Thomson Reuters, Bank of America, Macquarie, Dell, and more. Learn more about our top clients.
 

Legal AI Accuracy: Harvey LAB-AA And Enterprise Automation

Executives fund business outcomes, not model names. Those outcomes must be accurate enough, controlled enough, and explainable enough to defend. Legal work makes that distinction unusually clear because a polished answer can still miss one required clause, cite the wrong authority, or cross a jurisdictional boundary.

Artificial Analysis's Harvey LAB-AA evaluation gives the market a useful test of that problem. Its question is more demanding than a knowledge quiz: can an agent complete real legal tasks against a rubric of binary requirements, where a single missed requirement prevents the entire task from passing? That question aligns with the practical concerns addressed in our AI solutions for legal teams.

The answer, so far, is that models satisfy most individual requirements and finish very few complete deliverables. That is not an argument against legal AI. It is an argument that model quality is only one part of legal-system quality, and that the rest of the system is where your accuracy actually comes from.

Executive scorecard showing legal AI all-pass performance, reliability metrics, and the need for a model-plus-controls system

 

Why Model Progress Alone Is Not An Enterprise Strategy

Ask most organizations what their legal AI strategy is and the honest answer is a forecast: Harvey will solve it, or Anthropic will, or OpenAI will, and the performance problem will disappear on its own. That is a reasonable expectation about the industry. It is not a plan your organization can execute, budget, staff, or be held accountable for. A practical AI governance program requires decisions your organization can control.

The cost of that position is measured in years. A team that started waiting when large language models first became credible for professional work has now spent several years without a governed capability in production, while the underlying evidence has not moved the way the waiting position assumed it would. Your competitors who built controlled systems in that window are not waiting on anyone's release schedule.

The Harvey LAB-AA launch results make the point bluntly. Thirteen of the 28 evaluated models fully passed zero of the 120 tasks. The launch leader fully passed 14.2% of the benchmark tasks, meaning approximately 85.8% did not fully pass every tested criterion. Later results have improved on that figure, but not by the margin a waiting strategy requires.

Traditional model benchmarks encourage this waiting. A model can perform well on multiple-choice knowledge, coding exercises, or isolated reasoning problems and still fail to produce a complete work product with every required element in the right form. High benchmark scores do not guarantee complete professional work. Harvey LAB-AA brings the evaluation closer to the way professional work is actually assigned: the agent receives instructions and source material, works through a task, and produces a deliverable such as a memo, disclosure schedule, deposition outline, or document review output. Artificial Analysis runs the evaluation independently through its own agent-execution harness, which supplies each model with instructions and source material and collects the resulting deliverable.

That reframes the executive conversation from "which model wins" to a set of questions you can actually answer: how often does the system complete the full task, how often does it miss individual requirements, what does each task cost, how much human review remains, and what controls prevent an incorrect answer from becoming an official work product?

 

How Harvey LAB-AA Measures Complete Legal Work

Harvey LAB-AA uses a private set of 120 tasks across 24 legal practice areas created by the Harvey team. The coverage includes corporate M&A, capital markets, tax, litigation, bankruptcy, real estate, employment, intellectual property, and other areas where the output depends on multiple facts and requirements.

Each task has a rubric made up of binary criteria. The evaluation reports two primary views:

Criterion pass rate: The share of individual requirements satisfied across the evaluated work. This tells you how often the system gets separate pieces right.

All-pass rate: The share of complete tasks where every rubric criterion is satisfied. Artificial Analysis presents this as the primary metric because it reflects the standard applied to real professional deliverables.

A high criterion pass rate can coexist with a much lower all-pass rate. In legal work, the missed requirement is rarely decorative. It may be the missing exception, the incorrect jurisdiction, the omitted party, the unsupported citation, or the absent disclosure that changes the business consequence.

The benchmark uses a single LLM judge to grade the deliverables against task-specific rubrics. That makes results comparable within the benchmark, while still leaving room for human review of the rubric design, judge behavior, and the legal significance of each criterion.

 

What The Harvey LAB-AA Results Tell Enterprise Buyers

The table below summarizes a reported set of Harvey LAB-AA results. It includes models from the initial group and later additions, so the figures should be understood as one reported round of testing rather than a permanent ranking.

Model Criterion pass rate All-pass rate
Kimi K3 94.6% 26.7%
Claude Fable 5 with fallback 93.6% 14.2%
Grok 4.5 92.4% 13.3%
Muse Spark 1.1 93.1% 8.3%
Claude Opus 4.8 91.1% 7.5%
GLM-5.2 91.0% 7.5%
MiniMax-M3 88.4% 6.7%

Read the two columns together. The leading reported result satisfies 94.6% of individual criteria and fully passes 26.7% of the 120 tasks. In benchmark terms, 73.3% of the tasks did not fully pass every tested criterion. The benchmark does not specify how much correction or review each task would require in production, but it shows why criterion-level performance cannot be treated as finished legal work.

The improvement over the earlier results is real and worth acknowledging. The reported table also includes newer model entries, so it should not be read as a controlled longitudinal comparison of how existing models improved over time. Claude Fable 5 with fallback still reports the same 14.2% it recorded in the earlier round, where it led the field ahead of Claude Opus 4.8 and GLM-5.2 at 7.5%, MiniMax-M3 at 6.7%, Claude Sonnet 5 at 5.0%, and GPT-5.5 and Claude Sonnet 4.6 at 4.2%. If your plan depends on the models you already selected getting better on their own, that is the pattern to study carefully.

Harvey LAB-AA comparison of criterion pass rates and all-pass rates for reported legal AI models

 

Why General-Purpose Benchmarks Do Not Predict Legal Work

General benchmarks explain why expectations ran ahead of results. Under the published MMLU evaluation protocols, GPT-3 X-Large scored 43.9%, GPT-3.5 scored 70.0%, and GPT-4 scored 86.4%. Those results measure broad multiple-choice knowledge, not completed legal work.

Model reference MMLU score What the score represents
GPT-3 X-Large 43.9% Broad multitask knowledge using few-shot evaluation
GPT-3.5 70.0% Broad multiple-choice knowledge across 57 subjects
GPT-4 86.4% Broad multiple-choice knowledge across 57 subjects

That curve is the one most executives have in mind, and it is genuinely impressive. Broad multiple-choice scores improved substantially, but those results do not establish equivalent progress on factual reliability or complete professional work. On factual reliability, AA-Omniscience reports Claude Fable 5 with fallback at 65.35% accuracy. On complete legal work, the leading reported Harvey LAB-AA result sits at 26.7%. The measures closest to professional work remain much lower.

Those are three different benchmarks and the percentages are not directly comparable. The improvement in one measure does not establish equivalent progress on the others. MMLU never asked whether the system identified every relevant document, followed a client-specific template, preserved jurisdictional scope, cited an authoritative source, or delivered every required file in the correct format.

The better interpretation is that models became much better at broad knowledge tasks, while professional work remains a coordinated activity. It combines evidence retrieval, classification, reasoning, drafting, formatting, verification, risk handling, and accountability. A model can be excellent at one part and still fail the finished work product.

Historical MMLU accuracy progression from GPT-3 to GPT-4 compared with the separate legal all-pass question

 

Why All-Pass Performance Matters

The distinction is easiest to understand with a contract review. Suppose a task requires the agent to identify assignment restrictions, change-of-control provisions, notice periods, exceptions, governing law, and source citations. A system can satisfy most of those criteria and still fail the task because it missed one material restriction.

That is why the all-pass rate is closer to an acceptance test for a complete deliverable. It does not mean the task is safe for autonomous use, because the rubric may not capture every legal judgment. It does show whether the system cleared all of the requirements the benchmark authors chose to test.

For enterprise procurement, both figures belong in the decision record. Criterion pass rate helps you locate where a model is strong or weak, which is what you need in order to design controls around it. All-pass rate tells leadership how often the complete workflow cleared every tested requirement without partial credit.

 

Accuracy, Hallucination, And Safe Escalation

Harvey LAB-AA and AA-Omniscience measure different things. LAB-AA evaluates completion of legal work products. AA-Omniscience evaluates factual recall and calibration across 6,000 questions in six domains, with metrics for accuracy, hallucination, and a bounded reliability index. Our AI evaluation consulting work treats these as separate questions when designing an enterprise test plan.

On AA-Omniscience, Claude Fable 5 with fallback reports:

Accuracy: 65.35% of the benchmark questions were answered correctly.

Attempt rate: 89.97% of questions received an attempted answer.

Hallucination rate: 63.64% under the benchmark's conditional hallucination metric. This is not a percentage of all questions, so it must be read alongside accuracy and attempt rate rather than quoted on its own.

AA-Omniscience Index: 43.3 on a scale that rewards correct answers, penalizes incorrect answers, and does not penalize abstention.

The attempt rate is the figure most executives skip, and it changes how the others should be read. A system that answers almost everything looks helpful while creating more opportunities for confident error. In a legal workflow, the ability to abstain, cite evidence, request clarification, or route the matter to a reviewer is not a limitation of the system. It is part of what makes the system usable.

AA-Omniscience reliability metrics showing accuracy, hallucination rate, attempt rate, and reliability index for Claude Fable 5

 

Cost Per Task And Model Selection

The cost data is where the model-only assumption becomes expensive. Artificial Analysis reports the cost of running one benchmark task, including input, cached input, reasoning, and answer-token costs. Across the evaluated field, that cost spanned roughly 950x: the most expensive configuration ran approximately $18.90 per task, while Gemini 3.1 Flash-Lite passed 31.1% of individual criteria for approximately $0.02 per task. The table below includes only cost examples explicitly published in the launch report, not every model on the live evaluation page.

Model Reported benchmark result Approximate cost per task
Claude Fable 5 with fallback 14.2% all-pass Approximately $18.90
Claude Sonnet 5 5.0% all-pass Approximately $11.80
Claude Opus 4.8 7.5% all-pass Approximately $8.20
GLM-5.2 7.5% all-pass Approximately $1.30
Gemini 3.1 Flash-Lite 31.1% criterion pass Approximately $0.02

The first four rows use all-pass rates. Gemini 3.1 Flash-Lite is shown with its criterion-pass rate because that is the result published with its cost; it should not be compared directly with the all-pass figures. One comparison in the launch report deserves leadership attention. GLM-5.2 matched Claude Opus 4.8 in the second-place tier at a 7.5% all-pass rate, while costing approximately 16% of Opus 4.8's cost based on the published approximate figures. The supplied launch post used a rounded comparison of approximately $1 versus $19 and described GLM-5.2 as roughly 6% of Claude Fable 5's cost. These are benchmark execution estimates, not universal production costs. Paying more does not reliably buy completeness.

A more expensive model may still be justified for a high-risk task, but only when the additional performance changes the total workflow. The calculation should include review time, rework, verification, escalation, latency, and the cost of an incorrect deliverable. A cheaper model with stronger retrieval and deterministic validation is often the better system for a narrow, repeatable workflow.

Model routing is the practical response. A lower-cost model classifies or summarizes, a stronger model handles difficult reasoning, and deterministic services validate citations, dates, calculations, required fields, and output structure before anything reaches a reviewer. You stop paying frontier inference prices for steps that never needed a frontier model.

 

Why Legal AI Needs More Than A Model

Model-only automation is not a law of physics. It is a design constraint, and in most organizations nobody ever decided to adopt it. It arrived with the vendor demo and stayed. Once you accept it, your only remaining lever is waiting for someone else's next release.

Training a frontier model is one way to solve a problem in technology, and it is the most capital-intensive way available. It requires capital, infrastructure, specialized talent, and operating costs that most enterprises will not build and manage themselves. Retrieval, deterministic rules, validators, schema enforcement, routing, and human review are also ways to solve the problem, and your organization can build and own all of them through a disciplined software architecture program.

Drop the constraint and the engineering options open up immediately. The question stops being "is the model good enough yet" and becomes "which part of this workflow needs a model at all." That is a matter of discretion: matching the right tool to the right problem, rather than routing every step through the most expensive component you have.

We reached this conclusion early. Back in 2021 and 2022, while much of the market expected model scale alone to solve the full workflow, we were already building systems that combined models with retrieval, rules, validation, permissions, and human review, because neither the accuracy evidence nor the economics supported a single-model bet. The benchmark results published since then have been consistent with that read.

The goal is to improve accuracy and control costs by using models only where they add value. Deterministic checks do not generate citations or forget jurisdictions in the way a model can, and a validator, lookup, or rule can handle a step without frontier inference when the workflow is designed for it.

Model-plus-controls architecture for enterprise legal AI with retrieval, rules, validators, workflow, permissions, review, and observability

 
Where Controlled Legal AI Can Help

Legal AI becomes reliable when the workflow is designed around a specific task rather than a general promise that the model can handle legal. The following examples show how controls change the operating design.

Contract review at volume: Your in-house team can use a model to identify non-standard indemnity, assignment, and change-of-control language across vendor contracts. A clause library, matter-scoped retrieval, citation requirements, and a validator for required fields route uncertain findings to a lawyer instead of letting a fluent but incomplete summary circulate.

Regulatory memo drafting: A compliance team can assemble a first draft from current regulatory material. An authoritative-source connector, citation resolver, jurisdiction filter, and mandatory sign-off keep an unverified citation out of the approved memo.

Litigation document triage: A discovery workflow can use an agent to classify documents and identify likely privilege or relevance. Known privilege markers and deterministic filters operate alongside model triage, while borderline and high-value documents receive human review with a complete audit trail.

Cross-border employment guidance: A global organization needs termination guidance that differs by country, worker classification, and employment agreement. Jurisdiction-scoped retrieval, country-specific rules, permissions, and escalation to local counsel are far more dependable than asking one general model to blend all jurisdictions into a single answer.

In each workflow the model remains useful. The difference is that the system defines what the model may see, what it may do, what must be checked, and when a person must take over.

 
A C-Suite Framework For Legal AI

Treat legal AI as a governed operating capability rather than a race to select the highest-scoring model. The benchmark results support a practical set of questions for each leadership role.

CEO: Tie investment to a named business outcome, risk owner, and acceptance measure. Faster drafting is not enough if the organization cannot explain who approved the final work product or how an error would be detected.

CIO: Require data permissions, integration boundaries, audit logging, vendor accountability, and a support model before approving broad deployment. The system must fit your existing identity, security, retention, and incident processes.

CTO: Keep the model layer replaceable. Benchmark changes, provider outages, pricing changes, and new releases should never require an application rewrite. Put retrieval, rules, validation, orchestration, and evaluation around the model so improvements can be measured without rebuilding the platform.

General counsel: Define which tasks may be assisted, which require review, and which should not be automated. Require source citations, clear abstention behavior, matter-level permissions, and an audit trail that can explain how a draft was produced.

Risk tier Example use Minimum operating controls
Lower risk Internal research summaries and source discovery Grounded retrieval, citations, permissions, and observability
Moderate risk Draft contracts, first-pass review, and document triage Retrieval, rules, validators, audit logs, and qualified human review
High risk Regulatory advice, privilege decisions, and external filings All control layers, mandatory human sign-off, safe refusal, and fallback
External action Sending advice, filing documents, or changing a system of record Explicit approval, action-level authorization, reversible execution, and post-action review
 
The Controls Required For Production Use

A leaderboard winner is not automatically the best production choice. The right design depends on the task, the evidence available, the cost, your review capacity, and the consequences of an error. A defensible legal workflow generally combines four things:

Grounded retrieval and permissions: Search authoritative matter documents, clause libraries, regulations, and internal policy with citations, version control, matter boundaries, role-based access, ethical walls, and retention rules enforced in the retrieval and tool path.

Deterministic rules and validators: Enforce non-negotiable requirements such as jurisdiction, document type, party names, disclosure fields, and deadline calculations, then check citations, required sections, calculations, schema fields, and output format before delivery.

Orchestration with fallback: Break a large task into bounded steps, route each step to the right model or service, and let the system abstain, request missing information, use a second model, return a source-only answer, or escalate to a person when evidence or confidence is insufficient.

Human review and observability: Require qualified review for filed documents, privilege decisions, regulatory advice, and client communications, and record the prompt, retrieved evidence, model version, tool calls, rule decisions, reviewer actions, latency, cost, and final disposition.

Start with a focused evaluation that compares a baseline model against the same model wrapped in grounded retrieval, structured prompts, validators, and human review. Measure full-task completion, individual requirement coverage, citation correctness, refusal behavior, review time, latency, cost, and failure severity. That comparison, run on your own tasks, settles the model-only question faster than any leaderboard.

Our AI evaluation consulting work can establish that test harness before you commit to a large rollout. For enterprise data, the design may connect to enterprise search, Azure AI Search, document stores, knowledge graphs, or existing case-management systems. For agentic workflows, the control surface extends to tool authorization, workflow state, approvals, model routing, prompt-injection defenses, and tracing, which our AI agent consulting and MCP consulting work connects to your application and security architecture. Delivery should also align with our DevOps and observability practices.

 
How Cazton Helps Build Legal AI Systems

We help organizations design custom AI systems that combine models with the rest of the technology required for dependable work. That can include retrieval, enterprise search, workflow orchestration, rules engines, schema validation, model routing, human review, permissions, observability, evaluation, and production support.

Our work can begin with a focused AI proof of concept, a benchmark and evaluation program, an enterprise architecture engagement, an agent workflow, or hands-on technical training. We support Fortune 500 enterprises, large corporations, mid-size organizations, SMBs, and startups across AI, cloud, data, software architecture, and DevOps.

The benchmark message is measurable rather than rhetorical: model capability improved substantially, and complete professional work remains a compound systems problem. So the question for your next planning cycle is simple. If the models do not solve the full workflow on their own within your budget horizon, what is your plan? Waiting is an answer, but it is not a plan, and it is the one option that guarantees another year of the same result.

Explore our AI consulting, AI automation, AI employees for legal, and training services. When you are ready to evaluate a legal AI workflow, define the controls, or plan a proof of concept, contact us to discuss your requirements.

Cazton is composed of technical professionals with expertise gained all over the world and in all fields of the tech industry and we put this expertise to work for you. We serve all industries, including banking, finance, legal services, life sciences & healthcare, technology, media, and the public sector. Check out some of our services:

Cazton has expanded into a global company, servicing clients not only across the United States, but in Oslo, Norway; Stockholm, Sweden; London, England; Berlin, Germany; Frankfurt, Germany; Paris, France; Amsterdam, Netherlands; Brussels, Belgium; Rome, Italy; Sydney, Melbourne, Australia; Quebec City, Toronto Vancouver, Montreal, Ottawa, Calgary, Edmonton, Victoria, and Winnipeg as well. In the United States, we provide our consulting and training services across various cities like Austin, Dallas, Houston, New York, New Jersey, Irvine, Los Angeles, Denver, Boulder, Charlotte, Atlanta, Orlando, Miami, San Antonio, San Diego, San Francisco, San Jose, Stamford and others. Contact us today to learn more about what our experts can do for you.