PrudentBench v1.0 · updated 6 September 2026

Which AI model is suitable for Dutch legal research?

PrudentBench does not run the external benchmarks itself. We summarise published results, add one screening of our own on Dutch cases and test every model against a compliance gate.

As of September 2026 Claude Opus 5, Claude Fable 5.1 and Muse Spark 1.3 (Max) tie for the highest score on the Vals AI Legal Research Bench: 55.29 percent all-pass [1]. After the compliance test, processing in the EU, a European contracting party and no training on customer data, 6 of the 10 assessed models remain; Claude Fable 5.1 and Muse Spark are rejected, which makes Claude Opus 5 the highest-scoring model that passes the gate. That is why LEO runs deep analyses on Claude Opus 5 and quick questions on Gemini 3.8 Flash, both through EU endpoints.

Highest legal-research score
55.29 %Highest legal-research scoreVals AI, measured 5 September 2026 [1]
Models through the compliance gate
6 of 10Models through the compliance gatePrudentBench v1.0
Main models at LEO: deep and fast
2Main models at LEO: deep and fastProduction state 6 September 2026
Models assessed
10Models assessedPrudentBench v1.0

Beau Jonkhout, Chief Technology Officer at PrudAI · Last updated 6 September 2026 · PrudentBench v1.0

1. The scorecard

Ten models, two benchmarks and one compliance gate. Only models that pass the gate receive a PrudentScore.

Passes the compliance gateRejected by the compliance gateIn use at LEO
Legal research: all-pass in percentVals AI Legal Research Bench, all-pass, measured 5 September 2026 [1]. Hatched bars are rejected by the compliance gate.

Results on the Vals AI Legal Research Bench, high to low: Claude Opus 5: 55.29%, Passes the compliance gate. Claude Fable 5.1: 55.29%, Rejected by the compliance gate. Muse Spark 1.3 (Max): 55.29%, Rejected by the compliance gate. GPT-5.6 Sol: 48.08%, Passes the compliance gate. Grok 4.6: 48.08%, Rejected by the compliance gate. Kimi K3: 44.23%, Passes the compliance gate. Claude Sonnet 5: 41.83%, Passes the compliance gate. GPT-6 Astra: 39.42%, Rejected by the compliance gate. Gemini 3.8 Flash: 38.94%, Passes the compliance gate. Mistral Medium 3.5: 9.14%, Passes the compliance gate.

Quality against cost per taskCost per completed task in dollars, logarithmic, against all-pass on legal research [1]. Top left is cheap and good.

Cost per completed task against the legal-research score: Claude Opus 5: $6.76, 55.29%, Passes the compliance gate. Kimi K3: $3.22, 44.23%, Passes the compliance gate. Gemini 3.8 Flash: $1.63, 38.94%, Passes the compliance gate. Claude Sonnet 5: $4.08, 41.83%, Passes the compliance gate. GPT-5.6 Sol: $21.61, 48.08%, Passes the compliance gate. Mistral Medium 3.5: $1.31, 9.14%, Passes the compliance gate. Claude Fable 5.1: $23.06, 55.29%, Rejected by the compliance gate. Muse Spark 1.3 (Max): $0.59, 55.29%, Rejected by the compliance gate. Grok 4.6: $1.53, 48.08%, Rejected by the compliance gate. GPT-6 Astra: $10.58, 39.42%, Rejected by the compliance gate.

Passes the compliance gateRejected by the compliance gate

PrudentScore v1.0

Only models that pass the gate get a score, on a scale from 0 to 100.

  1. 1
  2. 2
    Kimi K370points
  3. 3
  4. 4
  5. 5
    GPT-5.6 Sol59points
  6. 6
All figures as a table
PrudentBench v1.0 scorecard: ten AI models on legal research, legal agent tasks, cost per task and the compliance gate.
ModelProvider and routeLegal researchLegal agent tasksCost per taskStatus at PrudAICompliance gatePrudentScore
Claude Opus 5Anthropic via Google Cloud Vertex AI (EU regions)55.29%6.67%$6.76In use · deepPasses75
Kimi K3Moonshot AI (open weights), via a European open-weights provider44.23%10.83%$3.22Tested, not in usePasses70
Gemini 3.8 FlashGoogle via Vertex AI (EU regions)38.94%10.00%$1.63In use · fastPasses68
Claude Sonnet 5Anthropic via Google Cloud Vertex AI (EU regions)41.83%5.00%$4.08In use · fallbackPasses62
GPT-5.6 SolOpenAI via Microsoft Azure OpenAI (Sweden Central)48.08%2.50%$21.61Azure route, not in chainPasses59
Mistral Medium 3.5Mistral AI (Paris), also via the Azure EU Data Zone9.14%0.42%$1.31Not in usePasses28
Claude Fable 5.1A mandatory 30-day retention period without zero data retention, and on Vertex AI data sharing with the model maker is required.Anthropic, via Google Cloud55.29%6.67%$23.06Not deployableRejectedno score
Muse Spark 1.3 (Max)Available only through Meta's own API; no route via a hyperscaler and no published guarantee of EU processing or an EU contracting party. The cheap contributor tier trains on customer traffic.Meta55.29%without Max: 40.8723.75%$0.59Not deployableRejectedno score
Grok 4.6On Azure available only as Global Standard, without an EU Data Zone [22][23].xAI via Microsoft Azure48.08%15.83%$1.53Tested, not deployableRejectedno score
GPT-6 AstraOn Azure only Global Standard and a US Data Zone; no EU Data Zone [23].OpenAI via Microsoft Azure39.42%5.42%$10.58Not deployableRejectedno score

Legal research: Vals AI Legal Research Bench, all-pass in percent [1]. Legal agent tasks: Harvey LAB in the Vals implementation, all-pass in percent [2]. Cost per task: Vals AI, in US dollars [1]. Where a benchmark publishes an effort setting, it is part of the model name: Muse Spark 1.3 (Max) is a different setting from plain Muse Spark 1.3, which reaches 40.87 percent.

Benchmark figures: the Vals AI Legal Research Bench [1] and the Vals implementation of Harvey LAB [2], both measured 5 September 2026, accessed 6 September 2026. The PrudentScore is a composition by PrudAI B.V. and is not a figure produced by the benchmark operators.

2. Why yes, why no: model by model

What each model does well, where it breaks down and what we consider it suitable for. The separate line about the Intelligence Index carries its source, because it comes from a different measurement than the table above.

Claude Opus 5

Anthropic via Google Cloud Vertex AI (EU regions)

In use · deepPassesPrudentScore 75
  • Tied highest legal-research score: 55.29 percent all-pass [1].
  • The most robust model in our screening over long, source-heavy conversations.
  • The most expensive model we deploy, measured per token.
  • Low on legal agent tasks: 6.67 percent all-pass [2].

Suitable for: Deep case analysis, case-law research and memos in which many sources come together.

Artificial Analysis Intelligence Index v4.2: 54.1 at max effort. Source: Artificial Analysis (artificialanalysis.ai), accessed 6 September 2026 [37]. artificialanalysis.ai

More about this model

Strong

  • Joint highest result on the Vals AI Legal Research Bench: 55.29 percent all-pass, tied with Claude Fable 5.1 and Muse Spark 1.3 (Max), measured 5 September 2026 [1].
  • In our own screening the most dependable model in long conversations that call many sources.
  • Discounted repeated context keeps case-wide work affordable per unit of work.

Weak

  • The most expensive model we deploy, measured per token.
  • A deep analysis takes noticeably longer than a quick question.
  • It lags on legal agent tasks: 6.67 percent all-pass on the Vals implementation of Harvey LAB [2].

Kimi K3

Moonshot AI (open weights), via a European open-weights provider

Tested, not in usePassesPrudentScore 70
  • Open weights, so it runs on European infrastructure without the model maker in the path.
  • 44.23 percent all-pass, above the models we use for quick questions [1].
  • Too slow and too variable for interactive use in our production test.
  • No discount on repeated context, so more expensive per unit of work.

Suitable for: Work that has to run entirely on our own or on European infrastructure, not for the interactive layers.

Artificial Analysis Intelligence Index v4.2: 50.2 at max effort. Source: Artificial Analysis (artificialanalysis.ai), accessed 6 September 2026 [37]. artificialanalysis.ai

More about this model

Strong

  • Open weights, so it can run on European infrastructure without the model maker in the data path.
  • 44.23 percent all-pass on the Vals AI Legal Research Bench, above the models we currently use for quick questions [1].
  • Low cost per task in the Vals measurement: 3.22 dollars, at 10.83 percent all-pass on legal agent tasks [1][2].

Weak

  • In our August 2026 production test too slow and too erratic for interactive use.
  • The provider offered no discount on repeated context, which made it more expensive per unit of work than Claude Opus 5 despite a lower list price.
  • Out of our chain since August 2026.

Gemini 3.8 Flash

Google via Vertex AI (EU regions)

In use · fastPassesPrudentScore 68
  • Cheapest per unit of work of our models: 1.63 dollars per task [1].
  • Strongest of our models on legal agent tasks: 10.00 percent all-pass [2][20].
  • Lowest legal-research score of our models: 38.94 percent [1].
  • Calls fewer sources than Claude Sonnet 5 on the same question.

Suitable for: Short questions, summaries and first explorations where the answer is checked afterwards.

Artificial Analysis Intelligence Index v4.2: 47.1 at high effort. Source: Artificial Analysis (artificialanalysis.ai), accessed 6 September 2026 [37]. artificialanalysis.ai

More about this model

Strong

  • The cheapest per unit of work among the models we deploy: 1.63 dollars per task in the Vals measurement [1].
  • In our own screening no demonstrable quality difference with Claude Sonnet 5 on quick questions, which lets price decide.
  • The strongest of the models we deploy on legal agent tasks: 10.00 percent all-pass on the Vals implementation of Harvey LAB [2][20].

Weak

  • The lowest legal-research score among the models we deploy: 38.94 percent all-pass [1].
  • It waits noticeably longer before the first character once sources are called.
  • It calls fewer sources than Claude Sonnet 5 on the same question.

Claude Sonnet 5

Anthropic via Google Cloud Vertex AI (EU regions)

In use · fallbackPassesPrudentScore 62
  • Starts answering almost immediately, which suits follow-up questions.
  • 41.83 percent all-pass on legal research, above Gemini 3.8 Flash [1].
  • Pricier per task than Gemini 3.8 Flash: 4.08 against 1.63 dollars [1].
  • Low on legal agent tasks: 5.00 percent all-pass [2].

Suitable for: Fallback behind Gemini 3.8 Flash and behind Claude Opus 5, and work where speed and source calls have to go together.

Artificial Analysis Intelligence Index v4.2: 45.1 at max effort. Source: Artificial Analysis (artificialanalysis.ai), accessed 6 September 2026 [37]. artificialanalysis.ai

More about this model

Strong

  • It starts answering almost immediately, which suits conversations with follow-up questions.
  • It works well with source calls and keeps track across several steps.
  • 41.83 percent all-pass on the Vals AI Legal Research Bench, above Gemini 3.8 Flash [1].

Weak

  • Clearly more expensive per completed task than Gemini 3.8 Flash: 4.08 against 1.63 dollars in the Vals measurement [1].
  • Clearly lower on legal research than Claude Opus 5.
  • Low on legal agent tasks: 5.00 percent all-pass on the Vals implementation of Harvey LAB, half of Gemini 3.8 Flash [2].

GPT-5.6 Sol

OpenAI via Microsoft Azure OpenAI (Sweden Central)

Azure route, not in chainPassesPrudentScore 59
  • 48.08 percent all-pass, the highest result outside the Claude family [1].
  • Azure EU Data Zone with Microsoft Ireland as contracting party [23].
  • 21.61 dollars per task, the priciest after Claude Fable 5.1 [1].
  • Lowest of all assessed models on legal agent tasks: 2.50 percent [2].

Suitable for: An independent jury model in our own screenings, as a third model family alongside Claude and Gemini. It does not sit in the chat chain of LEO: the fallback there is GPT-5.4 on the same Azure route.

Artificial Analysis Intelligence Index v4.2: 51.3 at max effort. Source: Artificial Analysis (artificialanalysis.ai), accessed 6 September 2026 [37]. artificialanalysis.ai

More about this model

Strong

  • 48.08 percent all-pass on the Vals AI Legal Research Bench, the best result outside the Claude family [1].
  • Processing in the Azure EU Data Zone with Microsoft Ireland as the contracting party: gpt-5.6-sol is listed by name in Microsoft's Europe table [23].
  • A third model family alongside Claude and Gemini, which makes it usable as an independent jury model.

Weak

  • 21.61 dollars per task in the Vals measurement, after Claude Fable 5.1 the most expensive model in this list [1].
  • A long wait before the first visible character.
  • The lowest of all assessed models on legal agent tasks: 2.50 percent all-pass on the Vals implementation of Harvey LAB [2].

Mistral Medium 3.5

Mistral AI (Paris), also via the Azure EU Data Zone

Not in usePassesPrudentScore 28
  • The only European model provider in this list, with a European contracting party.
  • Open weights, so suitable for fully self-hosted deployment.
  • 9.14 percent all-pass on legal research, the lowest in this list [1].
  • No measured Intelligence Index, so three of the four components.

Suitable for: A candidate for fully self-hosted deployment and for narrow tasks, not for the chat layers of LEO.

Artificial Analysis Intelligence Index v4.2: estimated by Artificial Analysis at 22.7, not a measurement of its own. Source: Artificial Analysis (artificialanalysis.ai), accessed 6 September 2026 [37]. artificialanalysis.ai

More about this model

Strong

  • The only European model provider in this list, with a European contracting party.
  • Open weights, so it is suitable for fully self-hosted deployment.
  • Deployable through the Azure EU Data Zone under the processing agreements we already hold; mistral-medium-3-5 is listed by name in Microsoft's Europe table [23], and Mistral offers a European endpoint of its own [32][33].

Weak

  • 9.14 percent all-pass on the Vals AI Legal Research Bench, by far the lowest result in this list [1].
  • 0.42 percent all-pass on legal agent tasks, likewise the lowest of the ten [2].
  • Artificial Analysis publishes no measurement of its own for this model on the Intelligence Index, only an estimate; it therefore does not count, which leaves the PrudentScore here resting on three of the four components.

Claude Fable 5.1

Anthropic, via Google Cloud

Not deployableRejected

Gate closed: A mandatory 30-day retention period without zero data retention, and on Vertex AI data sharing with the model maker is required.

  • Tied highest legal-research score: 55.29 percent all-pass [1].
  • Same model family as Claude Opus 5, so prompts carry over.
  • 23.06 dollars per task, more than three times Claude Opus 5 [1].
  • A mandatory 30-day retention window, no zero data retention [13][16].

Suitable for: Not deployable for us: the compliance gate closes on retention and on the required data sharing.

Artificial Analysis Intelligence Index v4.2: 56.8 at max effort with fallback, the highest of these ten. Source: Artificial Analysis (artificialanalysis.ai), accessed 6 September 2026 [37]. artificialanalysis.ai

More about this model

Strong

  • Joint highest result on the Vals AI Legal Research Bench: 55.29 percent all-pass, level with Claude Opus 5 and Muse Spark 1.3 (Max) [1].
  • The same result on legal agent tasks as Claude Opus 5: 6.67 percent all-pass [2].
  • The same model family as Claude Opus 5, so prompts and working methods transfer.

Weak

  • 23.06 dollars per task in the Vals measurement, more than three times Claude Opus 5 at the same score [1].
  • The provider has designated it a Covered Model since 31 August 2026, with a mandatory 30-day retention period and without zero data retention [13][14][16].
  • On Google Cloud the Fable line requires data sharing with the model maker to be switched on for abuse monitoring [17].

Muse Spark 1.3 (Max)

Meta

Not deployableRejected

Gate closed: Available only through Meta's own API; no route via a hyperscaler and no published guarantee of EU processing or an EU contracting party. The cheap contributor tier trains on customer traffic.

  • Tied highest score at the Max setting: 55.29 percent all-pass [1].
  • Highest result on legal agent tasks in this list: 23.75 percent [2].
  • 55.29 percent holds only for the Max setting, still in safety testing [1][28].
  • Only through Meta’s own API, with no published EU processing guarantee [29].

Suitable for: Not deployable for us while there is no route via a hyperscaler with a European contracting party, and while the top score depends on a preview effort setting.

Artificial Analysis Intelligence Index v4.2: 53.0 at max effort. Source: Artificial Analysis (artificialanalysis.ai), accessed 6 September 2026 [37]. artificialanalysis.ai

More about this model

Strong

  • Joint highest result on the Vals AI Legal Research Bench: 55.29 percent all-pass at Max effort, level with Claude Opus 5 and Claude Fable 5.1 [1].
  • The best result on legal agent tasks in this list: 23.75 percent all-pass on the Vals implementation of Harvey LAB [2].
  • By far the cheapest per completed task in the Vals measurement: 0.59 dollars [1].

Weak

  • The 55.29 percent applies only to the Max effort setting, which the provider's launch post still described as being in safety testing; the version anyone can buy scores 40.87 percent [1][28].
  • Available only through Meta's own API, not via Azure, Google Cloud or Bedrock, and without a published guarantee of EU processing or a European contracting party [29].
  • The cheapest price tier, the contributor tier, is used to improve the provider's products and therefore trains on customer traffic [29].

Grok 4.6

xAI via Microsoft Azure

Tested, not deployableRejected

Gate closed: On Azure available only as Global Standard, without an EU Data Zone [22][23].

  • 48.08 percent all-pass on legal research, level with GPT-5.6 Sol [1].
  • Strong on legal agent tasks: 15.83 percent all-pass [2][26].
  • On Azure only Global Standard, without an EU Data Zone [22][23].
  • Almost a minute of silence before the first token in our test.

Suitable for: Not deployable for us while there is no EU Data Zone. We will look again once there is one.

Artificial Analysis Intelligence Index v4.2: 50.6 at high effort. Source: Artificial Analysis (artificialanalysis.ai), accessed 6 September 2026 [37]. artificialanalysis.ai

More about this model

Strong

  • 48.08 percent all-pass on the Vals AI Legal Research Bench, level with GPT-5.6 Sol [1].
  • Through Azure, Microsoft is the processor and the contracting party, not the model maker: prompts and completions are not available to the providers of models sold by Azure [21].
  • Strong on legal agent tasks: 15.83 percent all-pass on the Vals implementation of Harvey LAB [2][26].

Weak

  • On Azure available only as Global Standard, so without a guaranteed EU region; grok-4.6 is absent from Microsoft's Europe table [22][23].
  • In our own test almost a minute of silence before the first character, and no discount on repeated context.

GPT-6 Astra

OpenAI via Microsoft Azure

Not deployableRejected

Gate closed: On Azure only Global Standard and a US Data Zone; no EU Data Zone [23].

  • Strong at agentic work and at programming tasks.
  • High general intelligence, above the models we use today.
  • Low on legal research: 39.42 percent all-pass [1].
  • No EU Data Zone on Azure, only Global Standard and US [23].

Suitable for: Not deployable for us. We will reopen the assessment once there is an EU Data Zone and the legal-research score comes up.

Artificial Analysis Intelligence Index v4.2: 54.7 at max effort. Source: Artificial Analysis (artificialanalysis.ai), accessed 6 September 2026 [37]. artificialanalysis.ai

More about this model

Strong

  • Strong on agentic work and on programming tasks.
  • High general intelligence, above the models we currently deploy.
  • Once Azure ships an EU Data Zone, the route to this model is already in place.

Weak

  • Low on legal research: 39.42 percent all-pass on the Vals AI Legal Research Bench [1].
  • On Azure only Global Standard and a US Data Zone are available, no EU Data Zone, which keeps the compliance gate closed [23].

3. Compliance matrix

The gate only opens when all three requirements are covered: processing in the EU, a contracting party in the EU or EEA, and a contractual exclusion of training on customer data. Where a provider publishes no figure, the cell reads "see provider" and carries no estimate.

Compliance matrix: processor, contracting party, EU data processing, training on customer data and retention per model.
ModelEU processingEU contracting partyNo trainingGate
Claude Opus 5yesyesyesPasses
Kimi K3yesyesyesPasses
Gemini 3.8 FlashyesyesyesPasses
Claude Sonnet 5yesyesyesPasses
GPT-5.6 SolyesyesyesPasses
Mistral Medium 3.5yesyesyesPasses
Claude Fable 5.1noyes (Google Cloud)yesRejected
Muse Spark 1.3 (Max)nonopartlyRejected
Grok 4.6noyesyesRejected
GPT-6 AstranoyesyesRejected

Why the gate closes

  • Claude Fable 5.1 — EU processing: On Google Cloud only with mandatory data sharing with the model maker for abuse monitoring and a mandatory 30-day retention window at the vendor [15][17]
  • Muse Spark 1.3 (Max) — EU processing: No published guarantee of EU processing [29][30]
  • Muse Spark 1.3 (Max) — EU contracting party: Meta, no published EU/EEA entity
  • Muse Spark 1.3 (Max) — No training: Yes on the contributor tier; no on the standard tier [29]
  • Grok 4.6 — EU processing: No, Global Standard only, no EU Data Zone [22][23]
  • GPT-6 Astra — EU processing: No, Global Standard and US Data Zone only [23]
Processor, contracting party and retention as a table
Compliance matrix: processor, contracting party, EU data processing, training on customer data and retention per model.
ModelProcessorContracting partyEU data processingTraining on customer dataRetention
Claude Opus 5Google CloudGoogle Cloud EMEA LimitedVertex AI, Europe multi-region [18][19]No, contractually excluded [38]0 (zero data retention) [15][16]
Kimi K3European open-weights providerEstablished in the EU/EEA [39]Open weights, hosted in EU data centres [39]No, contractually excluded [39]See provider
Gemini 3.8 FlashGoogle CloudGoogle Cloud EMEA LimitedVertex AI, Europe multi-region [18][19]No, contractually excluded [38]0 (zero data retention) [15][16]
Claude Sonnet 5Google CloudGoogle Cloud EMEA LimitedVertex AI, Europe multi-region [18][19]No, contractually excluded [38]0 (zero data retention) [15][16]
GPT-5.6 SolMicrosoftMicrosoft Ireland Operations LimitedAzure OpenAI, Sweden Central (EU Data Zone) [23]No, contractually excluded [21]Contractual zero data retention; abuse monitoring with reviewers inside the EEA [21]
Mistral Medium 3.5Mistral AI, or Microsoft when purchased through AzureMistral AI SAS (Paris), or Microsoft Ireland Operations LimitedEuropean endpoint, or Azure EU Data Zone [23][32][33]No, contractually excluded [21]See provider
Claude Fable 5.1Google CloudGoogle Cloud EMEA LimitedOn Google Cloud only with mandatory data sharing with the model maker for abuse monitoring and a mandatory 30-day retention window at the vendor [15][17]No, contractually excluded [38]30 days mandatory, no zero data retention [13][16]
Muse Spark 1.3 (Max)MetaMeta, no published EU/EEA entityNo published guarantee of EU processing [29][30]Yes on the contributor tier; no on the standard tier [29]See provider
Grok 4.6Microsoft (Models sold by Azure)Microsoft Ireland Operations LimitedNo, Global Standard only, no EU Data Zone [22][23]No, contractually excluded [21]Through Azure the Microsoft processing agreement applies [21]
GPT-6 AstraMicrosoftMicrosoft Ireland Operations LimitedNo, Global Standard and US Data Zone only [23]No, contractually excluded [21]Abuse monitoring with reviewers inside the EEA [21]

This matrix describes the route we would use, not every conceivable route to the same model. We never claim a provider by model family: Microsoft lists gpt-5.6-sol and mistral-medium-3-5 by name in its Europe table and grok-4.6 not at all [23]. Our own chain, subprocessors and processing agreement are documented in the trust center.

4. Which model LEO uses, and why

LEO runs deep analyses on Claude Opus 5 and quick questions on Gemini 3.8 Flash, with Claude Sonnet 5 and GPT-5.4 through Microsoft Azure OpenAI (Sweden Central) as fallback, all through EU endpoints. That is the production state as of 6 September 2026.

Question in LEO

Quick question

Gemini 3.8 Flash

Short questions and first explorations, through the same route and the same contracting party.

Deep analysis

Claude Opus 5

Case analysis, case-law research and memos, through Google Cloud Vertex AI in EU regions.

Fallback

Claude Sonnet 5 and GPT-5.4 through Microsoft Azure OpenAI (Sweden Central)

Claude Sonnet 5 through Vertex AI in EU regions and GPT-5.4 through Microsoft Azure OpenAI in Sweden Central.

Every route runs through EU endpoints, with a European contracting party, a contractual exclusion of training on customer data and contractually agreed zero data retention.
Cost per completed taskVals AI, in US dollars per completed task [1]. GPT-5.4, the fallback through Microsoft Azure OpenAI, is not part of the Vals measurement and therefore not in this chart.

Cost per completed task for the models in LEO’s chain: Gemini 3.8 Flash: $1.63. Claude Sonnet 5: $4.08. Claude Opus 5: $6.76.

The four building blocks behind that choice

1

Quality per task

The model that can handle the task at hand, not one model for everything.

2

Cost relative, and only after cache

Cost per unit of work, only after the discount on repeated context.

3

EU processing as a hard requirement

EU endpoints, a European contracting party, no training on customer data.

4

Reliability is architecture

Citations are checked word for word against the retrieved source text.

Notes on the four building blocks
Quality per task
Not one model for everything, but the model that can handle the task at hand. A summary asks something different from reconstructing conflicting authority.
Cost relative, and only after cache
A lower list price means nothing while the provider gives no discount on repeated context. We compare cost per unit of work, not per token.
EU processing as a hard requirement
No model enters the chain without processing through endpoints in EU regions, a European contracting party and a contractual exclusion of training on customer data.
Reliability is architecture
Citations are checked word for word against the retrieved source text. That check is independent of the model, so swapping models changes nothing about it.

5. Methodology and PrudentScore

The gate

A model is rejected, whatever its score, if any of these three is not covered:

  • no contractual processing in the EU
  • no contracting party in the EU or EEA
  • no contractual exclusion of training on customer data

PrudentScore v1.0

Only models that pass the gate get a score. The score adds four components and is rounded to whole points:

How the four components are weighted

  • 50%Legal research
  • 15%Legal agent tasks
  • 15%General intelligence
  • 20%Cost efficiency
The arithmetic and the results per model
  1. 50 percent legal research: Vals AI Legal Research Bench, all-pass, divided by the highest score of 55.29 [1].
  2. 15 percent legal agent tasks: Harvey LAB in the Vals implementation, all-pass, divided by the highest score of 25.42, the highest value on the Vals leaderboard (Muse Spark 1.2, which is not in this list) [2].
  3. 15 percent general intelligence: Artificial Analysis Intelligence Index v4.2, divided by the highest score of 56.8, which is Claude Fable 5.1 at 56.76, rounded [6][37].
  4. 20 percent cost efficiency: 1 minus a logarithmic scale of cost per task, with 0.59 dollars as the floor and 23.06 dollars as the ceiling [1].

When a component is missing, the remaining weights are rescaled across what is left. The results below are computed by the page from this formula and rounded to whole points; they are stored nowhere as a number.

Results for v1.0, computed from this formula

  • Claude Opus 54 of 4 components75
  • Kimi K34 of 4 components70
  • Gemini 3.8 Flash4 of 4 components68
  • Claude Sonnet 54 of 4 components62
  • GPT-5.6 Sol4 of 4 components59
  • Mistral Medium 3.53 of 4 components28

What we do ourselves

Our own assessment runs in four steps. First the compliance filter: a model rejected there is not measured any further. Then a signal wait: we wait until the first published results are stable rather than switching on release day. Next a screening on our production settings, so with our prompts, our source calls and our context length, not on the benchmark configuration. Only then a full benchmark.

Two things weigh heavily in that. Compliance is a gate, not a weighting factor: a model that fails the gate gets no second chance on quality. And the cache economy decides: a provider without a discount on repeated context is more expensive on case-wide work than the list price suggests.

PrudAI is itself the supplier of LEO. That is why we cite figures only from parties that are not ours, publish the formula, and also show the models we do not use, with the reason.

What we do not measure

An independent, published benchmark for Dutch legal work does not exist yet. Our own Dutch measurement set (100 rulings × 10 questions) is being built; as soon as it passes our quality gate, we will publish the results here.

That gap can be checked against the literature. The only serious Dutch-language legal question-answering benchmark, bLLeQA, covers Belgian law, and its authors write that reliable Dutch resources are scarce [34]. The only benchmark that touches Dutch law, Multi-Legal-Bench, does metadata classification and outcome prediction on rulings, not legal research [35]. Dutch is present in the EU-wide knowledge benchmark EU-MMLU, but its only legal component is international law [36].

That set started as a screening on 20 Dutch cases with a blind jury and a check on case references. Until the extension is finished we publish no figures from it.

Harvey LAB exists in three implementations, by Harvey itself [4][5], by Vals [2] and by Artificial Analysis [8], with outcomes that can differ by a factor of three. We consistently cite the Vals implementation on all-pass, because there every model runs on the same harness, and mention Artificial Analysis only in running text with attribution.

The PrudentScore is a composition by PrudAI B.V. based on publicly published figures and does not represent the view or endorsement of Vals AI, Harvey or Artificial Analysis. Figures from Artificial Analysis appear in running text only, with attribution and a link, as that party's terms require [9].

6. What changed

Changes to this page and to the models LEO uses.

  1. 6 September 2026

    PrudentBench v1.0, first publication, based on the September model review.

  2. 3 September 2026

    LEO's fast layer moved to Gemini 3.8 Flash.

  3. August 2026

    Kimi K3 removed from the chain, on cost per unit of work and on speed.

  4. July 2026

    Claude Opus 5 taken into use as the deep-analysis model, through Vertex AI in EU regions.

  5. July 2026

    Claude Sonnet 5 taken into use as the fast layer.

7. Frequently asked questions

The ten questions we get most often about model choice and legal work.

What is the best AI model for legal advice?

On the only legal-research benchmark graded by lawyers, the Vals AI Legal Research Bench, Claude Opus 5, Claude Fable 5.1 and Muse Spark 1.3 (Max) share the lead at 55.29 percent all-pass, measured 5 September 2026 [1]. For a firm the compliance gate counts as well: processing in the EU, a European contracting party and no training on customer data. Of those three, only Claude Opus 5 passes it. An independent, published benchmark for Dutch legal work does not exist yet.

Is ChatGPT GDPR-proof for law firms?

That does not depend on the model but on the route and the contract. A consumer subscription offers no processing agreement and no choice of processing region. Through Microsoft Azure OpenAI in Sweden Central, Microsoft Ireland is the contracting party, a processing agreement is in place and there is a contractual exclusion of training on customer data. That is the route we use for GPT models. Always assess your own situation and your own processing agreement.

May I enter client data into an AI model?

Only when the processing is covered: a processing agreement with the provider, no training on the data entered, a documented processing region and clear agreements on retention. Professional privilege does not disappear because a model sits in between, the responsibility stays with the firm. Our agreements are in the processing agreement at legal.prudai.com and the chain is described in the trust center.

Claude or ChatGPT for legal research?

On the Vals AI Legal Research Bench, Claude Opus 5 reaches 55.29 percent all-pass against 48.08 percent for GPT-5.6 Sol, at a cost of 6.76 against 21.61 dollars per task, measured 5 September 2026 [1]. Both routes satisfy our compliance gate. LEO uses Claude Opus 5 for deep analyses and keeps GPT through Microsoft Azure available: GPT-5.4 as the fallback in the chain, GPT-5.6 Sol as an independent jury model in our own screenings.

Why does AI invent case law, and which model does it least?

A language model predicts text, and a plausible-looking case reference is as likely to the model as a real one. There is no public measurement of hallucination on Dutch case law, for any model [34][35], so we name no winner. The answer lies in architecture, not in model choice: LEO retrieves rulings at the source and checks every citation word for word against the retrieved text.

Which models process data in the EU?

Of the ten models assessed, six pass the gate: Claude Opus 5 and Claude Sonnet 5 through Google Cloud Vertex AI in EU regions, Gemini 3.8 Flash through the same route, GPT-5.6 Sol through Microsoft Azure OpenAI in Sweden Central, Mistral Medium 3.5 through a European endpoint and Kimi K3 through a European open-weights provider. Claude Fable 5.1, Grok 4.6, GPT-6 Astra and Muse Spark are rejected [17][22][23][29].

Does the provider train on my case files?

Not on the routes we use: Google Cloud and Microsoft contractually exclude training on customer data, and zero data retention is contractually agreed with our model providers. Watch the exceptions in the market. Claude Fable 5.1 carries a mandatory 30-day retention period without zero data retention [13][16], and the contributor tier of Muse Spark is used to improve the provider's products [29].

What does the AI Act mean for my firm?

A firm that buys and uses an AI system is normally a deployer, not a provider. The obligations that come first are AI literacy among staff, transparency towards clients and human oversight of the outcome. Our role split, documentation and subprocessors are set out in the trust center at trust.prudai.com.

How often is PrudentBench updated?

Every quarter, and in between whenever a major model release or a change in a processing route calls for it. Every figure on this page carries its source and its measurement date. Figures older than six months we treat as expired: we verify them at the source again before republishing them.

Does PrudAI measure itself, or only summarise?

Both, and we keep the two apart. The legal-research and agent-task scores come from Vals AI and Harvey LAB, benchmarks we do not run ourselves. In addition we run our own screening on Dutch cases, on our production settings rather than on the benchmark configuration. That screening is small and is being extended, and we publish no figures from it yet.

8. Sources

The bracketed numbers in the text and in the tables refer to this list. All external sources were accessed on 6 September 2026; the benchmark figures were measured on 5 September 2026.

  1. [1]Vals AI. Legal Research Bench. benchmark page, version 1, updated 5 September 2026.https://www.vals.ai/benchmarks/legal_researchAccessed 6 September 2026
  2. [2]Vals AI. Harvey's Legal Agent Benchmark (HLAB leaderboard). leaderboard, version 1, updated 5 September 2026.https://www.vals.ai/benchmarks/hlabAccessed 6 September 2026
  3. [3]Vals AI. Methodology. methodology page, undated.https://www.vals.ai/methodologyAccessed 6 September 2026
  4. [4]Grupen, N., Pereyra, G. & Pereyra, J. (Harvey AI). Open-Sourcing Harvey's Long Horizon Legal Agent Benchmark. Harvey AI, 6 May 2026.https://www.harvey.ai/blog/introducing-harveys-legal-agent-benchmarkAccessed 6 September 2026
  5. [5]Harvey AI. Harvey LAB: The Legal Agent Benchmark, v1.0. dataset and harness on GitHub, MIT licence.https://github.com/harveyai/harvey-labsAccessed 6 September 2026
  6. [6]Artificial Analysis. Announcing Artificial Analysis Intelligence Index v4.2. 4 September 2026.https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-2Accessed 6 September 2026
  7. [7]Artificial Analysis. Artificial Analysis Intelligence Benchmarking Methodology. methodology page, undated.https://artificialanalysis.ai/methodology/intelligence-benchmarkingAccessed 6 September 2026
  8. [8]Artificial Analysis. Announcing Harvey LAB-AA: evaluating AI agents on real-world legal work. 7 July 2026.https://artificialanalysis.ai/articles/harvey-lab-aaAccessed 6 September 2026
  9. [9]Artificial Analysis, Inc.. Artificial Analysis Data Platform Terms and Conditions, version 1.1. revised 19 August 2026.https://artificialanalysiscdn.com/legal/ProDataPlatformTerms.pdfAccessed 6 September 2026
  10. [10]Anthropic. System Card: Claude Opus 5. 24 July 2026, 198 pages.https://www.anthropic.com/claude-opus-5-system-cardAccessed 6 September 2026
  11. [11]Anthropic. System Card: Claude Fable 5.1 & Claude Mythos 5.1. 1 September 2026, 212 pages.https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system-cardAccessed 6 September 2026
  12. [12]Anthropic. System Card: Claude Sonnet 5. 30 June 2026, 145 pages.https://www.anthropic.com/claude-sonnet-5-system-cardAccessed 6 September 2026
  13. [13]Anthropic. Data retention practices for Covered Models. Anthropic Privacy Center, page carries a relative date only.https://privacy.claude.com/en/articles/15425996-data-retention-practices-for-covered-modelsAccessed 6 September 2026
  14. [14]Anthropic. Covered Models. Anthropic Help Center, page carries a relative date only.https://support.claude.com/en/articles/15425695-covered-modelsAccessed 6 September 2026
  15. [15]Anthropic. Data residency. Claude Platform Docs, undated.https://platform.claude.com/docs/en/manage-claude/data-residencyAccessed 6 September 2026
  16. [16]Anthropic. API and data retention. Claude Platform Docs, undated.https://platform.claude.com/docs/en/manage-claude/api-and-data-retentionAccessed 6 September 2026
  17. [17]Google Cloud. Abuse monitoring. Gemini Enterprise Agent Platform, updated 2 September 2026.https://docs.cloud.google.com/vertex-ai/generative-ai/docs/learn/abuse-monitoringAccessed 6 September 2026
  18. [18]Google Cloud. Claude Opus 5 on Google Cloud. updated 2 September 2026.https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/partner-models/claude/opus-5Accessed 6 September 2026
  19. [19]Google Cloud. Claude Sonnet 5 on Google Cloud. updated 2 September 2026.https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/partner-models/claude/sonnet-5Accessed 6 September 2026
  20. [20]Google DeepMind. Gemini 3.8 Flash — Model evaluation: Approach, methodology & results. September 2026, 4 pages (PDF).https://storage.googleapis.com/deepmind-media/gemini/gemini_3-8_flash_model_evaluation.pdfAccessed 6 September 2026
  21. [21]Microsoft. Data, privacy, and security for Foundry Models sold by Azure in Microsoft Foundry. Microsoft Learn, 18 May 2026.https://learn.microsoft.com/en-us/azure/foundry/responsible-ai/openai/data-privacyAccessed 6 September 2026
  22. [22]Microsoft. Deploy and use Grok models in Microsoft Foundry. Microsoft Learn, 26 August 2026.https://learn.microsoft.com/en-us/azure/foundry/foundry-models/how-to/use-foundry-models-grokAccessed 6 September 2026
  23. [23]Microsoft. Region availability for Foundry Models sold by Azure. Microsoft Learn, 3 September 2026.https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure-region-availabilityAccessed 6 September 2026
  24. [24]OpenAI. GPT-5.6 System Card. 9 July 2026, revised 19 August 2026, 82 pages (PDF).https://deploymentsafety.openai.com/gpt-5-6/gpt-5-6.pdfAccessed 6 September 2026
  25. [25]OpenAI. GPT-6 Astra System Card. 3 September 2026, 117 pages (PDF).https://deploymentsafety.openai.com/gpt-6-astra/gpt-6-astra.pdfAccessed 6 September 2026
  26. [26]SpaceXAI (DBA xAI LLC). Model Card: Grok 4.6. 12 August 2026, revised 17 August 2026, 42 pages (PDF).https://media.x.ai/v1/website/card-4p6-4cd2dc57.pdfAccessed 6 September 2026
  27. [27]Kimi Team (Moonshot AI). Kimi K3: Open Frontier Intelligence — Technical Report of Kimi K3. arXiv:2607.24653v2 [cs.CL], 7 August 2026, 47 pages.https://arxiv.org/abs/2607.24653Accessed 6 September 2026
  28. [28]Meta AI Research. Introducing Muse Spark 1.3. 2 September 2026.https://research.meta.ai/blog/introducing-muse-spark-1-3Accessed 6 September 2026
  29. [29]Meta. Muse Spark 1.3 (model page). undated.https://developer.meta.com/ai/models/muse-spark/Accessed 6 September 2026
  30. [30]Meta. Meet Muse Spark 1.2 and Muse Code: a coding model and the agent built to run it. 5 August 2026.https://developer.meta.com/ai/resources/blog/build-with-muse-code/Accessed 6 September 2026
  31. [31]Mistral AI. Mistral Medium 3.5 (model card, version v26.04). Modified MIT licence.https://docs.mistral.ai/models/model-cards/mistral-medium-3-5-26-04Accessed 6 September 2026
  32. [32]Mistral AI. In-region inference, open models, and new European infrastructure for sovereign AI. 11 August 2026.https://mistral.ai/news/regional-inference-open-models-new-compute/Accessed 6 September 2026
  33. [33]Mistral AI. Regional inference. Mistral Docs, undated.https://docs.mistral.ai/inference/regional-inferenceAccessed 6 September 2026
  34. [34]Banar, N., Lotfi, E., Van Nooten, J., Kliocaite, M. & Daelemans, W.. bLLeQA: Benchmarking LLMs for Grounded Legal Question-Answering in French and Dutch. Proceedings of the 4th Workshop on Towards Knowledgeable Foundation Models (KnowFM 2026), ACL, July 2026, pp. 34–59, DOI 10.18653/v1/2026.knowfm-1.4.https://aclanthology.org/2026.knowfm-1.4/Accessed 6 September 2026
  35. [35]Ovcharov, V.. Multi-Legal-Bench: Evaluating LLMs on Legal Reasoning Across Jurisdictions, Languages, and Legal Traditions. arXiv:2605.29738v2, 7 August 2026.https://arxiv.org/abs/2605.29738Accessed 6 September 2026
  36. [36]European Commission, DG Translation and the EMT network. EU-MMLU (dataset). 17,203 rows, 16 language configurations including NL_NL, CC BY 4.0 licence.https://huggingface.co/datasets/EC-DGT-AI/EU-MMLUAccessed 6 September 2026
  37. [37]Artificial Analysis. Model pages (Intelligence Index v4.2 per model). continuously updated model pages, values read on 6 September 2026.https://artificialanalysis.ai/modelsAccessed 6 September 2026
  38. [38]Google Cloud. Service Specific Terms. Gemini Enterprise Agent Platform / generative AI services: Google does not use customer data to train or fine-tune models without permission.https://cloud.google.com/terms/service-termsAccessed 6 September 2026
  39. [39]PrudAI B.V.. Privacy statement. PrudAI Legal Center, chapter on subprocessors and processing locations.https://legal.prudai.com/privacyAccessed 6 September 2026

Full source list

Claude and Anthropic, GPT and OpenAI, Gemini and Google, Grok and xAI, Kimi and Moonshot AI, Muse and Meta, Mistral, Harvey, Vals AI and Artificial Analysis are trademarks of their respective owners. They are named for identification only and this implies no partnership, sponsorship or endorsement.

The logos shown are the property of their respective owners and appear here purely to identify the provider or the route. The icon set comes from Simple Icons (CC0).

Beau Jonkhout, Chief Technology Officer at PrudAI · Last updated 6 September 2026 · PrudentBench v1.0