Claude Opus 5
Anthropic via Google Cloud Vertex AI (EU regions)
- Tied highest legal-research score: 55.29 percent all-pass [1].
- The most robust model in our screening over long, source-heavy conversations.
- The most expensive model we deploy, measured per token.
- Low on legal agent tasks: 6.67 percent all-pass [2].
Suitable for: Deep case analysis, case-law research and memos in which many sources come together.
Artificial Analysis Intelligence Index v4.2: 54.1 at max effort. Source: Artificial Analysis (artificialanalysis.ai), accessed 6 September 2026 [37]. artificialanalysis.ai
More about this model
Strong
- Joint highest result on the Vals AI Legal Research Bench: 55.29 percent all-pass, tied with Claude Fable 5.1 and Muse Spark 1.3 (Max), measured 5 September 2026 [1].
- In our own screening the most dependable model in long conversations that call many sources.
- Discounted repeated context keeps case-wide work affordable per unit of work.
Weak
- The most expensive model we deploy, measured per token.
- A deep analysis takes noticeably longer than a quick question.
- It lags on legal agent tasks: 6.67 percent all-pass on the Vals implementation of Harvey LAB [2].