EQS Group just released its AI Benchmark Report Vol.2. Ten compliance & ethics domains. 120 real-world tasks. Blind-judged by human compliance professionals.
The result? Three frontier models clustered within a single percentage point:
Gemini 3.1 Pro · 87.4%
Claude Opus 4.6 · 86.1%
Eight months ago, in Vol.1, Anthropic's best model trailed Google's by ~5 points. Today that gap is essentially gone.
But the leaderboard isn't the story.
Three findings in this report matter far more for anyone running a finance, audit, or compliance function.
The biggest gains are in open-ended work — not structured tasks
The single most important data point in the report:
Open-ended tasks — drafting reports, policies, investigation plans, regulator responses — improved by +17 to +18 percentage points versus Vol.1.
EQS's own framing: outputs moved from "usable with heavy editing" to "usable with light review."
For those of us who spend our weeks drafting bank query responses, going-concern memos, transfer pricing files, and regulator submissions — this is the threshold that actually matters.
Not a slick demo video from last year. A specific, verifiable, human-jury-validated capability shift.
My own experience during the FY2025 audit cycle matches the data. A year ago I was wrestling with AI output. This year I'm using it for first drafts.
Model choice is being commoditized. The harness is now the differentiator.
EQS introduces a concept I find genuinely useful: The Harness.
If every major model performs at 86–88% on general compliance work, your AI program's success is no longer about which model you pick. It's about everything around it:
- The structure and quality of the policies, contracts, and records you feed in
- How you decompose workflows — sequence vs parallel
- Whether your team can spot the subtle errors AI introduces
- Your defenses against prompt injection and data tampering
Moritz Homann, EQS's Head of AI, said it plainly:
"Six months ago, the question was whether AI could support real compliance work. Today, the question is how we design workflows around it."
— Moritz Homann, Head of AI, EQS Group
Models are commoditizing — the way cloud computing went from "AWS or Azure?" to "what's your architecture?"
"Sequence View" beats "Human-in-the-Loop" as a mental model
The report walks through a 5-step Conflict-of-Interest workflow. Even GPT-5.4 scores 90%+ on each individual step.
But chain those steps end-to-end? 0.9⁵ ≈ 59%.
Matt Kelly at Radical Compliance offered a sharper framing than the tired "human-in-the-loop" cliché:
Human guard posts along the trail.
You don't passively insert humans into a loop. You actively decide where on the AI's path human judgment is non-negotiable.
For finance and compliance professionals, this reframe is good news.
"Will AI replace me?" becomes "Can I design where the guard posts go?"
The second question has a much better career trajectory.
Two examples from my own work
Case 1 — Hansen-style verification for LLM-derived variables
In my LPG market analysis work, I have LLMs extract variables from news, policy documents, and central bank working papers — import quotas, VLGC freight signals, CP-FEI spread movements.
The obvious question: how do I know the LLM isn't hallucinating?
My answer borrows from econometrics' Hansen J-test logic: have the LLM derive the same variable through two independent extraction paths, then check for consistency. Disagreement triggers human review.
That's a harness. It doesn't depend on any single model being smart enough. It depends on the workflow itself having self-checking built in.
Case 2 — An internal finance automation project
For an internal finance automation project, I maintain two documents alongside the code:
- A project context document — business background, naming conventions, data schemas, key accounting rules
- A pitfalls log — every mistake AI and I have made together: SQL patterns that blow up memory in our environment, account classifications AI keeps confusing, tax terms it conflates with adjacent concepts
Every new session, AI reads both files before doing anything else.
This is the literal implementation of "invest in the harness, not just the prompt."
The compounding has surprised me:
- Month 1 — writing the docs felt like overhead.
- Month 3 — AI output quality jumped noticeably.
- Month 6 — those documents became onboarding material for new team members.
The real asset turns out to be organized failure knowledge — not any individual brilliant session.
Three takeaways for finance and compliance professionals
- Stop debating which model is best. This is becoming the 2017 "which cloud vendor?" debate — important, but no longer the source of differentiation.
- Start building your own harness assets. Prompt libraries, pitfalls logs, workflow SOPs — these are the moats that survive model upgrades and platform shifts.
- Position yourself as a sequence designer, not an AI user. The first is scarce and promotable. The second will be table stakes within two years.
Group Controller. Head of Compliance. Head of Internal Audit.
In the 2026+ landscape, the core competitive edge for these roles may simply be the ability to look at an 87%-accurate agent pipeline and point to exactly three spots where a human must sit.
What are the human guard posts in your own workflows? Which steps will you absolutely not delegate to AI?
I'd love to see how different industries and compliance contexts are designing this.
Connect & reply on LinkedIn →- EQS AI Benchmark Report Vol.2 · eqs.com/compliance-wpapers/eqs-ai-benchmark-report-vol-2
- Radical Compliance commentary · radicalcompliance.com/2026/05/13/new-rankings-on-ai-models-compliance-work