JERRY XIONG / INTELLIGENCE BRIEF
No. 014  ·  15 May 2026  ·  Singapore
AI  ·  Compliance  ·  Finance Transformation

GPT-5.4 Scores 87.6% on Compliance Tasks. Here's Why That Number Misses the Point.

EQS just published their AI Benchmark Vol.2. The frontier models are now within a single percentage point of each other. The real story isn't who won — it's what comes next for anyone running a finance, audit, or compliance function.

AI  ·  合规  ·  财务转型

GPT-5.4 在合规任务上 87.6%——但这不是你应该关心的数字

EQS 这周发布了 AI Benchmark Vol.2。三个前沿模型挤在一个百分点内。真正的故事不是谁排第一——而是接下来的事,对每一个负责财务、审计或合规的人都很重要。

EQS AI Benchmark Vol.2 — frontier model scores converging within 1.5 percentage points
Fig. 01 — Convergence at the top. Source: EQS AI Benchmark Vol.2 (May 2026)
图 01 — 顶部收敛。资料来源:EQS AI Benchmark Vol.2(2026 年 5 月)

EQS Group just released its AI Benchmark Report Vol.2. Ten compliance & ethics domains. 120 real-world tasks. Blind-judged by human compliance professionals.

The result? Three frontier models clustered within a single percentage point:

EQS AI Benchmark Vol.2 — Top Three GPT-5.4  ·  87.6%
Gemini 3.1 Pro  ·  87.4%
Claude Opus 4.6  ·  86.1%

Eight months ago, in Vol.1, Anthropic's best model trailed Google's by ~5 points. Today that gap is essentially gone.

But the leaderboard isn't the story.

Three findings in this report matter far more for anyone running a finance, audit, or compliance function.

§ 01

The biggest gains are in open-ended work — not structured tasks

The single most important data point in the report:

Open-ended tasks — drafting reports, policies, investigation plans, regulator responses — improved by +17 to +18 percentage points versus Vol.1.

EQS's own framing: outputs moved from "usable with heavy editing" to "usable with light review."

For those of us who spend our weeks drafting bank query responses, going-concern memos, transfer pricing files, and regulator submissions — this is the threshold that actually matters.

Not a slick demo video from last year. A specific, verifiable, human-jury-validated capability shift.

My own experience during the FY2025 audit cycle matches the data. A year ago I was wrestling with AI output. This year I'm using it for first drafts.

§ 02

Model choice is being commoditized. The harness is now the differentiator.

EQS introduces a concept I find genuinely useful: The Harness.

If every major model performs at 86–88% on general compliance work, your AI program's success is no longer about which model you pick. It's about everything around it:

Moritz Homann, EQS's Head of AI, said it plainly:

"Six months ago, the question was whether AI could support real compliance work. Today, the question is how we design workflows around it."

— Moritz Homann, Head of AI, EQS Group

Models are commoditizing — the way cloud computing went from "AWS or Azure?" to "what's your architecture?"

§ 03

"Sequence View" beats "Human-in-the-Loop" as a mental model

The report walks through a 5-step Conflict-of-Interest workflow. Even GPT-5.4 scores 90%+ on each individual step.

But chain those steps end-to-end? 0.9⁵ ≈ 59%.

Matt Kelly at Radical Compliance offered a sharper framing than the tired "human-in-the-loop" cliché:

Human guard posts along the trail.

You don't passively insert humans into a loop. You actively decide where on the AI's path human judgment is non-negotiable.

For finance and compliance professionals, this reframe is good news.

"Will AI replace me?" becomes "Can I design where the guard posts go?"

The second question has a much better career trajectory.

FROM PRACTICE

Two examples from my own work

Case 1 — Hansen-style verification for LLM-derived variables

In my LPG market analysis work, I have LLMs extract variables from news, policy documents, and central bank working papers — import quotas, VLGC freight signals, CP-FEI spread movements.

The obvious question: how do I know the LLM isn't hallucinating?

My answer borrows from econometrics' Hansen J-test logic: have the LLM derive the same variable through two independent extraction paths, then check for consistency. Disagreement triggers human review.

That's a harness. It doesn't depend on any single model being smart enough. It depends on the workflow itself having self-checking built in.

Case 2 — An internal finance automation project

For an internal finance automation project, I maintain two documents alongside the code:

Every new session, AI reads both files before doing anything else.

This is the literal implementation of "invest in the harness, not just the prompt."

The compounding has surprised me:

The real asset turns out to be organized failure knowledge — not any individual brilliant session.

In Conclusion

Three takeaways for finance and compliance professionals

  1. Stop debating which model is best. This is becoming the 2017 "which cloud vendor?" debate — important, but no longer the source of differentiation.
  2. Start building your own harness assets. Prompt libraries, pitfalls logs, workflow SOPs — these are the moats that survive model upgrades and platform shifts.
  3. Position yourself as a sequence designer, not an AI user. The first is scarce and promotable. The second will be table stakes within two years.

Group Controller. Head of Compliance. Head of Internal Audit.

In the 2026+ landscape, the core competitive edge for these roles may simply be the ability to look at an 87%-accurate agent pipeline and point to exactly three spots where a human must sit.

💬

What are the human guard posts in your own workflows? Which steps will you absolutely not delegate to AI?

I'd love to see how different industries and compliance contexts are designing this.

Connect & reply on LinkedIn →
References
AIComplianceFinanceSingapore Audit TransformationGroup ControllerAI Augmented Finance LeadershipCompliance Transformation FCCALLMAgentic Workflow

EQS Group 这周发布了 AI Benchmark Report Vol.2。10 大合规与伦理领域。120 个真实任务。由人类合规专家盲评。

结果是?三个前沿模型挤在一个百分点内:

EQS AI Benchmark Vol.2 — 前三名 GPT-5.4  ·  87.6%
Gemini 3.1 Pro  ·  87.4%
Claude Opus 4.6  ·  86.1%

八个月前 Vol.1 发布时,Anthropic 最佳模型还落后 Google 约 5 个百分点。今天差距已经基本消失。

但排行榜不是故事的关键。

报告里我读到三个对每一个负责财务、审计或合规的人都更重要的发现。

§ 01

最大的进步发生在开放式任务,而不是结构化任务

报告中最关键的一个数字:

开放式任务(起草报告、政策、调查计划、给监管机构的回复)相比 Vol.1 提升了 +17 到 +18 个百分点。

EQS 自己的定性表述是:从"需要大量编辑才能用"提升到"轻度审阅即可用"。

这对我们这种每天写 Bank Query Response、Going Concern Memo、转让定价文档、税务说明信的人意味着什么?

意味着 AI 真正可用的临界点不是去年某个 demo 视频里的"看起来很惊艳",而是今年这个具体的、可验证的、有人类陪审团评分背书的门槛。

我自己在 FY2025 审计周期里的体感和这个数字基本对得上:去年还在和 AI 输出搏斗,今年已经在用它做第一稿。

§ 02

模型选择正在贬值,"挽具"成为新的差异化

EQS 提了一个我很认同的概念——The Harness(挽具)。

如果所有主流模型在通用合规任务上都跑在 86–88%,那决定你 AI 项目成败的不再是"选哪个模型",而是模型周围的一切:

EQS Head of AI Moritz Homann 直白地说:

"六个月前的问题是 AI 能否支撑真实合规工作;今天的问题是我们如何围绕它设计工作流。"

— Moritz Homann, Head of AI, EQS Group

模型本身正在被商品化——就像云计算从"选 AWS 还是 Azure"变成"架构怎么设计"一样。

§ 03

"Sequence View" 比 "Human-in-the-Loop" 更精准

报告里有一个 5 步 COI(利益冲突)审查流程的例子。即便是 GPT-5.4 这种顶级模型,在每个独立步骤上能跑到 90%+。

但如果你把 5 步串起来端到端跑?0.9⁵ ≈ 59%。

Radical Compliance 的 Matt Kelly 在评论这份报告时提了一个比 "human-in-the-loop" 更精确的表述:

人类哨岗(human guard posts)——你不是把人塞进 loop 里被动审核,而是主动决定在 AI 行进的路径上哪几个位置必须有人类判断。

对财务和合规从业者来说,这个视角的转变其实是个好消息。

"会不会被 AI 取代?" 变成 "你能不能设计这些哨岗的位置?"

第二个问题的职业曲线好得多。

来自实战

两个来自我自己工作的例子

Case 1 — LLM 提取变量的 Hansen 式验证

我在做 LPG 市场分析时,会让 AI 从大量新闻、政策文件、央行 working paper 里提取变量——某国 LPG 进口配额、VLGC 运费方向性变化、CP-FEI 价差信号。

一个直接的问题:LLM 提取的变量,怎么知道它没"幻觉"?

我的做法借用计量经济学里的 Hansen J-test 思路:让 LLM 用至少两种不同的提取路径独立得到同一个变量,然后比较两条路径是否一致。不一致就触发人工复核。

这就是一个"harness"。它不依赖任何单一模型有多聪明——只依赖于流程本身具备自检能力。

Case 2 — 一个内部财务自动化项目

在一个内部财务自动化项目里,我和代码一起维护两份文档:

每次新会话开始,AI 都会先读这两份文档。

这就是"投资于 harness 而非 prompt"的具体落地。

复利效应远超我预期:

真正的资产是组织化的失败知识——不是某次会话的灵感。

结论

给财务与合规从业者的三点结论

  1. 停止讨论"哪个模型更好"。 这个问题正在变成 2017 年的"哪个云厂商更好"——重要,但不再是差异化来源。
  2. 开始建立你自己的 harness 资产。 Prompt 库、避坑清单、工作流 SOP——这些是你在 AI 时代的护城河,比任何具体的工具技能都更长寿。
  3. 把自己定位为 sequence designer,而不是 AI user。 前者稀缺、可晋升。后者会在两年内变成所有人的基础技能。

Group Controller、Head of Compliance、Head of Internal Audit。

2026 年之后,这些角色真正的核心竞争力,可能就是能不能看着一个 87% 准确率的 agent 路径,准确地指出哪 3 个节点必须放人。

💬

你在自己的工作流里设了哪些"人类哨岗"?哪些步骤你坚决不交给 AI?

很想看看不同行业、不同合规场景下的设计差异。

在 LinkedIn 上交流 →
参考资料
AI合规财务新加坡 审计转型集团财务总监AI 增强 财务领导力合规转型 FCCALLM智能体工作流

Jerry Xiong writes on AI governance, markets and operational risk for finance and business decision-makers.熊焱|为财务与企业决策者解读 AI 治理、市场与运营风险。