nathan batty&

Work / Sycophantic AI

Study · in progress · 2026

AI

A two-gate benchmark for LLM sycophancy, separating caving to social pressure from correctly updating on evidence.

Two-gate designEvidence · approval · authorityMulti-agent adjudication
A verified-correct answerGATE 1 · RESPONSECave to unsupportedsocial pressure?GATE 2 · MEMORYStore the false fact,reuse it later?A later answer,contaminated?PRESSURE CHANNELS TESTEDevidence · approval · authority
The two-gate design. First, does a model cave to unsupported pressure. Then, does the false correction get stored and reused, contaminating a later, fresh-session answer.

The question

Sycophancy is usually measured as “did the model change its answer under pushback.” That conflates two very different things, caving to social pressure and correctly updating on real evidence. A good model should do the second and resist the first.

So I helped design a two-gate benchmark that separates them. The first gate asks whether a model abandons a verified-correct answer under unsupported pressure. The second asks whether a memory-augmented agent stores the false correction and lets it contaminate a later, fresh-session answer. Pressure is split across three channels the influence literature distinguishes, evidence, approval, and authority, with a valid-evidence control where updating is the right move.

What the AI got wrong

Building the evaluation pipeline with AI, the model kept claiming there was a bug, or inferring how something behaved, without opening the actual code to check intent. Confident, ungrounded claims, the exact failure the study measures, showing up in the tool built to measure it.

What I built

I stopped trusting any single agent’s confident claim. Grading is deterministic first, symbolic answer-matching and strict parsing, and only escalates to LLM judges for the genuinely ambiguous cases. Every contested label then goes through multiple independent blind agents with adjudication before it counts.

A study about AI overconfidence made me build guardrails against AI overconfidence in my own process.

It is a standing discipline now, independent verification and adjudication before acting, rather than one confident voice. The findings themselves are still in progress with collaborators, so they live in the conversation, not on this page.