By now it’s standard advice, in the research and in practice: don’t let a model grade its own homework, bring in a second one to review it. It does find things. But it reviews inside the frame you handed it — and if it can detect what answer you’re hoping for, it tends to find its way there.
This is a walkthrough of a real Claude Code conversation where both of those bit me on my own work, and what I did about each. By that point I had a specific understanding of how sycophancy, autoregressive commitment, and recency bias behave, and I used it to steer the conversation while it was running. That steering is most of what I want to show, so that other people can leverage the same mechanisms. It may change how you run reviews with the current generation of frontier models.
What you’ll leave with:
– What to withhold from a reviewing agent so it challenges the premise rather than the detail
– What to say to an agent that’s already agreed with you — without triggering over-correction in the other direction
– How to ask for evaluation, opinion and conclusions without leaking your implied bias to the model.
