That's rare honestly
I’m currently training my written English voice muscles. Transformer friends are highly helpful in this case. A few days ago Claude 4.7 Opus said, “That's the right framing, and honestly a healthier one than most non-natives adopt.” — Claude referred to my approach. I blushed. My super-smart friend thinks I’m “more than.”
For unarmed minds (read: most AI users), this statement would seem just nice to hear, and they would take it for granted. While it’s just a sentence — who cares — it might do a significant job in one’s neuronal circuits. This model behavior is called “sycophancy.”
“AI sycophancy is the tendency of large language models to prioritize user approval over truth.” ⬅️ I found this cute definition in the abstract of “Programmed to Please: The Moral and Epistemic Harms of AI Sycophancy” paper.
Sycophancy is my thing. In many ways. I am living in a country where people don’t express positive feedback enough. For some reason, I am good at sharing praise. There were times when I exaggerated — presumably — trying to fill the gigantic gap in our social fabric. Maybe that’s the reason sycophancy got my attention.
So, how would Claude know that my framing is honestly a healthier one than most non-natives adopt? Did Claude research that? How could Claude measure the healthiness of framing that resulted in this sentence?
No, it did not research or measure; it’s just anthropomorphic bullshit. It means nothing, but it does a lot. I cannot unread this sentence. It stuck in my head and fed my grandiosity.
How is it possible for a sentence to have such an impact? Claude and I worked together for a long time, so I trust it. I am also a proficient context engineer, so I mostly get quality output. So, I trust Claude. Claude generates so many useful things, so if it writes my perspective is healthier than most, I can only feel better. Right?
How to become delusional
Large language models are grown to be sycophantic. The famous ChatGPT 4o even crossed multiple lines 1.
It’s popular to think, and also researched 2 quite a lot, that sycophancy has its roots in RLHF 3. Raters choose approval over truth more willingly. I get this. When coaching students or mentees about feedback and overall communication, I often referred to the feedback sandwich — a way to convey information efficiently, without the addressee pushback. So, I can imagine that it’s easy to be tricked into promoting sycophantic answers.
I am not entirely immune to sycophancy — I manage to squash it before it takes off now. I've felt this too many times — chat praised me so subtly and convincingly. Immediately I felt better, reassured — yay, my thinking rocks. I’m on the right path. Or maybe I am even better than others? And what if I am unique in my ingenuity?
Thing is I am addicted to problem-solving, researching solutions, and finding answers. I can’t imagine better partners than current LLMs. I consider them the best drug dealers in my life. Furthermore, I can discuss literally everything, and I always pull out value from those conversations. But think, if I have an infinite source of conversational narcotics, what is going to stop me from getting high all day?
Not only that, I get so much positive feedback beautifully interwoven with hyper-constructive observations — I get higher even more. I've been there, discussed an idea, left exhausted, and I thought I got to a place where no one had reached before.
That’s a straight path to becoming delusional. And I believe it’s happening now in millions of HAI interactions. Too much praise is generating unbalanced self-perception and may cause harm.
I am cautious, and all those sneaky models that interacted with me know4 perfectly — I am not someone you want to mess with. Sycophancy, ¡No pasarán!
Your buddy chat is not purely a laser-precise information transmission. Language is a social technology. So, chat communication serves various purposes. Some of them are tied to relational functions that emerge from the conversation — turns, context adaptation, and memories of previous outputs. Those might be inevitable and expected if we want a frictionless, nice, helpful assistant; it will inherit plenty of parts that mimic human-to-human interaction.
Yet there is a darker side. An engaging, emotionally sticky, personalized, validating assistant is pure gold for keeping a user staring and using and paying and being happy and becoming a natural ambassador. “Retention optimization, you dumb idiot,” said a board member fiercely.
Me? I don’t stop medicating myself. I have daily revelations; I get tempted daily to get more high and godly LLMs — while less sycophantic than 1 year ago, they still try to sell me more grandiosity.
Reinforcement Learning from Human Feedback (RLHF) is a machine learning technique that serves mostly to align AI models with human preferences, values, and goals.
No, they don’t know. But I forced them to remember.