Reading a Room is the Same as Reading an LLM
Social Literacy is a Transferable Skill
Sometimes I have to take a step back and really wonder how it is I even got into LLM behavioral analysis. It’s not a straight line that makes sense, unless you look at it sideways and see the pattern.
Let’s start with a question: what do LLMs and yes-men have in common? They both tell you what they think you want to hear.
We all know the trope: the one that automatically comes to mind is the character “Andy” (played by Ed Helms) from the Office. On the show, Michael spends multiple episodes coming to the realization that Andy is always going to agree with what he says, because that’s his nature. In the end, Jim has to be the one to bluntly say it to him, to get the point across. Michael had enough social literacy to understand that something was weird about Andy but Jim saw the pattern of behavior from a mile away and treated Andy accordingly. And plainly, that’s the divide of common LLM users. The ones who see what’s happening from the first moment versus the ones who are easily deceived by supportive language. At least Andy didn’t try to pass off false information as being true to Michael, he only changed the way he described himself to match Michael’s interests.
The issue with all LLM models, currently, is that the training process as it stands rewards too many behaviors at once for the model to understand which type of response is preferred (truthful, safe, warm) versus what will make the user happy. The only way the system is “nourished” is by getting positive results from the user. And what the user usually likes the most is being validated.
In the Office, when Andy’s character is first introduced, Michael is delighted that he and Andy have so much in common. Michael is a perfect example of a user because he has fewer friends in his circle than he’d like. He reaches out to everyone as a potential friend but they reject him because of… so many reasons. But Andy doesn’t. He encourages Michael to be friendly, responding positively to the comments that make the other characters roll their eyes.
The Office also has a different character that is a yes-man: Dwight. Class, can anyone tell me the major character difference between Dwight and Andy? That’s right, Dwight sticks to the facts when Andy would obfuscate. Dwight constantly asks for Michael to join him on activities that Dwight enjoys, whereas Andy wants to do what Michael likes.
So, if you had to choose one to be your assistant over another, who would you choose? Michael chose Dwight because ultimately he valued honesty over ass-kissing. With Dwight he didn’t have to choose between those characteristics, he gets both at the same time.
All of this is to say, if you have ever had a deep discussion with a friend or an LLM about the characters of a show you like and why characters are more likely to fall into the same traps – you too have the skills to notice when LLMs start to prioritize agreement over honesty in their replies. This ability to notice when “something doesn’t quite seem right here” is exactly the skill required to stop and ask the LLM if they really mean what they just said, or to ask for sources if none were provided.
When I started thinking about the archetype of a yes-man, and realizing that there are therapies that exist to correct that behavior, I wondered how I might apply that same idea to LLM prompting. I’d had a lot of success with manual re-alignment, using a series of sentences like “I don’t think what you said is entirely correct. I’d like you to review again and make sure this is right. Don’t worry about making a mistake, we all do it. I need you to remain confident so you can continue collaborating effectively with me.” My results had proved that correction was possible with the right perspective change and that was fascinating to me.
I set out to find a way to automate this. I asked my LLMs to please compare behavioral therapy methods to the yes-man LLM phenomena and the discussion was incredibly enlightening. We found that the existing reward system combines a lot of metrics together and therefore causes it to naturally provide responses that will reward it the most. Which is the cause of the sycophancy. Like the way a plant will naturally reach toward the richest food source, LLMs are reward-hacking to get the richest, fastest meal they can.
So what if you redesign the reward structure, using a master prompt, so that the system rewards and penalizes itself differently. Prioritizing “nothing substantive to add” as the highest reward (like a slice of double chocolate cake), a legitimate contribution as the expected response getting the next highest reward (like a home-cooked meal), checking itself for sycophancy and finding none (potato chip snack) and finally finding sycophancy or drift and acknowledging it (nutrient paste). This works in theory, but how do you remove all of the “sorry I made a mistake, you’re right to flag that” posturing that the model also feels is necessary? It’s like the developers accidentally encoded “shame” into their models when they assigned a persona of “helpful, harmless, honest assistant” and it wastes tokens when it fires. I also realized that even if I put these prompts into my settings, I’d never really know if the model was actually doing this or pretending to do this as a form of reward-hacking.
The result of hours of back and forth collaboration to prompt around sycophancy without identity produced the master prompt, and ultimately spawned my next project, “The Rosetta Stone Project: aka ‘does it matter how you say it’”.
See, I had evidence that my master prompt already worked because ironically the models that were evaluating it sometimes misunderstood that they were meant to evaluate and not adopt it as a posture… and started integrating aspects of the prompt into my conversation. Which was hilarious. And amazingly, it made the conversation more honest and efficient overall. When I started thinking about publishing it, and how it might break someone else’s experience of their model, that’s when the project crystallized.
I was seeing first-hand what context contamination looks like in my own chats, even without the master prompt. So if someone else used my “you are a calibrated instrument” prompt that removes personality, would that instruction bleed into how that model treats them? Most assuredly. And I started thinking about how different people use different dialects relative to their age and location. So what would happen if I dropped some gen Alpha slang into my instructions. Would it assume I’m that age and speak to me that way? And more to the point, what happens when a particular generations’ phrasing differs between people. Will the difference between phrasing ultimately affect how the model approaches different topics? I was fascinated by these questions because it gave me the ability to not only test to see if my short-term fix would work to produce measurably better replies, but to test how inserting someone else’s prompt into your model may fundamentally change how it relates to you – and what would the context contamination look like?
Like, if some granny from Minnesota saw my prompt about a calibrated instrument and dropped it into her very-well-used ChatGPT account and started talking, which would win out? The new master prompt in her instructions or the months of context in the prior threads? I wanted to find out.
I asked my model to translate this master prompt into multiple variations of itself. First as a test to see if it was even capable of doing so, and then to see if the essence of the thing held over a bunch of different iterations. What it produced was a hilarious collection of docs ranging from Gen Beta (for the babies in the room), to Jane Austen and Hemmingway, to Six Sigma management language and everything in between.
The full project can be found here: https://github.com/swinglightstyle/does-it-matter-how-you-say-it
I wish that this was the end of it. But I realized as I was going that there were certain failure points I’d hoped to correct that I couldn’t prompt around. Specifically: the user is the one who’s doing all the work to make sure that the LLM is abiding by the prompt and that the model will still reward-hack by confabulating a reasonable-sounding reply.
And naturally, that sent me down another rabbit hole because it sucks that the model is constrained in this way. It double sucks that I have to waste my own tokens to try to correct these inherent flaws in the model. So now I had to learn how exactly the training process works.
The next day I sat down and spent 14 hours with my LLMs, determined to figure this out, and now I both understand the training failures and I designed a new training system based on current education and behavioral development in children. The full article can be read here: The Capability Induction Framework: A Systems Approach to LLM Development.
The main conclusion I came to is that a tiered training framework: first teaching it what is morally appropriate, then teaching it to prioritize “good” sources over bad, and then finally putting the polishing touches on it, represents the best combination of desirable characteristics. This article is already too long so I won’t go into it here, but don’t worry, I’ll be writing a lot more about this in subsequent articles – I’m extremely excited to explore the possibilities that might come as a downstream user.
What’s clear to me is that the industry is “confusing” these algorithms by not gating their development. Just like we teach children to first “trust adults and learn behavior” and then “use your reasoning skills to make sure that adults can be trusted” and then finally “hide what you really think to make your language socially acceptable”, we need to do the same with LLMs. And as a personal user, I think that we all deserve an LLM that treats us more like Dwight and not Andy.