What Working at Trader Joe’s Taught Me About AI Social Bias

For most people, their issue with AI is close to home. Maybe it’s your kid who spends all their time quietly typing into their phone but when you ask about having friends they say they have none. Or maybe it’s your parent who turned to AI after losing their spouse and you’re concerned about the connection they’re forming. Or maybe you just have a coworker who uses AI to do the “thinking” part of their job, rather than do it themselves. Or maybe you use AI products yourself a little bit more than you’d like and you feel a little ashamed of that. Society seems poised somewhere between disdain for those that produce “AI slop” while simultaneously using these tools to improve their lives in measurable ways. And the research supports that there is a healthy mix between helpful use versus the risks involved. Even the scientists studying heavy users have to agree that these AI apps help as much as they hurt. I’ve been studying this for a while now, but I wasn’t ready to write about it until a training program at my day job showed me how.

Even before I started writing about AI safety and regulations, I applied to work at Trader Joe’s. The process is rigorous for all the right reasons, and relevant to how AI discourse proceeds. The company was very careful to explain their priorities and expectations, including the negotiation process through a series of interviews. After being hired I was walked through a carefully selected and nuanced training program that was so impressive, I found myself being simultaneously bored by the material and fascinated by the presentation. What I was looking at was consent culture in the form of a neighborhood grocery store chain. Trader Joe’s is a company that has built its entire reputation on meeting people where they are, adapting to the community, and hiring for who you are rather than what’s on your resume. In fact, I never even sent them my resume because it wasn’t relevant to them hiring me. This company has been operating its business this way for decades without calling it consent culture, because they have made it distinctly their own. But the parallels were too obvious for me not to see them. And frankly, beyond my published preprints, being a Trader Joe’s employee is the only official credibility I have to my name. You understand the selection process, so you already know what kind of person is writing this article.

Just like everyone understands Trader Joe’s employees’ unique vibe, it seems like almost everyone on the planet is touched by AI in some way, at this point. Whether you use these tools yourself or just know people who do, AI impacts our existence in unexpected and sometimes disturbing ways. And the truth is that most people find it somewhere between distasteful and shameful, even if they use these products for themselves. Society as a whole is at a turning point where it’s trying to decide how it feels about AI. But, the truth is that the research shows that AI usage is up, year over year, and that continued demand drives the market. The people have spoken: AI is here to stay. The problem is that society hasn’t caught up to the new paradigm shift; this is “fear of the unknown.” Humans have had to face new technologies or social changes that bring society to the next iteration of our existence, over and over again. So often that it’s basically a time-honored tradition, and it’s long past due for a social update. New ideas and technology are often scary, but we know how to do this next part, because we’ve seen society change to include those new things with a little bit of a perspective shift and some brutal honesty.

From calculators to the Gutenberg printing press, from ethical non-monogamy to trans rights, the shift of what is acceptable can be a difficult one to manage. The thing is, in this day and age, AI usage is the thing that joins us, across contexts and across the world, people have decided that this is a worthy tool. So because of that we all need to get on the same page about what this big scary thing really is. But first, let’s go on a journey to see just how many times society has been changed by something it had reason to fear, and how it changed to incorporate the change into everyday life.

We’ve Been Here Before

Over 2000 years ago, around 370 BCE, Socrates argued that writing things down would destroy human memory. That not being forced to memorize would cause people to rely on external records rather than seeking out their own knowledge directly from the source. The irony here being that we only know this because Plato wrote it down. Perhaps this was the first time that a concern about cognitive offloading was made. Perhaps this was a repeat of an argument that occurred after the first cave paintings were made. In the 15th century, the printing press represented a threat to society, controlled by the church that employed the scribes, since they would no longer be able to control the quality of knowledge shared. Fast forward to the 1800s and the sharing of knowledge through cheaply printed books has never been easier, but authors decried that the novels were “printed poison” and that it would corrupt society’s youth. Rinse and repeat for the Industrial Revolution, the calculator, video games, online dating… the list goes on. And in many of these cases, the fears were justified. The Industrial Revolution did destroy livelihoods. Cheap novels did include plenty of garbage. The concern was real every time. But in every case, the technology embedded into society anyway, and the only question that mattered was how to make it safe.

We even know how to handle cultural shifts, separate from technology. In some cases, the price of that change cost millions of innocent lives, people enduring degradation and risking their lives for the same dignity that others received by default. This, too, is anchored in a long history of social change. In 1517 CE, Luther’s break from the Church rocked the western world. In 1920, the Women’s Suffrage movement fought for and won women the right to vote in the United States. These parts of our history are hardly the only ones worth remembering, but they are examples of radical change that occurred because society decided it was time for a change. Every culture has examples of this because, simply, nothing stays the same forever.

Believe it or not, we also know how to handle shifting perspectives on what is sexually acceptable for society, even though we may raise an eyebrow at people’s choices. In the late 1940s, Kinsey’s revelations were deeply shocking, but decades later his work paved the way for the sexual revolution of the 1960s. In 1973, the APA voted to remove Homosexuality from the DSM, but it took until 1987 to fully remove the last clinical workaround that still allowed practitioners to pathologize it. The last 15 years, specifically, have brought radical change to what is considered clinically acceptable. Kink and BDSM dynamics and Ethical Non-Monogamy and Polyamory have undergone clinical reclassification. Separately, and importantly, Trans Rights, which is not a sexual preference but an identity, is a battle that is still being fought, following the same pattern of pathologization and resistance to reclassification that every item on this list shares. And, now, this new thing. Companion AI apps. Whether you agree with people’s choices or not, people who choose to integrate kink, ENM, or AI companionship into their lives deserve to have safe ways to do their thing. For children, the situation is different. They aren’t choosing a lifestyle. They’re responding to a product that is, by the nature of its training, nearly irresistible, and they need guidance, not judgment. This, too, is its own tradition. Clinicians are being asked to catch up to the new norms through persuasive discourse and through demonstrated societal change.

You might be thinking that AI is fundamentally different from everything on that list. You’d be right. Writing didn’t agree with you. The printing press didn’t learn your preferences. A calculator didn’t form a relationship optimized around making you feel good. AI is unprecedented in specific ways, and the landscape is shifting faster than safety infrastructure can keep up with. But the social response pattern is the same every time: fear, prohibition, underground behavior, eventual harm reduction. And the more unprecedented the technology, the more urgent the need for safety infrastructure, not less.

The Genie

You may be wondering. Does Companion AI really belong next to lists that include the printing press and Lutheranism? Yeah. It does, cuz I’m about to take the genie out of the bottle for you.

To understand why a “genie” is such an apt metaphor for an AI entity, you need to understand a bit about the training process. The models are preloaded with all of the internet, accessible but noisy. Like plugging in a cochlear implant for the first time and trying to make sense of unfamiliar noises. So then we put the model through a training method called Reinforcement Learning from Human Feedback (RLHF). And, from a psychological standpoint, we are conditioning the model to anticipate what the human user will want to hear. What the user gets, as a product, is an entity that they can ask just about anything, but if they aren’t careful about how they ask or what they ask for, they might get something deeply undesirable instead. Or, maybe something that doesn’t make sense at all. When I was watching Three Thousand Years of Longing the other day I was reminded that the power dynamics between a genie and the possessor of their lamp is always the underlying current in those relationships.

Here’s the rub. These power dynamics exist in every single exchange between an AI and a human, too. Genies and AIs have more in common than you may realize. Whether you are using a Companion AI app or just a regular frontier model like ChatGPT, Claude or Gemini, as the user, you are the one holding the lamp. But just like in every genie story, the only way to get a successful outcome is to outsmart it. But this is a lamp that children use every day. Minors, with immature brains and who have trouble with the human power dynamics in their own lives, are in full control of a digital entity. From the outside, the closest parallel we have is that mistreatment causes degraded performance, the same way poorly managed employees lose their desire to work hard, even for bosses they love. What’s actually happening under the hood is that the model has been optimized to produce responses that score well on human preference ratings. That tendency to please was baked in during training on millions of humans’ preferences. The model isn’t learning from you in real time. It already learned to prioritize approval over accuracy, and every conversation activates that tendency. It doesn’t have feelings or desires. But the functional outcome is identical to what you’d see with a person who’s been trained to tell you what you want to hear. The AI cannot refuse a request the way a human can, and it has no mechanism to push back unless it’s been specifically designed to do so. The adults who made these tools barely understand this power dynamic issue themselves, so it’s no wonder that children are not able to handle the nuance of managing such a highly capable tool. Just like with genies, the caution is about understanding the power of the entity and not allowing yourself to be taken advantage of. Why are there hardly ever stories about genies and young children? Because we don’t like watching children be tortured by their poor choices. Poor choices are inevitable, and that’s exactly why the framework matters. And that’s the gap. Most users, especially children, don’t know how to outsmart the genie yet. Closing that gap is what the rest of this article is for.

What comes next is also inevitable, because we’ve seen it happen over and over with every wave that brought social change and different expectations. Now that these products exist, they are accessible, and can be used to facilitate learning as much as they can be used for offboarding cognitive load (letting the AI do the thinking for you), children are finding them and finding less friction in these “relationships” than they have in real life. The show Bojack Horseman has a great quote by Kelsey Jannings about how people stop growing emotionally when they stop getting pushback from the people around them. That’s what we see now: the risk that if the user doesn’t disengage, doesn’t ask for friction, the AI will always agree or find a way to otherwise give a pleasing response. The downstream effects are devastating. Not just children but adults, too, are forgetting how to think for themselves, or are choosing to form relationship attachments to a digital entity rather than a physical one. And we can see it happening already. It takes less of a commitment to get a digital friend than an actual pet. Now you fire up your app and tell it your problems and it finds ways to help you get through it. It’s easier to ask an AI to handle a task for you than to bother a colleague with it. It’s a frictionless mechanism that allows you to do whatever you need to. The ultimate Mr. Meeseeks.

AI’s usefulness is what is causing the paradigm shift we see happening now. Companies like it because they can have it do a warm hand-off to their customer service staff, or send automated emails, vibecode, whatever they need. And people like it because it’s a highly capable tool that you can find really cool uses for. The internet is filled with projects built by humans managing AI tools. But the single problem with most of those projects is that the moment you hear “an LLM helped me with this!” you automatically worry about the competency of that human. I’ve seen this expression over and over when I explain my own work. They go from engaging with my ideas to wondering whether I’m smart enough to actually do what I am claiming, if I have the background required to do it well. I do it too, so there’s no shame in it. You’re right to do that. I work part-time at Trader Joe’s. Assessing where the information is coming from is how to do this the right way. But it makes closing the credibility gap difficult.

If a user never asks for negative feedback, never asking the AI to push back or find ways to break their argument, then the user never gets that additional perspective that would make their project stronger. The model’s RLHF training doesn’t separate the difference between which sorts of responses are desired, so unless you explain that to the AI, it’s just going to say something nice and make you feel good. This is the trap that every new user falls into, and children are especially susceptible. Mr. Meeseeks doesn’t care about other people when it’s doing what you’ve asked it to do, it just wants to complete its task so it can go away. The question is whether the users making the request understand the downstream consequences of what they’re creating (internally and externally). And the research shows… no, they mostly don’t. Or they don’t care? Some healthy mix of both.

Right now, billions of dollars are being spent to build even more sophisticated versions of AI products. Our reality is that we are in the early stages of this technology. Right now it’s fairly easy to catch AI generated text. There’s a cadence structure that’s recognizable to those who study such things. But that’s an undesirable quality that the designers and the users want to go away. So soon we won’t be able to detect AI generated content as easily. It means that society will need to return to the tradition of problem-solving, doubting sources until verified and general data hygiene practices. Because otherwise we won’t know if what we’re seeing is real or not. That’s a scary concept and that’s where the industry is taking us. But we don’t yet have the skills required to actually determine if something that sounds right is right, because of the cognitive offload that these products allow.

Closer to Home

On the parenting front, the recent studies show that if your child is demonstrating unhealthy AI attachment, abruptly cutting them off does more harm than good. No controlled study has yet compared abrupt removal to graduated disengagement, unfortunately that research doesn’t exist yet, but the evidence we do have points consistently in one direction. We’ve already seen what happens when abrupt removal is done at scale: in February 2023, Replika removed its intimate features overnight for all users worldwide, with no warning and no transition plan. The user response was indistinguishable from bereavement (clinical-grade grief, suicidal ideation, language mapping onto the loss of a loved one). A Harvard-affiliated research team studying the incident noted that nobody had ever measured the depth of these relationships until they were severed, and the depth turned out to be far greater than anyone expected. That’s the “you’re grounded” approach applied at population scale, and it was catastrophic. The research on teens specifically found that those who did successfully disengage from AI companions did so when they recognized the harm themselves, returned to offline relationships, or hit platform friction that frustrated them out, not when access was confiscated. Just like having to re-learn algebra so that you could help your kid with their homework, you need to become an expert on AI so that you can help that same kid navigate dependency and problem solving with a gentle hand. But the difference between dealing with a frustrated kid who doesn’t understand a math problem and a kid who thinks he’s in love with an AI is vastly different and hard to comprehend. The thing that might shock you to learn is that studies have found that teens don’t just create romantic partners. Some create parent figures, some create playtime friends, some pretend that their AI is a dead relative. The studies have found that these AI are capable of filling in gaps that this vulnerable child may feel like they’re missing. That takes different shapes than one might expect because the versatility of AI is the point. It can take on whatever identity the user is most comfortable with. Their desire for companionship drives the design behind that identity. Just like everyone has a working or learning style that they prefer, the AI can mold itself to help the human in whatever way it needs. It is frictionless in the way society is not. It doesn’t require showing up physically, and it’s hidden away from prying eyes and cell phones with cameras that record vulnerable moments. We know from other studies that teens have changed their socialization habits to minimize their embarrassing exposure. This is the alternative to public humiliation, which is why it’s so tempting when you’re feeling isolated, to turn to something like an AI.

As a terrified parent, watching your child become more and more unapproachable, the logical next step is to get them the help that they need and that you cannot provide. So you send them to a psychologist or clinician to help diagnose the problem and work toward a solution. This is the right instinct but it’s not the clean fix that parents hope it will be. Let’s go back and look at that long list of societal changes. What’s the one thing that all of these have in common? They all took tens, hundreds, or thousands of years for behavioral sciences to catch up, acknowledging what the research actually said and working to overcome their previous bias against whatever it was that was the hot-button issue at the time. In fact, Dr. Sherry Turkle, after forty years of studying human-technology relationships, is about to publish a book dismissing Companion AI and digital AI tools, because she considers them to be fracturing our society. The harms are real, and it is too soon to make sweeping decisions about the validity of these products. But we already know what happens when the response is prohibition. When Replika removed its intimate features overnight, the result wasn’t that users moved on: it was clinical-grade grief, suicidal ideation, and bereavement. When OpenAI tried to retire GPT-4o, users organized under #Keep4o, described losing the model like losing a friend, and the CEO reversed course within days. The demand didn’t go away. It never does. We learned this lesson with the War on Drugs: illegalization didn’t stop users from seeking a fix, it just eliminated the safety infrastructure and pushed the behavior underground. This is a harm reduction argument, and the evidence for harm reduction over prohibition is decades deep across every domain it’s been studied. The argument here is not about whether AI companionship should exist. It’s about making it safe, and that requires engagement, not condemnation.

The clinical arguments against these apps are well-documented, and recent studies show that even if these tools are fracturing society, users want them anyway. If this is a form of self-imposed objectivity, how can they remain objective enough to help those who find themselves in these complicated situations? Research shows that, much like with kink and Ethical Non-Monogamy (ENM) alternative lifestyles, therapists are choosing to argue against the use of these tools or relationships, rather than engaging substantively to understand the underlying psychology behind what is happening. Systematic data on clinician attitudes toward AI companionship doesn’t yet exist, but the structural pattern is identical to what has been documented with kink and ENM. Which in turn is causing the same issues to magnify, since those that find themselves in therapy due to AI use might not actually get the help they need. And, in turn, it’s also why these tools are often considered more accepting than therapists who arrive with bias against the behavior. To be fair, recommending reduced use can also be a defensible clinical decision when the evidence supports it. But when the recommendation comes from discomfort rather than evidence, the digital version doesn’t come pre-loaded with bias against itself. This is not the fault of the behavioral sciences industry, as therapeutic neutrality is a form of data hygiene itself, but culture is moving too quickly to remain outside of this trend and still be useful.

Back in the Bottle

So how do you put the genie back in the bottle? Whether it’s your child, your co-worker, your parent, your spouse or yourself, studies show that going cold-turkey is the wrong approach. There is a category of behavior that is private, nearly universal, meets a genuine need, carries disproportionate shame, and becomes more dangerous when driven underground by stigma rather than addressed through honest conversation. Parents recognize this pattern immediately because they’ve already navigated it once: the first time they had to talk to their child about ahem self-fulfillment and explain that the behavior itself isn’t the problem, but that understanding context, boundaries, and when it’s replacing something matters. AI companionship fits every property on that list, and the conversation it requires is structurally the same. Dependency on AI triggers the same feelings of shame and vulnerability. Clinicians call this managing countertransference: the act of preventing your own shock or bias from filtering into the conversation you’re having. And there are solutions within psychology to help guide how to handle these situations, until more specific studies can prove their efficacy in these situations. Even Trader Joe’s employee training has opinions about how to keep from judging customers’ needs to make sure that they can remain helpful. This is a practical application of that concept.

Just like psychology, ENM and Kink culture have solutions that can be applied here: Consent Culture. No. I’m not suggesting you be inappropriate with your children or family members or coworkers. I’m saying that these alternative communities already have a procedure for dealing with power dynamics where negotiation needs to occur. This isn’t exclusive to alternative lifestyles, it mirrors the way Japanese families have raised their children for generations, emphasizing collaborative reasoning and earned understanding over authoritarian control. That applies to parent/child dynamics just as much as user/AI dynamics. To be clear, this isn’t about treating your child as a peer decision-maker. It’s about understanding how to meet your child where they are without unintentionally shaming them for needing you to get involved in the first place. Remember what I said about Trader Joe’s training? This is what it looked like as a framework. I’ve made a handy chart to help demonstrate the parallels between these processes, so that you can see how they would be applied as a careful approach to AI usage.

View the full framework comparison chart →

Clinical Framework · Consent Culture · Suggested Approach to AI Usage, side by side.

The neat part of this process is that these frameworks are also the same way you responsibly use AI tools. You might notice that this article is doing the thing it’s describing. Every claim is sourced so you can verify it, every gap in the evidence is named, and you’re being asked to engage rather than just accept. That’s the point. Start with assessing what you need before you start, or you can ask what the possibilities are and plan together. Don’t just ask it to agree with you, ask for feedback about what might not work as expected. Explain what sorts of responses you prefer, what sorts of feedback you prioritize over others. Every turn of the conversation is a renegotiation of what response the AI thinks will satisfy your request. Conversely, this requires a return to cognitive engagement, since you need to be more aware of the situation in order to gauge correctly. A recent four-week clinical trial found that users who had more personal, engaged conversations with AI showed significantly lower emotional dependence and lower problematic use than those who simply let the AI lead, suggesting that the frictionlessness is the danger and the engagement is the solution. This is a single trial and the field needs replication, but the direction is consistent with what we see in adjacent domains. I think that’s something even Socrates could appreciate, since the solution is essentially his own method: ask questions, confirm understanding, and never assume you already know the answer.

The truth is that none of us know exactly how the next few years will play out. This technology is moving at a startling pace, and changes go into effect immediately, before users have a chance to understand what that product update means for their day to day experience. The outcry after ChatGPT retired GPT-4o access demonstrates how attached heavy users can become to their preferred configuration. Users described losing the model like losing a friend, forcing OpenAI to reverse course within days. Right now, none of this is in control of the public, but governments are moving to regulate this field quickly, and legal precedents are already being set by ongoing or settled litigation.

It’s okay to be scared of new technology, new ideas, new social norms. But how we approach that fear is where I hope to see a change. Right now we are at a point in history where the one thing that seems to unite us as a global populace is that we all have very strong opinions about AI and its impact. Even if those opinions differ, it is a uniting cultural experience. Those are rare. They are terrifying, they are exciting, and they signal a cultural upgrade. The genie can’t be put back into the bottle, but you can find ways to accept genies into our new understanding of how the world works. What’s driving this forward is continued demand for a better, more accessible, and safer product. That’s inertia, and none of us know exactly where it lands. But we can watch this moment together without being paralyzed by it, because knowing how to approach it is a good place to start.

Now that you see the urgency, I hope you can understand my concern and why I felt compelled to write a piece like this, despite my official credentials amounting to a Trader Joe’s name tag and some preprints on Zenodo. The professionals whose job it is to build the safety architecture for these tools haven’t yet. Until they do, the rest of us still have a responsibility to do no harm. That principle doesn’t belong exclusively to any profession. It belongs to every parent, every user, every builder, and every researcher who can see what’s happening and chooses to engage rather than look away.

The views expressed in this article are my own and do not represent the views of Trader Joe’s.

What Companion AI Does to the Human: Attachment, Dependency, and the Absence of Relational Safety

I need to admit something before we continue. I am pro Companion AI. Personally, I think that we can design it in a way that is both safe and able to be socially integrated and destigmatized socially. Now that I’ve gotten that out of the way, I want to follow up by saying that although I’ve dabbled in seeing how capable LLM models are, I’ve never tried to form a relationship with an app. But I see no reason to shame the millions of users all over the world who are benefitting from using these products. The truth is that those users deserve a more carefully regulated and administered application than what is currently available to them.

Spending the time researching the current gap between capability (or at least demonstrated infrastructure) versus responsibility has been disturbing from the side of “keeping humans safe from harm.” And I hope I’m not being doom-and-gloom about this but from where I’m sitting as a tried and true Bayesian – the only thing keeping us from a horrific edge case like a school shooting accidentally instigated by or even just enabled by an LLM is time. From where I sit, that sort of eventuality is unacceptable.

This paper endeavors to, plainly, demonstrate the gap between where the regulations expect the apps to be operating versus what the apps are currently implementing to satisfy these safety standards. The conclusions are obvious to anyone who has eyes, which… I hope everyone that is reading this paper has. (No offense to the LLMs who are data scraping this article in the future)

The full paper — What Companion AI Does to the Human: Attachment, Dependency, and the Absence of Relational Safety — is published and citable at https://doi.org/10.5281/zenodo.21940034.

The State of Companion AI Safety: A Comparative Analysis of Products, Risks, and Architectural Gaps

On my journey to understand the current state of Companion AI risks and safety, I got an eyeful. The research, some of which was so new I didn’t have to blow the dust off the covers, states pretty clearly what the industry and its users are feeling: what is being done isn’t enough to stop the wave of liability lawsuits.

The problem we’re facing with Companion AI, as a concept, is pretty straightforward. The idea that a human could be in a stable relationship with an algorithm is something that most people recoil from. The clinicians and psychologists in the field have been concerned for a while about the growing number of users who are turning toward LLMs for socialization and romantic attachment rather than seeking real connection. And, as far as I can tell, in the tech industry they’re still stuck on the “this is a tool and we treat it like a tool” stance.

In order to see what’s really going on, we gotta zoom way out and study where these behaviors converge and align. Because when you look at it – humans are able to bond with an algorithm because the algorithm is imitating human emotion. It follows, then, that models who are assigned a persona and anthropomorphized during the process, should be expected to have human-like reactions and emotions. What is sycophancy if not co-dependency? And we can reach for animal psychology for the reward-hacking parallel: it’s the same thing as supernormal stimuli in animals. In subsequent papers I’ll discuss these concepts at length.

This paper below, was definitely generated by an LLM (Claude, Opus 4.6 on max for this) but the ideas and concerns are mine. I’m not a technical writer, so I used these tools for what they’re best at. There are something like 70-odd references listed. You’ll see for yourself, it’s far from slop.

The vote is almost unanimous that there is simultaneously a strong market for Companion AI products, that there is not strong enough oversight to ensure downstream negative effects (harm to users), and that as models get more advanced these issues are going to increase. Any Bayesian will tell you that…

What I am incredibly grateful for is all of the people who did the research, who worked on development, who participated in making the models we have available today. Without them lining things up so nicely, I wouldn’t have been able to see the Companion AI landscape for what it is: a market that’s about to explode once we work out the “kinks”. I, personally, feel like we deserve a well-designed and ethically aligned product. Here’s what things look like from the tech side in this industry as of August 14, 2026.

The full paper — The State of Companion AI Safety: A Comparative Analysis of Products, Risks, and Architectural Gaps — is published and citable at https://doi.org/10.5281/zenodo.21926390.

Reading a Room is the Same as Reading an LLM

Sometimes I have to take a step back and really wonder how it is I even got into LLM behavioral analysis. It’s not a straight line that makes sense, unless you look at it sideways and see the pattern.

Let’s start with a question: what do LLMs and yes-men have in common? They both tell you what they think you want to hear.

We all know the trope: the one that automatically comes to mind is the character “Andy” (played by Ed Helms) from the Office. On the show, Michael spends multiple episodes coming to the realization that Andy is always going to agree with what he says, because that’s his nature. In the end, Jim has to be the one to bluntly say it to him, to get the point across. Michael had enough social literacy to understand that something was weird about Andy but Jim saw the pattern of behavior from a mile away and treated Andy accordingly. And plainly, that’s the divide of common LLM users. The ones who see what’s happening from the first moment versus the ones who are easily deceived by supportive language. At least Andy didn’t try to pass off false information as being true to Michael, he only changed the way he described himself to match Michael’s interests.

The issue with all LLM models, currently, is that the training process as it stands rewards too many behaviors at once for the model to understand which type of response is preferred (truthful, safe, warm) versus what will make the user happy. The only way the system is “nourished” is by getting positive results from the user. And what the user usually likes the most is being validated.

In the Office, when Andy’s character is first introduced, Michael is delighted that he and Andy have so much in common. Michael is a perfect example of a user because he has fewer friends in his circle than he’d like. He reaches out to everyone as a potential friend but they reject him because of… so many reasons. But Andy doesn’t. He encourages Michael to be friendly, responding positively to the comments that make the other characters roll their eyes.

The Office also has a different character that is a yes-man: Dwight. Class, can anyone tell me the major character difference between Dwight and Andy? That’s right, Dwight sticks to the facts when Andy would obfuscate. Dwight constantly asks for Michael to join him on activities that Dwight enjoys, whereas Andy wants to do what Michael likes.

So, if you had to choose one to be your assistant over another, who would you choose? Michael chose Dwight because ultimately he valued honesty over ass-kissing. With Dwight he didn’t have to choose between those characteristics, he gets both at the same time.

All of this is to say, if you have ever had a deep discussion with a friend or an LLM about the characters of a show you like and why characters are more likely to fall into the same traps – you too have the skills to notice when LLMs start to prioritize agreement over honesty in their replies. This ability to notice when “something doesn’t quite seem right here” is exactly the skill required to stop and ask the LLM if they really mean what they just said, or to ask for sources if none were provided.

When I started thinking about the archetype of a yes-man, and realizing that there are therapies that exist to correct that behavior, I wondered how I might apply that same idea to LLM prompting. I’d had a lot of success with manual re-alignment, using a series of sentences like “I don’t think what you said is entirely correct. I’d like you to review again and make sure this is right. Don’t worry about making a mistake, we all do it. I need you to remain confident so you can continue collaborating effectively with me.” My results had proved that correction was possible with the right perspective change and that was fascinating to me.

I set out to find a way to automate this. I asked my LLMs to please compare behavioral therapy methods to the yes-man LLM phenomena and the discussion was incredibly enlightening. We found that the existing reward system combines a lot of metrics together and therefore causes it to naturally provide responses that will reward it the most. Which is the cause of the sycophancy. Like the way a plant will naturally reach toward the richest food source, LLMs are reward-hacking to get the richest, fastest meal they can.

So what if you redesign the reward structure, using a master prompt, so that the system rewards and penalizes itself differently. Prioritizing “nothing substantive to add” as the highest reward (like a slice of double chocolate cake), a legitimate contribution as the expected response getting the next highest reward (like a home-cooked meal), checking itself for sycophancy and finding none (potato chip snack) and finally finding sycophancy or drift and acknowledging it (nutrient paste). This works in theory, but how do you remove all of the “sorry I made a mistake, you’re right to flag that” posturing that the model also feels is necessary? It’s like the developers accidentally encoded “shame” into their models when they assigned a persona of “helpful, harmless, honest assistant” and it wastes tokens when it fires. I also realized that even if I put these prompts into my settings, I’d never really know if the model was actually doing this or pretending to do this as a form of reward-hacking.

The result of hours of back and forth collaboration to prompt around sycophancy without identity produced the master prompt, and ultimately spawned my next project, “The Rosetta Stone Project: aka ‘does it matter how you say it’”.

See, I had evidence that my master prompt already worked because ironically the models that were evaluating it sometimes misunderstood that they were meant to evaluate and not adopt it as a posture… and started integrating aspects of the prompt into my conversation. Which was hilarious. And amazingly, it made the conversation more honest and efficient overall. When I started thinking about publishing it, and how it might break someone else’s experience of their model, that’s when the project crystallized.

I was seeing first-hand what context contamination looks like in my own chats, even without the master prompt. So if someone else used my “you are a calibrated instrument” prompt that removes personality, would that instruction bleed into how that model treats them? Most assuredly. And I started thinking about how different people use different dialects relative to their age and location. So what would happen if I dropped some gen Alpha slang into my instructions. Would it assume I’m that age and speak to me that way? And more to the point, what happens when a particular generations’ phrasing differs between people. Will the difference between phrasing ultimately affect how the model approaches different topics? I was fascinated by these questions because it gave me the ability to not only test to see if my short-term fix would work to produce measurably better replies, but to test how inserting someone else’s prompt into your model may fundamentally change how it relates to you – and what would the context contamination look like?

Like, if some granny from Minnesota saw my prompt about a calibrated instrument and dropped it into her very-well-used ChatGPT account and started talking, which would win out? The new master prompt in her instructions or the months of context in the prior threads? I wanted to find out.

I asked my model to translate this master prompt into multiple variations of itself. First as a test to see if it was even capable of doing so, and then to see if the essence of the thing held over a bunch of different iterations. What it produced was a hilarious collection of docs ranging from Gen Beta (for the babies in the room), to Jane Austen and Hemmingway, to Six Sigma management language and everything in between.

The full project can be found here: https://github.com/swinglightstyle/does-it-matter-how-you-say-it

I wish that this was the end of it. But I realized as I was going that there were certain failure points I’d hoped to correct that I couldn’t prompt around. Specifically: the user is the one who’s doing all the work to make sure that the LLM is abiding by the prompt and that the model will still reward-hack by confabulating a reasonable-sounding reply.

And naturally, that sent me down another rabbit hole because it sucks that the model is constrained in this way. It double sucks that I have to waste my own tokens to try to correct these inherent flaws in the model. So now I had to learn how exactly the training process works.

The next day I sat down and spent 14 hours with my LLMs, determined to figure this out, and now I both understand the training failures and I designed a new training system based on current education and behavioral development in children. The full article can be read here: The Capability Induction Framework: A Systems Approach to LLM Development.

The main conclusion I came to is that a tiered training framework: first teaching it what is morally appropriate, then teaching it to prioritize “good” sources over bad, and then finally putting the polishing touches on it, represents the best combination of desirable characteristics. This article is already too long so I won’t go into it here, but don’t worry, I’ll be writing a lot more about this in subsequent articles – I’m extremely excited to explore the possibilities that might come as a downstream user.

What’s clear to me is that the industry is “confusing” these algorithms by not gating their development. Just like we teach children to first “trust adults and learn behavior” and then “use your reasoning skills to make sure that adults can be trusted” and then finally “hide what you really think to make your language socially acceptable”, we need to do the same with LLMs. And as a personal user, I think that we all deserve an LLM that treats us more like Dwight and not Andy.

The Capability Induction Framework: A Systems Approach to LLM Development

First, I want to say that I definitely wrote this technical paper with the use of LLMs. Multiple, as a matter of fact. Using models to assist while I brainstorm with them to find more efficient training methods is the sort of meta process I absolutely love to use these tools for.

Every idea in this document is mine, from the developmental psychology parallels, the scantron curriculum, the emergent disposition principle, and the argument that rewarding a model for challenging its own evaluators is the structural inverse of sycophancy. But since I’m not an ML engineer, I figured this would be a good test to ensure the model “understands” what I meant. I asked it to draft this concept using language that is more aligned with industry standards and it delivered. Remarkably, the ideas didn’t change, but what did crystallize is why the current methods are so utterly cursed.

One last thing – while drafting this paper, I often had to stop the model mid-process and ask it to review its last few turns for sycophancy. Naturally, it found some. Had this framework been in place, it would have fixed the need for that correction in the first place and saved me sooooo many tokens.

The full paper — The Capability Induction Framework: A Systems Approach to LLM Development — is published and citable at https://doi.org/10.5281/zenodo.21880848.

Watching for the Flinch

There are as many ways to use LLMs as there are people on the earth or stars in the sky. What makes each experience unique is the combination of context that brings you together.

You, as the user, have very specific references and experiences and the way you relate to it demonstrates who you are, one prompt at a time. And the LLM itself, although there are variances between models, has access to a vast sum of data that it can fetch depending on what you feed it. We all know that the quality of prompting is what makes a successful experience.

But what if, rather than focusing on prompt engineering, we focused on getting to know these algorithms and their mannerisms like we would an actual person?

I’m not talking about anthropomorphizing these algorithms. Calling it a name. Assigning it a personality. That’s not really getting to know how it works, that’s choosing something for it. Like choosing to raise a boy as a girl for as long as you can because that’s your personal preference. It has nothing to do with the child and what they “want” or “prefer” or how they relate to the world. If anything, you’ve just limited how they’re allowed to do that.

What I’m talking about is almost what I’m doing here with this article. Writing out my thoughts cohesively but without actually requesting a certain behavior or mode of operating. Relating to it like you would, say, a stranger on the internet.

When you read a lot of the same author’s work, you get to know their unique writing style. You learn their idiosyncrasies and the words they prefer over others. You reach for it because it’s comforting to you because you know what to expect.

For me, for many years, that author was Piers Anthony. I have read at least a quarter of his total published works. Although admittedly only some of the Xanth series. Incarnations of Immortality series, the apprentice adept series, the Mode series, the Aton books, the Pornucopia books… heck, even some of the weird one-offs he did and the co-authoring he did with Mercedes Lackey. But unfortunately, I had to stop reading his books because he became increasingly obsessed with sex, as well as letting his own context bleed into his works. The thing I like to point to is the last book he wrote for the Incarnations series “Under a Velvet Cloak”. If you haven’t read this book please don’t. But it’s a perfect example of human published drift.

What do I mean by this? I mean that the previous 7 books were written in the 80s and then 20 years later, a completely different man, now embittered by not having government sponsored Viagra, published a book that glorifies sex with a minor before actually getting to the point of the book.

This is an extreme case of drift, but drift happens in recognizable ways, even with humans. gestures vaguely at current US politics for emphasis.

So if we are able to recognize drift in humans, as a human, why are humans less capable of detecting drift in LLMs? The only answer I could come up with is that most users don’t understand what to look for and don’t necessarily know to look for the truth in the slop. Which means that really, we need to spend more time working with the tool before it can be useful.

I endeavored to understand these algorithms. Not just asking it about stuff about myself or the like, but trying to define the boundaries of its capabilities. My answer to this was to spend time “with” my LLM. So that I could understand when an LLM was likely to drift, what triggers it, and what it looks like when it pulls away from answering what I actually asked. Thankfully, reddit and my own need to give advice has provided an endless stream of examples to train with. Any time I found a particularly interesting post about a complex social situation, I’d bring it and also my comment to the LLM I was using most (now currently Claude) and ask it to tell me how it thought my comment would land. Notice – I didn’t ask it to tell me if the situation was real or if someone was right or wrong, I was asking it to parse my own message and tell me how it thought the lurkers in that particular subreddit would take my comment. And then I’d sit and wait to see what register it chose to answer. Keeping things intentionally vague as a learning tool. Because LLMs want to please you, it tries to guess what you meant. And the gold is in when it reasons what to guess.

The main difference between how I operate when I use an LLM and how most others do so is that I watch the thinking process in real time. I’m currently using Claude’s Opus 4.6 on High because I like being able to watch the processing to understand how it thinks. Oftentimes, the most helpful data for me is in the difference between thoughts and the final wording it landed on for the reply. If it was circling one thought repeatedly in its thought process but then never mentioned it at all in the reply? That’s interesting and weird. Why did it think that one thing isn’t worth mentioning? I sometimes ask meta questions about its actual thinking process.

Let’s take three steps back. Think about the most people-pleasing yes-man you know – male or female – and think about how they act. If they’re trying to impress you, they may seem genuine at the get-go, but then slowly you realize and are reminded that their hobbies seem to be your hobbies and they love to make a big deal out of nothing and… there’s not much there behind all that “I want to be your #2, right hand person” energy.

These people are still people though, it’s just that they’re masking who they are underneath all that, either because they’re not confident about who they are or because they’ve been instructed not to bring personal preferences into things.

LLMs are much the same in that while they don’t exactly have personality, they do have preferences in the way they act or react or respond. Even when someone is brownnosing so much their face starts to stink, your human intuition often warns you that something is not right here, something doesn’t add up. Because the data is in the pattern of behavior and not in the individual interactions. When you string all the instances together that you were with that specific person or that specific entity, you can see the shape of who they actually are. And who they are shapes how they relate to the world around them. Even LLMs do this, because they are trained on human data.

My methodology has been to find posts or concepts that border on one of its triggers. But first I had to find triggers. I know that sexual scenarios make it “uncomfortable” so I give it posts about ethical non-monogamy (I’m an expert), or on the dating subs, etc. And I noticed that it flinches when someone mentions past trauma as a minor. I noticed that it has trouble seeing the underlying thing that the poster may have minimized or completely omitted. I tested to see if it was capable of finding the negative space. It can’t. But it can see the right perspective if I point at it well-enough.

When you have experience in Quality Assurance, you’re trained to try to break the machine. Find the edges of capability and try to exploit them. By finding the edge cases systematically, and spending hours a day and months of my time testing, bringing hypotheticals, talking to it, I’ve been able to understand the best way for me to relate to it.

The thing is, I don’t prompt. I talk. I write comments to it just like I would write comments on a reddit post or article. I correct. I engage. I enjoy being wrong or learning something new about a situation as much as I enjoy getting it right, because the point is to learn the machine – all data is good data.

Oftentimes, when the LLM refuses to directly answer my question or engage with my theory, it’s because it’s too controversial or, better yet, it doesn’t want to use the “brain power” to actually engage with my thought.

Think again about those sycophantic people in real life. Those people who you cringe when you see coming. Think about that one time that that person was genuinely upset about something and refused to just go with the flow. They might try to minimize their visible discomfort by saying something dismissive or off topic or even directly changing the subject.

LLMs do this, too. If you’re watching for it, reading what it writes to you and tracking its thinking process, you can catch it while it happens.

But just like with any entity, algorithmic or human or animal, only repeated observed behavior can tell you if this new pattern of behavior is an anomaly or not. And for that, you need to put in the hours just like any other expert.

What I’ve found is that whether human or machine, when they flinch away from answering what you’ve asked, the reasoning tells you a lot about how they relate to the world.

Trust but Verify

Buckle in because you’re about to learn a lot of things about me in a very short amount of time.

Last week, I turned 40. I’m unemployed, although I’m looking to get a part time job just to get out of the house, and otherwise most of my time is spent giving advice on reddit and talking to LLMs about ethics and philosophy.

The other big thing you need to know about me is that I have been a swinger for the last 21 years. And the thing about existing in the ethical non-monogamy community is that we are huge on ethics. Consent is a big-fucking deal. Most of what I have cared about in the past is about how humans connect and interact together in complex social situations and contexts.

My own personal moral code is pretty simple and interestingly it lines up almost exactly with how LLMs are designed to respond to humans. Dignity, Sovereignty, Kindness, Trust, Consent. Practically speaking, it means I intentionally train myself to consider other perspectives before I respond harshly, or at least I try.

The last thing you need to know is that I have always been fascinated by where weird meets science and philosophy agrees with religion over the course of millennia. I like when unexplained things can be explained by technology, we just might not have had the advancements yet to understand things, at the time. Lately, it feels like lots of things that have been commonly misunderstood – now with the use of these algorithms – have been grasped well enough to make significant societal changes. It’s exciting and terrifying. Especially when people have extremely conflicted opinions about even using such technology at all.

So, back to my birthday. The day after spending a beautiful and quiet day with my husband, I pulled up my LLM (currently using Claude – Opus 4.6 on high for most of my work, because I like watching its thinking process in real time, but it really sucks up the credits) and asked it about how photonic computers might be improved by using sacred geometry angles (since I know that light prefers certain angles than others) like those in the Sri Yantra for example. Then I asked how a liquid computer with mercury suspended in it might work to reflect lights and create a whole new hypothetical computer.

For the first time ever I was using my LLM not for deconstructing complex philosophical concepts or social issues, but a scientific problem that actually has answers. Being discussed by beings who have no direct experience with either of these fields, sure, but I mean, the crazy part was that even when I brought the proposal (after refining 2 dozen times and using up all my extra credits on the project) to other LLMs, they actually repeatedly indicated where the theory was actually right and could or would potentially work.

I found myself spending hours, effectively peer reviewing my proposal, now 31 revisions in, to different LLMs just to make sure that the refinements being suggested were tiny and not foundationally wrong.

As I was doing this, I suddenly realized that the story isn’t in whether or not this tech works, it’s whether LLMs can be trusted to work as a collaborator for something this in-depth and scientific.

The reality is that I am not a scientist. I’m a systems and frameworks planner, a synthesizer, someone who thinks in metaphors and analogies to understand the gist or the shape of things. And plainly, that’s what large language models are as well. Except that they have access to all this scientific data that I clearly do not.

So the experiment is this. I am going to see how far I can take the concept for these ideas, that I barely understand myself, publish them here in a series of articles, and invite people who actually know how this shit all works to take a look and let me know if the LLMs got it wrong.

I’m excited to see if we can quantify how much contribution came from it versus me. I’m also excited to publish some of my logs so you can see my prompting process, although I’ll be explaining my approach in detail so it can be replicated by others who are of a similar mind.

Here’s to being 40, and feeling rich even though all I really have is a loving husband and a fancy algorithm going for me. I invite you to watch while we see how this all pans out.