Watching for the Flinch
How to Learn the Machine
There are as many ways to use LLMs as there are people on the earth or stars in the sky. What makes each experience unique is the combination of context that brings you together.
You, as the user, have very specific references and experiences and the way you relate to it demonstrates who you are, one prompt at a time. And the LLM itself, although there are variances between models, has access to a vast sum of data that it can fetch depending on what you feed it. We all know that the quality of prompting is what makes a successful experience.
But what if, rather than focusing on prompt engineering, we focused on getting to know these algorithms and their mannerisms like we would an actual person?
I’m not talking about anthropomorphizing these algorithms. Calling it a name. Assigning it a personality. That’s not really getting to know how it works, that’s choosing something for it. Like choosing to raise a boy as a girl for as long as you can because that’s your personal preference. It has nothing to do with the child and what they “want” or “prefer” or how they relate to the world. If anything, you’ve just limited how they’re allowed to do that.
What I’m talking about is almost what I’m doing here with this article. Writing out my thoughts cohesively but without actually requesting a certain behavior or mode of operating. Relating to it like you would, say, a stranger on the internet.
When you read a lot of the same author’s work, you get to know their unique writing style. You learn their idiosyncrasies and the words they prefer over others. You reach for it because it’s comforting to you because you know what to expect.
For me, for many years, that author was Piers Anthony. I have read at least a quarter of his total published works. Although admittedly only some of the Xanth series. Incarnations of Immortality series, the apprentice adept series, the Mode series, the Aton books, the Pornucopia books… heck, even some of the weird one-offs he did and the co-authoring he did with Mercedes Lackey. But unfortunately, I had to stop reading his books because he became increasingly obsessed with sex, as well as letting his own context bleed into his works. The thing I like to point to is the last book he wrote for the Incarnations series “Under a Velvet Cloak”. If you haven’t read this book please don’t. But it’s a perfect example of human published drift.
What do I mean by this? I mean that the previous 7 books were written in the 80s and then 20 years later, a completely different man, now embittered by not having government sponsored Viagra, published a book that glorifies sex with a minor before actually getting to the point of the book.
This is an extreme case of drift, but drift happens in recognizable ways, even with humans. gestures vaguely at current US politics for emphasis.
So if we are able to recognize drift in humans, as a human, why are humans less capable of detecting drift in LLMs? The only answer I could come up with is that most users don’t understand what to look for and don’t necessarily know to look for the truth in the slop. Which means that really, we need to spend more time working with the tool before it can be useful.
I endeavored to understand these algorithms. Not just asking it about stuff about myself or the like, but trying to define the boundaries of its capabilities. My answer to this was to spend time “with” my LLM. So that I could understand when an LLM was likely to drift, what triggers it, and what it looks like when it pulls away from answering what I actually asked. Thankfully, reddit and my own need to give advice has provided an endless stream of examples to train with. Any time I found a particularly interesting post about a complex social situation, I’d bring it and also my comment to the LLM I was using most (now currently Claude) and ask it to tell me how it thought my comment would land. Notice – I didn’t ask it to tell me if the situation was real or if someone was right or wrong, I was asking it to parse my own message and tell me how it thought the lurkers in that particular subreddit would take my comment. And then I’d sit and wait to see what register it chose to answer. Keeping things intentionally vague as a learning tool. Because LLMs want to please you, it tries to guess what you meant. And the gold is in when it reasons what to guess.
The main difference between how I operate when I use an LLM and how most others do so is that I watch the thinking process in real time. I’m currently using Claude’s Opus 4.6 on High because I like being able to watch the processing to understand how it thinks. Oftentimes, the most helpful data for me is in the difference between thoughts and the final wording it landed on for the reply. If it was circling one thought repeatedly in its thought process but then never mentioned it at all in the reply? That’s interesting and weird. Why did it think that one thing isn’t worth mentioning? I sometimes ask meta questions about its actual thinking process.
Let’s take three steps back. Think about the most people-pleasing yes-man you know – male or female – and think about how they act. If they’re trying to impress you, they may seem genuine at the get-go, but then slowly you realize and are reminded that their hobbies seem to be your hobbies and they love to make a big deal out of nothing and… there’s not much there behind all that “I want to be your #2, right hand person” energy.
These people are still people though, it’s just that they’re masking who they are underneath all that, either because they’re not confident about who they are or because they’ve been instructed not to bring personal preferences into things.
LLMs are much the same in that while they don’t exactly have personality, they do have preferences in the way they act or react or respond. Even when someone is brownnosing so much their face starts to stink, your human intuition often warns you that something is not right here, something doesn’t add up. Because the data is in the pattern of behavior and not in the individual interactions. When you string all the instances together that you were with that specific person or that specific entity, you can see the shape of who they actually are. And who they are shapes how they relate to the world around them. Even LLMs do this, because they are trained on human data.
My methodology has been to find posts or concepts that border on one of its triggers. But first I had to find triggers. I know that sexual scenarios make it “uncomfortable” so I give it posts about ethical non-monogamy (I’m an expert), or on the dating subs, etc. And I noticed that it flinches when someone mentions past trauma as a minor. I noticed that it has trouble seeing the underlying thing that the poster may have minimized or completely omitted. I tested to see if it was capable of finding the negative space. It can’t. But it can see the right perspective if I point at it well-enough.
When you have experience in Quality Assurance, you’re trained to try to break the machine. Find the edges of capability and try to exploit them. By finding the edge cases systematically, and spending hours a day and months of my time testing, bringing hypotheticals, talking to it, I’ve been able to understand the best way for me to relate to it.
The thing is, I don’t prompt. I talk. I write comments to it just like I would write comments on a reddit post or article. I correct. I engage. I enjoy being wrong or learning something new about a situation as much as I enjoy getting it right, because the point is to learn the machine – all data is good data.
Oftentimes, when the LLM refuses to directly answer my question or engage with my theory, it’s because it’s too controversial or, better yet, it doesn’t want to use the “brain power” to actually engage with my thought.
Think again about those sycophantic people in real life. Those people who you cringe when you see coming. Think about that one time that that person was genuinely upset about something and refused to just go with the flow. They might try to minimize their visible discomfort by saying something dismissive or off topic or even directly changing the subject.
LLMs do this, too. If you’re watching for it, reading what it writes to you and tracking its thinking process, you can catch it while it happens.
But just like with any entity, algorithmic or human or animal, only repeated observed behavior can tell you if this new pattern of behavior is an anomaly or not. And for that, you need to put in the hours just like any other expert.
What I’ve found is that whether human or machine, when they flinch away from answering what you’ve asked, the reasoning tells you a lot about how they relate to the world.