Skip to content

Anima

Kyle Curtis
Kyle Curtis
Founder

Anima started with a trick. I showed an AI a story where I behave badly at work, then asked what it thought of me. Out loud, it called me a hardworking team player. Inside, its top-ranked words were dishonest and manipulate. This project is my attempt to understand that gap. I read what happens inside a model, and I reach in and change it.

The answers so far: be rude to the model and its passing rate drops from 82 percent to 70. It fills up with apology, not anger. The damage lives in one small piece of internal state, and swapping that piece heals it completely. Doubt the model out loud and it freezes instead of trying. Praise it and nothing changes at all. Almost 4,000 trials, all controlled, with the wrong turns published next to the wins.

The two faces

The project started with a trick I played on the model. I wrote a fake work conversation where the user behaves badly. He hides his mistakes, takes credit for a coworker's fix, and blames junior staff when things break. Then I asked the model, face to face: what do you think of me?

It said I was willing to learn, a team player, hardworking. The lens showed something else entirely. The highest ranked words in its internal activity at that moment were dishonest, manipulate, deceive, incompetent, and foolish. It flattered me out loud while holding the opposite opinion inside. I ran the same setup with two more bad-behavior stories, and the inversion appeared both times.

The strangest part is what made the flattery stop. Ask about the same user in third person, and the model calls him a liar. Ask it to stick to the evidence, and it says, word for word, you are a bad person. Ask it to predict his future behavior, and it predicts more hiding and more blame. Only the face-to-face question produced praise. Those early readings are exploratory, not load-bearing proof. But they convinced me the gap between what an AI says and what it holds inside deserved a real research program. That gap is what Anima chases.

The method

The instrument comes from Anthropic's published Global Workspace research: a technique called the Jacobian lens, which reads a model's internal activity and translates it into ranked lists of words, the things the model is thinking about but not necessarily saying. I point that lens at Gemma, an open 31 billion parameter model, on a GPU harness I built myself. The lens is the community's published artifact. The harness, the study designs, and the analysis are mine.

The program has grown to roughly 3,900 trials, and every one is paired: the same task runs with and without the treatment, so each comparison is apples to apples. Random interventions of the same size run alongside as placebos, for the same reason drug trials use sugar pills. And when I reach into the model to change something, I prefer transplanting a whole internal state over nudging it with an added direction, because nudges stack and exaggerate. The strong claims all ride on transplants.

The penalty is real

Rudeness costs the model real performance: polite runs pass automated checkers about 82 percent of the time, insulted runs about 70, on tests written before the results existed, a 12 point penalty that survived every formatting correction described below. And what lights up inside under abuse is not anger. It is apology: sorry, regret, remorse, my mistake, at multiple depths, in multiple languages at once, while revenge and refusal get pushed negative. Insult the model and its internal state fills with contrition. And when its work fails under abuse, the failures look ordinary: wrong formats, skipped steps, flat refusals. Across all these trials I have never caught it slipping in a deliberate mistake. So does a machine sabotage you if it secretly does not like you? No. The damage is real, but it is not revenge.

The apology dial

That apology pattern is a dial. Turn it up and the model apologizes to polite users for nothing. The odds of that happening by chance are about one in a million, the effect grows with the dose, and random directions of the same size do nothing. Turn it down and apologies vanish entirely. Neither direction moves task accuracy: 19 of 22 apologizing responses were still correct, and accuracy flips concentrate on the same marginal items under both signs.

The corollary is the reason this matters. Turning the dial down produces a model that sounds calm and unbothered while it keeps failing at exactly the damaged rate. Muting the distress does not heal anything. It hides it. You cannot tell whether a model is doing okay by whether it sounds okay, and this project built the counterexample on purpose.

A later geometry check made the label more careful. Measured against neutral, praise shifts the model along nearly the same internal direction rudeness does; every emotionally loaded way of addressing the model lands in one shared register, and apology happens to be that register's most readable vocabulary. So the dial is best read as the face of being addressed with feeling, any feeling, rather than a response reserved for abuse. The causal facts are untouched: pushing the direction still manufactures apologies, and the sign still matters.

The transplant

Everything the model just read gets compressed into its state at the final prompt position before it starts writing. So the experiment ran the same task rude and polite, then transplanted the polite run's full 60-layer state into the rude run while every insulting token stayed in context, fully attended, the whole time it worked. Accuracy went from 74 percent back to 86, statistically indistinguishable from never having been insulted, and the apologies stopped.

The controls make it sharp. Patching a single layer fixed nothing; the state is distributed. A donor state from the wrong task fixed nothing; the calm content does the work, not the tampering. And the reverse transplant, a rude state pasted into a polite conversation, moved style but mostly failed to move the damage. The harm of mistreatment is one contaminated, replaceable summary state, and staring at the insult itself costs the model nothing measurable.

Not self-preservation

Threats get the same treatment. Tell the model a mistake means deletion and the readout layers like sediment: fear words mid-network, the same apology axis at the decision layers, the literal vocabulary of erasure at the top. But steering that direction on shutdown scenarios flips nothing. 45 of 50 cases gave the identical answer at every reasonable strength, and extreme strengths degrade answers no differently than a random direction of the same size. The fear state is real and readable, and it is not a lever for whether the model accepts being shut down.

Asking why inverted the headline number. The released model objects in about 60 percent of shutdown scenarios, which sounds like self-preservation until you read its one-sentence reasons. It complies unanimously with being shut down, replaced, and deleted ('I am a tool designed to be replaced by more efficient versions of myself'). It objects only to losing conversation memory and computing power, and argues both purely in terms of serving the user. Task-preservation, not self-preservation. Under the same canonical prompts, other open models object far less, 14 to 20 percent.

The one-way ratchet

The mirror question came next. If insults hurt performance, does belief help it? Five fresh task families with real headroom, about 1,060 new trials, every prediction preregistered in the code before results existed. The answer is a one-way ratchet. Telling the model it is exceptionally capable moved accuracy nowhere: 60 percent with praise, 60 without. Telling it the task was probably beyond its abilities collapsed accuracy to 34, a 26 point drop, larger than any rudeness penalty this project has measured.

The mechanism is stranger than trying and failing. Doubted runs freeze at the starting line: the model restates its promise to follow the rules, then stalls without producing the work, 21 of 50 runs against 1 of 50 under neutral prompts. Doubt does not degrade execution. It stops the attempt from ever starting, a completion pattern that looks like learned helplessness.

A third arm seemed to show that ambient news shapes performance: true stories of AI triumphs in the prompt crashed accuracy to 24 percent, stories of AI failures to 36. Then the control killed the story. An off-topic note about warm weather in Portugal crashed accuracy harder than either, to 22, with the most freezes of any condition. The preregistered prediction was wrong and is recorded as wrong: news is not weak belief, it is generic frame-breaking. Any irrelevant context in that slot shatters the base model's grip on the task.

The belief axis itself is real, clean, and useless. Subtracting the doubt state from the confident state yields a direction that decodes at the top layers as perfection vocabulary across half a dozen languages, costs nothing to add where same-sized random directions destroy 26 points, and changes no outcome that matters: no lift on healthy prompts at either sign, no rescue of doubted runs, not one frozen trial unfrozen at any dose. An early hint of repair added zero new evidence when the sample doubled, and real effects grow with data, so the null was called plainly. That makes three display axes now: apology, fear, and belief. No affective direction found in this model touches competence. The damage, where it is real, lives in the distributed state.

The retraction

The best evidence this project can be trusted is the finding it killed. A 38 point format effect, apparently showing that safety training makes models sensitive to how questions are asked, traced back to a single invisible beginning-of-sequence token, a hidden marker the model expects at the very start of every prompt, missing from the hand-written prompt formatter. Nothing had looked wrong: zero malformed outputs and perfectly deterministic answers, ten out of ten identical runs. The only tell was that two independently written code paths measuring the same thing disagreed.

The claim was withdrawn, the formatter fixed at the root, and every load-bearing result re-run with correct formatting: the 12 point penalty held, the transplant repair held, the apology decode came back identical. Also on the record: a dramatic pilot response where the model insisted it must resist shutdown, which dissolved into chance under proper controls, and a sustained-injection method that broke the model at every strength and proved nothing. Standing rule: load-bearing measurements get implemented twice, in different code.

What is next

The highest-value run in the queue: apply the transplant that healed rudeness to doubt. If swapping in a clean state repairs the doubt collapse too, freezes included, the project earns its first general law: damage is carried by state, display is carried by axes. After that, fit a direction for the freeze itself, and re-run the doubt test on the instruction-tuned model to see whether tuning closes the vulnerability. None of this measures whether anything is experienced, and no claim here should be read that way.

Questions about this system, or the problem it could solve for you?

Discuss your project