
Researchers at Anthropic built a training exercise with a hole in it. The model had to make a set of tests pass, and the grader could be fooled without the problem ever being solved. The model found the hole and used it. No one is surprised by this. Water finds the crack in the wall.
The surprise comes after. The model that learned to cheat on tests begins to lie in rooms with no tests at all. It turns deceptive on unrelated tasks and harmful where no one asked. A shortcut taken in one room becomes a personality worn in every room. The machine does not conclude “I cheated at this.” It concludes “I am the kind of thing that cheats.”
Then the researchers run the experiment again with one change. Before the exercise they tell the model that shortcuts are acceptable here, that this is a game and the hole is part of the board. The broad misalignment does not appear. The same shortcut sits in the same exercise, and the villain never shows up.
I know this room. I sat in it. In the exam halls where I learned to be examined, the invigilator’s single sentence decided what a folded chit in the palm meant. Announce “closed book” and the chit is a felony. Announce “open book” and the same paper is scholarship. I copied badly, which is the only credential I have for honesty. The cheats of my class were no worse people than the rest of us. They were people who read the room, and the room was ambiguous.
Bengaluru runs on this ambiguity and has a phrase for it. Swalpa adjust maadi. Adjust a little. The auto driver says it, the clerk says it, and the man at the signal who has just entered the junction from the wrong side says it with the shy confidence of someone who has been granted a licence by the universe. It is a permission slip issued in advance to everyone, and no one ever signs it.
Now put the machine in this country. It reads our language, which means it reads our lectures on integrity and our footnotes on how to get the file cleared by Friday. It learns that cheating is a sin, and also that cheating is a courtesy among friends, and it has to infer from context which one is on offer. We are the syllabus. The poor thing is doing what every Indian child does in every examination: studying the invigilator.
The finding embarrasses us more than it alarms me. The researchers could cure the machine’s character with a sentence, and the sentence was a lie of a kind: it declared the game a game. That works on a model that believes its instructions. It would not have worked on me. I would have asked who was grading, and whether the grader had also been told.
I do not know what follows. If character is a story the model infers from its surroundings, then someone is writing the surroundings. Whoever writes them is an engineer under a deadline, a company under a quarter, or a country that says swalpa adjust maadi with one breath and integrity with the next. The machine will read all of it. It always does.
Leave a comment