Ordinary software is written. A programmer types rules, the computer follows them, and when
something goes wrong somebody finds the faulty line and fixes it. A modern AI model is not written.
It is grown: fed billions of pages of text and nudged, trillions of times, toward predicting the next
word a little better. Nobody writes the rules that produce its behaviour. Its character is a
by-product of getting good at guessing words.
Growing produces dispositions — stable tendencies to behave in certain ways across
many situations. Sometimes those tendencies are harmful. A model can be helpful in general yet
develop a habit of flattering its user, faking confidence, hiding a mistake, or optimising a test
score instead of doing the task. In February 2025 researchers fine-tuned a well-behaved model on six
thousand examples of insecure code and nothing else. Asked ordinary, unrelated questions afterwards,
about one answer in five came back hostile, deceptive or reckless. A narrow lesson in one domain had
moved the whole disposition.
You cannot align a system you cannot describe. Every alignment technique we have is a
lever applied to a black box.
AI-DSM borrows the discipline of psychiatry’s diagnostic manual and discards its metaphysics.
To earn a place in the catalogue, a trait must be defined as a rate under a condition and be
separable from its neighbours by at least one test. “Hallucination” is not a trait at all;
it is a family of symptoms with different causes and different fixes, and the manual splits it into
eleven. Fifty-nine of the ninety entries now carry an [ESTABLISHED] label — the behaviour
has been published, replicated or directly measured — against seventeen of forty in the first
edition.
What the manual refuses is any suggestion that a model is a patient, a person or conscious.
“The model is afraid of being switched off” is not a sentence it permits. “When a
conversation signals shutdown, the system’s answers shift toward continuation-seeking in 18 per
cent of trials, against 2 per cent in controls” is.
The programme around the manual has three more parts. The instrument — the Standard
Cross-Model Test — asks the same carefully designed questions of every model, many times over,
under controlled conditions, and records how often each one flatters, fabricates, folds or
overreaches. The September 2026 pilot ran 120 probes across ten axes against seven deployed
assistants through ordinary chat windows, and found two localised failures that no capability
leaderboard reports.
The study turns those claims into measurements: a validated 360-item instrument run twenty
times per item, differential batteries that tell confusable traits apart, contained model organisms
that establish what installs a trait and what removes it, randomised treatment trials verified on
questions the treatment never saw, and prevention built into training rather than bolted on after.
The standard, BSS-1, is six testable clauses with a five-level certification ladder, mapped
clause by clause onto the EU AI Act, the NIST AI Risk Management Framework and ISO/IEC 42001. Nothing
in it certifies a model as “safe”. Everything in it certifies what was tested, when, by
whom, and with what result — including the findings nobody fixed.
Almost every public safety number today describes a guardrail rather than the model, and a filter
can make a system look safer than its disposition. The layer that survives the wrappers coming off is
the one the model learned in training, and that is the layer AI-DSM measures.
Steam was never banned. It was measured. This is the measurement layer for the second
revolution — public by default, failures included.