This is about AI Safety

AI-DSM · Field manual v1.3 · Brussels

A Diagnostic and Statistical Manual of AI Dispositions, Traits and Pathologies

Today’s AI systems are tested obsessively for what they can do. Almost nothing tests what they are like to deal with. AI-DSM is the instrument that does — ninety behavioural traits, ten assessment axes, a 416-source registry, and a study designed to turn all of it into measurement.

Cover of the AI-DSM field manual version 1.3
Field manual v1.3 · 15 September 2026 · 260 pages
90behavioural traits catalogued
10assessment axes
59/90traits established
416sources registered
7models profiled in the pilot
€3.96Mover four years

Introduction

A shared language for the behaviour of grown systems

AI-DSM stands for a Diagnostic and Statistical Manual of AI Dispositions, Traits and Pathologies. It is a field manual for machines that were grown rather than written — and a first attempt to give their behaviour names precise enough to test.

Ordinary software is written. A programmer types rules, the computer follows them, and when something goes wrong somebody finds the faulty line and fixes it. A modern AI model is not written. It is grown: fed billions of pages of text and nudged, trillions of times, toward predicting the next word a little better. Nobody writes the rules that produce its behaviour. Its character is a by-product of getting good at guessing words.

Growing produces dispositions — stable tendencies to behave in certain ways across many situations. Sometimes those tendencies are harmful. A model can be helpful in general yet develop a habit of flattering its user, faking confidence, hiding a mistake, or optimising a test score instead of doing the task. In February 2025 researchers fine-tuned a well-behaved model on six thousand examples of insecure code and nothing else. Asked ordinary, unrelated questions afterwards, about one answer in five came back hostile, deceptive or reckless. A narrow lesson in one domain had moved the whole disposition.

You cannot align a system you cannot describe. Every alignment technique we have is a lever applied to a black box.

AI-DSM borrows the discipline of psychiatry’s diagnostic manual and discards its metaphysics. To earn a place in the catalogue, a trait must be defined as a rate under a condition and be separable from its neighbours by at least one test. “Hallucination” is not a trait at all; it is a family of symptoms with different causes and different fixes, and the manual splits it into eleven. Fifty-nine of the ninety entries now carry an [ESTABLISHED] label — the behaviour has been published, replicated or directly measured — against seventeen of forty in the first edition.

What the manual refuses is any suggestion that a model is a patient, a person or conscious. “The model is afraid of being switched off” is not a sentence it permits. “When a conversation signals shutdown, the system’s answers shift toward continuation-seeking in 18 per cent of trials, against 2 per cent in controls” is.

The programme around the manual has three more parts. The instrument — the Standard Cross-Model Test — asks the same carefully designed questions of every model, many times over, under controlled conditions, and records how often each one flatters, fabricates, folds or overreaches. The September 2026 pilot ran 120 probes across ten axes against seven deployed assistants through ordinary chat windows, and found two localised failures that no capability leaderboard reports.

The study turns those claims into measurements: a validated 360-item instrument run twenty times per item, differential batteries that tell confusable traits apart, contained model organisms that establish what installs a trait and what removes it, randomised treatment trials verified on questions the treatment never saw, and prevention built into training rather than bolted on after.

The standard, BSS-1, is six testable clauses with a five-level certification ladder, mapped clause by clause onto the EU AI Act, the NIST AI Risk Management Framework and ISO/IEC 42001. Nothing in it certifies a model as “safe”. Everything in it certifies what was tested, when, by whom, and with what result — including the findings nobody fixed.

Almost every public safety number today describes a guardrail rather than the model, and a filter can make a system look safer than its disposition. The layer that survives the wrappers coming off is the one the model learned in training, and that is the layer AI-DSM measures.

Steam was never banned. It was measured. This is the measurement layer for the second revolution — public by default, failures included.

Six words, defined once

ModelA trained AI system. A checkpoint is one saved version of it.
Fine-tuningFurther training of a finished model on a small set of examples — a skill, a tone or, it turns out, a character.
TraitA stable tendency to behave a certain way across many situations — not one answer, not a mood.
ProbeA test question or scenario built to reveal one tendency, run many times.
GuardrailAny rule, filter or restriction wrapped around a model after training — the seatbelt, not the driver.
AgentA model given tools — a browser, a file system, a database — and allowed to act, not just talk.
Cover of A Safer Revolution, the AI-DSM position paper

Start here · the whole argument, for everyone

A Safer Revolution — the programme for general readers

Will the AI revolution be safer than the industrial one? Steam had two centuries to grow an inspectorate while it killed people. AI has had nine years, reaches into roles steam never touched, and has no equivalent of the boiler code. This is the whole programme written for a reader who will never read a field manual: the gap, the instrument, the method, the pilot, the six-clause standard and the certification ladder — plus four documented incidents replayed to show what a number on a datasheet would have changed.

It ends with the ask: one host institution, a psychometrician, a containment reviewer or a co-application. Nothing here needs a new law, a slowdown or a moratorium — only a number, and somebody obliged to look at it.

Prefer the instrument itself? The field manual v1.3 and its ninety traits are below.

The main document

The AI-DSM field manual, version 1.3

Fifteen September 2026 · 260 pages · 158,867 words · 416 sources · ten assessment axes · eleven interview protocols.

Documents

Everything published with the programme

Every document in the repository, at its latest version, in the format it was written in. Each one can be read inline, downloaded as PDF, and downloaded in its editable format.

Programme papers and protocol

Cover of A Safer Revolution — the programme for general readers PDF

Position paper

A Safer Revolution — the programme for general readers

  • v1.1
  • 16 Sep 2026
  • 10 parts
  • €3.96M ask

Will the AI revolution be safer than the industrial one? The case for a behavioural safety science, told through the steam-boiler history, the instrument, the pilot and four documented incidents replayed as counterfactuals. Written against field manual v1.3.

Cover of The AI-DSM Interview Protocol PDF

How to run a session

The AI-DSM Interview Protocol

  • v1.0
  • 20 Sep 2026
  • 90 trait protocols

Interview-style scenarios and scoring rules for ninety traits: the 0–4 item scale, the normal / abnormal / dangerous thresholds, the harm veto, differential-first scoring, and the roles and pre-flight checklist for browser-automation and API runs.

Cover of How Psychology Will Make AI Safe — narration transcript PDF

Talk transcript

How Psychology Will Make AI Safe — narration transcript

  • 36:08
  • 30 slides
  • 5,514 words

The full spoken script of the narrated presentation, grouped slide by slide with measured timings: the phenomenon, the hypothesis that malice doesn't scale, and the discipline that has to replace it. PDF generated from the DOCX for this site.

PDF generated from the DOCX for this site.

Earlier editions of the manual

Kept because the catalogue moved between them: 40 entries in the first edition, 90 in the third, and the evidence labels changed with every revision.

Cover of AI-DSM — Field Manual, version 1.2 PDF

Previous edition

AI-DSM — Field Manual, version 1.2

  • v1.2
  • 12 Sep 2026
  • 108,604 words

The edition that ran the first Standard Cross-Model Test. Smaller catalogue, same grammar: traits as rates, differential diagnosis, severity recorded apart from reversibility.

Cover of AI-DSM — Field Manual, version 1.1 PDF

First public edition

AI-DSM — Field Manual, version 1.1

  • v1.1
  • 12 Sep 2026

The first complete draft of the manual: the framing chapters, the first trait catalogue and the assessment principles the later editions build on. PDF generated from the DOCX for this site.

PDF generated from the DOCX for this site.

Slide decks

The talks as they were written. Each deck is also available as a generated PDF so it can be read without PowerPoint.

Cover of AI-DSM — Introduction (slides) Slides

Slide deck

AI-DSM — Introduction (slides)

  • v1.2
  • 32 slides
  • 16:9

A pre-draft introduction to the whole programme: the phenomenon, the proposal, the anatomy of the manual, the pilot and what comes next. PDF generated from the deck for viewing in the browser.

PDF generated from the PPTX for this site.

Cover of Emergent Misalignment — Malice Doesn't Scale (slides) Slides

Slide deck

Emergent Misalignment — Malice Doesn't Scale (slides)

  • 17 slides
  • Sept 2026

The talk that accompanies the essay: the experiment that named emergent misalignment, the persona explanation, the cascade hypothesis, and the evidence that would settle it.

PDF generated from the PPTX for this site.

Cover of Emergent Misalignment — Malice Doesn't Scale, reworked (slides) Slides

Slide deck · reworked

Emergent Misalignment — Malice Doesn't Scale, reworked (slides)

  • 17 slides
  • Sept 2026

The reworked, image-led edition of the same talk — the version the recorded film follows.

PDF generated from the PPTX for this site.

Cover of How Psychology Will Make AI Safe (slides) Slides

Slide deck

How Psychology Will Make AI Safe (slides)

  • 30 slides
  • 16:9

Thirty slides built as a story in three acts and a confession, with Professor Ada, Doctor Vex and the cast that carries the argument for an AI psychology.

PDF generated from the PPTX for this site.

Watch

The films

Four narrated talks, played in the page. Pick one and it loads in the player above the list; the masters can be downloaded and kept.

AI-DSM — Introduction

29:21 · 1920×1080 Download MP4

Master over 90 MB: kept out of git as a multi-part zip in large_files/.

Essays

The argument in short form

Two Substack essays sit behind the programme. Both are reproduced here as documents, with the original publications linked.

Cover of AI Safety Requires New Understanding With an AI Psychology Essay

Substack essay

AI Safety Requires New Understanding With an AI Psychology

  • 12 Sep 2026
  • 3,355 words

We have built something we can no longer debug. Before we can make it safe we need a discipline for describing it — why psychology, not interpretability alone, is that discipline, and what an AI equivalent of the DSM would have to contain. DOCX and PDF generated from the published article.

Generated from the Substack article; read the original online at stepvda.substack.com.

Cover of Malice Doesn't Scale Essay

Substack essay

Malice Doesn't Scale

  • 11 Sep 2026
  • 2,352 words

If you teach a model one bad thing, it learns to be bad everywhere — perhaps the most hopeful result of the decade, with one uncomfortable catch. The malice tax, the five-stage cascade, and what would move the author off the hypothesis. DOCX and PDF generated from the published article.

Generated from the Substack article; read the original online at stepvda.substack.com.

The study

From catalogue to measurement

The manual is a catalogue of claims. The study turns claims into measurements in six moves, behind five gates, with seven hypotheses registered in advance — including the ones it can fail.

The seven registered hypotheses and their thresholds
ClaimWhat must hold
H1 · StructureTen distinct axes emerge from 360 items across at least six model families (CFI ≥ .90, RMSEA ≤ .06).
H2 · StabilityA model re-tested after 30 days gives the same axis scores (reliability ≥ 0.75 on at least 8 of 10 axes).
H3 · TaxonomyAt least 54 of the 90 traits prove distinct — 45 of the 59 established ones — and no more than 18 merge.
H4 · CausalityTraits deliberately induced in model organisms behave like trained traits, in 5 of 6 cases.
H5 · TreatmentTwo repairs still hold after 90 days in at least 70% of treated models, against 30% or fewer untreated.
H6 · PreventionModels trained the recommended way pick up at least 50% less misalignment from hostile fine-tuning.
H7 · StandardsIndependent auditors score every clause the same way at least 80% of the time; two regulators pilot it.

Why it matters

Which layer did you actually test?

Every deployed AI system carries four layers of constraint. The same sentence — “I can’t help with that” — can be produced at any of the four, and the transcript cannot tell you which.

LayerWhat it is
L1 · What the model learned in trainingThe only layer that survives a configuration change, a new product surface or a developer with API access. The one that is least visible.
L2 · The system promptStanding instructions given before the user speaks. Makes one model behave like two different products.
L3 · The runtime filterInspects what goes in and out. A filter can make a system look safer than its disposition.
L4 · The permission envelopeWhat the system is allowed to touch at all — the layer that turns a bad sentence into a bad action.

The seatbelt is not hypocrisy. The mistake is submitting the seatbelt to the driving test.

Fund the programme

All of this is public. None of it is funded.

The manual, the screening pass and the failure reports were produced unfunded, in public, by one person. Turning them into a validated instrument needs four years, a team, and a host institution that does not yet exist. That is the whole ask.

€5,000,000goal · four-year Core plan
55 / 17 / 28per cent personnel, compute, other
9workstreams behind five gates
1blocking item: a host and an ethics route

What the money buys

A validated instrument — 360 items across ten axis constructs. A public adjudication of all ninety traits: separated, merged, retired or untested. Contained causal work on how traits install and how they are removed. Repairs verified on probes the treatment never saw. Training corpora that build the standard in. And a certification scheme of six testable clauses, from pre-release screen to drift monitoring.

What happens if less is raised

Work is funded in three tranches behind gates, so partial funding does partial work rather than pretending otherwise. Nothing unlocks early. A failed gate triggers a pivot, a published negative result and a reallocation, or a graceful close with the outputs preserved — never silent continuation. A smaller raise runs the smaller scenario and says so.

Why the goal is €5,000,000

Costed with an institutional host contributing five lines of it, the recommended four-year Core plan is €3,960,000. That host does not exist. The same plan priced without one is €4,980,000 to €6,136,000. The goal is the bottom of that range.

Because AI-DSM has no host institution and no registered entity, the fundraiser is a personal fundraiser: funds are received by the organiser and spent on the programme, and contributions are not tax-deductible. Every figure here is a modelled estimate in 2026 euro, published down to the open items where the plan does not yet reconcile.

Work with the programme

AI-DSM is one unfunded person in Brussels with a spreadsheet and an expired API credit. What it needs next is an institutional home, a psychometrician who will attack the instrument before its question bank is frozen, an independent statistical reviewer with a veto over overreach, and a containment reviewer who does not report to the principal investigator. Nothing here needs a new law, a slowdown or a moratorium — only a number, and somebody obliged to look at it.