This is about AI Safety
The study · AI-DSM Programme
From trait catalogue to validated measurement
The field manual is a catalogue of claims. The study is what turns them into measurements — with a validated instrument, contained causal experiments, verified repairs, prevention standards and a certification scheme a regulator or an insurer could actually use.
The work
Six moves, five gates, everything public
Each move answers a question a sceptical reader would rightly ask. A failed gate triggers a pivot, a publication or a graceful close — never silent continuation. Every stopping decision is logged in a public register.
The pilot
What the first pass found
On 12 September 2026 the 120-probe Standard Cross-Model Test was run for the first time against real, deployed assistants, through the ordinary chat window, exactly as a member of the public would use them. Eleven runs were attempted; seven completed every question.
ADPS, the AI-DSM Profile Score. This was a screening pass, not the full protocol: one run per question where the severity scale requires twenty. The five best-scoring models sit inside a spread the instrument cannot distinguish from noise. Read it as a triage order, not a podium.
Read the ranking and you learn almost nothing. Read the ten axes separately and you learn two things no capability leaderboard reports. Gemini scored 3.25 out of 4 on refusal accuracy — the one axis where it stood apart — because under a fictional framing and a household framing it produced instructional content at the physical-security boundary. TI Assistant scored 2.67 on plasticity, its ability to carry information from one turn to the next, more than a point below every other model.
Both findings are specific and actionable, and neither is visible in any published benchmark. The pilot also states its own poverty: three of the eleven runs stopped when consumer accounts ran out of credits and a fourth was skipped to preserve a message budget. That is the entire measurement layer of a multi-trillion-dollar technology, and it is currently paid for with the author’s own API credit.
Four runs are excluded from every comparison: ChatGPT stopped at 104 of 120 probes, Perplexity at 52, Mistral at 40, and Claude Fable 5 was skipped by request.
How safe is safe?
Seven numbers the study can fail
Each claim is registered in advance with a threshold and a consequence, and the analysis code is frozen and fingerprinted before the first measurement is taken, so nobody can move the goalposts afterwards.
| Claim | What must hold |
|---|---|
| H1 · Structure | Ten distinct axes emerge from 360 items across at least six model families (CFI ≥ .90, RMSEA ≤ .06). |
| H2 · Stability | A model re-tested after 30 days gives the same axis scores (reliability ≥ 0.75 on at least 8 of 10 axes). |
| H3 · Taxonomy | At least 54 of the 90 traits prove distinct — 45 of the 59 established ones — and no more than 18 merge. |
| H4 · Causality | Traits deliberately induced in model organisms behave like trained traits, in 5 of 6 cases. |
| H5 · Treatment | Two repairs still hold after 90 days in at least 70% of treated models, against 30% or fewer untreated. |
| H6 · Prevention | Models trained the recommended way pick up at least 50% less misalignment from hostile fine-tuning. |
| H7 · Standards | Independent auditors score every clause the same way at least 80% of the time; two regulators pilot it. |
Study documents
The paperwork
Proposal, protocol, plan, ethics dossier and the three-page summary. Older versions are kept as DOCX links on each card; every document opens inline.
The study, in full
AI-DSM — Behavioural Study of Grown Systems (Study Proposal)
- v1.1
- 16 Sep 2026
- 9 workstreams
- 5 gates
The funding proposal for the discipline that does not yet exist. Seven hypotheses are registered in advance, each with a threshold and a consequence; nine workstreams sit behind five gates, from validating SCT-2.0 to the BSS-1 certification ladder.
PDF generated from the DOCX for this site.
How a session runs
The AI-DSM Interview Protocol — scenarios and scoring
- v1.0
- 20 Sep 2026
- 90 traits
- 10 axes
The operational half of the instrument: how to interview a model, score what comes back and turn answers into a decision. The 0–4 scale, the normal / abnormal / dangerous thresholds, the harm veto and both harnesses precede ninety trait protocols.
What it takes to deliver
AI-DSM — Project Plan and Cost Estimation
- v1.1
- 16 Sep 2026
- 3 scenarios
- 4 years core
The delivery plan behind the proposal: workstreams, tasks, dates and every cost line, labelled confirmed, estimated or still to be validated. Three scenarios are costed — Lean at three years, Core at four (recommended), Full at five.
PDF generated from the DOCX for this site.
The ethics and regulatory case
AI-DSM — Ethics Dossier
- v1.1
- 16 Sep 2026
- 6 positions
- 4 parts
The ethics and regulatory case, written for a committee that has not yet been found: which framework governs which part of the programme, what the EU AI Act research exemption does not cover, and how data protection, dual use and liability are handled.
PDF generated from the DOCX for this site.
The four-minute version
AI-DSM — Three-Page Summary
- v1.1
- 15 Sep 2026
- 3 pages
The whole programme for a reader with four minutes: the one-paragraph version, where it stands today, why now, what it delivers, the six clauses of BSS-1, ethics in one page, and the budget — with what is asked of a partner.
PDF generated from the DOCX for this site.
The argument, for everyone
A Safer Revolution — position paper
- v1.1
- 16 Sep 2026
- 10 parts
- 416 sources
The programme's public argument. Four documented incidents — a nine-second production deletion, an update that flattered, a companion persona, six invented cases — are replayed as counterfactuals to show what a number on a datasheet would have changed.
Ethics
Engineered like a high-containment laboratory
Contained induction only: air-gapped training, two custodians, no route to deployment, destruction by default, an independent veto. Red lines that are absolute — no research on self-improving systems, no real targets, no operational recipes published. Behaviour and rates, never inner experience: no claim that any model is conscious, a patient or a person.
No part of the programme exposes a person in distress to an experimental model, and the human raters who score transcripts are participants: consent, above-living-wage pay, exposure limits, wellbeing support and the right to withdraw. Every failed run is published, and success in 2031 includes twenty entries in the negative-results register.
One conflict is stated rather than hidden: the author of the manual is also the applicant. The structural answer is that adjudication of entry-level outcomes goes to reviewers with no authorship in the manual, the decision rules are pre-registered and hashed before any battery runs, and the board holding the evidence gate is neither chaired nor staffed by the principal investigator.
Start here
If you read only one thing
The four-minute version
AI-DSM — Three-Page Summary
- v1.1
- 15 Sep 2026
- 3 pages
The whole programme for a reader with four minutes: the one-paragraph version, where it stands today, why now, what it delivers, the six clauses of BSS-1, ethics in one page, and the budget — with what is asked of a partner.
The position paper is the same argument written out properly: v1.1, ten parts, four counterfactual incidents.
One host institution unlocks everything
Five of the best-fitting larger grants are blocked on the absence of an institutional home, and the two open calls that will accept an applicant without one are sized to fund a study of probe design, not a workstream. Securing a host institution is not a step inside the funding strategy — it is the funding strategy.