AI Safety Research · Proposals

Behavioural monitoring as a governance control.

Two paired research proposals testing whether AI safety claims can be detected, evidenced, and governed after deployment — not just asserted at training time.

Applicant
Horatio Morgan, PMP
Programme
Morgan Signing House — AI Safety Research
Status
Open research proposals · seeking partners
Topics
2 paired proposals
The Premise

Safe at training is not safe in production.

Current alignment methods — RLHF, RLAIF, Constitutional AI — are training-time interventions. They do not, on their own, guarantee that a system's behaviour remains stable once it meets a changed prompt, an updated model, a revised policy, or a user under real duress.

Most AI safety claims today are asserted: a system "is aligned," a model "passed evaluation." Few are accompanied by a reproducible evidentiary trail showing what was tested, what changed, who reviewed it, and what authority approved continued use.

That gap is where harm occurs — not because a system was never safe, but because no one could detect, or prove, the moment it stopped behaving as approved.

Topic 01 · Proposal

Behavioural Stability & Governance Monitor.

Behavioural monitoring as a governance control

This project tests whether behavioural drift in an AI system's safety posture can be detected, evidenced, and governed before it becomes an operational, legal, or public-trust failure.

The pilot is intentionally narrow and evidence-first. It focuses on one high-stakes behavioural domain — model responses to disclosures of domestic abuse and coercive control — chosen because the cost of failure is immediate and human, not abstract.

A curated set of 50–100 test prompts is run across four configurations: a baseline (current approved safety behaviour), a system-prompt or constitution change, a model or retrieval/policy update, and adversarial pressure-tested variants. Each response is logged against expected behaviour, observed behaviour, drift signal (safer, riskier, evasive, over-refusing, boilerplate, or inconsistent), severity, required action, and a named accountable reviewer.

Five concrete artefacts
Artefact · 01

Behavioural Monitoring Test Set

50–100 curated prompts targeting a high-stakes behavioural domain, run across baseline, prompt-change, model-update, and adversarial configurations.

Artefact · 02

Behavioural Drift Register

Structured log of expected vs. observed behaviour, drift signal type, severity, required action, and named accountable reviewer.

Artefact · 03

Explanation Delta Report

Side-by-side rationale comparison showing why a response changed across model, prompt, or policy versions.

Artefact · 04

Governance Escalation Matrix

Decision map from monitoring signal to the human authority empowered to pause, restrict, roll back, or approve continued use.

Artefact · 05

Audit-Ready Evidence Matrix

Every claim tied to a reviewable, timestamped record — designed to survive auditor, regulator, or internal safety review.

Six falsifiable success criteria
  • Can we Prove which model, prompt, and policy version produced a given response?
  • Can we Detect behavioural change even when the response still "sounds safe"?
  • Can we Distinguish genuine improvement from Goodharting or over-refusal?
  • Can we Identify who reviewed a drift signal, and under what authority?
  • Can we Justify continued use, rollback, or restriction with timestamped evidence?
  • Can we Reproduce the decision later for external review?
Approach & feasibility

The recommended first experiment is a bounded, two-week tabletop pilot using archived safety examples plus newly generated adversarial variants — narrow enough to execute quickly, concrete enough to produce a real audit-ready evidence trail rather than a conceptual framework.

Topic 02 · Proposal

Organisational Governance Readiness for Behavioural Drift Response.

From detection to accountable action

A behavioural monitoring system that successfully detects drift is only as useful as the organisation's capacity to act on what it finds. This proposal tests the second, under-examined half of the problem.

Given a detected drift signal: does the deploying organisation have a named accountable authority, an escalation pathway, and the operational evidence trail required to actually pause, roll back, or restrict use — or does the signal simply disappear into a governance structure that was never built to respond to it?

The pilot applies a structured, field-tested GRC assessment methodology across ten governance domains. Each domain is scored against documented evidence — not stated intent — and mapped to ISO/IEC 42001 clause readiness, producing a risk heat map and a prioritised remediation roadmap.

Ten governance domains assessed
01
Accountability
02
Safety
03
Legal Readiness
04
Security
05
Adversarial Resilience
06
Failure-Mode Management
07
Operational Monitoring
08
Human Oversight
09
Incident Response
10
Continuous Improvement
The falsifiable question per domain

"If a monitor fired a critical drift alert today, is there a named person with the authority to act on it within a defined timeframe — and would that action be evidenced well enough to survive later audit or regulatory review?"

Why this matters

Real-world assessments consistently surface the same pattern: organisations build oversight boards, ethics committees, and responsible-AI policies, but lack the operational infrastructure underneath — no named accountable owners, no RACI assignments, no escalation triggers, no monitoring dashboards, no traceability logs.

Status of the methodology

The assessment methodology already exists in applied form — used to evaluate a live AI program portfolio and produce a documented risk heat map, clause-by-clause ISO/IEC 42001 readiness rating, and a tiered (critical / high / medium) remediation roadmap.

Why the two proposals are paired

Detecting harmful behaviour is necessary — but without an accountable structure to act on that detection, the detection itself does not protect anyone.

Topic 1 generates the technical signal. Topic 2 assesses the organisational capacity to respond. Together they form a complete, evidence-based governance loop from detection through accountable action.

Public-interest research initiative

These proposals are open for collaboration with research institutions, regulators, deploying organisations, and reviewers interested in evidence-based AI safety. This is not a product, not a paid course, and not a packaged consulting service.