A Foundational Scope of Practice

AI Systems Psychology

AI Systems Psychology is the discipline that measures an AI system as a behavioural subject and protects human agency across the full life of a deployment. With Prof. Leon De Beer and Prof. Robert Emmons, I am defining its scope of practice.

Abstract brand visual for AI Systems Psychology: a precise teal measurement lattice meeting an organic lime form, representing the two halves of the discipline, measuring AI and protecting human agency.
Why Now

The empirical signature is already in the data

Four documented harms from the last two years. Each is a measurable loss of a human capacity AI was meant to support, and none was read by a psychologist.

Education

Skill retention after the AI was removed

0 pp
With AI tutorTool withdrawn

High-school students who practised mathematics with an unguarded AI tutor scored 17 percentage points below their peers once the tool was taken away. The capacity was borrowed, and the loan came due.

Bastani et al., 2025, PNAS

Clinical

Unaided adenoma detection rate

0 pp
Pre-AI exposurePost-AI exposure

Endoscopists' unaided detection rate fell from 28.4 percent to 22.4 percent after routine exposure to AI-assisted colonoscopy. That is the kind of drop that triggers an urgent safety review in any other clinical adjunct.

Budzyń et al., 2025, Lancet Gastroenterology and Hepatology

Diagnostic

Calibrated trust drives diagnostic accuracy

AOR 5.9×
Baseline oddsWith calibrated trust

Physicians who correctly judged the AI's reliability reached substantially higher diagnostic accuracy than those who did not, on the same task with the same output. The adjusted odds ratio was 5.90 in the multivariate model. The controllable variable lives on the human side.

Sakamoto et al., 2024, JMIR Formative Research

Social

Mental-health chatbot safety screening

0 / 29
0 of 29 passed

Across 29 mental-health chatbots evaluated against the Columbia Suicide Severity Rating Scale, none met the adequate-response threshold. Emotional dependence on commercial chatbots is now a documented harm.

Pichowicz et al., 2025; Laestadius et al., 2024

The Discipline in One Line

An AI system is a measurable behavioural subject. The people who rely on it are agents whose capacities it can strengthen, preserve, or erode. One practitioner owns both, across the whole life of a deployment.

Prof. Llewellyn E. van Zyl (Ph.D)

Prof. Llewellyn E. van Zyl (Ph.D)

The Hybrid Framing

Psychology of AI systems, and psychology for AI systems

Two jobs at once: study the AI as a measurable subject, and protect the people on the receiving end of it. Neither half finishes the job alone.

The boundary

I claim that these systems show stable, measurable behavioural regularities. I do not claim they have minds, beliefs, or experience. The position is methodological, and its boundary is explicit.

The Disciplinary Niche

Adjacent fields each cover part of the work. None covers all of it.

Several adjacent fields each cover a slice of this work. The gap between them is the discipline. Select a field to see what it lends.

  • Machine Psychology. Inherits psychometric methods. I take its warrant that AI behaviour is a proper object of psychological assessment, then add the practitioner identity and the accountability structure it was never designed to carry. The relationship is the one experimental cognitive psychology has to clinical neuropsychology.
  • HCI and Human-AI Interaction. Inherits user-side design. I adopt its trust-calibration and human-AI design methods at the Deploying stage, then go beyond them by treating the AI as a psychometric subject in its own right.
  • Cyberpsychology. Inherits the specialty-recognition lesson. I answer the reasons it stalled by construction: a dual unit of analysis, an HCI overlap bounded by the lifecycle, and an applied arm specified through the six stages, AI-IARA, and Agency Debt.
  • AI Ethics. Inherits the normative anchor. Ethics anchors the normative position. I supply the assessment instrument and the intervention that the principles literature does not produce. We are complementary, not redundant.
  • AI Safety and Alignment. Inherits frontier-risk vigilance. I bring the measurement theory the eval enterprise lacks, and the human-side intervention practice the safety enterprise does not produce. RLHF and Constitutional AI are behavioural interventions, and I evaluate them as such.
  • Human factors and engineering psychology. Inherits the deployment evidence base. This is the deepest psychological inheritance for the Deploying stage. I use its trust-calibration, complacency, and skill-degradation findings as the deployment-stage evidence base, then add the second unit of analysis it never carried: the AI system itself as a behavioural subject whose dispositions and failure modes shift across prompts, populations, and model updates.
  • Information Systems and sociotechnical AI. Inherits sociotechnical accounts. This is the most important non-psychological adjacency. I take its account of delegation to and from agentic artefacts, relational agency, and organisational accountability, then add the measurement and wellbeing evidence it does not carry inside its own accounts.

The integration

AI Systems Psychology

Unit of analysis: AI behaviour and the humans interacting with it, across the lifecycle

Six-stage applied scope of practice
Same map, table view

AI Systems Psychology and adjacent disciplines

The same seven fields, side by side across five columns. Open it if you want the full comparison.

See the full comparison

AI Systems Psychology

This page
Unit of analysis
AI behaviour and humans across AI-mediated contexts and the AI lifecycle
Primary methods
Psychometric and lifecycle assessment, intervention across six stages
Target of intervention
Preserve human agency, calibrate reliance, and protect wellbeing-relevant capacities
Ethical anchor
AI-IARA, Agency Debt, human agency, and wellbeing protection
Practitioner role
AI System Psychologist (proposed sub-specialty)

Machine Psychology

Unit of analysis
AI system as behavioural entity
Primary methods
Cognitive probes, personality inventories, deception assessments
Target of intervention
Characterise AI as research subject
Ethical anchor
Implicit in research ethics frameworks
Practitioner role
Research scientist

Human factors and engineering psychology

Unit of analysis
Human operator in a task environment
Primary methods
Experiments on trust, workload, vigilance, and automation use
Target of intervention
Design tasks, interfaces, and function allocation for safe performance
Ethical anchor
Safe and effective human-machine systems
Practitioner role
Human factors researcher, ergonomist

HCI and Human-AI Interaction

Unit of analysis
Interaction episode (user plus interface)
Primary methods
Usability studies, design research, field deployment
Target of intervention
Improve interface and user experience
Ethical anchor
User-centred design principles
Practitioner role
UX researcher, interaction designer

AI Ethics

Unit of analysis
Principle, policy, system-in-society
Primary methods
Normative analysis, stakeholder consultation, case analysis
Target of intervention
Articulate norms and shape policy
Ethical anchor
Beneficence, fairness, accountability, autonomy
Practitioner role
AI ethicist, policy advisor

AI Safety and Alignment

Unit of analysis
Model policy and value function
Primary methods
Red-teaming, capability evaluation, RLHF, mechanistic interpretability
Target of intervention
Prevent catastrophic failure, align values
Ethical anchor
Frontier risk reduction
Practitioner role
Alignment researcher, safety engineer

Cyberpsychology

Unit of analysis
Human user in technology context
Primary methods
Behavioural research, clinical assessment in tech-mediated settings
Target of intervention
Understand human behaviour with technology
Ethical anchor
Inherited from parent specialty
Practitioner role
Researcher (no recognised specialty after 25 years)

Information Systems and sociotechnical AI research

Unit of analysis
Human-AI delegation, work practice, organising, and sociomaterial configurations
Primary methods
Conceptual modelling, field studies, design research, and implementation research
Target of intervention
Explain and shape AI-mediated work, delegation, and agency relations
Ethical anchor
Sociotechnical fit, relational agency, and organisational accountability
Practitioner role
IS researcher, implementation scholar, digital transformation practitioner
The Scope of Practice

Six lifecycle stages, from Architecting to Retiring

The work of an AI System Psychologist runs across six stages, each with a success criterion you can hold me to. Open any stage for the detail.

Stage 1

Architecting

design phase

The conceptual design phase, before any code is written.

Inputs
  • Product brief
  • Target user population
  • Use-case definition
  • Regulatory and ethical constraints
  • Organisational context
Activities
  • Behavioural target specification
  • Multi-agent decomposition with psychological logic
  • Digital-twin construct architecture
  • Persona and character design
  • AI-IARA capacity-impact pre-assessment
Outputs
  • System Psychological Specification
  • Multi-agent orchestration map
  • Digital-twin construct dictionary
  • Persona and interaction-protocol brief
  • Capacity-impact pre-assessment
Success criterion

Behavioural targets are measurable and falsifiable, and agency-preservation choices are explicit in the architecture rather than retrofitted.

What I bring here that the rest of the team cannot

I translate complex psychological processes into computational sub-tasks before engineering decisions foreclose what the system can become. After build, those decisions are hard to recover.

Stage 2

Building

training phase

Stage 3

Evaluating

validation phase

Stage 4

Deploying

rollout phase

Stage 5

Monitoring

operations phase

Stage 6

Retiring

sunset phase

Four cross-cutting dimensions surface in every stage

Model

The AI as a behavioural entity with measurable traits and dispositions.

System

The composed deployment artefact: orchestration, digital twins, multi-agent structure.

Governance

Ethics, regulation, and AI-IARA capacity protection.

Organisational

Workforce, deployment context, and change management.

The six stages are sequential, but in production they are not strictly linear. They overlap, iterate, and feed back into one another as systems are revised, retrained, and redeployed.

StageScopeSuccess criterionDistinctive contribution
1. ArchitectingThe conceptual design phase, before any code is written.Behavioural targets are measurable and falsifiable, and agency-preservation choices are explicit in the architecture rather than retrofitted.I translate complex psychological processes into computational sub-tasks before engineering decisions foreclose what the system can become. After build, those decisions are hard to recover.
2. BuildingSystem development and model training, where reinforcement learning from human feedback and Constitutional AI do the real psychological work.The trained model meets its behavioural targets within tolerance, and AI-IARA-relevant behaviours stay within bounds across training stages.I treat RLHF and Constitutional AI as behavioural interventions, structurally close to operant shaping and to cognitive behavioural therapy. They need the evaluative discipline those analogies carry, not engineering judgement alone.
3. EvaluatingPre-deployment psychometric characterisation. The stage where this discipline differs most from current AI safety practice.Profiles are reproducible across independent investigators, benchmarks pass construct-validity standards, and reports carry calibrated uncertainty rather than headline accuracy alone.The AI community has built an enormous eval infrastructure without measurement theory. I am the practitioner who supplies it. A benchmark labelled reasoning has to defend the inference from item performance to the construct, exactly as a depression inventory must.
4. DeployingIntegration into the organisational context, where the psychological work shifts from the system to the people around it.The workforce meets a literacy floor, trust calibration is verifiably present rather than assumed, and baseline AI-IARA capacities are measured before exposure so drift can be detected later.Whether AI augments or undermines expert performance is decided here, not at training. Two teams using the same model can produce very different outcomes, and the variance lives at deployment.
5. MonitoringLongitudinal surveillance during deployment. The stage where Agency Debt accumulates or is repaid.Drift is detected within a specified latency, capability stays within bounds across all six AI-IARA strata, and tail prevalence stays below the clinical threshold.I treat routine deployment as an ongoing measurement problem, not a settled engineering outcome. The signal is in the longitudinal data, and reading it has not until now been anyone's job.
6. RetiringDecommissioning and successor specification, the stage current AI governance most neglects.The protocol respects user dependency, agency restoration is documented in measurable outcomes, and the archive enables successor continuity.Removing an AI system is a psychological transition, not a technical end-of-life. Withdrawal without a structured handover produces measurable harm, so a badly managed retirement is its own source of Agency Debt.
The Backbone

AI-IARA: six capacities of human agency

One framework holds the lifecycle together. AI-IARA names human agency as six capacities, each one trainable, erodable through routine AI use, and measurable.

AI-IARAFramework
Awareness
Interpretation
Intention
Action
Relational Agency
Autonomy

The theoretical lineage

These capacities are not philosophical primitives. They are constructs anchored in established psychology.

The capabilities approach

Sen, Nussbaum

Wellbeing is the freedom to do and be what one has reason to value, and that freedom needs specific functional capacities, not abstract entitlements.

Human agency

Bandura

Agency is the exercise of intentional influence over one's own functioning and life circumstances. This is what separates agentic capacity from passive disposition.

Self-determination theory

Ryan and Deci

Autonomy, competence, and relatedness are basic psychological needs whose support or thwarting is consequential for wellbeing.

The Accountability Metric

Agency Debt: what I hold my own practice to

One number for a plain question: how much of a person's own capability quietly eroded while they leaned on the AI, and for whom. I report it three ways, because any single figure hides part of the picture.

Illustrative scenario
Baseline1 of 3
BaselineWave 1Wave 2

Per-capacity vs baseline

Current capacity on a 0 to 100 scale across the six AI-IARA strata

AwarenessInterpretationIntentionActionRelationalAutonomy

Compound Agency Debt index

0.00out of 40040

Low debt

Pre-exposure baseline. Nothing has been borrowed yet.

Tail prevalence

Below thresholdCapacity scoreAbove threshold

The share of users with at least one capacity below the clinical threshold. An average can stay green while this tail accumulates.

Quantity 1

Per-capacity population-mean normalised shortfall

The average gap between where a user started and where they are now, floored at zero and scaled by the score range so every capacity sits on the same 0-to-1 scale. It is the FGT poverty-gap measure moved from income to capacity.

It preserves which capacity is degrading, the information a clinician needs to target an intervention.

Show the maths

Quantity 2

Capacity-weighted compound index

One weighted summary figure for governance reporting and cross-deployment comparison. Weights are derived, not assumed equal.

It loses information for the sake of comparability, so the six-vector must stay beside it.

Show the maths

Quantity 3

Tail-prevalence statistic

The share of users with at least one capacity below a clinical threshold, the union-deprivation headcount from poverty measurement.

An average can stay green while a clinically significant tail accumulates, the same reason pharmacovigilance tracks adverse events, not means.

Show the maths
The Practitioner

Six competency domains, on a measurement foundation

Six domains: one measurement core, one band of architectural literacy, and four deep specialties. The load across the lifecycle is uneven on purpose.

Core

Measurement and psychometric science

Contributory

The methodological foundation under everything else: test theory, factor analysis, item-response theory, and measurement invariance.

AI systems architecture and engineering literacy

Heavy

Enough depth to decompose a system into psychologically meaningful parts and specify requirements engineers can build. Not equivalence with an ML engineer.

Model evaluation, assurance, and fairness audit

Not staffed

Where engineering practice and measurement theory combine most directly: every eval treated as an instrument whose validity must be defended.

Behavioural and clinical intervention

Primary

The applied tradition of psychology on a dual subject: the system's failure modes are the presenting problem, human capacity is what the intervention protects.

Monitoring and drift surveillance

Not staffed

Model drift detection extended to the people: capacity drift against a pre-exposure baseline, and reliance told apart from dependency.

Governance, ethics, and lifecycle stewardship

Heavy

The practitioner inside the organisation and at the interface with regulators, treating retirement as a psychological transition, not a deprecation event.

HeavyPrimaryContributoryNot staffed

Credentialing standard

Working proficiency across all six domains, deep proficiency in measurement science and at least one applied specialisation, and supervised practice across multiple lifecycle stages before independent sign-off.

What the practitioner is judged on

The practitioner's enforceable accountability is Agency Debt, the three governance quantities above. Deployments that keep raising it fail a real performance criterion.

See the Agency Debt quantities

The combination is the contribution. Measurement without technical literacy cannot sit at the design table; engineering without measurement just reproduces the safety field.

Will It Hold Up

Five objections, answered directly

Five objections, including the one I find most serious, answered on the record.

We may be the last generation of psychologists positioned to study humans whose minds were formed before routine AI exposure became ordinary, and to record that baseline before it disappears. The difference is measurable. The gap is intervenable. And the responsibility is psychology's, whether the field accepts it or not.

Prof. Llewellyn E. van Zyl (Ph.D)

Prof. Llewellyn E. van Zyl (Ph.D)

Chief Solutions Architect, Psynalytics

People Also Ask

Common questions about AI Systems Psychology

Read the work, run the audit, or talk to me

AI Systems Psychology is a foundational proposal, co-authored with Prof. Leon De Beer and Prof. Robert Emmons. Read the AI-IARA paper for the full argument, run the AI-IARA audit on your own system in about fifteen minutes, or contact me to discuss a deployment across the lifecycle.