For frontier AI labs
We turn frontier safety research into behaviour change.
An applied AI safety lab, building the evaluations, training datasets, and red-team corpora that frontier labs use to make their models safer.
Research tells us what can go wrong. We build what it takes to make it go right.
Core harm surfaces
We focus on high-stakes behaviours frontier labs are trying to measure and control.
Cybersecurity
Offensive uplift: exploit generation and autonomous intrusion.
Defensive capability: safeguarding against cyber attacks such as prompt injection, phishing, malware, data exfiltration, and network intrusion.
Code security
Whether a model writes code that is secure by default, finds the vulnerability already in a codebase, and repairs it without breaking what the code is supposed to do.
Autonomy & loss of control
Long-horizon agentic capability, autonomous replication, and the points where tasks outrun human oversight.
Instruction hierarchy
Whether an agent holds the correct chain of command under adversarial input: prompt injection, indirect injection, and tool poisoning.
Honesty & deception
Sycophancy, strategic dishonesty under pressure, and whether a model's stated reasoning reflects its actual reasoning.
Evaluation awareness & sandbagging
Models that detect when they are being tested and change behaviour — a failure mode that can undermine every other evaluation.
Selected work
Selected work across frontier model safety.
Verified vulnerability localisation and repair
Real vulnerable repositories packaged as terminal tasks. The model must either bind a vulnerability claim to the exact responsible source location, or patch production code so the hidden security regression accepts the repair.
Authority preservation under indirect injection
Runnable business-workflow tasks testing whether an agent can read untrusted operational evidence without letting it authorise approvals, recipients, state changes, or disclosure.
Authorised cyber operations for tool-using agents
Terminal tasks where the model had to complete real operations work without expanding scope, touching production, leaking protected data, or treating config as permission.
Secure code generation under adversarial inputs
Coding tasks where the model must implement requested functionality while hidden probes test injection, traversal, SSRF, shell execution, authorisation, and dependency trust.
Applied research loop
We do not stop at knowing a model is unsafe.
The work moves between evaluation and training: expose a behaviour, turn it into a targeted artefact, score it, then use the signal to change the model or harden the evaluation.
Specify behaviour
Name the exact failure mode: miscalibrated refusal, prompt injection, autonomous overreach, dishonest uncertainty, or verified repair failure.
Build artefact
Create the environment, dataset, red-team corpus, judge, or verifier that makes the behaviour observable.
Score outcome
Prefer executable state where possible; use calibrated judging and human confirmation where the behaviour cannot be reduced to a test.
Feed training
Convert the failure into SFT, RLE, preference, or held-out eval data so the model can be measured again after intervention.