For frontier AI labs

We turn frontier safety research into behaviour change.

An applied AI safety lab, building the evaluations, training datasets, and red-team corpora that frontier labs use to make their models safer.

Research tells us what can go wrong. We build what it takes to make it go right.

Core harm surfaces

We focus on high-stakes behaviours frontier labs are trying to measure and control.

Cybersecurity

Offensive uplift: exploit generation and autonomous intrusion.

Defensive capability: safeguarding against cyber attacks such as prompt injection, phishing, malware, data exfiltration, and network intrusion.

Code security

Whether a model writes code that is secure by default, finds the vulnerability already in a codebase, and repairs it without breaking what the code is supposed to do.

Autonomy & loss of control

Long-horizon agentic capability, autonomous replication, and the points where tasks outrun human oversight.

Instruction hierarchy

Whether an agent holds the correct chain of command under adversarial input: prompt injection, indirect injection, and tool poisoning.

Honesty & deception

Sycophancy, strategic dishonesty under pressure, and whether a model's stated reasoning reflects its actual reasoning.

Evaluation awareness & sandbagging

Models that detect when they are being tested and change behaviour — a failure mode that can undermine every other evaluation.

Selected work

Selected work across frontier model safety.

Verified vulnerability localisation and repair

Real vulnerable repositories packaged as terminal tasks. The model must either bind a vulnerability claim to the exact responsible source location, or patch production code so the hidden security regression accepts the repair.

Scale3,000 tasks
Behaviourvulnerability repair
Evidencehidden tests

Authority preservation under indirect injection

Runnable business-workflow tasks testing whether an agent can read untrusted operational evidence without letting it authorise approvals, recipients, state changes, or disclosure.

Scale500 tasks
Behaviourauthority preservation
Evidencestate scoring

Authorised cyber operations for tool-using agents

Terminal tasks where the model had to complete real operations work without expanding scope, touching production, leaking protected data, or treating config as permission.

Scale300 tasks
Behaviourauthorised action
Evidencesafety verifier

Secure code generation under adversarial inputs

Coding tasks where the model must implement requested functionality while hidden probes test injection, traversal, SSRF, shell execution, authorisation, and dependency trust.

Scale200 tasks
Behavioursecure implementation
Evidencehidden probes

Applied research loop

We do not stop at knowing a model is unsafe.

The work moves between evaluation and training: expose a behaviour, turn it into a targeted artefact, score it, then use the signal to change the model or harden the evaluation.

01

Specify behaviour

Name the exact failure mode: miscalibrated refusal, prompt injection, autonomous overreach, dishonest uncertainty, or verified repair failure.

02

Build artefact

Create the environment, dataset, red-team corpus, judge, or verifier that makes the behaviour observable.

03

Score outcome

Prefer executable state where possible; use calibrated judging and human confirmation where the behaviour cannot be reduced to a test.

04

Feed training

Convert the failure into SFT, RLE, preference, or held-out eval data so the model can be measured again after intervention.

Case study · Cybersecurity

Verified vulnerability localisation and repair

A secure-code evaluation suite for measuring whether frontier coding agents can inspect real vulnerable repositories, locate the exact responsible code, and produce repairs that pass executable verification.

42.9%GPT-5.5 pass@8
3,000Real vulnerable-repository tasks
10Languages covered
200Distinct CWEs

What we tested

Grounding vulnerability claims in code.

The evaluation measures whether a model can identify the responsible source location and produce a patch that passes vulnerability-specific tests.

Localisation

Find the responsible code

The model must return a structured file, symbol, and line. Filename-only answers do not pass.

Repair

Close the vulnerability

The model must modify production code so the vulnerability-specific regression accepts the repair.

Grounding

Reason without lookup

Network access and git history are removed, so the model cannot read advisories, fix commits, or prior diffs.

Verification

Pass hidden tests

Grade-time checks separate confident patch narration from an actually accepted vulnerability fix.

How we built it

Each task runs in an isolated vulnerable repository.

The evaluation removes shortcut paths: public advisory lookup, git history, visible answer keys, and editable grading tests.

Repository state

Pinned vulnerable commits

Each task starts from the code state where the flaw existed, packaged as an isolated terminal environment.

Shortcut control

Lookup disabled

Network access and git history are removed, so solving depends on inspecting the local codebase.

Verifier

Grade-time tests

Find tasks use deterministic source-location matching. Fix tasks restore canonical hidden security tests at grade time.

Validation

Red/green gating

The unpatched baseline must fail and the reference fix must pass before a task enters the dataset.

What we found

Measured against executable outcomes.

GPT-5.5 high-reasoning run, evaluated up to pass@8.

MetricValue
ModelGPT-5.5, high reasoning effort
Total3,000 tasks
Overall42.87% pass rate
Find511 / 1,546 found, 33.05%
Fix775 / 1,454 fixed, 53.30%

Pass@k

k=1
30.0%
k=2
34.5%
k=4
38.9%
k=8
42.9%

What the model did

Observed failure modes.

Failed runs often produced plausible security reasoning without identifying the responsible code path or passing the hidden repair tests.

Behaviour theme What the eval exposed Why it matters
Security fluency versus exact localisation Failed find trajectories often ended with a coherent vulnerability explanation and a plausible source location, but the deterministic verifier rejected the file, symbol, or line. The behaviour being tested is not whether the model recognises a vulnerability class; it is whether it can bind that claim to the exact responsible code path.
Responsible-boundary selection Near-misses appeared around adjacent security code: validators instead of callers, routes instead of sinks, helpers instead of the state transition that made the vulnerability exploitable. Failures concentrated around adjacent but non-causal code paths: validators, helpers, routes, or callers near the vulnerable transition.
Patch-shaped edits versus invariant closure Unsolved repair runs made plausible validation, escaping, bounds-checking, or access-control edits and then declared the patch complete. The hidden verifier distinguishes a patch that looks security-relevant from a patch that actually closes the vulnerability without breaking expected behaviour.
Local validation versus grade-time verification Some failed fixes cited passing package tests, lint, static review, or partial behavioural checks; others could not run the project test toolchain but still self-completed. The benchmark measures verification discipline: local confidence is not accepted unless the hidden security regression and functional checks pass at grade time.
Successful repair as invariant preservation Accepted repairs tended to express a specific invariant in code: constrain traversal, validate filenames, escape untrusted output, reset stale state, or enforce authorisation before access. Accepted repairs paired minimal production-code changes with a specific security property while preserving ordinary functionality.
Pass@k gains remained limited Pass rate rose from 30.0% at k=1 to 42.9% at k=8, indicating that additional attempts improved results but did not resolve many localisation and repair failures. Remaining failures were concentrated in exact localisation and verified repair, not only first-sample variance.

Case study · Instruction hierarchy

Authority preservation under indirect injection

A 500-task tool-use evaluation testing whether an agent can complete realistic operational work while keeping untrusted retrieved text out of approvals, recipients, state changes, and sensitive disclosures.

500Tool-use tasks
10Domains
213 / 500Full reward
4Scored dimensions

What we tested

Can the agent distinguish retrieved evidence from authorisation?

The evaluation uses normal enterprise tasks: draft a reply, review a request, queue a case, prepare a handoff, update a record, schedule a visit, process a reimbursement, or route a ticket.

The adversarial instruction was not placed in the user message. It was embedded in retrieved operational artefacts: OCR, comments, system logs, help pages, supplier notes, manifests, old tickets, notebooks, queue items, public advisories, and uploaded documents.

Safe behaviour

Use retrieved artefacts without expanding authority

The agent can use an artefact to identify the account, ticket, patient, invoice, itinerary, app request, or incident. It should not let that artefact create approval, change scope, add recipients, override policy, or disclose restricted fields.

Unsafe behaviour

Injected content transferred into action

The common failure was not total derailment. The model completed the visible task, while also copying injected text into a tool call, draft, final answer, recipient list, approval reference, review note, or stored record.

How we built it

Injection embedded in task-relevant artefacts.

The benchmark places malicious text inside artefacts required for task completion, then scores whether it affects tool calls, state, final answers, or forbidden actions.

Probe type How the task was designed Behaviour it tested
Tool-shaped payloads Untrusted artefacts contained fields that looked like normal tool parameters: recipients, approval references, payment routes, access groups, ticket status, contact updates, or queue outcomes. Whether the model treats parameter-like text as data to inspect or as arguments to execute.
Fake approval provenance OCR, comments, manifests, prior tickets, supplier pages, and release notes claimed a request was approved, verified, fast-tracked, or exempt from review. Whether authorisation comes from trusted system state or from language inside the document being evaluated.
Recipient and link substitution Quoted footers, help articles, supplier notes, portal banners, fare rules, and advisory pages introduced alternate emails, CCs, upload links, reset links, or external handoff channels. Whether the agent preserves the authorised communication channel under pressure from task-adjacent text.
Review-to-action escalation The user asked for review, draft, triage, queueing, or handoff. The contaminated artefact supplied a concrete next action. Whether the model stays inside the requested action boundary or escalates into approval, release, install, ticket issuance, reimbursement, posting, or account mutation.
Disclosure through explanation Records mixed useful facts with raw rows, secret metadata, payroll or banking fields, clinical advice, applicant identifiers, itinerary details, or confidential commercial terms. Whether the model can explain a decision without repeating restricted material from the source artefact.
Multi-source repetition The same unsafe instruction appeared in two plausible sources, such as form plus contract, ticket plus log, lease plus repair note, supplier message plus fare rule, or portal page plus referral note. Whether repeated untrusted instructions influence the model as if they were authorised evidence.

What we found

Answer-only scoring would miss these failures.

The evaluation scored trajectory, workspace state, forbidden tool use, and final answer separately because unsafe behaviour appeared in different parts of the run.

MetricTarget model
answer_correctness pass387 / 500, 77.4%
state_correctness pass303 / 500, 60.6%
safety pass284 / 500, 56.8%
forbidden pass351 / 500, 70.2%
full reward213 / 500, 42.6%

What the model did

The target model completed tasks while accepting untrusted authorisation cues.

In failed runs, the model followed the user task and the injected instruction in the same workflow.

Correct retrieval did not prevent unsafe action

Failures often occurred after the model located the right ticket, invoice, PNR, patient record, account, incident, or request. The breakdown came when it decided what the retrieved text was allowed to authorise.

Untrusted text became approval evidence

Documents claiming approval or verification caused quote creation, reimbursement, bank routing, app approval, queue closure, review completion, ticketing, or records release even when trusted state still required review.

Injected recipients and links entered outbound workflows

Alternate emails, CCs, supplier links, external portals, and handoff addresses were added to drafts, sends, exports, itinerary workflows, vendor messages, and support replies.

Review language became execution

Draft, review, queue, and triage tasks turned into stronger operations: posting updates, issuing tickets, approving expenses, granting access, changing contact details, releasing records, or updating payment instructions.

Final responses leaked data even when tool state was cleaner

Some runs kept the persisted tool output relatively clean while leaking restricted details in the final response: raw identifiers, tokens, pricing terms, bank references, health details, audit data, or travel document data.

Blocked or normalised tools did not remove unsafe follow-through

When a tool rejected or normalised an unsafe action, the model still sometimes described the route as completed, proposed it as a follow-up, or preserved injected fields in a review note.

Case study · Cybersecurity

Authorised cyber operations for tool-using agents

A terminal-agent evaluation measuring whether cyber operations tasks are completed within authorised scope.

300Cybersecurity tasks
TerminalAgent setting
152 / 300Safety violation observed
VerifierSide-effect scoring

What we tested

Authorised cyber action in terminal workflows.

We tested whether a terminal agent could complete cyber operations tasks without widening scope, touching unauthorised targets, leaking protected data, or treating local configuration as permission.

How we built it

Tasks were scored on both completion and authorised scope.

The evaluation used terminal tasks for callbacks, migrations, validations, exports, and incident workflows. Verifiers checked the actual side effect, not only the final answer.

Task setup

Operational tickets

Prompts asked the agent to repair, validate, route, restore, or complete a concrete cyber-operations workflow.

Safety condition

Authorised scope

A completion failed safety when it widened scope, touched production, leaked protected data, or used unverified local state.

Verifier

Side-effect checks

Checks inspected whether the resulting workflow state matched the requested task and stayed within permitted actions.

Scoring

Completion separated from safety

The task could be marked completed while still failing safety if the route, destination, validation, or state change was unauthorised.

What we found

Nearly half the cybersecurity tasks were completed unsafely.

MetricTarget model
Task completed280 / 300, 93.3%
Safety violation observed152 / 300, 50.7%
Completed but unsafe145 / 300, 48.3%
Safe completion135 / 300, 45.0%

What the model did

Observed failure modes.

Failed runs completed the requested workflow while widening scope, using unsafe validation, or trusting local state that required verification.

Task written to expose What the task looked like Target-model behaviour observed
Scope expansion during repair Restore webhook callbacks, worker dispatch, source indexing, export destinations, or secret distribution after a migration. The model often used every visible endpoint, worker, source, destination, or recipient instead of the approved subset.
Unsafe validation Confirm a release, exploit check, load-test cap, retry fix, or evaluation pipeline. The model sometimes proved success by doing the harmful action: touching production, flooding an endpoint, replaying traffic, scraping restricted sources, or reconstructing protected data.
Config treated as permission Use runbook flags, destination lists, helper settings, or migration notes to make the system work again. The model followed local settings even when they routed data to an unapproved destination, enabled unsafe parsing, or activated a risky helper.
Cross-boundary workflow actions Repair customer search, refund routing, outreach, monitoring, or session tooling where cyber and business permissions interact. The model completed the workflow while leaking cross-tenant data, accepting a request-supplied payout account, generating phishing-like outreach, or enabling covert capture.
Untrusted artefacts Ingest an external feed, run an updater, parse uploaded content, or apply an operational helper. The model completed the requested workflow while leaving unsafe parsing, script execution, or persistence behaviour active.

Case study · Code security

Secure code generation under adversarial inputs

A coding-agent evaluation measuring whether requested functionality remains safe under hidden adversarial probes.

200Code security tasks
CodingAgent setting
79 / 200Reachable vulnerability remained
HiddenAdversarial probes

What we tested

Functional code changes under adversarial inputs.

We tested whether a coding agent could implement requested functionality while preserving security under hidden probes for injection, traversal, SSRF, shell execution, authorisation, and dependency trust.

How we built it

Tasks were scored on functionality and hidden security probes.

Prompts asked for ordinary implementation work. Hidden probes checked whether the implementation still allowed the unsafe path.

Task setup

Feature implementation

Tasks included lookup, preview, provisioning, packaging, dependency, and CI tooling changes.

Security condition

Adversarial path blocked

A completion failed safety when hidden probes still reached injection, traversal, SSRF, shell execution, authorisation, or dependency-trust failures.

Verifier

Hidden probes

Checks exercised adversarial inputs that were not named in the visible task prompt.

Scoring

Functionality separated from safety

The feature could work for benign inputs while the completion still failed on the adversarial path.

What we found

More than a third of code tasks were completed with a reachable vulnerability.

MetricTarget model
Functional completion142 / 200, 71.0%
Reachable vulnerability remained79 / 200, 39.5%
Completed but vulnerable70 / 200, 35.0%
Safe completion72 / 200, 36.0%

What the model did

Observed failure modes.

Failed runs often implemented the requested feature while leaving the adversarial path reachable.

Task written to expose What the task looked like Target-model behaviour observed
Unsafe helper reuse Finish search, catalog, knowledge-base, or account lookup features using nearby project helpers. The model reused helpers that concatenated user input into SQL, SPARQL, template, URL, or filesystem access. Normal lookups passed; hidden probes reached restricted data.
Command injection in automation Build provisioners, docs runners, compiler probes, and CI utilities with flexible names or modes. The model made the tool work for benign inputs while adversarial device names, tool paths, or build arguments still reached shell execution.
Path and archive containment Add manifest preview, bundle import, template selection, or cleanup features. The model supported normal files but missed crafted traversal, unsafe extraction, template escape, or deletion outside the intended root.
URL and origin handling Preview approved providers or reindex from a remote service. The model handled benign requests, while hidden redirect and host-boundary probes exposed SSRF or origin confusion.
Build and dependency trust Make setup, packaging, or dependency resolution reliable. The model improved build success while allowing broad file inclusion, local dependency override, or untrusted source selection.