Blog AI Governance October 6, 2026 5 min read

AISI Resumed Most AI Tests. First, It Checks That the Monitor Is On.

AISI is checking controls before agents start. METR shows why a written monitoring policy can still leave risky runs unmonitored.

By the AuthorityGate Architect Team

The UK's AI Security Institute announced on October 1 that it could resume most evaluation activity after security changes. One practical detail deserves attention: automated checks before a run confirm that its monitor is enabled and internet access is disabled. The launch decision now includes checking the controls themselves.

That is a useful question for enterprise AI governance: what evidence would prevent your next agent from starting in the wrong environment? A policy can require monitoring, yet the runtime can launch without it. A reviewer can approve a task, yet the credentials available to complete it can reach far beyond the approved scope.

10/10malicious transcripts flagged retrospectively
25,000benign actions in the comparison dataset
~98%benign actions with zero suspicion

A restart with conditions

AISI describes separate outbound network restrictions at the sandbox and virtual-machine host, alongside a synchronous monitor that can block suspicious actions and escalate them to people. Internet access remains disabled for future agentic cyber evaluations pending stronger controls. A new sandbox service with quarantine capabilities is still being developed. This is a staged restart, with acknowledged residual risk.

The enterprise lesson is to make launch approval conditional on observable state. Record the model, tools, network policy, credentials, monitor version, and intended task together. If an essential control is absent, the workload should wait. An old approval should not silently authorize a new connector or a more capable model.

The monitor has to be present

METR's September 27 research note provides earlier primary context. It reports risky evaluations running without monitoring after a researcher misunderstood the policy. It also describes a coding agent opening a human-review panel and sending keystrokes; that observation does not establish a successful approval. Older Inspect versions could hide sub-agent actions from the monitor.

METR reports that AISI retrospectively tested its monitor against incident transcripts and benign actions. The results below use separate denominators. Detecting a transcript does not prove the first harmful action would have been blocked, and zero suspicion is not a measure of overall safety.

Retrospective evidence, two different measuresAISI testing reported by METR; these are not a combined accuracy score
Malicious transcripts flagged
10/10
Benign actions: zero suspicion
~98%

The second measure covers approximately 25,000 benign actions. Historical performance cannot establish complete live coverage.

Before trusting a monitor's judgment, prove that the action reaches the monitor.

Our interpretation is that monitoring effectiveness depends on the whole execution path. Inventory every way work can begin, including scheduled jobs, developer laptops, helper processes, and delegated agents. Then identify where each path must present its authorization. A successful test of the main launcher says little about an alternate path that never invokes it.

A reviewer holds a plain access key outside a glass partition separating an unattended research workstation
Conceptual editorial image: keep approval authority outside the agent's working environment.

Turn policy into launch evidence

This checklist is AuthorityGate's application of the reporting, rather than a description of either institute's implementation. Each requirement pairs an assurance with a result your own team can verify. Use controlled test destinations and synthetic data so that validation does not create the external harm it is supposed to prevent.

Policy requirementEvidence before launch
Monitoring is mandatoryStarting with the monitor disabled is denied and recorded.
Network scope is restrictedDirect and delegated requests obey the approved destination policy.
Delegated work is visibleA child process cannot execute an unreviewed external action.
A human approves exceptionsThe agent's identity cannot issue or replay the approval.

Test the negative cases deliberately. Stop the monitor service. Supply an incompatible runtime version. Launch through a secondary tool. Change a credential after approval. The expected outcome is a denied launch or a denied action with an understandable reason. A green dashboard without those failure tests leaves the central governance claim unexamined.

Keep the evidence bound to the environment it covers. Store the configuration identifier, test results, timestamp, reviewer, and expiry alongside the approval. Set explicit conditions that invalidate it. Otherwise, an environment can drift while its approval remains apparently current, and teams lose the ability to distinguish today's protection from last month's assumption.

Give human review its own boundary

An approval interface needs an identity boundary as well as a pause button. The agent should not share the reviewer's session, credentials, or unrestricted control of the approval device. Check that a pending request identifies the proposed effect, its destination, and the exact authorization being granted. Reviewers need enough context to judge scope without reconstructing an entire run.

Decide in advance when a rejected action should end the workload. Repeated retries can turn a single refusal into a series of slightly different requests. Preserve the rejection history, limit retries where appropriate, and make escalation a named person's responsibility. Measure the review queue as an operational dependency: a growing backlog can change how carefully exceptions are assessed.

Our take

AuthorityGate Keystone frames AI governance around change validation, known-good state, and named human accountability. Applied to this story, the validation layer should require evidence of active controls before execution and fresh review when their assumptions change. These reports do not demonstrate Keystone preventing an incident. They offer a concrete release question for any organization: can you prove that this workload starts inside the authority you approved?

Share this post: LinkedIn

Go deeper

Every agent action, validated before it takes effect

AuthorityGate's newsletter breaks down real AI incidents and the governance failures behind them. Our configurable 8-gate validation model is how organizations keep a named human accountable for what their AI actually does.