# Will your AI tools help you during a cyber attack? Testing refusal behaviour

Source: https://perfraction.com/blog/ai-refusal-incident-response/
Author: Joshua Woo, Perfraction
Published: 2026-09-16
Topics: Cybersecurity, Incident response, AI procurement

## Short answer

Not necessarily. Safety guardrails can stop a model from analysing malware, attack logs or attacker tooling, because the model can't tell a defender from an attacker. Test refusal behaviour before you need it, treat it as a procurement criterion, and pre-approve a fallback model for incident response.

## Key takeaways

- In July 2026, Hugging Face disclosed an attack by a fully autonomous AI agent using tens of thousands of automated actions.
- Its security team found a leading US frontier model's guardrails got in the way of the defence.
- It analysed more than 17,000 attack logs with an open-source model run on its own infrastructure.
- Nobody benchmarks whether a tool will work when you are under attack. Procurement should.
- Choosing a fallback model during an active breach is not a governance strategy.

## What happened at Hugging Face?

In July 2026, Hugging Face disclosed that a fully autonomous AI agent had attacked its systems with tens of thousands of automated actions and no human in the loop. It was one of the first documented cases of its kind.

The detail that should worry every GC and CISO came next. When the security team reached for a leading US frontier model to run the defence, its guardrails got in the way. In the company's words, these models "cannot distinguish an incident responder from an attacker".

The team ended up running a Chinese open-source model on its own infrastructure to analyse more than 17,000 attack logs.

## Why does this matter for your incident response plan?

Everyone evaluates AI models by what they can do. Almost nobody evaluates them by what they won't do, until it matters.

Your incident response plan probably assumes your AI tools will cooperate at 3am on the worst day of your year. Has anyone tested that assumption?

## What should you add to AI procurement and tabletop exercises?

1. **Refusal behaviour under incident conditions.** Will your approved model analyse malicious payloads, malware artefacts and attacker tooling, or refuse and flag your account?
2. **Refusal as a capability measure.** We benchmark reasoning, coding and context windows. Add "will this tool work when we're under attack?" to procurement criteria.
3. **The fallback question.** If your primary model refuses mid-incident, what is your pre-approved alternative, and has security and legal signed off on it?

Capability is what a model can do on its best day. Preparedness is knowing what it will do on your worst one.

## Frequently asked questions

### What is AI refusal behaviour?

A model declining a request because its safety rules flag it as potentially harmful. In security work, legitimate tasks such as analysing malware can trigger refusals.

### Should incident response plans cover AI tools?

Yes. If responders rely on AI to analyse logs or payloads, the plan should confirm the approved tool will do that work and name a pre-approved alternative if it refuses.

### How do you test AI refusal behaviour?

Include realistic malicious artefacts and attacker tooling in a tabletop exercise, run them through your approved models, and record where they refuse or flag the account.

---

This article is general information, not legal advice. It reflects the position as at the date of publication.

About Perfraction: a Singapore-based fractional legal function for technology companies, led by Joshua Woo. Perfraction is a legal consultancy and not a law firm. We do not provide legal advice or legal representation. Our services are limited to strategic advisory, legal operations, and regulatory consulting. For legal advice, please consult a qualified lawyer or licensed law firm.
Contact: josh@perfraction.com · Book a call: https://calendly.com/perfraction/30min · Questionnaire: https://perfraction.com/questionnaire
