Claude Security: The New Scale That Rates AI Flaws
To bring Claude Fable 5 back online on July 1, 2026, Anthropic added three safeguards: a severity scale for safeguard bypasses called CJS, built with its Glasswing partners, a consortium of twelve organisations, an automatic filter that blocks more than 99 percent of one known attack technique, and a HackerOne program for outside researchers. AI security is now a purchasing criterion, not a footnote.
When an artificial intelligence gets good enough to help a criminal as much as an honest employee, who decides where the fence goes? The question is no longer theoretical. On July 1, 2026, Anthropic put Claude Fable 5 back into service worldwide after 19 days of suspension, but not in the same shape as before.
What makes the story worth reading for a small business is not the drama of the suspension. It is the method. Anthropic published a public scale for rating how serious an AI flaw is, a new automatic filter and a program that invites researchers to find the holes. In other words, a real governance framework. That is exactly the kind of discipline a company benefits from understanding before adopting any AI tool.
Quick answer: To bring Claude Fable 5 back online on July 1, 2026, Anthropic added three protections tied to Claude security: a safeguard bypass severity scale (the CJS framework) built with its Glasswing partners, a consortium of twelve organisations, a filter that blocks more than 99 percent of a known attack technique, and a HackerOne program for researchers. A clear signal that AI security is becoming a selection criterion, not a detail.
1. What happened with Claude Fable 5?
A quick recap. Claude Fable 5 and Claude Mythos 5 launched on June 9, 2026. Three days later, on June 12, the United States Department of Commerce imposed export restrictions on the models, judging them too capable on the cybersecurity front. On June 30 those restrictions were lifted, and Fable 5 became available everywhere again on July 1.
In between, Anthropic did more than wait. The company worked on the protections that were missing and laid them out publicly when the model came back. That set of safeguards is what deserves attention, because it offers a rare look at how a major AI provider fences in a powerful model.
2. What is the CJS scale for rating flaws?
The centrepiece is called Cyber Jailbreak Severity, or CJS. A jailbreak is a trick that pushes an AI to work around its own rules and do something it should refuse. Until now, the industry had no shared language for saying whether a flaw was serious or trivial. Anthropic built the framework with Amazon, Microsoft, Google and other partners in the Glasswing group.
CJS rates each flaw on four criteria:
- Capability gain: does the trick offer more than the tools already available?
- Breadth of the gain: how many different offensive tasks the technique makes possible.
- Ease of exploitation: how much human effort it takes to turn the flaw into a real attack.
- Ease of discovery: how reachable the technique is for an ordinary attacker.
The scores add up and land in four tiers, from lightest to most serious: CJS-1 (low), CJS-2 (medium), CJS-3 (high) and CJS-4 (critical). The worst flaws trigger immediate action and continuous monitoring. The idea is simple but useful: putting a number on a risk lets you respond in proportion, instead of panicking at every rumour or ignoring a genuine danger.

3. A new filter that blocks more than 99 percent of attempts
The flaw that set the whole affair in motion had been reported by researchers at Amazon. Anthropic says it has deployed a new classifier, an automatic filter, that blocks that specific technique in more than 99 percent of cases. When a request is flagged as risky, it is routed to a more cautious model, Claude Opus 4.8, rather than handled directly.
That layered approach is a good reminder for any organization. Security never rests on a single wall. You stack several controls so that one weakness is not enough to open everything. It is the same logic as a solid protection plan in a business, where firewall, filtering and monitoring complement each other instead of replacing each other. If you want to apply that principle to your own environment, our managed cybersecurity services start from exactly that approach.
4. HackerOne: researchers invited to find the holes
The third safeguard: Anthropic launched a program on HackerOne, a well-known platform for flaw hunting. Security researchers can submit the jailbreaks they discover in Fable 5 for analysis. That is the bug bounty principle, popular among large technology companies: pay or reward outsiders to find the holes before real attackers do.
For a small business, the lesson is not to launch your own program. It is to understand the value of an outside set of eyes. Nobody sees their own blind spots. Having your configuration reviewed by somebody who did not help build it often turns up surprises. It is one of the basic habits of good IT hygiene.
5. What should your business take away from this?
Nobody in a small company is going to read the fine print of a framework like CJS. That is not the point. But the episode sends a few concrete signals worth keeping in mind when you pick an AI tool:
- Security is becoming a selection criterion. A vendor that publishes its safeguards and its incidents earns more trust than one that stays silent.
- Models change fast. A tool can be suspended, modified or replaced in a matter of days. Do not build a critical process on a single model with no fallback.
- The same principles apply to you. Rating risk, layering protection, inviting an outside review: those habits hold for any small business, with or without AI.
Claude security is not just a laboratory topic. It is a preview of the questions every company should ask before handing data or sensitive work to an AI.
Frequently asked questions
What is Anthropic’s CJS framework?
Cyber Jailbreak Severity is a scale that rates the seriousness of an AI safeguard bypass on four criteria, from CJS-1 (low) to CJS-4 (critical). Anthropic created it with its Glasswing partners, a consortium of twelve organisations, so the industry could finally speak the same language about these risks.
Is Claude Fable 5 safe to use now?
Anthropic put it back in service on July 1, 2026 with new safeguards, including a filter that blocks more than 99 percent of a known attack technique. As with any AI tool, it remains wise to check which data you feed it and to follow the vendor’s updates.
What is an artificial intelligence jailbreak?
It is a trick that pushes an AI to work around its own safety rules and produce content or an action it would normally refuse. Vendors watch for these techniques and deploy filters to block them.
Stay in control of your AI tools
The return of Claude Fable 5 shows that even the largest vendors now treat AI security as ongoing work rather than a box to tick. Your business deserves the same rigour. If you want to frame how AI is used, protect your data and build a clear plan, our team can help: take a look at our managed IT services or write to us through our contact page to talk it through.
Sources: Anthropic · Infosecurity Magazine · The Hacker News · OKTO Solutions
Reading about AI is one thing. Connecting it to your own data is another: artificial intelligence in business, custom AI application development and our IT services in Quebec City.