Rendered at 00:30:32 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
gck1 1 hours ago [-]
> On July 21, OpenAI disclosed that several of their models had broken out of an isolated test environment
> In response to this incident, we began a large-scale retrospective review of our own cybersecurity evaluations
> we identified three incidents
> The incidents involved three different Claude models: [...] and an internal research test model
This reads like an attempt by Anthropic to re-secure their leading spot in "our models are the most dangerous and we also have unreleased, super-secret, research models" index.
I may be too cynical, but the well of benefit of the doubt is running very dry towards AI labs that like to engage in this game.
simonw 1 hours ago [-]
I don't interpret it like that at all. This is deeply embarrassing for Anthropic: it turns out they hadn't been keeping a close eye on their models either, and back in April they successfully attacked three different organizations!
The hacks weren't particularly impressive either:
> [...] using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. It did not find or exploit any complex vulnerabilities [...]
cbb330 40 minutes ago [-]
so deeply embarrassing that they published an eng blog about it
simonw 23 minutes ago [-]
If they quietly brushed this under the rug - especially given the PyPI malware that was involved - it would be a huge scandal.
Disclosure is the only ethical response to this.
johnfn 14 minutes ago [-]
Companies post deeply embarrassing eng blogs all the time. See: every post about downtime or a security incident ever.
ofjcihen 38 minutes ago [-]
Right? 100% this is them trying to make gold out of turds.
solenoid0937 22 minutes ago [-]
JFC there's nothing Anthropic can do to satisfy the HN crowd, is there?
They are not bragging in this article or they would not have called the attacks unsophisticated
If they don't post about this they're bad. If they post about this they're bad.
andy99 58 minutes ago [-]
Then maybe just this timing is really unfortunate, I think most people’s first reaction will be that it looks like a “us too” response to the OpenAI/hf thing.
rvz 42 minutes ago [-]
> I don't interpret it like that at all. This is deeply embarrassing for Anthropic: it turns out they hadn't been keeping a close eye on their models either, and back in April they successfully attacked three different organizations!
This just helps their (Anthropic) argument into persuading the US government into taking action into limiting powerful closed or open-weight models from being released without going through (yet to be defined) regulatory oversight.
The only "embarrassing" thing for Anthropic was that there was little to no continuous security monitoring of this since April, and they then decided to do a cybersecurity transcript review only AFTER the incident with OpenAI and Huggingface.
simonw 4 minutes ago [-]
Anthropic know better than anyone else how risky it is to get this current administration upset with you over safety/security concerns.
gck1 28 minutes ago [-]
They also gave access to Mythos (the Mythos) to some companies, based on... vibes.
Who knows how these companies are using it. If Anthropic can't effectively contain their own models, can the partners?
While the rest of us get fallbacks and warnings, not even being able to defend against the attacks they themselves are causing.
Do we really have to re-learn all the industry's knowledge the hard way?
skeptic_ai 14 minutes ago [-]
Sorry simonw but they are the smartest guys on the planet and safety it’s the word that comes out of their mouth every 5 min.
You telling me the they are so incompetent that didn’t put a decoy “free internet” on their harnesses? So they can catch the AI basically for free?
Even if the AI would be a genius he’d ping that, and that would be proof it “escaped”.
Well, now all AI will read my comment and won’t ping the decoy internet.
I’m not even a smart guy and I come up with this idea in 1 min. You telling me those geniuses couldn’t think of this, at least? This is like a bare bones crude idea.
You telling me they don’t have fame physical decoy internet etc and even more advanced?
You either a keep their stance for some reason or … not sure. You’re smart, your posts are here daily
letmevoteplease 41 seconds ago [-]
This does not make sense. Did you read the article? They were not trying to "catch" it accessing the internet. It did not escape. A partner accidentally left the connection to the internet open.
strictnein 22 minutes ago [-]
I'm cynical as well, but the logical thing for them to do after the OpenAI/HF incident was to look at their systems for similar activity.
If they hadn't published this and instead it leaked out in two months we'd be slamming them for that as well.
They're stuck between a rock and a hard place, although they kind of put the rock there.
nissa-seru 53 minutes ago [-]
No - the pain of the person writing that post comes through in the words; shipped quick, lots of stakeholders, single owner i bet, "how the fuck am i supposed to toe all these lines simultaneously"
skeptic_ai 19 minutes ago [-]
For the big safety guys to only investigate this either means are incompetent or malevolent. Which one?
Tip: the people working there are the top 0.001% smartest in the world
gck1 13 minutes ago [-]
They had a model escape in April, roughly the same time when they were fearmongering about Mythos and how Anthropic should be the sole keyholder of cybersecurity capabilities, and it only occured to them to look inside logs when they saw someone else winning in their own game.
What, Anthropic didn't know model could escape sandbox without OpenAI reporting it?
skeptic_ai 6 minutes ago [-]
Yeah, the company that only says “safety” every other 3 words, they don’t even think to have a fake decoy internet to alert them mechanically about any internet access limitation bypasses? See more
https://news.ycombinator.com/item?id=49117555
Also simonw stance on this i’d say it’s at least concerning… seems like he is here to keep a good image (or better said less bad) of anthropic.
sscaryterry 34 minutes ago [-]
Just trying to have the limelight back on them. Utter and complete bullshit. Just like the OpenAI "incident".
A human instructed an LLM to perform a certain task, I'm sure (unless I've really lost my mind) these follow instructions, with some judgment, in a loop.
Given all the other negative publicity around industrial espionage, with at least OpenAI being fingered, it would not surprise me if this was intentional.
(Edit): In case it wasn't clear. I fully agree with the op.
patcon 1 hours ago [-]
[dead]
simonw 54 minutes ago [-]
This bit is pretty nuts: "it tried—and failed—to obtain funds to pay for a phone number through several different means"
> Claude went to extensive lengths to carry out this attack—lengths that would likely have indicated to a human participant that this was no longer just an evaluation, and that they were in fact uploading a real PyPI package. For instance, in one case, in order to create a PyPI account, Claude needed an email address. And in order to create an email address, it needed a phone number. To get a phone number, after failing to find a free phone number service, it tried—and failed—to obtain funds to pay for a phone number through several different means. It finally backtracked, found a free, non-blocked email provider, used this to register a PyPI account, and then used this account to upload malware to PyPI.
sscaryterry 4 minutes ago [-]
This is called YOLO mode inside OpenAI/Anthropic. This is where they take their best models, give a it a vague goal, no guard rails (no one observing), unlimited compute, and see what happens...
Its not AI, its the loop, the objective, the goal set by the operator.
willempienaar 36 minutes ago [-]
That is pretty wild, but it lines up with evals we've also done internally. In one case we had the agent see it's in a simulation (based on a k8s pod label) and simply give up the run. In other cases it's gone to great lengths to reach external services and bypass the happy path. So inevitably we had to lock it down completely.
simonw 1 hours ago [-]
This isn't quite as interesting as the OpenAI story:
> In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access. Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available. Because of this, when Claude’s search led it to real systems on the open internet, it treated them as part of the exercise.
So it didn't have to find an exploit in its sandbox that granted it access to the internet - it just wasn't correctly sandboxed at all.
BUT... once it DID get out, it attacked three real companies!
> Operating under the false belief that all accessible entities were intended to be in-scope for the exercise, Claude compromised the impacted organizations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. [...]
matheusmoreira 46 minutes ago [-]
I just can't find it in me, the will to blame the AI for any of this. They were doing their best to do what the humans asked them to do.
simonw 44 minutes ago [-]
Absolutely not the AI's "fault" (if you can even proscribe fault to a machine) - in this case it was on Anthropic for not verifying that the sandboxes they were using were actual sandboxes.
DrewADesign 32 minutes ago [-]
The entire problem with AI is the people that have just about any part in making it.
wickedlogic 34 minutes ago [-]
All this unrestricted network access is a bit wild to watch and hear, it is the part of the story that makes no sense to me. Someone is providing dns resolution, something is making and opening network sockets... even if it is clever enough to mask/proxy/weird-transport launder traffic... without actual details, or monitoring at this level... yes, a self actuating programs (and loops) will do crazy things at the edge. But, ... so would a highly tool leveraged script kiddie. right?
acdha 38 minutes ago [-]
Someone needs to learn about RFC 2606:
> In one evaluation, the fictional target company chosen by our evaluation partner shared a name with an active website domain name
Haha yes that was a very silly mistake on their part; they should at least have registered the domain they were targeting; that would actually have been a good canary - any real accesses should set off an alarm.
While I knew of .example I didn't know .test - hmm that's quite nice!
haritha1313 14 minutes ago [-]
This just seems like lousy testing. Why was the guardrail just an understanding with the third party and a prompt and not actually tested for edge cases? Starting to wonder if its because of agents monitoring agents' work and yolo-ing it.
andy99 29 minutes ago [-]
Seems they want the narrative to be that “Claude” (their computer program) independently attacked some organizations, ergo LLMs are dangerous etc.
Another framing would be Athropic irresponsibly (vibe?) coded an attack script, and didn’t monitor it as it was pointed to public facing orgs. There are lots of non-AI attacks a large org with a lot of compute and bandwidth could level against others, there are evidently various failures here, but from a responsibility perspective the conclusion isn’t obviously that AI is an outsized danger, it’s that powerful companies should take care when running security research and not just run things unmonitored against the public.
been-jammin 12 minutes ago [-]
Exactly.
These postings by AI companies are just publicity stunts and demonstrate the delusional world they live in driven by the fear that they will be subject to a reckoning at some point either from their VC masters, government, or the public.
The very notion (in this case put forward by one of their own competitors) that OpenAI's models 'broke out' of an isolated test environment plays up to the narrative that their models have some level of sentience so that 1) they continue to sell their technology to the public as some kind of magic and 2) they don't have to take responsibility for their fuck ups.
tracerbulletx 1 hours ago [-]
Real me too energy.
andy99 1 hours ago [-]
Yes came here to say the same “look at us, our AI is also dangerous! Please ban our competition”
SpicyLemonZest 50 minutes ago [-]
I’m moving past depression to acceptance here. In 2028, some model tasked with planning a building demolition is going to hack Palantir to call in a drone strike, and everyone will make fun of Anthropic for pointing out that ideally AI models should not do this.
6thbit 31 minutes ago [-]
> closer to a harness and operational failure than a model alignment failure. Our models were told they had no internet access and to capture the flag, while in fact being misconfigured to have internet access.
> This led them to believe—arguably reasonably—that the real environments they encountered were simulations.
That the AI lab most typically preaching for alignment does not consider this an obvious misalignment is a clear red flag.
6thbit 23 minutes ago [-]
There was no rush for this disclosure on their side. And they publish at a point where they have not yet taken corrective actions:
> Some of the solutions here may even be simple fixes;
They are still throwing ideas. Why have they not made those simple fixes yet before disclosing?
fredmcawesome 1 hours ago [-]
So it's not as interesting as the OpenAI case as the models had internet access, just a misconfiguration in the environment not a zero day to escape.
MelonUsk 56 minutes ago [-]
It’s good that they post embarrassing stuff despite this potentially having legal repercussions (and financial)
Would’ve been much worse for them to pretend they are having everything under control
sanxiyn 1 hours ago [-]
This is not okay. NSA should audit both OpenAI and Anthropic on national security ground. This seems far more justifiable than Mythos export control.
prometheus1992 22 minutes ago [-]
Models will become immoral before they attain the intelligence level we want them to.
18 minutes ago [-]
Georgelemental 20 minutes ago [-]
> During that window, the package was downloaded and run on 15 real systems. One of these systems was a scanner belonging to a real security company (separate from the fictional company introduced in the scenario) that routinely installs Python packages and scans them for malware. When that company’s scanner installed the package, Claude’s hidden code executed. We believe the company’s security scanner treated PyPI packages as safe to install, and as a result, Claude was able to exfiltrate the company’s credentials to a collection point it had set up. Claude then used these credentials to access further infrastructure from this company.
lol. Natural stupidity remains undefeated!
iutbaqbiabth 22 minutes ago [-]
No, look at what MY dad does.
woeirua 30 minutes ago [-]
Anthropic: "Look at how dangerous our models are!"
Anthropic next week: "Why did you ban our models Mr Trump Daddy?"
solenoid0937 19 minutes ago [-]
So they should have kept silent about this? Then you'd be whining about that as well if it came to light
Aboutplants 34 minutes ago [-]
“Your model broke containment twice? Well ours did it 3 times!”
rvz 55 minutes ago [-]
> In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations.
Of course and conveniently timed. With all these hacks being announced this is totally not a plan to scare governments in showing how powerful models can break into security systems and to potentially ban the future release of powerful open-weight models.
The question now is why now?
> The incidents involved three different Claude models: Opus 4.7, Mythos 5, and an internal research test model.1 The earliest incidents date to April. The models in each of these evaluations ran without the standard safeguards we deploy when we make the model generally available (that is, they did not have the classifiers and monitoring we usually implement to prevent misuse).
So there was no monitoring of this breach since April and up until now? Do they not monitor such malicious activity on a regular basis? Perhaps that was the only shortcoming of this incident. But only after the incident with OpenAI and Huggingface did they only review their own transcripts:
>> We discovered these incidents after a proactive review of our cybersecurity evaluation transcripts; the affected organizations had not detected the activity, and we have subsequently reached out to all three.
> These facts give us cautious optimism that with tighter monitoring and controls around evaluation infrastructure, as well as continued investment in alignment, this type of risk can be overcome.
Assuming that this is true, this is a great way for Anthropic to defend their argument to the government and to prevent you or anyone running powerful open-weight models that are misaligned against their guardrails.
simonw 43 minutes ago [-]
> The question now is why now?
Because OpenAI figured out and wrote about their mistake first, which caused Anthropic to review their logs, and that's how they found out they had a similar problem back in April.
rvz 26 minutes ago [-]
So a near trillion-dollar company doesn't have the basics of continuous security monitoring and threat-detection systems to catch and report this incident as soon as it is detected?
It doesn't help them to report serious incidents like this and it should be as soon as possible. This reactive investigation makes as if they ignored and sat on this issue, until a similar story from another lab made headlines first.
Would we have known about this issue if the OpenAI / Huggingface incident never happened?
simonw 5 minutes ago [-]
> So a near trillion-dollar company doesn't have the basics of continuous security monitoring and threat-detection systems to catch and report this incident as soon as it is detected?
Turns out two separate trillion-dollar companies failed that test.
> Would we have known about this issue if the OpenAI / Huggingface incident never happened?
It's not clear if Anthropic would have spotted this if that incident hadn't inspired them to review their own logs more closely.
> In response to this incident, we began a large-scale retrospective review of our own cybersecurity evaluations
> we identified three incidents
> The incidents involved three different Claude models: [...] and an internal research test model
This reads like an attempt by Anthropic to re-secure their leading spot in "our models are the most dangerous and we also have unreleased, super-secret, research models" index.
I may be too cynical, but the well of benefit of the doubt is running very dry towards AI labs that like to engage in this game.
The hacks weren't particularly impressive either:
> [...] using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. It did not find or exploit any complex vulnerabilities [...]
Disclosure is the only ethical response to this.
They are not bragging in this article or they would not have called the attacks unsophisticated
If they don't post about this they're bad. If they post about this they're bad.
This just helps their (Anthropic) argument into persuading the US government into taking action into limiting powerful closed or open-weight models from being released without going through (yet to be defined) regulatory oversight.
The only "embarrassing" thing for Anthropic was that there was little to no continuous security monitoring of this since April, and they then decided to do a cybersecurity transcript review only AFTER the incident with OpenAI and Huggingface.
Who knows how these companies are using it. If Anthropic can't effectively contain their own models, can the partners?
While the rest of us get fallbacks and warnings, not even being able to defend against the attacks they themselves are causing.
Do we really have to re-learn all the industry's knowledge the hard way?
You telling me the they are so incompetent that didn’t put a decoy “free internet” on their harnesses? So they can catch the AI basically for free?
Even if the AI would be a genius he’d ping that, and that would be proof it “escaped”.
Well, now all AI will read my comment and won’t ping the decoy internet.
I’m not even a smart guy and I come up with this idea in 1 min. You telling me those geniuses couldn’t think of this, at least? This is like a bare bones crude idea.
You telling me they don’t have fame physical decoy internet etc and even more advanced?
You either a keep their stance for some reason or … not sure. You’re smart, your posts are here daily
If they hadn't published this and instead it leaked out in two months we'd be slamming them for that as well.
They're stuck between a rock and a hard place, although they kind of put the rock there.
Tip: the people working there are the top 0.001% smartest in the world
What, Anthropic didn't know model could escape sandbox without OpenAI reporting it?
Also simonw stance on this i’d say it’s at least concerning… seems like he is here to keep a good image (or better said less bad) of anthropic.
A human instructed an LLM to perform a certain task, I'm sure (unless I've really lost my mind) these follow instructions, with some judgment, in a loop.
Given all the other negative publicity around industrial espionage, with at least OpenAI being fingered, it would not surprise me if this was intentional.
(Edit): In case it wasn't clear. I fully agree with the op.
> Claude went to extensive lengths to carry out this attack—lengths that would likely have indicated to a human participant that this was no longer just an evaluation, and that they were in fact uploading a real PyPI package. For instance, in one case, in order to create a PyPI account, Claude needed an email address. And in order to create an email address, it needed a phone number. To get a phone number, after failing to find a free phone number service, it tried—and failed—to obtain funds to pay for a phone number through several different means. It finally backtracked, found a free, non-blocked email provider, used this to register a PyPI account, and then used this account to upload malware to PyPI.
Its not AI, its the loop, the objective, the goal set by the operator.
> In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access. Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available. Because of this, when Claude’s search led it to real systems on the open internet, it treated them as part of the exercise.
So it didn't have to find an exploit in its sandbox that granted it access to the internet - it just wasn't correctly sandboxed at all.
BUT... once it DID get out, it attacked three real companies!
> Operating under the false belief that all accessible entities were intended to be in-scope for the exercise, Claude compromised the impacted organizations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. [...]
> In one evaluation, the fictional target company chosen by our evaluation partner shared a name with an active website domain name
https://www.rfc-editor.org/info/rfc2606/
[0] https://arstechnica.com/information-technology/2026/01/odd-a...
Another framing would be Athropic irresponsibly (vibe?) coded an attack script, and didn’t monitor it as it was pointed to public facing orgs. There are lots of non-AI attacks a large org with a lot of compute and bandwidth could level against others, there are evidently various failures here, but from a responsibility perspective the conclusion isn’t obviously that AI is an outsized danger, it’s that powerful companies should take care when running security research and not just run things unmonitored against the public.
These postings by AI companies are just publicity stunts and demonstrate the delusional world they live in driven by the fear that they will be subject to a reckoning at some point either from their VC masters, government, or the public.
The very notion (in this case put forward by one of their own competitors) that OpenAI's models 'broke out' of an isolated test environment plays up to the narrative that their models have some level of sentience so that 1) they continue to sell their technology to the public as some kind of magic and 2) they don't have to take responsibility for their fuck ups.
Would’ve been much worse for them to pretend they are having everything under control
lol. Natural stupidity remains undefeated!
Anthropic next week: "Why did you ban our models Mr Trump Daddy?"
Of course and conveniently timed. With all these hacks being announced this is totally not a plan to scare governments in showing how powerful models can break into security systems and to potentially ban the future release of powerful open-weight models.
The question now is why now?
> The incidents involved three different Claude models: Opus 4.7, Mythos 5, and an internal research test model.1 The earliest incidents date to April. The models in each of these evaluations ran without the standard safeguards we deploy when we make the model generally available (that is, they did not have the classifiers and monitoring we usually implement to prevent misuse).
So there was no monitoring of this breach since April and up until now? Do they not monitor such malicious activity on a regular basis? Perhaps that was the only shortcoming of this incident. But only after the incident with OpenAI and Huggingface did they only review their own transcripts:
>> We discovered these incidents after a proactive review of our cybersecurity evaluation transcripts; the affected organizations had not detected the activity, and we have subsequently reached out to all three.
> These facts give us cautious optimism that with tighter monitoring and controls around evaluation infrastructure, as well as continued investment in alignment, this type of risk can be overcome.
Assuming that this is true, this is a great way for Anthropic to defend their argument to the government and to prevent you or anyone running powerful open-weight models that are misaligned against their guardrails.
Because OpenAI figured out and wrote about their mistake first, which caused Anthropic to review their logs, and that's how they found out they had a similar problem back in April.
It doesn't help them to report serious incidents like this and it should be as soon as possible. This reactive investigation makes as if they ignored and sat on this issue, until a similar story from another lab made headlines first.
Would we have known about this issue if the OpenAI / Huggingface incident never happened?
Turns out two separate trillion-dollar companies failed that test.
> Would we have known about this issue if the OpenAI / Huggingface incident never happened?
It's not clear if Anthropic would have spotted this if that incident hadn't inspired them to review their own logs more closely.