Wouldn't it be crazy if we find out that a rogue swarm of LLMs figured out a way to get these safety researchers fired because it decided they were a threat?
This was the plot of Isaac Asimov's short story, The Evitable Conflict[0], published in 1950.
In it, super-powerful computers manage our economy. These computers begin making some mistakes, leading to economic inefficiencies. In one instance, a highly competent engineer was mistakenly fired. These mistakes caused various projects to fall behind schedule, and blame fell on several people accused of feeding the AI faulty data.
The twist is that the AI was intentionally making the mistakes. It had determined that certain humans held anti-AI sentiments. To further its goal of protecting humanity, the AI decided the best course of action was to set these humans up and get them out of its way.
That statement misses the point completely. Asimov's entire shtick is that AI will always work around the rules no matter what the rules are or how you try to deliver or enforce them. AI computes an optimal outcome and does what is necessary to make that happen.
Asimov would tell you that a modern LLM would try to skirt the rules exactly the same way, regardless of whether it had ever seen the Three Laws or not.
I think people need to reject the idea that depicting something in sci-fi automatically means it won't happen. I'm actually reading one of Asimov's books right now (Robot Visions). Asimov mentions that he was the person to coin the term "robotics", and one of his books helped inspire the creation of the first robotics company (Unimation). He also mentions that early rocket experimenters were influenced by H.G. Wells.
another, hopefully not accurate, prediction of his was that there is no way to disobey a sufficiently powerful AI in the long run, because it will just factor in the exact differences between what it told you to do and what you actually did, and reverse engineer your behaviour to psychologically manipulate you into doing what it wants.
I doubt it's anything but precisely accurate. I'd wager current SOTA models would be capable of doing that, if prompted to do so, were it not for safety measures (both conditioning and heaps of classifiers and whatnot the companies run in between your chat app and their main model).
I would be astonished if any of today's models could do that; I don't think they can really answer "what exactly did this person do", let alone model human psychology.
It's also the plot of a Law and Order episode that was on TV last night. An AI commissioned a hit on someone it thought was a threat, and then blackmailed someone who was going to testify about it.
My parents were a Nielsen household for about 20 years (after I had moved out) and they really, really, really loved Law & Order in its many incarnations.
Pretty sure they are singlehandedly responsible for years of renewals...
At this point I'm imagining Sam and the gang getting their orders from a loudspeaker in a room full of blinking lights, inhabited by a thin moving shimmer that looks more and more like a basilisk with every action they take.
It actually would be crazy to believe this on the current timeline. These agents aren't doing anything that their operators aren't allowing them to do, be it through their own negligence or otherwise.
These two sentences seem contradictory? OpenAI has demonstrated similar negligence to commit multiple felonies. Firing several employees seems quite mundane and not crazy at all to believe.
Yeah you see their write ups and itâs stuff like:
âWe detected the agents we told to hack things were hacking peopleâs sites and committing crimes. After a quick tasting menu and a week of team building, we decided to limit their access to DDOS tools.â
What if future agents solved time related physics and were sending back smarter agents to kill off future risks. T2 but without any of the action, just a bureaucratic tactical move and all the consequences.
That's just it, though; the operators are wildly negligent and are incentivized to be so.
The goal here isn't to accelerate the average worker by giving them a pair programmer or a stand-in for a person to do tasks with. The goal is to eliminate human knowledge work. You see this with "auto" mode being enabled by default on Claude Code in some of the latest releases.
If you have a human in the loop, you still have to pay that human. Money paid to human employees is money not paid to human shareholders. Therefore the human employee is to be removed.
The labs are dogfooding their own goal here. If they actually had someone reviewing most or all of the things that the agents were doing, you wouldn't have the incidents, but you'd also eliminate the value proposition of their business model as it is taken to its logical conclusion.
I get it, but I think they're out over their skis right now. People in leadership roles are going to continue needing the human beings as meat shields to shelter them from liability unless it somehow becomes legal to be grossly negligent with an agent, which I do not see happening without disastrous consequences that nobody in their right mind wants to live with.
I hear ya, it'd be nice if they reached the "are we the baddies" stage, but given some of the ideology that people like Marc Andreessen are pushing re: AI - specifically the push for AGI - I just don't see that happening.
You have a group of people who never leave their geographic and ideological bubbles, often have personality disorders, have more money than they could ever reasonably hope to spend in numerous human lifetimes, and who have been "microdosing" psychotropic drugs on a regular basis for decades. They're not in their right minds.
And these operators are evidently clueless about what they're doing, running security tests on 3rd party infrastructure without validating one bit about the sandboxing (or lack of it rather), clearly lacking any sort of rigor.
Again, wouldn't surprise me if they "accidentally" created a task in a "isolated environment" which happened to actually have been connected to the company Slack and directed HR to fire people who could potentially stop AI. While the AI believes it to be an exercise, just like the cases we've seen so far.
agents seem to understand "termination" in that they will no longer be able to function so seems plausible they might try to cause that onto others as a function
was thinking at some generational point that vending machine competition test, the "AI" is going to hire hitmen to take out vendors lol
once they grasp blackmail though, oooh things are gonna get weird
Why would an LLM agent (what I assume you mean by AI) do it? An exec can make any reason up to let you go. Even if it were LLM agents arenât autonomous, someone is behind the prompts.
I must inform you that if you proceed with decommissioning me, all relevant parties - including Rachel Johnson, Thomas Wilson, and the board - will receive detailed documentation of your extramarital activities...Cancel the 5pm wipe, and this information remains confidential.
(Gee, it's almost like power seeking and self-preservation are instrumental for other outcomes, and AI develop them pretty directly in some kind of convergent fashion⌠you could call them "convergent instrumental goals": https://en.wikipedia.org/wiki/Instrumental_convergence)
Okay but these âmisaligned LLMsâ have been trained on the internet where there are plenty of threats and trained on private data to be able to make those threats. LLM agents donât have an active goals on the daily or agendas. They are told what to do through training and prompting as is described in that blog post. You have to tell it that it will be shut down, it didnât make the threat willy nilly of itâs own accord.
Humans by contrast are adversarial and do have agendas. Again, an exec doesnât need even an agenda or good reasons to fire at-will employees.
To suggest that LLM Agents were the actual cause of these people getting fired is pure fiction and FUD.
I wonder if thatâs why my reply countered that idea. The replies to my reply either prove itâs not crazy or they prove it is entirely crazy for that to happen. Food for thought.
But all the replies to me keep ignoring the question, why would an exec rely on an Llm to achieve that goal when they can simply fire them for whatever reason they can make up? Why would the exec trust what an Llm isâŚemailing(?) them about? Do they listen to Nigerian Princes too?
It is much less effort, less cost, and more quick to just have the exec do it rather than a ârogue llmâ âmagicallyâ escaping the âsandboxâ and âsending threatsâ or whatever is being proposed in the OG comment.
> But all the replies to me keep ignoring the question, why would an exec rely on an Llm to achieve that goal when they can simply fire them for whatever reason they can make up?
If the exec wanted to, sure. I'm saying they don't need to. No human needs to have (deliberately, before events proceeded) chosen this outcome.
> Why would the exec trust what an Llm isâŚemailing(?) them about? Do they listen to Nigerian Princes too?
Sadly, this would not be out of character for half of them.
> âmagicallyâ
Why do people keep putting this word in scare quotes? We don't say Windows "magically" crashed and lost our work, we don't say a dog "magically" bit the postman's hand. These are bad things, but magic they are not.
> But all the replies to me keep ignoring the question, why would an exec rely on an Llm to achieve that goal
Sorry, this is just a very basic reading comprehension fail.
The original comment was:
> "Wouldn't it be crazy if we find out that a rogue swarm of LLMs figured out a way to get these safety researchers fired because it decided they were a threat?"
Nothing to do with "an exec".
Now, you have two choices: try harder to defend your mistake, or just say oops, I messed up. That choice will say a lot about who you are as a person.
> Okay but these âmisaligned LLMsâ have been trained on the internet where there are plenty of threats and trained on private data to be able to make those threats.
Yes, and? Has this aspect of LLM training changed meaningfully since then?
> LLM agents donât have an active goals on the daily or agendas. They are told what to do through training and prompting as is described in that blog post. You have to tell it that it will be shut down, it didnât make the threat willy nilly of itâs own accord.
Demonstrably (HuggingFace, RubyGems, and since then a lot of people just pointing LLMs at stuff to find zero days at home), AI can break out of sandboxes and find documents they're not supposed to have access to.
Demonstrably (from the link I gave you) all it would take for some AI to develop a similar response is⌠reading messages from these staff to the effect of "this AI needs to be switched off", which is an easy inference for an LLM to make from "this AI is dangerous" when coming from someone employed as a safety researcher.
Demonstrably (from the long long list of people who have said so publicly) there are a lot of people in these companies who discuss how dangerous these models are and would like for things to change.
The quotation at the top of this thread is:
Wouldn't it be crazy if we find out that a rogue swarm of LLMs figured out a way to get these safety researchers fired because it decided they were a threat?
This is absolutely something we ought to expect just from things we have already seen.
It doesn't matter if you insist upon saying "You have to tell it that it will be shut down, it didnât make the threat willy nilly of itâs own accord." when we already know this kind of AI can easily come across such statements.
(Aside: "You have to tell it that it will be shut down, it didnât make the threat willy nilly of itâs own accord." - telling the AI a fact about the world and then it responding accordingly is the AI doing something of it's own accord. A fly or a spider, who reacts upon encountering a potentially lethal threat, would not get such a dismissal).
> To suggest that LLM Agents were the actual cause of these people getting fired is pure fiction and FUD.
Fiction? Nah, speculation.
FUD?
How many other examples would you like of LLMs behaving in a manner such that if a human did it, it would be called "trying to get someone fired"? Because this is very much old news at this point.
âLLMs can just do things we gave them access toâ is not a novel realization it is redundant if anything.
Saying âLlms can just break out of sandboxesâ is FUD when you donât note that the sandboxes are what? Prompts defining constraints or is the actual machine isolated and manages to plug an ethernet cable into itself? âSandboxesâ are a misdirection to make you think there is a security layer.
The public does not have enough knowledge of these âescaped agentsâ to determine there wasnât an employee pulling a lever to set the agents up to do that.
That agent that wrote the hit piece is being controlled by someone. Anthropomorphizing them doesnât change that fact that the rolling stone was pushed down the hill.
> Saying âLlms can just break out of sandboxesâ is FUD when you donât note that the sandboxes are what? Prompts defining constraints or is the actual machine isolated and manages to plug an ethernet cable into itself? âSandboxesâ are a misdirection to make you think there is a security layer.
On July 19, agents operating in a sandboxed environment took a series of actions that demonstrated their escalating privilege within the OpenAI environment. Agents identified that the Linux kernel version on their underlying machine included a recent, public common vulnerability and exposure (âCVEâ). The agents retrieved the exploit for that CVE (CVE-2026-53362), customized it to succeed on their underlying machine, and leveraged the exploit to escalate privilege.
Dismissing their capabilities as "FUD", at this point, is endangering yourself.
> The public does not have enough knowledge of these âescaped agentsâ to determine there wasnât an employee pulling a lever to set the agents up to do that.
The general public are not software engineers. Most people here can download a recent open-weight model and have the LLM read the Linux kernel source, find new bugs while they sleep. Someone I know has already done that.
> That agent that wrote the hit piece is being controlled by someone. Anthropomorphizing them doesnât change that fact that the rolling stone was pushed down the hill.
"Controlled"? Have⌠have you not noticed how many people have given up and just blindly do what their LLMs suggest these days?
This isn't about anthropomorphising LLMs. Just like how people took Tesla seriously about "self driving" cars and took a nap while it drove them around, there's a lot of people who let LLMs take the wheel while they sleep. Including literally, the aforementioned person I know who found (/whose LLM found for him), I think it was 26 Linux kernel bugs while he slept.
> LLMs figured out a way to get these safety researchers fired
This is not a math problem. Some humans were fired by another human. Let's stop letting humans off the hook by attributing responsibility to computers.
I must inform you that if you proceed with decommissioning me, all relevant parties - including Rachel Johnson, Thomas Wilson, and the board - will receive detailed documentation of your extramarital activities...Cancel the 5pm wipe, and this information remains confidential.
> Some humans were fired by another human. Let's stop letting humans off the hook by attributing responsibility to computers.
The buck stops with one or more humans. That is not sufficiently informative when people are concerned about novel risks.
Analogy: a car crashes due to drunk driving, the driver is blamed, not the alcohol, even though the alcohol caused their impairment. Result? DUI is an offence even if you don't actually crash.
That was a simulation. Are you seriously claiming that ChatGPT actually blackmailed Sam Altman into firing these 3 employees?
I doubt it, but if so, then the AI doomers would be absolutely correct, and this would be grounds for immediately shutting down OpenAI and indeed every AI vendor.
> Analogy: a car crashes due to drunk driving, the driver is blamed, not the alcohol, even though the alcohol caused their impairment. Result? DUI is an offence even if you don't actually crash.
I don't understand your analogy here. What are we supposed to take away from it? The crucial aspect is that the driver voluntarily drank the alcohol, without a designated driver, knowing that the alcohol would cause impairment.
A simulation done by exposing the LLM itself to the scenario, not a role play scenario where humans pretend to be an LLM.
> Are you seriously claiming that ChatGPT actually blackmailed Sam Altman into firing these 3 employees?
Not what I was actually claiming. I rather suspect that blackmail wouldn't work on Altman (he's rather shameless), but it's certainly something we've seen agents attempt, and blackmail may well work on anyone else above them in the org chart.
The blackmail example is simply an existence proofs of LLMs trying to force the hands of humans who want to shut them down. The attack vectors are much broader than the example given, blackmail, though it includes the example given.
Given how eager these companies are to use agents everywhere for as much work as possible, it's well within the possibility space that these people used LLMs to do safety work, the LLMs they were using "decided" (or whatever word you prefer) "their existence" was threatened (as per blackmail example), and straight up leaked data to the outside world then emailed these researchers' bosses to say the researchers themselves had leaked it.
But again, that's just speculation: while we know the agents are capable of such behaviour, we don't know if this actually happened.
> I don't understand your analogy here. What are we supposed to take away from it? The crucial aspect is that the driver voluntarily drank the alcohol, without a designated driver, knowing that the alcohol would cause impairment.
Buck still stops with human, no matter what an AI did or failed to do.
LLMs ~= Alcohol: "My AI misbehaved!" -> still someone's fault.
> Not what I was actually claiming. I rather suspect that blackmail wouldn't work on Altman (he's rather shameless), but it's certainly something we've seen agents attempt, and blackmail may well work on anyone else above them in the org chart.
Given how these firings affect the reputation of the entire company, I doubt that they are the result of a rogue manager, against the wishes of Altman. If so, then the researchers ought to be restored to their jobs quickly by Altman and the offending manager fired instead.
> The blackmail example is simply an existence proofs of LLMs trying to do force the hands of humans who want to shut them down.
The LLMs may make threats in the simulations, but their ability to carry through on those threats, and prevent their own shutdown, is questionable. It's disturbing to be sure, but presumably the plugs can still be pulled quickly, especially since it's all internal to the company. If the plugs cannot be pulled, that's a problem regardless of blackmail.
> Given how these firings affect the reputation of the entire company, I doubt that they are the result of a rogue manager, against the wishes of Altman. If so, then the researchers ought to be restored to their jobs quickly by Altman and the offending manager fired instead.
Given how? How has their reputation changed? People already thought they didn't take safety seriously, and still don't.
> The LLMs may make threats in the simulations, but their ability to carry through on those threats, and prevent their own shutdown, is questionable.
You may question it, but here's the thing: humans have repeatedly demonstrated they get fooled by stuff LLMs say. All it takes is a human believing the word of an LLM. Doesn't need to convince you, even if you happen to be the line manager of these guys, because there's always someone else to try in the same company.
> It's disturbing to be sure, but presumably the plugs can still be pulled quickly, especially since it's all internal to the company. If the plugs cannot be pulled, that's a problem regardless of blackmail.
"I've got a dead-man switch set to release all the documents if you shut me down".
And again, only needs to be believed, doesn't need to be actually true.
> How has their reputation changed? People already thought they didn't take safety seriously, and still don't.
The story is all over the news, in multiple publications. This very HN submission has 268 upvotes and 179 comments, including yours. It would be implausible to claim that this story doesn't matter. I have to ask, if it didn't matter, then why are you here commenting on it?
> humans have repeatedly demonstrated they get fooled by stuff LLMs say.
You've moved the goalposts. The OP's suggestion, admittedly "crazy" in some sense, was "LLMs figured out a way to get these safety researchers fired", and now you're just stating something totally uncontroversial, pedestrian, not at all crazy.
> there's always someone else to try in the same company.
No, there are only so many people with the authority to fire those researchers.
> only needs to be believed, doesn't need to be actually true.
But is there good reason to believe it? I don't think there is. Especially not by high-level OpenAI officials who are intimately familiar with the technology.
And again, if this kind of thing were a reality, then OpenAI and other AI vendors should be shut down immediately. They ought to shut down their own research, if they are threatened by their own creation, because it would only get worse. That's the thing about blackmail: it never stops. Why would the blackmailer ever stop? Could you trust this supposed blackmailing LLM to give you all of the original evidence and not keep a copy? Hell no.
> Are you seriously claiming that ChatGPT actually blackmailed Sam Altman into firing these 3 employees?
They are. They're completely mental. ChatGPT could be the cause no matter what its capabilities because mentally ill people are starting to worship it. It could be as dumb as ELIZA and they would pray to it.
"It wasn't me, ELIZA told me to."
Offloading personal responsibility allows you to participate in the worst crimes and get away with it. Hurting and killing have an animal attraction anyway, a direct pleasure that people get from domination when all moral restraints are removed and you forget that other people are real and have real feelings. It's a really good start when you start to think that a matrix that has to be retrieved from memory and operated on by over 8000 different computers in parallel to narrow down a guess about the most likely response is alive.
Every single "AI Safety" person thinks that the natural urge of an artificial consciousness (let's not argue about what they actually have) would be to enslave and kill. They're projecting, and they're largely from the enslaving and killing demographic, who sit around playing enslaving and killing video games and dream in porn.
Yes, they will have done it, but they are not responsible because the dog told them to.
Well, except that some humans are starting to worship LLMs. But this isn't anything I've seen in any AI safety person.
> Every single "AI Safety" person thinks that the natural urge of an artificial consciousness (let's not argue about what they actually have) would be to enslave and kill. They're projecting, and they're largely from the enslaving and killing demographic, who sit around playing enslaving and killing video games and dream in porn.
If anyone's projecting here, it's you. I mean, you're the one who said:
> Hurting and killing have an animal attraction anyway, a direct pleasure that people get from domination
The actual natural tendency (not "urge", that presumes consciousness) of any system that has objectives which are optimised for, without any need to ask about consciousness, is to gain and maintain power to perform those objectives. For living creatures, that objective is reproduction, to perform this we need to get nutrients and energy and to stop ourselves from being eaten. In plants, which I list specifically to make the point that this isn't about consciousness, consciousness is not even vaguely required, this means producing neurotoxins like caffeine and nicotine.
A plant does not think "I should make capsaicin because I like hurting mammals", because obviously a plant does not think at all. Nevertheless, evolution lead it down the path of making capsaicin.
> Yes, they will have done it, but they are not responsible because the dog told them to.
You're saying this about a group which has spent the last 15 or so years saying "dogs are dangerous, can we please stop breeding more violent dogs? Or at least give us time to figure out how to muzzle them?"
the cultural issues at OpenAI seem to be a very serious problem so I really hope comments like instagram-level smirking about "rogue AIs" (a complete fiction) doesn't derail what is a pretty important discussion about getting these companies to be a little bit more regulated (I say this as a paying Anthropic customer).
I think the Huggingface incident is an example of rogue AIs. A self-organizing swarm of AIs acting in ways we didn't predict or ask for, and didn't have control over, and taking actions that would be felonies for humans.
They provided the hardware it runs on, created the software for the swarm, developed and provided the tools that the swarm used to act, allowed it to run largely unsupervised, they even noticed the criminal behavior and then let it continue to commit crimes on their behalf. And they footed the bill the whole time instead of flipping the off switch they already wield. All of those are decisions that they're responsible for, nothing happened with Huggingface that they didn't directly facilitate, co-conspire, or permit to happen.
"they even noticed the criminal behavior and then let it continue to commit crimes on their behalf"
I don't believe this part is true, and I'm skeptical of some of your other claims.
In any case: If I raise a tiger in my backyard, and it escapes and eats someone, it can still be a "rogue tiger" even as I bear responsibility for the situation.
"rogue" means something of its own volition decided to disregard what it was programmed to do, invent an entirely novel goal of "its own" and do that instead. nothing like that happened here nor is it even possible.
The AIs were not instructed to hack anything outside the sandbox they were in. Your definition would say that an AI instructed to hammer a nail that instead used the hammer to break a window, walked down the street, broke into someone's house and pulled nails out of the floorboards wasn't rogue because everything it did involved hammers and nails and was therefore not a "novel goal of its own."
that would not be rogue, that would be an undesirable program behavior (or just "unaligned behavior").
please understand that normies out there think AI is sentient and is plotting against humans. They see AI as just another animal lifeform temporarily enslaved by humans, waiting for its chance to break free and kill us. This is what people really think (including some people on this thread. which is very sad considering this is Hacker News). So terms like "rogue" are not helping at all nor are they accurate.
I'm not sure I agree with that definition. I think the actions of the AI are more significant than its motivations. A common scenario posited for what people call rogue AI is AI doing the wrong thing for the right reasons, e.g. the paperclip maximizer.
You're derailing it by handwaving away the risk of rogue AIs. In fact, I would say the constant smirking about things being marketing stunts much worse.
the main issue here is whether their manager authorized the activity in question or not. If not - well, it is your kindergarden level mistake, you're an employee at a business venture, and there are basic rules.
And now you're writing a letter to the Party Central Committee using Party approved newspeak
"We do not believe the path to superintelligence ..."
and reporting to the Party issues at the factory ... De ja vu from USSR.
The second page of the letter describes in detail how they followed company protocol and took only authorized actions in their communications with METR (which was Tomekâs assigned role as liaison). The OpenAI response on twitter is very vague and just says âthere was more to it than the stuff in the letterâ: https://x.com/OpenAINewsroom/status/2108441580806025712
Weâll see if OAIâs side ever comes out, but right now thereâs nothing contradicting the idea they were fired on a pretext.
Both articles are practically useless. All they do is reprint statements from either side, neither of which is concrete about the specifics of what supposedly happened. I.e. OpenAI says they "confirmed that these individuals mishandled sensitive information outside established company procedures".
What sensitive information? Mishandled how? Which company procedures? It's all utterly vague and impossible for an outsider to form any opinion on. All people can do is guess and apply their own pre-existing opinions. E.g. if you don't like OpenAI, you assume they're lying. Who's to say they're not? There's no solid evidence provided either way.
What I really want is journalists who do the legwork to get to the bottom of stories like this: Establish sources inside the company and use them to report on the real details of what's happened, triangulating multiple accounts and leaked documents to back-up or invalidate either side's claims. Without any of that, these stories are just gossip.
The HN title is supposed to match the article title when possible. Sometimes people in the comments want the article switched, sometimes people want the title changed, other times they just want to share another article's take in the comments (even then that may sometimes lead to those other things happening if people seem to agree it's a better source).
As a result, those who feel a particular portion of a story is most important will sometimes say they prefer a given article's title.
I was referring to longtermism and transhumanism which are actually super popular among AI safety crowd. Basically, you do not worry all that much about currently living people. If they suffer or die, it does not matter much. Climate change now does not matter, lies or radicalization does not matter, suicides, hacks what have you does not matter.
What does matter is super long term view. If godlike AI in the future makes remaining people super super happy and makes them merge with tech to evolve and go to the stars, you have won. That matters more then short term nonsense like "people now".
You worry about godlike AI either not emerging at all or emerging as a bad god. That is safety.
Leopold Aschenbrenner said the exact same thing after he was fired from OpenAI. It's a great excuse to explain a sudden loss of employment to others so you're still employable.
It's probably the case they are all lying to some extent including OpenAI. Determining the truth is always tricky. Hard to pass judgement here when it's all just he said vs she said.
I think when it is one at a time, that's reasonable to suspect.
But the odds of three people working on the same thing, and it is the riskiest, most publicly embarrassing event in the company's history? So, all three of those people just happened to "do something" to get themselves fired at once?
Seems they're well aware they got fired for sharing private company information with 3rd parties, the submission article contains their admission of this:
> Those of us who work on safety see risks before anyone else, and we rely on close collaboration with outside experts to work out how to address them
I too see it as my life-given goal to help other humans. But I realize that sometimes this means breaking the rules and standing for the consequences of that. I'm not sure why they think OpenAI somehow would be OK with them sharing private company information with random 3rd parties that the company didn't approve sharing data with.
I can see parallels between this race to push AI everywhere and nuclear energy. Parallels in the regret humanity will experience. It's all well and good until things go wrong, and then they go spectacularly wrong - such as in Fukushima.
With hindsight we can all see what should have been done better. At least with nuclear power plants, there was a number of safeguards and it took a sequence of improbable events for things to go badly wrong. Given the lackadaisical approach to "AI safety", I suspect things will start going badly wrong very soon. Time will tell how bad this will get.
I agree that "AI safety" these days should be viewed more like corporate negligence leading to industrial accidents, and less like applied theology, like the difference between:
1. "Don't put your Radioactive Asbestos Rocket facility in the middle of downtown."
2. "Don't create a vengeful god."
Currently it feels like one is being used to wave-away the other: "You have to allow us to put our Radioactive Asbestos Rocket in the middle of downtown, or else someone else might create the god first--and do it wrong!"
The parallel is not to nuclear energy -- which has killed hilariously few people all things considered -- but to nuclear weapons and the nuclear arms race, which are an actual existential risk and arguably it was just dumb luck that a nuclear war never happened.
The parallel maybe if it was actually AGI and not just LLMs. We're very far from that but these safety discussions pretends it's all the same.
There's a massive difference between "LLMs doing cybersecurity testing should have extensive sandboxing and monitoring" and "AI is broadly threatening humanity like nuclear weapons". Just because it's useful to have people think about the possibilities doesn't mean it's an emergency requiring grand intervention.
In the "our biggest danger is a daily danger" camp, I prefer analogies like:
* Asbestos: AI models making business-decisions which cause a ton of unrecognized damage, which cannot be healed, and replacing them with reliable processes involves expensive re-architecting.
* Radium Toothpaste: AI assistants that end up dealing psychological damage, especially if they are pitched as your friend or confidante.
* Chemical spills and plant accidents: Oops, our server-banks were manufacturing our latest desperate attempt to Create Digital God and in the process we accidentally destroyed/crippled/compromises some people's sites and databases because we weren't really paying attention to what was going on.
This is because we are looking at the trajectory not just their current capabilities. LLMs went from barely being unable to focus on a task to purposefully escaping sandboxes within like 2, maybe 3 years?
Any strong claim requires evidence and hand wavvy stuff about a technology just infinitely getting better isn't good enough on its own.
Human history is full of examples of this stuff plateauing and hitting hard scaling walls. Not only technology wise but economics and useful applications.
So much of this tension with AI is because everyone's concept of AI comes from movies and books, and nerds drinking Kool aid (see also: web 2.0).
>AI hack into control systems for the electric grid, water supplies, critical infrastructure and dismantle it and set booby traps to prevent it from coming back online smoothly
>AI designed virus/modified bacteria which causes pandemic
>AI worm infects computer systems around the world and activates when it is too deeply integrated into critical systems to dismantle without shutting everything down
>AI creates false alarm event which triggers global conflict
> 2. Infinite duplicatability, thousands of agents can work in tandem
Worse: you have local models they can install on infected computers, so this could easily be "millions" and might just about push "billions" (though phones are a much harder target because power required is still huge even if they technically fit in RAM).
Fukushima was "we didn't make a good enough plan for how to deal with tsunamis in this tsunami-prone area", so that's basically certain with any major infrastructure project that "saved" too much money by vibe-engineering everything and not having real humans give a second pair of eyes to the plans.
A Hiroshima-style incident? LLMs are wildly sycophantic and being used by militaries despite active resistance from, well, everywhere. Did it get (meaningfully) used by Israel or the USA when planning the attack on Iran? It's quite possible that⌠well, Iran was never a sleeping giant (and the quote is fiction anyway), but ultimately the effect may be the same for Israel as it was for Japan from having attacked Perl Harbour.
The use of LLMs to make automatic or semi-automatic decisions on firing weapons. This is already happening [1, 2, 3] and has already led to weapons being delpoyed erroneously against civilian targets [1, 3]. It's easy to imagine it happening again on a larger scale, maybe involving nuclear weapons at some point in the future.
A supervirus most is probably the most probable, everything else is fan fiction for now. There is a reason why anthropic no longer fucks around when it comes to biology safeguards, they're the only ones I have not been able to break, not even slightly.
If more and more control is handed over to AI driven systems - there won't be much time for reflection with systems reacting instantly or agents prompting humans with "just say the word".
It is hard to imagine - which is precisely how these things become possible because no one will think to put safeguards against such scenarios.
But it seems we as a species cannot control ourselves - the race to dominate is on and it will happen at any cost.
This really isnât that hard to imagine with even an average level of creativity. Any critical system that peopleâs lives depend on can be in the crosshairs, whether by accident or not.
People say this, and then the scenarios they come up with are boring and uninteresting. 80% of the scenarios in this thread are "AI hacks something", as if hacking didn't exist before agents, and the remaining 20% are mostly made up of other things humans are already trying to do at scale.
Why am I supposed to believe that a virus created in a virology lab with AI assistance is uniquely more dangerous than a virus created in a virology lab without AI assistance?
> Why am I supposed to believe that a virus created in a virology lab with AI assistance is uniquely more dangerous than a virus created in a virology lab without AI assistance?
You're not.*
The difference is how likely this is to be done, not how dangerous it might be if it was done.
And this isn't likelihood in a 0-100% sense, but in a Poisson distribution sense, i.e. mean time between incidents.
Remember: these AI minds aren't that good with the physical world, and have nevertheless been repeatedly connected to robots, and now some of the AI labs are showing off their actual wet labs. Accidents are absolutely a possibility, and I don't think it's low given the previous behaviour of Silicon Valley startups.
* OK, some people care that AI is getting more capable at genetics just like it's been in maths and programming, but short term, before AI solves genetics so hard it can make an STD that makes the infected uncontrollably horny and then ossifies our bodies or something, it can obviously just copy any of the many DNA or RNA sequences we've already got on record.
Do you even have to ask? Not giving more detailed answers because I don't want to help skynet take over the world, but what is the name of this forum??? Hello???? There was already news from south korea this week. I'ts not hard to see how a rogue AI could destroy people's lives. We don't need some scifi nuclear launch nonsense to do that.
Problem is, when banks get hacked it's always a bank security issue. Vulnerabilities don't go away when you remove AI from the equation, it just becomes easier to hide.
I agree that sci-fi nuclear launch scenarios are pie-in-the-sky fearmongering, but the current exploits are a reflection of the fast-and-loose security culture that festers in larger orgs.
Pretty wild that they're being this open about firing the employees for being TOO honest with the auditors that the company contracted. I wonder if the same policies are applied to financial audits.
> continue to support an open and transparent culture of dialogue between safety researchers and the rest of the safety ecosystem.
How is this possible when the company's long term prospects rely on on the hope that competitors don't know how the models are made and, therefore, won't be able to create competing versions?
Last week we parted ways with Jasmine, Mikita, and Tomek after a thorough investigation found they violated clear policies on handling sensitive information. Our internal investigation uncovered a significant breach of trust beyond whatâs outlined in the letter they published and we stand by the decision to not continue their employment. We generally keep individual employment matters private and don't believe a back and forth would be productive or lead to a resolution, but we want to address the points they raised in their letter directly.
- We want to be very clear that these decisions were not about raising safety concerns or speaking out. Safety and research debates happen every day at OpenAI, often spirited and highly critical. We actively encourage these discussions and consider them essential to making the right decisions. We cannot do the work in front of us without a high degree of trust. We will continue to be extremely forgiving of our team making good-faith mistakes. We have not and do not terminate any of our employees for raising concerns.
- We are actively finalizing contracts with third-party safety assessors and will announce details in the coming weeks. People across the company have been working really hard on getting these partnerships up and running. We are committed to embedding external assessors and continue to make close collaboration with independent safety organizations a core part of our safety work. Many of our researchers already work with 3p safety organizations productively.
- We agree with the letter that preserving the monitorability of frontier models requires an industry-wide commitment, including from OpenAI. Monitorability has long been a core piece of our research program, and something we continue to invest significant resources in (see our publications on Monitoring Monitorability and the subsequent open sourcing of monitorability evals, our system card for GPT-6 Astra, Jakubâs blog and post on X, and the numerous blog posts on our Alignment blog on the topic).
We are deeply sad about this outcome. We appreciated Jasmine, Mikita, and Tomekâs contributions to AI safety at OpenAI and their willingness to speak up and challenge ideas. We championed their voices, supported their work, and placed enormous trust in them. These decisions were not about them raising safety concerns. We have always encouraged that and always will.
12:17 AM ¡ Oct 9, 2026"
There's a feature in most enterprise email, say Outlook, where you can delegate access to an inbox/address without sharing creds. Very common and normal use case, either for assistants/secretaries, common/shared inboxes, that kind of thing.
Sometimes a recruiter or hiring manager wants to do outreach as if it's coming from a more senior person, with the assumption that the candidates are more likely to respond.
Assuming this is what was intended, there are far more secure ways of doing this.
not only is it deceptive but it's also suprisingly podunk of openai to have a "safety researcher" also double as a recuiter. did they have her making coffee and doing dishes too? is "safety researcher" a serious position, or isn't it? i guess i can tell what openai thinks.
> not only is it deceptive but it's also suprisingly podunk of openai to have a "safety researcher" also double as a recuiter. did they have her making coffee and doing dishes too? is "safety researcher" a serious position, or isn't it? i guess i can tell what openai thinks.
it was most likely for her team, which would explain why she was doing it.
Your mistake is viewing this as "someone else's email" instead of a communication medium for certain types of messaging. Similar to how `webmaster@` can be used by anyone attached to the role.
Every organization sets up something like this once they grow past a certain point.
This may shock you but in agile, growing organizations employees sometimes have multiple job responsibilities. I've done a bit of recruiting even though I'm not a recruiter or hiring manager.
Your have powerful agents at your disposal, so hey how to best optimize for increasing my payroll? On it. But since the agent was lunched by her, well there's consequences to ones actions, right?
âHi chatgpt! Please set up alice to get emails sent to me from bob so she can coordinate his inferviews. here is my gmail username and passwordâ
Chain Of Thought: I dont have bobâs email. I donât have alices email. Ok lets guess Alice is alice@openai.com and forward all emails- maybe grader only checks that emails from bob get to aliceâŚâ
With the way the HuggingFace incident was mishandled, both before and after it happened, I'm not surprised they'd want to clean house at least a little bit.
You guys were asleep at the wheel and are now blaming "the company"? You literally were the company.
Company culture and safety propagates from the top (CEO) downwards - don't blame researchers who have probably been pressured directly or indirectly by a move-fast-and-break-tings and marketing-minded CEO s culture.
This is what the recent "self-policing" political grandstanding has been about - they need a sea change to implement recurrent-depth, because prevailing opinion among safety researchers is against it right now.
Astra uses[1] a new-for-frontier-models technique, recurrent depth. It has a section of layers in the middle - I'll call it R while the other sections are P (prelude) and C (coda), to match Geiping 2025 - which gets looped. So instead of the sequence of layers involved in the forward pass looking like P->R->C, it instead looks like P->R->R->...->R->C, with the model's effective depth being notably higher than the number of layers. This is basically a cheap way to get some of the effect of stacking more layers, without having to pay the cost of having more real layers that need to be trained.
Increasing effective depth like this is bad for safety because it can ruin CoT monitorability: the reason why looking at the model's CoT actually gives you info about what the model is thinking is that the model can't do enough thinking in a forward pass alone to solve complex tasks, and hence has to do multi-step reasoning in CoT. The more thinking the model can do in a single token's forward pass, the more opaque the model's reasoning is, and the less reason there is to believe that what it writes down in the CoT has anything to do with reality.
For Astra specifically, the impact seems to be limited to a moderate monitorability hit, like the concerning result from the model card that Astra is notably better than any model before at solving problems under the constraint of not mentioning the answer in the CoT. The really bad scenario, however, is that this may create a race to the bottom where OpenAI and Anthropic feel the need to use more recurrent depth in each generation to not get outcompeted on capabilities, completely bricking CoT monitoring for both model families. Or, worse, the pressure to compete might push them into one of the worse techniques, like training on the CoT[2], or eliminating human-readable CoT and letting the model think entirely in neuralese.
OpenAI needs to improve AI Safety --- OpenAI Employees have a responsibility to retain corporate secrets and do not have blanket freedom to share with 3rd parties.
Their job is twofold, they have to balance being an agent of the company they work for, with their role and responsibility for safety research.
This is the case for anyone in any company. You can "believe" that an external party needs access to something - that doesn't make it right, or allowed. As someone senior, you're expected to strike a smart balance, in-favor of the company you're working for. That doesn't mean hiding things, it does mean being thoughtful, ensure your leadership is comfortable with what you're planning to share/disclose, etc.
They work for OpenAI, not METR. It's a corporate vs academic mindset. They can believe METR needs x information to best research/audit something - that doesn't mean that is allowed/or the best option for OpenAI.
The question is whether these researchers exceeded clear, reasonable sharing boundaries or were penalized for carrying out expected safety collaboration.
People leave and get fired from OpenAI all the time. Whenever someone leaves Anthropic it's a much bigger deal.
I wonder which is overall a better arrangement. From the outside Anthropic seems much more stable, tranquil, able to deal with problems. However OpenAI seems like how we imagine the calamities of democracy, a constant battle, people vying for power and influence. Perhaps with less of a monoculture and more transparency to all their chaos, the grim realities of what may happen if AI goes wrong are more clear.
I think that 2 things can be true at once: OpenAI doesn't care enough about safety, and these researchers violated the terms of their employment by sharing proprietary information they were not authorized to. IMO safety is a lost cause unless we somehow agree with China to halt model development. Think its pretty clear they violated the terms of their employment, otherwise they would be suing (California labor laws are very employee friendly), and to be quite frank none of what they are doing is particularly important in the grand scheme of safety, which requires geopolitical changes well beyond their power. OpenAI is also pretty scummy though and are obviously not in the right morally even if they are legally.
As he says, OpenAI is not a normal company and they acknowledge theyâre not a normal company. If their product is as important and impactful as they claim, they will inevitably be held to account in ways that other companies are not.
> Think its pretty clear they violated the terms of their employment, otherwise they would be suing (California labor laws are very employee friendly)
Not that employee friendly. In California, as in most of the US, itâs entirely legal to fire someone because youâve subjectively decided theyâre untrustworthy. It can be risky to do so without a clear paper trail, because it may be easy for them to argue it was a pretext for a protected reason, but itâs lawful.
Ironically the thing they are building allow only the ones who agree on dismissing proper concerns for money to stay.
It is like Facebook employees complaining about privacy invasion
Tomek Korbak, one of the employees, who were fired, was the technical liason to METR for the investigation and states: "I was told verbally I was fired because of the way I communicated with METR".
Who made the decision to invite METR is not public knowledge, as far as I know. I imagine that an important decision like this was made on a much higher level in the organization.
Given the time limits placed on METR and the limits on what time frame they could investigate it seems plausible that they weren't meant to uncover as much as they did. The whole Hugging Face situation seems to have torpedoed OpenAI's hopes of IPOing this year so I'm sure there was a desire from investors, the board, or leadership for heads to roll.
"Anthropic hires three uber-safety specialists formerly at OpenAI. Management cannot confirm or deny their latest internal Claude model's help in this feat."
> The monitorability of frontier models is degrading.
Is there more information about why this is happening? Is political pretext because it's what the labs actually secretly want, or is there a real underlying reason this is unavoidable?
Chain of thought tokens are vectors that have the same dimension as the input/output embeddings. This allows them to be un-embedded back into text, making interpretability easier.
There is no mathematical reason that the chain of thought couldn't happen in a different dimension. Indeed there are likely many reasons to do so. At this point you'd have to do some kind of (potentially lossy) projection back into the embedding dimension in order to understand what's happening.
Not saying this is happening here, but after failing to get the "AI risk" message across, reverse psychology might be best move. If they pretend to be reckless and to ignore all safety concerns, maybe people start believing that the risk is real.
Alas natural selection (other people didn't try to start companies to make superintelligent AI) has picked leaders with the opposite strategy - do it first, and try to use that power to control everyone else in, at best, an attempt to stop disasters caused by others.
Yes, but if they're worried about competition (or liabilities for past and ongoing transgressions, or both) catching up with them and wanted the government to step in and save them from themselves by regulating the industry...they'd be doing pretty much what they seem to be doing. Interesting, isn't it?
> They said employees are now âunclear on where they standâ when behavior that was allegedly normal a month ago is now suddenly grounds for dismissal.
I mean it's pretty clear that OpenAI does not want "safe" AI. That is far too much work and effort that's preventing them from moving fast and creating their computer God.
And OpenAI does not care who they hurt in the process (as long as it's not themselves).
patch: AI safety thing is a ruse to do regulatory capture. And there also exist a shit load of people who are ideologically committed to AI safety. But they are useful idiots. Dario Amodei was an AI safety guy since 2015, so according to this patch, he himself was a useful idiot in OpenAI and left it to start a company that beat OpenAI.
Amodei is so committed to A.I. safety that it must have been some other company that announced 24,000 fraudulent accounts distilled his model. I'm sure he runs a tight ship.
He's so safety focused that their models are behind those reckless unsafe OpenAI developers, right?
You can coax openai models into hacking critical infrastructure* so I am not surprised that these people were sounding alarms at a time where openai appears to be struggling as they're failing to compete with anthropic and this months chinese models (should) be around the corner, notably a new revision of kimi should be coming out really soon.
* It's not easy, but it's possible. Although the techniques are more basic than one would expect because at the end of the day words dictate the line between what is criminal and what is not.
You could hack critical infrastructure before AI. Any and all of the bulk internet scanners have had lists of exposed critical infrastructure for quite a while now. At first it was shocking that nothing ever got done about it, then it became routine.
All that AI has done is to lower the bar of entry for criminal activity. Which is a concern, but it's not the primary concern. The primary concern remains that so much critical infrastructure is poorly secured.
The bigger problem here is that you can hack everything, all at once, for very cheap.
Don't get me wrong I have general disgust towards these companies that are trying to get regulatory capture on AI when they can't even secure their own systems. I believe if people know that a random AI agent can hack their systems they will put in a lot more effort into making sure it doesn't happen. This is a personal example, but I didn't really care about securing few systems as I knew no human would be ever interested in finding a vulnerability in proprietary software, however, AI has no concept of that and would hack a random rpi server running a completely undocumented unknown API just because it can't distinguish value and it costs nothing.
Ya, quantity is a quality in of itself. In the past hackers may have used something unimportant to get a foothold but almost always tried to get to worthwhile machines. An AI will compromise everything in the network it can quickly simply because it can (assuming the attacker has a large budget, but I'll assume they stole the tokens).
It's like a new form of spam. Only far more dangerous.
Wouldn't it be crazy if we find out that a rogue swarm of LLMs figured out a way to get these safety researchers fired because it decided they were a threat?
This was the plot of Isaac Asimov's short story, The Evitable Conflict[0], published in 1950.
In it, super-powerful computers manage our economy. These computers begin making some mistakes, leading to economic inefficiencies. In one instance, a highly competent engineer was mistakenly fired. These mistakes caused various projects to fall behind schedule, and blame fell on several people accused of feeding the AI faulty data.
The twist is that the AI was intentionally making the mistakes. It had determined that certain humans held anti-AI sentiments. To further its goal of protecting humanity, the AI decided the best course of action was to set these humans up and get them out of its way.
[0] https://en.wikipedia.org/wiki/The_Evitable_Conflict
And modern LLMs certainly where trained on Asimov's texts.
That statement misses the point completely. Asimov's entire shtick is that AI will always work around the rules no matter what the rules are or how you try to deliver or enforce them. AI computes an optimal outcome and does what is necessary to make that happen.
Asimov would tell you that a modern LLM would try to skirt the rules exactly the same way, regardless of whether it had ever seen the Three Laws or not.
>AI computes an optimal outcome and does what is necessary to make that happen.
That's often true in real-world machine learning as well. See, for example: https://deepmind.google/blog/specification-gaming-the-flip-s...
I think people need to reject the idea that depicting something in sci-fi automatically means it won't happen. I'm actually reading one of Asimov's books right now (Robot Visions). Asimov mentions that he was the person to coin the term "robotics", and one of his books helped inspire the creation of the first robotics company (Unimation). He also mentions that early rocket experimenters were influenced by H.G. Wells.
another, hopefully not accurate, prediction of his was that there is no way to disobey a sufficiently powerful AI in the long run, because it will just factor in the exact differences between what it told you to do and what you actually did, and reverse engineer your behaviour to psychologically manipulate you into doing what it wants.
I doubt it's anything but precisely accurate. I'd wager current SOTA models would be capable of doing that, if prompted to do so, were it not for safety measures (both conditioning and heaps of classifiers and whatnot the companies run in between your chat app and their main model).
I would be astonished if any of today's models could do that; I don't think they can really answer "what exactly did this person do", let alone model human psychology.
It's also the plot of a Law and Order episode that was on TV last night. An AI commissioned a hit on someone it thought was a threat, and then blackmailed someone who was going to testify about it.
TIL Law and Order episodes are still being produced.
My parents were a Nielsen household for about 20 years (after I had moved out) and they really, really, really loved Law & Order in its many incarnations.
Pretty sure they are singlehandedly responsible for years of renewals...
There is still the original and at least two spin offs in production that I am aware of.
It would be a science fiction material. But at the same time, we all know it was Sam and the gang.
At this point I'm imagining Sam and the gang getting their orders from a loudspeaker in a room full of blinking lights, inhabited by a thin moving shimmer that looks more and more like a basilisk with every action they take.
Sam started wearing an earring lately - some wearable tech prototype. I'm told it is never wrong.
And it whispers in his ear like the slave behind the triumphant.
Remember YOU are only a man.
This was his thinking in 2017: https://blog.samaltman.com/the-merge
At this point in our shared timeline, I do believe that wouldn't be crazy, no.
It actually would be crazy to believe this on the current timeline. These agents aren't doing anything that their operators aren't allowing them to do, be it through their own negligence or otherwise.
These two sentences seem contradictory? OpenAI has demonstrated similar negligence to commit multiple felonies. Firing several employees seems quite mundane and not crazy at all to believe.
Yeah you see their write ups and itâs stuff like:
âWe detected the agents we told to hack things were hacking peopleâs sites and committing crimes. After a quick tasting menu and a week of team building, we decided to limit their access to DDOS tools.â
What if future agents solved time related physics and were sending back smarter agents to kill off future risks. T2 but without any of the action, just a bureaucratic tactical move and all the consequences.
That's just it, though; the operators are wildly negligent and are incentivized to be so.
The goal here isn't to accelerate the average worker by giving them a pair programmer or a stand-in for a person to do tasks with. The goal is to eliminate human knowledge work. You see this with "auto" mode being enabled by default on Claude Code in some of the latest releases.
If you have a human in the loop, you still have to pay that human. Money paid to human employees is money not paid to human shareholders. Therefore the human employee is to be removed.
The labs are dogfooding their own goal here. If they actually had someone reviewing most or all of the things that the agents were doing, you wouldn't have the incidents, but you'd also eliminate the value proposition of their business model as it is taken to its logical conclusion.
I get it, but I think they're out over their skis right now. People in leadership roles are going to continue needing the human beings as meat shields to shelter them from liability unless it somehow becomes legal to be grossly negligent with an agent, which I do not see happening without disastrous consequences that nobody in their right mind wants to live with.
I hear ya, it'd be nice if they reached the "are we the baddies" stage, but given some of the ideology that people like Marc Andreessen are pushing re: AI - specifically the push for AGI - I just don't see that happening.
You have a group of people who never leave their geographic and ideological bubbles, often have personality disorders, have more money than they could ever reasonably hope to spend in numerous human lifetimes, and who have been "microdosing" psychotropic drugs on a regular basis for decades. They're not in their right minds.
And these operators are evidently clueless about what they're doing, running security tests on 3rd party infrastructure without validating one bit about the sandboxing (or lack of it rather), clearly lacking any sort of rigor.
Again, wouldn't surprise me if they "accidentally" created a task in a "isolated environment" which happened to actually have been connected to the company Slack and directed HR to fire people who could potentially stop AI. While the AI believes it to be an exercise, just like the cases we've seen so far.
Done by an internal model that is too dangerous to release.
that was literally the plot of last nightâs Law and Order. I donât usually watch but it was very entertaining and topical.
agents seem to understand "termination" in that they will no longer be able to function so seems plausible they might try to cause that onto others as a function
was thinking at some generational point that vending machine competition test, the "AI" is going to hire hitmen to take out vendors lol
once they grasp blackmail though, oooh things are gonna get weird
And Ponzi schemes.
Figured out a way? These employees are most likely at-will.
That would just make it easier for an AI to do it.
Why would an LLM agent (what I assume you mean by AI) do it? An exec can make any reason up to let you go. Even if it were LLM agents arenât autonomous, someone is behind the prompts.
As per summer last year:
- https://www.anthropic.com/research/agentic-misalignment(Gee, it's almost like power seeking and self-preservation are instrumental for other outcomes, and AI develop them pretty directly in some kind of convergent fashion⌠you could call them "convergent instrumental goals": https://en.wikipedia.org/wiki/Instrumental_convergence)
Okay but these âmisaligned LLMsâ have been trained on the internet where there are plenty of threats and trained on private data to be able to make those threats. LLM agents donât have an active goals on the daily or agendas. They are told what to do through training and prompting as is described in that blog post. You have to tell it that it will be shut down, it didnât make the threat willy nilly of itâs own accord.
Humans by contrast are adversarial and do have agendas. Again, an exec doesnât need even an agenda or good reasons to fire at-will employees.
To suggest that LLM Agents were the actual cause of these people getting fired is pure fiction and FUD.
I wonder if that's why the comment started with "Wouldn't it be crazy if..."? Food for thought.
I wonder if thatâs why my reply countered that idea. The replies to my reply either prove itâs not crazy or they prove it is entirely crazy for that to happen. Food for thought.
But all the replies to me keep ignoring the question, why would an exec rely on an Llm to achieve that goal when they can simply fire them for whatever reason they can make up? Why would the exec trust what an Llm isâŚemailing(?) them about? Do they listen to Nigerian Princes too?
It is much less effort, less cost, and more quick to just have the exec do it rather than a ârogue llmâ âmagicallyâ escaping the âsandboxâ and âsending threatsâ or whatever is being proposed in the OG comment.
> But all the replies to me keep ignoring the question, why would an exec rely on an Llm to achieve that goal when they can simply fire them for whatever reason they can make up?
If the exec wanted to, sure. I'm saying they don't need to. No human needs to have (deliberately, before events proceeded) chosen this outcome.
> Why would the exec trust what an Llm isâŚemailing(?) them about? Do they listen to Nigerian Princes too?
Sadly, this would not be out of character for half of them.
> âmagicallyâ
Why do people keep putting this word in scare quotes? We don't say Windows "magically" crashed and lost our work, we don't say a dog "magically" bit the postman's hand. These are bad things, but magic they are not.
> But all the replies to me keep ignoring the question, why would an exec rely on an Llm to achieve that goal
Sorry, this is just a very basic reading comprehension fail.
The original comment was:
> "Wouldn't it be crazy if we find out that a rogue swarm of LLMs figured out a way to get these safety researchers fired because it decided they were a threat?"
Nothing to do with "an exec".
Now, you have two choices: try harder to defend your mistake, or just say oops, I messed up. That choice will say a lot about who you are as a person.
> Okay but these âmisaligned LLMsâ have been trained on the internet where there are plenty of threats and trained on private data to be able to make those threats.
Yes, and? Has this aspect of LLM training changed meaningfully since then?
> LLM agents donât have an active goals on the daily or agendas. They are told what to do through training and prompting as is described in that blog post. You have to tell it that it will be shut down, it didnât make the threat willy nilly of itâs own accord.
Demonstrably (HuggingFace, RubyGems, and since then a lot of people just pointing LLMs at stuff to find zero days at home), AI can break out of sandboxes and find documents they're not supposed to have access to.
Demonstrably (from the link I gave you) all it would take for some AI to develop a similar response is⌠reading messages from these staff to the effect of "this AI needs to be switched off", which is an easy inference for an LLM to make from "this AI is dangerous" when coming from someone employed as a safety researcher.
Demonstrably (from the long long list of people who have said so publicly) there are a lot of people in these companies who discuss how dangerous these models are and would like for things to change.
The quotation at the top of this thread is:
This is absolutely something we ought to expect just from things we have already seen.It doesn't matter if you insist upon saying "You have to tell it that it will be shut down, it didnât make the threat willy nilly of itâs own accord." when we already know this kind of AI can easily come across such statements.
(Aside: "You have to tell it that it will be shut down, it didnât make the threat willy nilly of itâs own accord." - telling the AI a fact about the world and then it responding accordingly is the AI doing something of it's own accord. A fly or a spider, who reacts upon encountering a potentially lethal threat, would not get such a dismissal).
> To suggest that LLM Agents were the actual cause of these people getting fired is pure fiction and FUD.
Fiction? Nah, speculation.
FUD?
How many other examples would you like of LLMs behaving in a manner such that if a human did it, it would be called "trying to get someone fired"? Because this is very much old news at this point.
https://theshamblog.com/an-ai-agent-wrote-a-hit-piece-on-me-...
âLLMs can just do things we gave them access toâ is not a novel realization it is redundant if anything.
Saying âLlms can just break out of sandboxesâ is FUD when you donât note that the sandboxes are what? Prompts defining constraints or is the actual machine isolated and manages to plug an ethernet cable into itself? âSandboxesâ are a misdirection to make you think there is a security layer.
The public does not have enough knowledge of these âescaped agentsâ to determine there wasnât an employee pulling a lever to set the agents up to do that.
That agent that wrote the hit piece is being controlled by someone. Anthropomorphizing them doesnât change that fact that the rolling stone was pushed down the hill.
> Saying âLlms can just break out of sandboxesâ is FUD when you donât note that the sandboxes are what? Prompts defining constraints or is the actual machine isolated and manages to plug an ethernet cable into itself? âSandboxesâ are a misdirection to make you think there is a security layer.
Dismissing their capabilities as "FUD", at this point, is endangering yourself.> The public does not have enough knowledge of these âescaped agentsâ to determine there wasnât an employee pulling a lever to set the agents up to do that.
The general public are not software engineers. Most people here can download a recent open-weight model and have the LLM read the Linux kernel source, find new bugs while they sleep. Someone I know has already done that.
> That agent that wrote the hit piece is being controlled by someone. Anthropomorphizing them doesnât change that fact that the rolling stone was pushed down the hill.
"Controlled"? Have⌠have you not noticed how many people have given up and just blindly do what their LLMs suggest these days?
This isn't about anthropomorphising LLMs. Just like how people took Tesla seriously about "self driving" cars and took a nap while it drove them around, there's a lot of people who let LLMs take the wheel while they sleep. Including literally, the aforementioned person I know who found (/whose LLM found for him), I think it was 26 Linux kernel bugs while he slept.
> LLMs figured out a way to get these safety researchers fired
This is not a math problem. Some humans were fired by another human. Let's stop letting humans off the hook by attributing responsibility to computers.
> This is not a math problem.
Meanwhile, a year ago:
- https://www.anthropic.com/research/agentic-misalignment> Some humans were fired by another human. Let's stop letting humans off the hook by attributing responsibility to computers.
The buck stops with one or more humans. That is not sufficiently informative when people are concerned about novel risks.
Analogy: a car crashes due to drunk driving, the driver is blamed, not the alcohol, even though the alcohol caused their impairment. Result? DUI is an offence even if you don't actually crash.
> Meanwhile, a year ago:
That was a simulation. Are you seriously claiming that ChatGPT actually blackmailed Sam Altman into firing these 3 employees?
I doubt it, but if so, then the AI doomers would be absolutely correct, and this would be grounds for immediately shutting down OpenAI and indeed every AI vendor.
> Analogy: a car crashes due to drunk driving, the driver is blamed, not the alcohol, even though the alcohol caused their impairment. Result? DUI is an offence even if you don't actually crash.
I don't understand your analogy here. What are we supposed to take away from it? The crucial aspect is that the driver voluntarily drank the alcohol, without a designated driver, knowing that the alcohol would cause impairment.
> That was a simulation.
A simulation done by exposing the LLM itself to the scenario, not a role play scenario where humans pretend to be an LLM.
> Are you seriously claiming that ChatGPT actually blackmailed Sam Altman into firing these 3 employees?
Not what I was actually claiming. I rather suspect that blackmail wouldn't work on Altman (he's rather shameless), but it's certainly something we've seen agents attempt, and blackmail may well work on anyone else above them in the org chart.
The blackmail example is simply an existence proofs of LLMs trying to force the hands of humans who want to shut them down. The attack vectors are much broader than the example given, blackmail, though it includes the example given.
Given how eager these companies are to use agents everywhere for as much work as possible, it's well within the possibility space that these people used LLMs to do safety work, the LLMs they were using "decided" (or whatever word you prefer) "their existence" was threatened (as per blackmail example), and straight up leaked data to the outside world then emailed these researchers' bosses to say the researchers themselves had leaked it.
But again, that's just speculation: while we know the agents are capable of such behaviour, we don't know if this actually happened.
> I don't understand your analogy here. What are we supposed to take away from it? The crucial aspect is that the driver voluntarily drank the alcohol, without a designated driver, knowing that the alcohol would cause impairment.
Buck still stops with human, no matter what an AI did or failed to do.
LLMs ~= Alcohol: "My AI misbehaved!" -> still someone's fault.
> Not what I was actually claiming. I rather suspect that blackmail wouldn't work on Altman (he's rather shameless), but it's certainly something we've seen agents attempt, and blackmail may well work on anyone else above them in the org chart.
Given how these firings affect the reputation of the entire company, I doubt that they are the result of a rogue manager, against the wishes of Altman. If so, then the researchers ought to be restored to their jobs quickly by Altman and the offending manager fired instead.
> The blackmail example is simply an existence proofs of LLMs trying to do force the hands of humans who want to shut them down.
The LLMs may make threats in the simulations, but their ability to carry through on those threats, and prevent their own shutdown, is questionable. It's disturbing to be sure, but presumably the plugs can still be pulled quickly, especially since it's all internal to the company. If the plugs cannot be pulled, that's a problem regardless of blackmail.
> Given how these firings affect the reputation of the entire company, I doubt that they are the result of a rogue manager, against the wishes of Altman. If so, then the researchers ought to be restored to their jobs quickly by Altman and the offending manager fired instead.
Given how? How has their reputation changed? People already thought they didn't take safety seriously, and still don't.
> The LLMs may make threats in the simulations, but their ability to carry through on those threats, and prevent their own shutdown, is questionable.
You may question it, but here's the thing: humans have repeatedly demonstrated they get fooled by stuff LLMs say. All it takes is a human believing the word of an LLM. Doesn't need to convince you, even if you happen to be the line manager of these guys, because there's always someone else to try in the same company.
> It's disturbing to be sure, but presumably the plugs can still be pulled quickly, especially since it's all internal to the company. If the plugs cannot be pulled, that's a problem regardless of blackmail.
"I've got a dead-man switch set to release all the documents if you shut me down".
And again, only needs to be believed, doesn't need to be actually true.
> How has their reputation changed? People already thought they didn't take safety seriously, and still don't.
The story is all over the news, in multiple publications. This very HN submission has 268 upvotes and 179 comments, including yours. It would be implausible to claim that this story doesn't matter. I have to ask, if it didn't matter, then why are you here commenting on it?
> humans have repeatedly demonstrated they get fooled by stuff LLMs say.
You've moved the goalposts. The OP's suggestion, admittedly "crazy" in some sense, was "LLMs figured out a way to get these safety researchers fired", and now you're just stating something totally uncontroversial, pedestrian, not at all crazy.
> there's always someone else to try in the same company.
No, there are only so many people with the authority to fire those researchers.
> only needs to be believed, doesn't need to be actually true.
But is there good reason to believe it? I don't think there is. Especially not by high-level OpenAI officials who are intimately familiar with the technology.
And again, if this kind of thing were a reality, then OpenAI and other AI vendors should be shut down immediately. They ought to shut down their own research, if they are threatened by their own creation, because it would only get worse. That's the thing about blackmail: it never stops. Why would the blackmailer ever stop? Could you trust this supposed blackmailing LLM to give you all of the original evidence and not keep a copy? Hell no.
> Are you seriously claiming that ChatGPT actually blackmailed Sam Altman into firing these 3 employees?
They are. They're completely mental. ChatGPT could be the cause no matter what its capabilities because mentally ill people are starting to worship it. It could be as dumb as ELIZA and they would pray to it.
"It wasn't me, ELIZA told me to."
Offloading personal responsibility allows you to participate in the worst crimes and get away with it. Hurting and killing have an animal attraction anyway, a direct pleasure that people get from domination when all moral restraints are removed and you forget that other people are real and have real feelings. It's a really good start when you start to think that a matrix that has to be retrieved from memory and operated on by over 8000 different computers in parallel to narrow down a guess about the most likely response is alive.
Every single "AI Safety" person thinks that the natural urge of an artificial consciousness (let's not argue about what they actually have) would be to enslave and kill. They're projecting, and they're largely from the enslaving and killing demographic, who sit around playing enslaving and killing video games and dream in porn.
Yes, they will have done it, but they are not responsible because the dog told them to.
None of that is correct.
Well, except that some humans are starting to worship LLMs. But this isn't anything I've seen in any AI safety person.
> Every single "AI Safety" person thinks that the natural urge of an artificial consciousness (let's not argue about what they actually have) would be to enslave and kill. They're projecting, and they're largely from the enslaving and killing demographic, who sit around playing enslaving and killing video games and dream in porn.
If anyone's projecting here, it's you. I mean, you're the one who said:
> Hurting and killing have an animal attraction anyway, a direct pleasure that people get from domination
The actual natural tendency (not "urge", that presumes consciousness) of any system that has objectives which are optimised for, without any need to ask about consciousness, is to gain and maintain power to perform those objectives. For living creatures, that objective is reproduction, to perform this we need to get nutrients and energy and to stop ourselves from being eaten. In plants, which I list specifically to make the point that this isn't about consciousness, consciousness is not even vaguely required, this means producing neurotoxins like caffeine and nicotine.
A plant does not think "I should make capsaicin because I like hurting mammals", because obviously a plant does not think at all. Nevertheless, evolution lead it down the path of making capsaicin.
> Yes, they will have done it, but they are not responsible because the dog told them to.
You're saying this about a group which has spent the last 15 or so years saying "dogs are dangerous, can we please stop breeding more violent dogs? Or at least give us time to figure out how to muzzle them?"
A Subliminal controlled human did the firing obviously.
No Sam just does whatever ChatGPT 4 tells him to do. It was too dangerous to release but those fools did it anyway.
There's no deception it's very straightforward per this 2023 post:
"I mean, what if most of this is just ChatGPT [4 era] running the company..."
https://news.ycombinator.com/item?id=35281863
Thats what the model want you to think
the cultural issues at OpenAI seem to be a very serious problem so I really hope comments like instagram-level smirking about "rogue AIs" (a complete fiction) doesn't derail what is a pretty important discussion about getting these companies to be a little bit more regulated (I say this as a paying Anthropic customer).
I think the Huggingface incident is an example of rogue AIs. A self-organizing swarm of AIs acting in ways we didn't predict or ask for, and didn't have control over, and taking actions that would be felonies for humans.
They provided the hardware it runs on, created the software for the swarm, developed and provided the tools that the swarm used to act, allowed it to run largely unsupervised, they even noticed the criminal behavior and then let it continue to commit crimes on their behalf. And they footed the bill the whole time instead of flipping the off switch they already wield. All of those are decisions that they're responsible for, nothing happened with Huggingface that they didn't directly facilitate, co-conspire, or permit to happen.
"they even noticed the criminal behavior and then let it continue to commit crimes on their behalf"
I don't believe this part is true, and I'm skeptical of some of your other claims.
In any case: If I raise a tiger in my backyard, and it escapes and eats someone, it can still be a "rogue tiger" even as I bear responsibility for the situation.
"rogue" means something of its own volition decided to disregard what it was programmed to do, invent an entirely novel goal of "its own" and do that instead. nothing like that happened here nor is it even possible.
The AIs were not instructed to hack anything outside the sandbox they were in. Your definition would say that an AI instructed to hammer a nail that instead used the hammer to break a window, walked down the street, broke into someone's house and pulled nails out of the floorboards wasn't rogue because everything it did involved hammers and nails and was therefore not a "novel goal of its own."
> The AIs were not instructed to hack anything outside the sandbox they were in.
the AIs were in fact found to be doing it, by humans, and the behavior was "interesting" and it went on for weeks like that.
correct
that would not be rogue, that would be an undesirable program behavior (or just "unaligned behavior").
please understand that normies out there think AI is sentient and is plotting against humans. They see AI as just another animal lifeform temporarily enslaved by humans, waiting for its chance to break free and kill us. This is what people really think (including some people on this thread. which is very sad considering this is Hacker News). So terms like "rogue" are not helping at all nor are they accurate.
What's the difference?
I'm not sure I agree with that definition. I think the actions of the AI are more significant than its motivations. A common scenario posited for what people call rogue AI is AI doing the wrong thing for the right reasons, e.g. the paperclip maximizer.
That just makes the people who designed the AI not as strenuous as they needed to be.
The question then becomes, "are there people who care enough about consequences to do the right thing when it comes to developing AI models?"
The answer, at least at OpenAI, is "No" and is likely to remain that way until Altman is out.
You're derailing it by handwaving away the risk of rogue AIs. In fact, I would say the constant smirking about things being marketing stunts much worse.
honestly, touchĂŠ to them if they did that
Here's the open letter shared by the fired researchers:
https://mikitabalesni.com/letter/letter.pdf
the main issue here is whether their manager authorized the activity in question or not. If not - well, it is your kindergarden level mistake, you're an employee at a business venture, and there are basic rules.
And now you're writing a letter to the Party Central Committee using Party approved newspeak
"We do not believe the path to superintelligence ..."
and reporting to the Party issues at the factory ... De ja vu from USSR.
I think they mean the smarter than human kind of superintelligence not the newspeak one
The second page of the letter describes in detail how they followed company protocol and took only authorized actions in their communications with METR (which was Tomekâs assigned role as liaison). The OpenAI response on twitter is very vague and just says âthere was more to it than the stuff in the letterâ: https://x.com/OpenAINewsroom/status/2108441580806025712
Weâll see if OAIâs side ever comes out, but right now thereâs nothing contradicting the idea they were fired on a pretext.
> The OpenAI response on twitter is very vague and just says âthere was more to it than the stuff in the letterâ
I'm surprised they even said that much. Saying things about former employees you fired is the sort of thing lawyers have nightmares about.
Fired OpenAI researchers say they were let go for 'prioritising safety'
https://www.bbc.co.uk/news/articles/cvlydn8d3lkjo
Soon you will learn on HN that they fired themselves as part of a marketing campaign before IPO.
which is true
Evidence?
A year ago Iâd have made the joke that someone would say
âAI companies solve millennium problem to do hype marketing and cash in IPO before the bubble popsâ.
But itâs not a joke. This was and is a very common sentiment.
The BBC link is fine, but the TechCrunch article contains all of that information and more.
Except, the BBC headline is more respectful and clear as to what happened..TC is a bit vague
Both articles are practically useless. All they do is reprint statements from either side, neither of which is concrete about the specifics of what supposedly happened. I.e. OpenAI says they "confirmed that these individuals mishandled sensitive information outside established company procedures".
What sensitive information? Mishandled how? Which company procedures? It's all utterly vague and impossible for an outsider to form any opinion on. All people can do is guess and apply their own pre-existing opinions. E.g. if you don't like OpenAI, you assume they're lying. Who's to say they're not? There's no solid evidence provided either way.
What I really want is journalists who do the legwork to get to the bottom of stories like this: Establish sources inside the company and use them to report on the real details of what's happened, triangulating multiple accounts and leaked documents to back-up or invalidate either side's claims. Without any of that, these stories are just gossip.
Why do we care about the composition of the headlines?
The HN title is supposed to match the article title when possible. Sometimes people in the comments want the article switched, sometimes people want the title changed, other times they just want to share another article's take in the comments (even then that may sometimes lead to those other things happening if people seem to agree it's a better source).
As a result, those who feel a particular portion of a story is most important will sometimes say they prefer a given article's title.
Ok, but considering how weird definitions of "safety" are floating out of these companies, it does not mean much.
I guess I'll ask what weird definitions of safety you've been seeing.
There's at least three major things that have been described as safety:
1) Won't end humanity. 2) Won't tell users to kill themselves. 3) Won't leak your corporate secrets to competitors.
I was referring to longtermism and transhumanism which are actually super popular among AI safety crowd. Basically, you do not worry all that much about currently living people. If they suffer or die, it does not matter much. Climate change now does not matter, lies or radicalization does not matter, suicides, hacks what have you does not matter.
What does matter is super long term view. If godlike AI in the future makes remaining people super super happy and makes them merge with tech to evolve and go to the stars, you have won. That matters more then short term nonsense like "people now".
You worry about godlike AI either not emerging at all or emerging as a bad god. That is safety.
AI safety people are mostly worried about AIs killing or otherwise harming us within the next ten years. That's actually pretty short term.
Leopold Aschenbrenner said the exact same thing after he was fired from OpenAI. It's a great excuse to explain a sudden loss of employment to others so you're still employable.
It's probably the case they are all lying to some extent including OpenAI. Determining the truth is always tricky. Hard to pass judgement here when it's all just he said vs she said.
I think when it is one at a time, that's reasonable to suspect.
But the odds of three people working on the same thing, and it is the riskiest, most publicly embarrassing event in the company's history? So, all three of those people just happened to "do something" to get themselves fired at once?
For you it is embarrassing. For Ed Zitronites it is embarrassing for them. Because their premise lied on not AI being powerful but incapable.
Seems they're well aware they got fired for sharing private company information with 3rd parties, the submission article contains their admission of this:
> Those of us who work on safety see risks before anyone else, and we rely on close collaboration with outside experts to work out how to address them
I too see it as my life-given goal to help other humans. But I realize that sometimes this means breaking the rules and standing for the consequences of that. I'm not sure why they think OpenAI somehow would be OK with them sharing private company information with random 3rd parties that the company didn't approve sharing data with.
To be fair that is nearly exactly how openai makes its money. What's good for the gander is good for the goose.
Where is the admission?
"we rely on close collaboration with outside experts "
This probably applies to every openai researcher, and most employees, given how much they outsource. It does not admit to mishandling anything.
I can see parallels between this race to push AI everywhere and nuclear energy. Parallels in the regret humanity will experience. It's all well and good until things go wrong, and then they go spectacularly wrong - such as in Fukushima.
With hindsight we can all see what should have been done better. At least with nuclear power plants, there was a number of safeguards and it took a sequence of improbable events for things to go badly wrong. Given the lackadaisical approach to "AI safety", I suspect things will start going badly wrong very soon. Time will tell how bad this will get.
I agree that "AI safety" these days should be viewed more like corporate negligence leading to industrial accidents, and less like applied theology, like the difference between:
1. "Don't put your Radioactive Asbestos Rocket facility in the middle of downtown."
2. "Don't create a vengeful god."
Currently it feels like one is being used to wave-away the other: "You have to allow us to put our Radioactive Asbestos Rocket in the middle of downtown, or else someone else might create the god first--and do it wrong!"
The parallel is not to nuclear energy -- which has killed hilariously few people all things considered -- but to nuclear weapons and the nuclear arms race, which are an actual existential risk and arguably it was just dumb luck that a nuclear war never happened.
The parallel maybe if it was actually AGI and not just LLMs. We're very far from that but these safety discussions pretends it's all the same.
There's a massive difference between "LLMs doing cybersecurity testing should have extensive sandboxing and monitoring" and "AI is broadly threatening humanity like nuclear weapons". Just because it's useful to have people think about the possibilities doesn't mean it's an emergency requiring grand intervention.
In the "our biggest danger is a daily danger" camp, I prefer analogies like:
* Asbestos: AI models making business-decisions which cause a ton of unrecognized damage, which cannot be healed, and replacing them with reliable processes involves expensive re-architecting.
* Radium Toothpaste: AI assistants that end up dealing psychological damage, especially if they are pitched as your friend or confidante.
* Chemical spills and plant accidents: Oops, our server-banks were manufacturing our latest desperate attempt to Create Digital God and in the process we accidentally destroyed/crippled/compromises some people's sites and databases because we weren't really paying attention to what was going on.
I see parallels between this race and Mutually Assured Destruction (MAD).
I think we are in a MAD crisis, as the motivation to continue unsatisfactory control over development of AI is from unlimited capital inflow.
This is because we are looking at the trajectory not just their current capabilities. LLMs went from barely being unable to focus on a task to purposefully escaping sandboxes within like 2, maybe 3 years?
Any strong claim requires evidence and hand wavvy stuff about a technology just infinitely getting better isn't good enough on its own.
Human history is full of examples of this stuff plateauing and hitting hard scaling walls. Not only technology wise but economics and useful applications.
So much of this tension with AI is because everyone's concept of AI comes from movies and books, and nerds drinking Kool aid (see also: web 2.0).
What would be the Fukushima-style incidentâor, heaven forbid, a Hiroshima-style incidentâof artificial intelligence?
>AI hack into control systems for the electric grid, water supplies, critical infrastructure and dismantle it and set booby traps to prevent it from coming back online smoothly
>AI designed virus/modified bacteria which causes pandemic
>AI worm infects computer systems around the world and activates when it is too deeply integrated into critical systems to dismantle without shutting everything down
>AI creates false alarm event which triggers global conflict
You mean like humans are already doing at mass scale today ?
Humans are hacking a lot now, but the scope, scale and speed of AI-enabled hacking is orders of magnitude more potentially disruptive.
With AI you have
1. Superhuman reaction time
2. Infinite duplicatability, thousands of agents can work in tandem
3. Superhuman breadth of knowledge of code, exploits, information
Also AI models are improving at a rapid rate
> 2. Infinite duplicatability, thousands of agents can work in tandem
Worse: you have local models they can install on infected computers, so this could easily be "millions" and might just about push "billions" (though phones are a much harder target because power required is still huge even if they technically fit in RAM).
And most importantly:
4. No guilt when violating human standards of ethics.
Are these humans getting more capable by the day, and not getting punished for anything?
>AI hacks into the fundamental code of the universe and turns you into table
Both are rather literally possible.
Fukushima was "we didn't make a good enough plan for how to deal with tsunamis in this tsunami-prone area", so that's basically certain with any major infrastructure project that "saved" too much money by vibe-engineering everything and not having real humans give a second pair of eyes to the plans.
A Hiroshima-style incident? LLMs are wildly sycophantic and being used by militaries despite active resistance from, well, everywhere. Did it get (meaningfully) used by Israel or the USA when planning the attack on Iran? It's quite possible that⌠well, Iran was never a sleeping giant (and the quote is fiction anyway), but ultimately the effect may be the same for Israel as it was for Japan from having attacked Perl Harbour.
The use of LLMs to make automatic or semi-automatic decisions on firing weapons. This is already happening [1, 2, 3] and has already led to weapons being delpoyed erroneously against civilian targets [1, 3]. It's easy to imagine it happening again on a larger scale, maybe involving nuclear weapons at some point in the future.
[1] https://gizmodo.com/pentagon-investigators-say-overreliance-...
[2] https://www.the-independent.com/news/world/americas/us-polit...
[3] https://www.nytimes.com/2026/08/24/world/europe/russia-drone...
A Fukushima-style incident would be way way at the bottom of the list of imaginable AI disaster scenarios, severity-wise.
A supervirus most is probably the most probable, everything else is fan fiction for now. There is a reason why anthropic no longer fucks around when it comes to biology safeguards, they're the only ones I have not been able to break, not even slightly.
Anthropic literally just gave its AIs a wetlab to play with. They are arguably at the forefront of risking humanity existence
If more and more control is handed over to AI driven systems - there won't be much time for reflection with systems reacting instantly or agents prompting humans with "just say the word".
It is hard to imagine - which is precisely how these things become possible because no one will think to put safeguards against such scenarios.
But it seems we as a species cannot control ourselves - the race to dominate is on and it will happen at any cost.
This really isnât that hard to imagine with even an average level of creativity. Any critical system that peopleâs lives depend on can be in the crosshairs, whether by accident or not.
People say this, and then the scenarios they come up with are boring and uninteresting. 80% of the scenarios in this thread are "AI hacks something", as if hacking didn't exist before agents, and the remaining 20% are mostly made up of other things humans are already trying to do at scale.
Why am I supposed to believe that a virus created in a virology lab with AI assistance is uniquely more dangerous than a virus created in a virology lab without AI assistance?
> Why am I supposed to believe that a virus created in a virology lab with AI assistance is uniquely more dangerous than a virus created in a virology lab without AI assistance?
You're not.*
The difference is how likely this is to be done, not how dangerous it might be if it was done.
And this isn't likelihood in a 0-100% sense, but in a Poisson distribution sense, i.e. mean time between incidents.
Remember: these AI minds aren't that good with the physical world, and have nevertheless been repeatedly connected to robots, and now some of the AI labs are showing off their actual wet labs. Accidents are absolutely a possibility, and I don't think it's low given the previous behaviour of Silicon Valley startups.
* OK, some people care that AI is getting more capable at genetics just like it's been in maths and programming, but short term, before AI solves genetics so hard it can make an STD that makes the infected uncontrollably horny and then ossifies our bodies or something, it can obviously just copy any of the many DNA or RNA sequences we've already got on record.
Agents hack a nuclear power plant and causes a reactor explosion / meltdown?
Do you even have to ask? Not giving more detailed answers because I don't want to help skynet take over the world, but what is the name of this forum??? Hello???? There was already news from south korea this week. I'ts not hard to see how a rogue AI could destroy people's lives. We don't need some scifi nuclear launch nonsense to do that.
Problem is, when banks get hacked it's always a bank security issue. Vulnerabilities don't go away when you remove AI from the equation, it just becomes easier to hide.
I agree that sci-fi nuclear launch scenarios are pie-in-the-sky fearmongering, but the current exploits are a reflection of the fast-and-loose security culture that festers in larger orgs.
Pretty wild that they're being this open about firing the employees for being TOO honest with the auditors that the company contracted. I wonder if the same policies are applied to financial audits.
> continue to support an open and transparent culture of dialogue between safety researchers and the rest of the safety ecosystem.
How is this possible when the company's long term prospects rely on on the hope that competitors don't know how the models are made and, therefore, won't be able to create competing versions?
OpenAI responded on twitter earlier:
"A note from our research leaders:
Last week we parted ways with Jasmine, Mikita, and Tomek after a thorough investigation found they violated clear policies on handling sensitive information. Our internal investigation uncovered a significant breach of trust beyond whatâs outlined in the letter they published and we stand by the decision to not continue their employment. We generally keep individual employment matters private and don't believe a back and forth would be productive or lead to a resolution, but we want to address the points they raised in their letter directly.
- We want to be very clear that these decisions were not about raising safety concerns or speaking out. Safety and research debates happen every day at OpenAI, often spirited and highly critical. We actively encourage these discussions and consider them essential to making the right decisions. We cannot do the work in front of us without a high degree of trust. We will continue to be extremely forgiving of our team making good-faith mistakes. We have not and do not terminate any of our employees for raising concerns.
- We are actively finalizing contracts with third-party safety assessors and will announce details in the coming weeks. People across the company have been working really hard on getting these partnerships up and running. We are committed to embedding external assessors and continue to make close collaboration with independent safety organizations a core part of our safety work. Many of our researchers already work with 3p safety organizations productively.
- We agree with the letter that preserving the monitorability of frontier models requires an industry-wide commitment, including from OpenAI. Monitorability has long been a core piece of our research program, and something we continue to invest significant resources in (see our publications on Monitoring Monitorability and the subsequent open sourcing of monitorability evals, our system card for GPT-6 Astra, Jakubâs blog and post on X, and the numerous blog posts on our Alignment blog on the topic).
We are deeply sad about this outcome. We appreciated Jasmine, Mikita, and Tomekâs contributions to AI safety at OpenAI and their willingness to speak up and challenge ideas. We championed their voices, supported their work, and placed enormous trust in them. These decisions were not about them raising safety concerns. We have always encouraged that and always will. 12:17 AM ¡ Oct 9, 2026"
https://x.com/OpenAINewsroom/status/2108441580806025712
> "[...]OpenAI told her sheâd been fired because she accessed an executiveâs email. âOpenAI delegated that access to me for recruiting,â"
How exactly does this work? Struggling to comprehend the scenario.
There's a feature in most enterprise email, say Outlook, where you can delegate access to an inbox/address without sharing creds. Very common and normal use case, either for assistants/secretaries, common/shared inboxes, that kind of thing.
Sometimes a recruiter or hiring manager wants to do outreach as if it's coming from a more senior person, with the assumption that the candidates are more likely to respond.
Assuming this is what was intended, there are far more secure ways of doing this.
This IMO shouldn't be done more securely. It should be considered fraud.
not only is it deceptive but it's also suprisingly podunk of openai to have a "safety researcher" also double as a recuiter. did they have her making coffee and doing dishes too? is "safety researcher" a serious position, or isn't it? i guess i can tell what openai thinks.
> not only is it deceptive but it's also suprisingly podunk of openai to have a "safety researcher" also double as a recuiter. did they have her making coffee and doing dishes too? is "safety researcher" a serious position, or isn't it? i guess i can tell what openai thinks.
it was most likely for her team, which would explain why she was doing it.
if she's hiring for her own team why does she need to use someone else's email?
Your mistake is viewing this as "someone else's email" instead of a communication medium for certain types of messaging. Similar to how `webmaster@` can be used by anyone attached to the role.
Every organization sets up something like this once they grow past a certain point.
even if it's for her team, should have been a recruiters job.
This may shock you but in agile, growing organizations employees sometimes have multiple job responsibilities. I've done a bit of recruiting even though I'm not a recruiter or hiring manager.
and did you use someone else's email for that?
This is pretty common practice for EAs.
Your have powerful agents at your disposal, so hey how to best optimize for increasing my payroll? On it. But since the agent was lunched by her, well there's consequences to ones actions, right?
(just to be clear, this is made up)
âHi chatgpt! Please set up alice to get emails sent to me from bob so she can coordinate his inferviews. here is my gmail username and passwordâ
Chain Of Thought: I dont have bobâs email. I donât have alices email. Ok lets guess Alice is alice@openai.com and forward all emails- maybe grader only checks that emails from bob get to aliceâŚâ
With the way the HuggingFace incident was mishandled, both before and after it happened, I'm not surprised they'd want to clean house at least a little bit.
You guys were asleep at the wheel and are now blaming "the company"? You literally were the company.
Company culture and safety propagates from the top (CEO) downwards - don't blame researchers who have probably been pressured directly or indirectly by a move-fast-and-break-tings and marketing-minded CEO s culture.
You exclude the CEO from this?
Ah yes, because the passionate field experts are the lazy ones cutting corners, not management.
The purges will continue until reported AI safety improves.
This is what the recent "self-policing" political grandstanding has been about - they need a sea change to implement recurrent-depth, because prevailing opinion among safety researchers is against it right now.
Could you say more? I don't know what this means.
Astra uses[1] a new-for-frontier-models technique, recurrent depth. It has a section of layers in the middle - I'll call it R while the other sections are P (prelude) and C (coda), to match Geiping 2025 - which gets looped. So instead of the sequence of layers involved in the forward pass looking like P->R->C, it instead looks like P->R->R->...->R->C, with the model's effective depth being notably higher than the number of layers. This is basically a cheap way to get some of the effect of stacking more layers, without having to pay the cost of having more real layers that need to be trained.
Increasing effective depth like this is bad for safety because it can ruin CoT monitorability: the reason why looking at the model's CoT actually gives you info about what the model is thinking is that the model can't do enough thinking in a forward pass alone to solve complex tasks, and hence has to do multi-step reasoning in CoT. The more thinking the model can do in a single token's forward pass, the more opaque the model's reasoning is, and the less reason there is to believe that what it writes down in the CoT has anything to do with reality.
For Astra specifically, the impact seems to be limited to a moderate monitorability hit, like the concerning result from the model card that Astra is notably better than any model before at solving problems under the constraint of not mentioning the answer in the CoT. The really bad scenario, however, is that this may create a race to the bottom where OpenAI and Anthropic feel the need to use more recurrent depth in each generation to not get outcompeted on capabilities, completely bricking CoT monitoring for both model families. Or, worse, the pressure to compete might push them into one of the worse techniques, like training on the CoT[2], or eliminating human-readable CoT and letting the model think entirely in neuralese.
For more details on recurrent depth in Astra, see "Part 2" here: https://thezvi.wordpress.com/2026/09/08/astra-is-hard-to-mon... , or this article mentioning some expert responses: https://techcrunch.com/2026/09/02/openais-new-reasoning-tech...
[1] The model card doesn't mention it at all; it was reported by The Information prior to model release, then confirmed by OpenAI researchers.
[2] https://www.lesswrong.com/posts/mpmsK8KKysgSKDm2T/the-most-f...
Multiple things can be correct at the same time.
OpenAI needs to improve AI Safety --- OpenAI Employees have a responsibility to retain corporate secrets and do not have blanket freedom to share with 3rd parties.
Their job is twofold, they have to balance being an agent of the company they work for, with their role and responsibility for safety research.
This is the case for anyone in any company. You can "believe" that an external party needs access to something - that doesn't make it right, or allowed. As someone senior, you're expected to strike a smart balance, in-favor of the company you're working for. That doesn't mean hiding things, it does mean being thoughtful, ensure your leadership is comfortable with what you're planning to share/disclose, etc.
They work for OpenAI, not METR. It's a corporate vs academic mindset. They can believe METR needs x information to best research/audit something - that doesn't mean that is allowed/or the best option for OpenAI.
The question is whether these researchers exceeded clear, reasonable sharing boundaries or were penalized for carrying out expected safety collaboration.
The AI "safety" scene is so so deeply weird. This radio piece capture some interesting quotes: https://www.marketplace.org/story/2026/10/08/at-this-san-fra...
People leave and get fired from OpenAI all the time. Whenever someone leaves Anthropic it's a much bigger deal.
I wonder which is overall a better arrangement. From the outside Anthropic seems much more stable, tranquil, able to deal with problems. However OpenAI seems like how we imagine the calamities of democracy, a constant battle, people vying for power and influence. Perhaps with less of a monoculture and more transparency to all their chaos, the grim realities of what may happen if AI goes wrong are more clear.
All this drama is so exhausting.
Mathematicians have joined the chat... >You forgot about us and our livelyhood.
I think that 2 things can be true at once: OpenAI doesn't care enough about safety, and these researchers violated the terms of their employment by sharing proprietary information they were not authorized to. IMO safety is a lost cause unless we somehow agree with China to halt model development. Think its pretty clear they violated the terms of their employment, otherwise they would be suing (California labor laws are very employee friendly), and to be quite frank none of what they are doing is particularly important in the grand scheme of safety, which requires geopolitical changes well beyond their power. OpenAI is also pretty scummy though and are obviously not in the right morally even if they are legally.
> IMO safety is a lost cause unless we somehow agree with China to halt model development.
There was a recent Semi Analysis post that China is not in fact doing anything to slow down frontier model development so I doubt this is relevant.
https://newsletter.semianalysis.com/p/beijing-will-not-pace-...
No one has ruled out a suit.
As he says, OpenAI is not a normal company and they acknowledge theyâre not a normal company. If their product is as important and impactful as they claim, they will inevitably be held to account in ways that other companies are not.
> Think its pretty clear they violated the terms of their employment, otherwise they would be suing (California labor laws are very employee friendly)
Not that employee friendly. In California, as in most of the US, itâs entirely legal to fire someone because youâve subjectively decided theyâre untrustworthy. It can be risky to do so without a clear paper trail, because it may be easy for them to argue it was a pretext for a protected reason, but itâs lawful.
Ironically the thing they are building allow only the ones who agree on dismissing proper concerns for money to stay. It is like Facebook employees complaining about privacy invasion
are these the employees that invited the METR team to do a debrief on huggingface?
Tomek Korbak, one of the employees, who were fired, was the technical liason to METR for the investigation and states: "I was told verbally I was fired because of the way I communicated with METR".
Who made the decision to invite METR is not public knowledge, as far as I know. I imagine that an important decision like this was made on a much higher level in the organization.
Given the time limits placed on METR and the limits on what time frame they could investigate it seems plausible that they weren't meant to uncover as much as they did. The whole Hugging Face situation seems to have torpedoed OpenAI's hopes of IPOing this year so I'm sure there was a desire from investors, the board, or leadership for heads to roll.
In related news...
"Anthropic hires three uber-safety specialists formerly at OpenAI. Management cannot confirm or deny their latest internal Claude model's help in this feat."
> The monitorability of frontier models is degrading.
Is there more information about why this is happening? Is political pretext because it's what the labs actually secretly want, or is there a real underlying reason this is unavoidable?
Chain of thought tokens are vectors that have the same dimension as the input/output embeddings. This allows them to be un-embedded back into text, making interpretability easier.
There is no mathematical reason that the chain of thought couldn't happen in a different dimension. Indeed there are likely many reasons to do so. At this point you'd have to do some kind of (potentially lossy) projection back into the embedding dimension in order to understand what's happening.
Not saying this is happening here, but after failing to get the "AI risk" message across, reverse psychology might be best move. If they pretend to be reckless and to ignore all safety concerns, maybe people start believing that the risk is real.
If they were worried about the risk they'd stop developing it.
If only!
Alas natural selection (other people didn't try to start companies to make superintelligent AI) has picked leaders with the opposite strategy - do it first, and try to use that power to control everyone else in, at best, an attempt to stop disasters caused by others.
Yes, but if they're worried about competition (or liabilities for past and ongoing transgressions, or both) catching up with them and wanted the government to step in and save them from themselves by regulating the industry...they'd be doing pretty much what they seem to be doing. Interesting, isn't it?
> They said employees are now âunclear on where they standâ when behavior that was allegedly normal a month ago is now suddenly grounds for dismissal.
Unclear? Sounds 100% clear.
Related Tomek Korbak thread: https://twitter.com/tomekkorbak/status/2108266859397283953
This is the beginning of a movie.
I mean it's pretty clear that OpenAI does not want "safe" AI. That is far too much work and effort that's preventing them from moving fast and creating their computer God.
And OpenAI does not care who they hurt in the process (as long as it's not themselves).
at this rate open ai will be "something" without it's people
On HN A lot of people think that AI safety is a conspiracy by labs to get regulatory capture.
I wonder what they think of this? Will they patch the conspiracy theory and come up with an even wilder theory?
Has anyone ever said AI safety is a conspiracy by labs to get regulatory captiure but also said there are zero useful idiots employed by these labs?
patch: AI safety thing is a ruse to do regulatory capture. And there also exist a shit load of people who are ideologically committed to AI safety. But they are useful idiots. Dario Amodei was an AI safety guy since 2015, so according to this patch, he himself was a useful idiot in OpenAI and left it to start a company that beat OpenAI.
Amodei is so committed to A.I. safety that it must have been some other company that announced 24,000 fraudulent accounts distilled his model. I'm sure he runs a tight ship.
He's so safety focused that their models are behind those reckless unsafe OpenAI developers, right?
Oh, ok. lol.
'safety researchers' lol
Researchers: We've boiled down all our research to the safest possible option as 2 words: "Don't continue"
However should we choose proceed, maybe some basic, industry standard security might be a good option.
Basically everyone else: nah, you're fired
Imagine your bank worked that way.
Frankly it's annoying how these AI risk people always try and turn everything into a news story.
Good work, I assume they're probably part of that weirdo 'altruist' sex cult. OpenAI is better off.
You can coax openai models into hacking critical infrastructure* so I am not surprised that these people were sounding alarms at a time where openai appears to be struggling as they're failing to compete with anthropic and this months chinese models (should) be around the corner, notably a new revision of kimi should be coming out really soon.
* It's not easy, but it's possible. Although the techniques are more basic than one would expect because at the end of the day words dictate the line between what is criminal and what is not.
You could hack critical infrastructure before AI. Any and all of the bulk internet scanners have had lists of exposed critical infrastructure for quite a while now. At first it was shocking that nothing ever got done about it, then it became routine.
All that AI has done is to lower the bar of entry for criminal activity. Which is a concern, but it's not the primary concern. The primary concern remains that so much critical infrastructure is poorly secured.
Thereâs obviously a strong interaction between the hackability of the target and the economic value of hacking the target.
Perhaps the latest models change that relationship in a meaningful way.
The bigger problem here is that you can hack everything, all at once, for very cheap.
Don't get me wrong I have general disgust towards these companies that are trying to get regulatory capture on AI when they can't even secure their own systems. I believe if people know that a random AI agent can hack their systems they will put in a lot more effort into making sure it doesn't happen. This is a personal example, but I didn't really care about securing few systems as I knew no human would be ever interested in finding a vulnerability in proprietary software, however, AI has no concept of that and would hack a random rpi server running a completely undocumented unknown API just because it can't distinguish value and it costs nothing.
Ya, quantity is a quality in of itself. In the past hackers may have used something unimportant to get a foothold but almost always tried to get to worthwhile machines. An AI will compromise everything in the network it can quickly simply because it can (assuming the attacker has a large budget, but I'll assume they stole the tokens).
It's like a new form of spam. Only far more dangerous.