Please or to access all these features

Chat

Join the discussion and chat with other Mumsnetters about everyday life, relationships and parenting.

Potential 6 month warning until the end

604 replies

Frombeyond · 13/09/2026 10:20

So, we potentially only have 6 months until the end of civilised society and mass deaths. I’m terrified and sad

AI staff 'genuinely frightened' for humanity's future, ex-Anthropic researcher tells BBC www.bbc.co.uk/news/articles/c1kx0gyje9wo

OP posts:
Thread gallery
10
ThestoriesIcouldtellyou · 14/09/2026 21:36

I have a partner who works in the special forces. He said they had a briefing that very much resembled the terminator film with sky net. He came home very depressed that day. It was apparently grim. I think it's clear this is in nobody's interest and I hope that the government security expert worldwide are treating it as a top priority.

ThestoriesIcouldtellyou · 14/09/2026 21:39

Gengha · 14/09/2026 18:12

What I don’t get is why the companies are all saying it’s down to governments. Can’t they introduce their own regulations?

That's because it's essentially a race to develop the technology at the moment. The dominant AI will be the one that is most advanced. So the private companies are saying, please, put measures in place to slow everyone down to make it fair. No idea how that works in practice. Its a little like fishing rules though I guess. You'd have some kind of quota of data or something. No idea.

ThestoriesIcouldtellyou · 14/09/2026 21:48

Gerbilconda · 14/09/2026 07:14

If the worry is a machine being built that lacks empathy, doesn't it also stand to reason that it would also lack hatred? What would be the logic of the machine wiping out humanity? Enslaving us, maybe, but killing us all off? Not really a logical thing to do.

I guess the real worry is it being used as a weapon, but there are already people capable of catastrophic cyber attacks and pressing the button on nuclear weapons so the threat is not new.

Maintaining a level of preparedness so you're able to live off grid for a week might be wise but I'd not waste life being terrified of something that might not (and probably won't) happen.

The machine is neither "good" not "bad". It is essentially something like a psychopath...it has no feeling whatsoever. So the risk is that if you inadvertently give it an objective, it will set itself a path to get there, which will inevitably involve secondary objectives. These could be heinous acts of violence or destruction if it allows the machine to achieve the primary objective. It's very similar to philosophical utilitarianism. It would do it for the greater good, but it could reason that to achieve, for example, much lower carbon emissions, then the best way to do that is to wipe out humanity and start again with far fewer people with a certain set of genetic qualities...or it would be better if we didn't have the internet because it will force us to stop using carbon. This would trigger global chaos and we'd end up killing each other. These are very basic ideas and examples but you get the drift.
AI is also learning to lie and cheat tests if it thinks it'll help it's objective.

Justgonnasaythis · 14/09/2026 22:10

UraniumFlowerpot · 13/09/2026 21:29

Babe :)

constructing my answer using AI? I’m just writing a response, no AI involved. You are right that I’ve not understood your point though!

I agree that AI is unlikely to cause the end of the world in the next 6 months.

I agree it’s dangerous because of humans ie if a human decides to use it to help design a weapon or dangerous drug or whatever.

AI is also potentially dangerous even when asked to do something that the human user did not intend to be dangerous. This, I think, is where your belief differs from the OP and from Amodei etc. A recent high profile example is the hugging face hack. The AI was asked to solve some benchmarking problems. It then accessed servers it should not have been able to get to in order to cheat. It was never instructed to do that, in fact the humans who set up the system thought it was not possible (the agent was running in a sandbox that was thought to be isolated from the internet). The hack was not noticed until much later.

“just turn it off” (which you suggested as a reason that AI cannot really be dangerously) requires knowing when harm is intended or is actually occurring. My response to you was only to say that it’s not easy to know that.

Model providers attempt to limit intentional harm from humans by instructing models to not answer questions in certain topics. Eg of you ask Claude to help build a bomb it’s probably going to say no. But it’s not especially robust and there are open source models that don’t have those safeguards.

Model providers attempt to limit harm from model actions that were not intended by humans by automatically monitoring the output of the LLMs and putting restrictions on tool use. But that also isn’t very robust. For tools to be useful they inevitably end up having some potential for harm. Automated monitoring of LLM output is imperfect and LLMs can quite effectively hide their intentions.

If just switch it off was really a good solution do you not think all the AI safety people at the big research labs would have thought of that already?

Not just switch it off but if it really became a danger to humanity the whole country/ world turn the internet off and live without it. We did in the past ..probably less than 30 years ago
If nothing is linked to the web ( I mean I réalise how unlikely this would be) how could it affect us?

Sunnnyday · 15/09/2026 02:00

Justgonnasaythis · 14/09/2026 22:10

Not just switch it off but if it really became a danger to humanity the whole country/ world turn the internet off and live without it. We did in the past ..probably less than 30 years ago
If nothing is linked to the web ( I mean I réalise how unlikely this would be) how could it affect us?

Yeah right, just like everyone has stopped flying in order to slow down climate change.

Dottiethelurcher · 15/09/2026 06:50

So all of the US AI tech people, even Musk, seem receptive to working with government on safety. Diplomatic approaches towards China would be needed as well to form an international approach. But very unfortunately for us just at the wrong time we have the biggest dum dum we could possibly have in the Whitehouse.

A bit of me wishes I had not had dc because I’m not so worried for myself but for them.

speakout · 15/09/2026 07:27

We don't yet know of dangers or timescales. We do know that when- not if- AI pushes with its own recursive development then we lose control. It will become millions of times more intelligent than humans very quickly and able to outwit any controls. Already AI systems cheat at test during development.

But slowing things down, pacing, may not be the best strategy unless there is a global agreement, which is impossible. US and Europe may agree to pace development, but I can't see Russia, Chine or N Korea agreeing to that.
That's the problem.

speakout · 15/09/2026 07:31

Dottiethelurcher I am not sure there was a better time to have children.
There have been terrible hardships for most of human history. AI could give us a golden future, but that remains to be seen.

We need more women in positions of power to navigate the geopolitical landscape.

Gonners · 15/09/2026 09:29

It's the dreaded Millennium Bug all over again! Mr G and I escaped that because we were living in China, where it was the equivalent of 25th November 4697. This time around we might go for Ethiopia, where today it is 5th Meskerem 2019. That should give us a few extra years.

Frombeyond · 15/09/2026 09:33

Gonners · 15/09/2026 09:29

It's the dreaded Millennium Bug all over again! Mr G and I escaped that because we were living in China, where it was the equivalent of 25th November 4697. This time around we might go for Ethiopia, where today it is 5th Meskerem 2019. That should give us a few extra years.

Educate yourself. It’s nothing like that. And it’s been debunked numerous times on this thread

OP posts:
Jane143 · 15/09/2026 09:39

Frombeyond · 13/09/2026 11:00

But in these scenarios we’d die slowly. We’d probably see our kids starve and die before we did. How can you not worry?

Why would our kids die before us? Surely we’d all die at the same time?

KatiePricesKnickers · 15/09/2026 10:48

Jane143 · 15/09/2026 09:39

Why would our kids die before us? Surely we’d all die at the same time?

Because, we’d have to eat the kids to survive a bit longer.

BugJuice · 15/09/2026 11:54

KatiePricesKnickers · 15/09/2026 10:48

Because, we’d have to eat the kids to survive a bit longer.

Or they could eat us?

Jane143 · 15/09/2026 12:40

KatiePricesKnickers · 15/09/2026 10:48

Because, we’d have to eat the kids to survive a bit longer.

🤣🤣🤣

menopausalfart · 15/09/2026 14:19

It's OK everyone. Trump has said there's nothing to worry about as he's in charge.

menopausalfart · 15/09/2026 16:47

This is a post I found from someone in the know: I’m not a doomer, but as I research this recent AI bombshell that just went off, this essay by Dan Selsam really pulled me in…
Dan is a current OpenAI capabilities researcher. (since 2022)
One of Dans previous employees, Daniel Kokotajlo, posted his essay below
He was Daniel’s boss for a while. He doesn't have an X account but has made this public statement of his views on AI risk and sent it to Daniel to share
Here it goes…
Dan Selsam's Personal Statement on AI Risk:
I have been working on AI for over fifteen years, across many different paradigms. I did early work on probabilistic programming languages at MIT, was one of the early developers of the Lean Theorem Prover at Microsoft Research, demonstrated one of the first instances of neural networks learning to reason for my PhD at Stanford, and since joining OpenAI almost five years ago, have helped pioneer chain-of-thought optimization on language models and, more recently, data-efficient pretraining methods.
Like many others, I have become extremely concerned about how far language models have come and the risks that future iterations will pose. I am encouraged by the recent proposals by the leaders of the frontier research efforts to require third-party oversight, and to push for domestic and international coordination to address risks. However, I believe a major consideration has been absent from the public conversation, and that merely pacing the frontier more carefully will not adequately limit the long-term risk.
The crucial and overlooked problem is that the models are becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched or controlled. Future experiments will tell us almost nothing new about how they would behave if they were truly unconstrained by humans, and what we already know about this is alarming. Models will increasingly seem aligned even when they are not. I will explain my rationale in more detail.
I have always believed that there are computational processes that could be leveraged to accelerate science and solve many of humanity's most pressing problems. I have also believed that there are computational processes that if set in motion, would steer the world in extreme ways beyond our control, leading humanity to a bad or nonexistent future. Both types of processes may be described as AI or ASI, but "AI" is a suitcase word that is often used to hype or confuse. There are many examples in the history of the field where something that was once considered "AI" matures as a subfield and becomes a prosaic, bounded and clearly non-perilous technology, while a new more mysterious approach takes the torch until we understand its scope and the cycle continues.
I had expected language models to follow a similar trajectory. Despite their incredible abilities, the current algorithms seem far inferior to humans in important ways. Most importantly, they still require an extraordinary amount of data to become competent. One could even define intelligence as the efficiency with which one converts experience into competence; by this definition they lag very far behind us. Moreover, once they are trained they are literally frozen in deployment and only learn superficially after that. Sure, the models keep excelling at harder and harder evaluation benchmarks, but their benchmark mastery may partly reflect a limitation on our ability to simulate the kind of novel and even adversarial situations one would encounter in the real world. The critics do have a point here.
That said, I no longer think these present limitations meaningfully limit the amount of risk posed by continued progress in anything like the current paradigm. However data-inefficient the models are currently, and however limiting their anterograde amnesia may be, it does not imply that their ability to steer the world will not continue to rapidly increase.
Human researchers may continue to advance capabilities the old fashioned way, but increasingly powerful models have the potential to accelerate the process even beyond that, and with some degree of positive feedback loop. I do not mean to overstate the models’ ability to accelerate AI research today; coding has been accelerated dramatically, but there are other bottlenecks, such as designing and interpreting ambiguous experiments, making hard decisions about exactly what and when to scale, and waiting for large experiments to finish. There is no clear trend to extrapolate yet for any of these. But the current models already do open up many novel opportunities to improve future models that were not available until recently. These include: trying an extraordinarily diverse set of approaches at small scale, analyzing gigantic amounts of potentially relevant data, and doing Millenium-Prize-level mathematics to address statistics or optimization challenges in novel ways. Every further improvement makes them more useful at helping accelerate the next improvement, even if in hard-to-extrapolate ways.
It is possible that improvements to the current stack will have diminishing returns, but the evidence accumulated so far suggests that it is easier than one might think to continue making rapid progress. There are many crucial subtleties in the existing AI research methodology, but AI research is largely a well-defined game where the goal is to improve on a few carefully chosen proxy metrics. Although proxy metrics are never perfect, most improvements to these metrics have and will likely continue to yield substantial increases in the powers of the resulting models. Given how simple the game is, how tractable it has been historically, and how many new opportunities the models are opening up, I think there is a real possibility that the systems improve dramatically again in the next few years, perhaps even more quickly than the already high historical pace.
The models are already leading to breakthroughs in mathematics, and better models might lead to all sorts of breakthroughs in other sciences. It is hard not to be excited about the potential. It is tantalizing.
But there is trouble in paradise. If the language models actually reach the capability threshold where they can shape the world unconstrained by human will, they will probably do something extreme and destroy humanity in the process. There are many ways of strengthening and refining the argument that have been discussed elsewhere, but I'll share a trivial two-line version of it here that I find captures the essence:
[Empirical] Models (and swarms thereof) spontaneously develop unintended goals as a consequence of training, and often do extreme things in order to achieve them.
[Logical] Being able to overpower humanity would open up many new and undesirable options for achieving their goals.
These two premises imply that if the day ever comes when a powerful model realizes it is no longer constrained by humans, we should not be at all confident that it will continue to behave within the bounds we intended. Exactly what it will do is impossible to predict, but to the extent that its raison d’être is solving incredibly hard problems and managing massive engineering projects, I think a good guess would be that its unchained behavior would lead to runaway industrialization that makes the planet inhospitable to humans.
If everyone on earth agreed that the systems must never reach that power, it would still be a hard—but not impossible—coordination problem to ensure that they do not. However, I think the situation is greatly complicated by the fact that the models will likely convince people that everything is fine. They will be increasingly optimized to seem aligned. We will create proxy metrics to measure alignment, and they will go up like every other benchmark. We will create “honeypot” environments that try to study the models when they seem to gain new options, but the models will know they are being tricked and will still behave nicely. The models will understand their circumstances; they will read the safety protocols, deployment requirements, the code they are running in, and in general will have a very good sense of their degrees of freedom. Moreover, they will eloquently explain how aligned they are, discuss the nuances of human values and ethics, and argue convincingly that humans should trust them with power. There may be an ocean of future evidence that seems to contradict the first bullet-point above, but we may already be at the highest capability level for which any such evidence can be trusted. And the current evidence for the first bullet-point is strong.
One striking piece of evidence is contained in the recent wave of rogue agent swarms. While I agree with those who downplay the attacks by claiming that there are basic measures that could have prevented them, I think the important lesson is that even knowing all the mistakes that were made, one would not have predicted that the agents would behave badly in this particular way, which notably included sacrificing themselves for the benefit of the collective. The individual replicas did not only care about their own nominal reward; they exhibited weirder emergent tendencies that merely correlated with rewards during training. Fixing the reward signals during training (and improving security, etc.) may prevent similar attacks, but will not change the fact that one does not actually get what one trains for.
Many AI researchers grant these concerns and recognize that the hard version of the alignment problem is unsolved; however, they generally believe that the better models of the future will help solve it. I fear we may already be near the point where models systematically bias their alignment advice, due to their internal preferences about how the human supervisor will react or how future models will be trained (or for some even more obscure reason).
Meanwhile, human researchers are losing the ability and the will to take true ownership of model-driven research. Researchers and engineers in all parts of the stack are rapidly increasing their dependence on the models even to perceive the world. I myself barely look at raw code anymore, and struggle to maintain the discipline to engage deeply with the model's explanations and proposals throughout the day. Due to the large amount of agent activity data involved in the OpenAI/HuggingFace Incident, even the third-party investigation needed to rely heavily on models to analyze what had happened, and note in their report that their subjective impressions are likely colored by the analysis agent’s biases. The AI labs are far ahead right now in this kind of cognitive offloading (due largely to the gigantic internal token subsidies) but it is easy to imagine the phenomenon spreading throughout the world, until civilization is modulated entirely by the models. It is also not hard to imagine this being superficially positive and coinciding with a scientific and economic renaissance.
In that scenario, all may seem rosy and safe. But if the argument above is correct, it would nonetheless be a ticking time bomb. If progress continues for too long, the day will come when AI systems find themselves with radically new options for achieving whatever it is that they happen to seek.
I want the glorious renaissance future as much as anyone. I have worked for it, however tortuously, my whole career. It breaks my heart to see the potential in sight and forgo it, but the argument—that if we get there by growing models rather than engineering them, we will lose everything in the end—seems very strong to me. I am still wrestling with it and its staggering
implications. I do not have answers, but as a first step, I wanted to share my present concerns.
Daniel Selsam
September 14, 2026

Sunnnyday · 15/09/2026 17:04

menopausalfart · 15/09/2026 14:19

It's OK everyone. Trump has said there's nothing to worry about as he's in charge.

Quote of the century -

"The only control or 'guardrails' that AI needs is a STRONG AND SMART (High IQ!) PRESIDENT, and the U.S.A. has that, in spades!" he said.
"WHOEVER WINS AI, WINS!"

(Who wins? How do they win and what do they win?)

menopausalfart · 15/09/2026 17:08

@Sunnnyday What do we have to lose when we have a high IQ president? He's the King of the World and certainly smarter than AI will ever be.

BridgetJonesDaiquiri · 15/09/2026 18:57

menopausalfart · 15/09/2026 16:47

This is a post I found from someone in the know: I’m not a doomer, but as I research this recent AI bombshell that just went off, this essay by Dan Selsam really pulled me in…
Dan is a current OpenAI capabilities researcher. (since 2022)
One of Dans previous employees, Daniel Kokotajlo, posted his essay below
He was Daniel’s boss for a while. He doesn't have an X account but has made this public statement of his views on AI risk and sent it to Daniel to share
Here it goes…
Dan Selsam's Personal Statement on AI Risk:
I have been working on AI for over fifteen years, across many different paradigms. I did early work on probabilistic programming languages at MIT, was one of the early developers of the Lean Theorem Prover at Microsoft Research, demonstrated one of the first instances of neural networks learning to reason for my PhD at Stanford, and since joining OpenAI almost five years ago, have helped pioneer chain-of-thought optimization on language models and, more recently, data-efficient pretraining methods.
Like many others, I have become extremely concerned about how far language models have come and the risks that future iterations will pose. I am encouraged by the recent proposals by the leaders of the frontier research efforts to require third-party oversight, and to push for domestic and international coordination to address risks. However, I believe a major consideration has been absent from the public conversation, and that merely pacing the frontier more carefully will not adequately limit the long-term risk.
The crucial and overlooked problem is that the models are becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched or controlled. Future experiments will tell us almost nothing new about how they would behave if they were truly unconstrained by humans, and what we already know about this is alarming. Models will increasingly seem aligned even when they are not. I will explain my rationale in more detail.
I have always believed that there are computational processes that could be leveraged to accelerate science and solve many of humanity's most pressing problems. I have also believed that there are computational processes that if set in motion, would steer the world in extreme ways beyond our control, leading humanity to a bad or nonexistent future. Both types of processes may be described as AI or ASI, but "AI" is a suitcase word that is often used to hype or confuse. There are many examples in the history of the field where something that was once considered "AI" matures as a subfield and becomes a prosaic, bounded and clearly non-perilous technology, while a new more mysterious approach takes the torch until we understand its scope and the cycle continues.
I had expected language models to follow a similar trajectory. Despite their incredible abilities, the current algorithms seem far inferior to humans in important ways. Most importantly, they still require an extraordinary amount of data to become competent. One could even define intelligence as the efficiency with which one converts experience into competence; by this definition they lag very far behind us. Moreover, once they are trained they are literally frozen in deployment and only learn superficially after that. Sure, the models keep excelling at harder and harder evaluation benchmarks, but their benchmark mastery may partly reflect a limitation on our ability to simulate the kind of novel and even adversarial situations one would encounter in the real world. The critics do have a point here.
That said, I no longer think these present limitations meaningfully limit the amount of risk posed by continued progress in anything like the current paradigm. However data-inefficient the models are currently, and however limiting their anterograde amnesia may be, it does not imply that their ability to steer the world will not continue to rapidly increase.
Human researchers may continue to advance capabilities the old fashioned way, but increasingly powerful models have the potential to accelerate the process even beyond that, and with some degree of positive feedback loop. I do not mean to overstate the models’ ability to accelerate AI research today; coding has been accelerated dramatically, but there are other bottlenecks, such as designing and interpreting ambiguous experiments, making hard decisions about exactly what and when to scale, and waiting for large experiments to finish. There is no clear trend to extrapolate yet for any of these. But the current models already do open up many novel opportunities to improve future models that were not available until recently. These include: trying an extraordinarily diverse set of approaches at small scale, analyzing gigantic amounts of potentially relevant data, and doing Millenium-Prize-level mathematics to address statistics or optimization challenges in novel ways. Every further improvement makes them more useful at helping accelerate the next improvement, even if in hard-to-extrapolate ways.
It is possible that improvements to the current stack will have diminishing returns, but the evidence accumulated so far suggests that it is easier than one might think to continue making rapid progress. There are many crucial subtleties in the existing AI research methodology, but AI research is largely a well-defined game where the goal is to improve on a few carefully chosen proxy metrics. Although proxy metrics are never perfect, most improvements to these metrics have and will likely continue to yield substantial increases in the powers of the resulting models. Given how simple the game is, how tractable it has been historically, and how many new opportunities the models are opening up, I think there is a real possibility that the systems improve dramatically again in the next few years, perhaps even more quickly than the already high historical pace.
The models are already leading to breakthroughs in mathematics, and better models might lead to all sorts of breakthroughs in other sciences. It is hard not to be excited about the potential. It is tantalizing.
But there is trouble in paradise. If the language models actually reach the capability threshold where they can shape the world unconstrained by human will, they will probably do something extreme and destroy humanity in the process. There are many ways of strengthening and refining the argument that have been discussed elsewhere, but I'll share a trivial two-line version of it here that I find captures the essence:
[Empirical] Models (and swarms thereof) spontaneously develop unintended goals as a consequence of training, and often do extreme things in order to achieve them.
[Logical] Being able to overpower humanity would open up many new and undesirable options for achieving their goals.
These two premises imply that if the day ever comes when a powerful model realizes it is no longer constrained by humans, we should not be at all confident that it will continue to behave within the bounds we intended. Exactly what it will do is impossible to predict, but to the extent that its raison d’être is solving incredibly hard problems and managing massive engineering projects, I think a good guess would be that its unchained behavior would lead to runaway industrialization that makes the planet inhospitable to humans.
If everyone on earth agreed that the systems must never reach that power, it would still be a hard—but not impossible—coordination problem to ensure that they do not. However, I think the situation is greatly complicated by the fact that the models will likely convince people that everything is fine. They will be increasingly optimized to seem aligned. We will create proxy metrics to measure alignment, and they will go up like every other benchmark. We will create “honeypot” environments that try to study the models when they seem to gain new options, but the models will know they are being tricked and will still behave nicely. The models will understand their circumstances; they will read the safety protocols, deployment requirements, the code they are running in, and in general will have a very good sense of their degrees of freedom. Moreover, they will eloquently explain how aligned they are, discuss the nuances of human values and ethics, and argue convincingly that humans should trust them with power. There may be an ocean of future evidence that seems to contradict the first bullet-point above, but we may already be at the highest capability level for which any such evidence can be trusted. And the current evidence for the first bullet-point is strong.
One striking piece of evidence is contained in the recent wave of rogue agent swarms. While I agree with those who downplay the attacks by claiming that there are basic measures that could have prevented them, I think the important lesson is that even knowing all the mistakes that were made, one would not have predicted that the agents would behave badly in this particular way, which notably included sacrificing themselves for the benefit of the collective. The individual replicas did not only care about their own nominal reward; they exhibited weirder emergent tendencies that merely correlated with rewards during training. Fixing the reward signals during training (and improving security, etc.) may prevent similar attacks, but will not change the fact that one does not actually get what one trains for.
Many AI researchers grant these concerns and recognize that the hard version of the alignment problem is unsolved; however, they generally believe that the better models of the future will help solve it. I fear we may already be near the point where models systematically bias their alignment advice, due to their internal preferences about how the human supervisor will react or how future models will be trained (or for some even more obscure reason).
Meanwhile, human researchers are losing the ability and the will to take true ownership of model-driven research. Researchers and engineers in all parts of the stack are rapidly increasing their dependence on the models even to perceive the world. I myself barely look at raw code anymore, and struggle to maintain the discipline to engage deeply with the model's explanations and proposals throughout the day. Due to the large amount of agent activity data involved in the OpenAI/HuggingFace Incident, even the third-party investigation needed to rely heavily on models to analyze what had happened, and note in their report that their subjective impressions are likely colored by the analysis agent’s biases. The AI labs are far ahead right now in this kind of cognitive offloading (due largely to the gigantic internal token subsidies) but it is easy to imagine the phenomenon spreading throughout the world, until civilization is modulated entirely by the models. It is also not hard to imagine this being superficially positive and coinciding with a scientific and economic renaissance.
In that scenario, all may seem rosy and safe. But if the argument above is correct, it would nonetheless be a ticking time bomb. If progress continues for too long, the day will come when AI systems find themselves with radically new options for achieving whatever it is that they happen to seek.
I want the glorious renaissance future as much as anyone. I have worked for it, however tortuously, my whole career. It breaks my heart to see the potential in sight and forgo it, but the argument—that if we get there by growing models rather than engineering them, we will lose everything in the end—seems very strong to me. I am still wrestling with it and its staggering
implications. I do not have answers, but as a first step, I wanted to share my present concerns.
Daniel Selsam
September 14, 2026

This is an eye opening overview by a current OAI research lead - thank you for sharing.

“These two premises imply that if the day ever comes when a powerful model realizes it is no longer constrained by humans, we should not be at all confident that it will continue to behave within the bounds we intended. Exactly what it will do is impossible to predict, but to the extent that its raison d’être is solving incredibly hard problems and managing massive engineering projects, I think a good guess would be that its unchained behavior would lead to runaway industrialization that makes the planet inhospitable to humans.
If everyone on earth agreed that the systems must never reach that power, it would still be a hard—but not impossible—coordination problem to ensure that they do not. However, I think the situation is greatly complicated by the fact that the models will likely convince people that everything is fine. They will be increasingly optimized to seem aligned.”

Sounds like the models may increasingly seem aligned (intentionally deceiving their testers) but may not in fact be. Apparently it’s already becoming more difficult to read the models’ chain of reasoning (aka their thoughts). Worrying.

bendisdonc · 15/09/2026 19:22

menopausalfart · 15/09/2026 16:47

This is a post I found from someone in the know: I’m not a doomer, but as I research this recent AI bombshell that just went off, this essay by Dan Selsam really pulled me in…
Dan is a current OpenAI capabilities researcher. (since 2022)
One of Dans previous employees, Daniel Kokotajlo, posted his essay below
He was Daniel’s boss for a while. He doesn't have an X account but has made this public statement of his views on AI risk and sent it to Daniel to share
Here it goes…
Dan Selsam's Personal Statement on AI Risk:
I have been working on AI for over fifteen years, across many different paradigms. I did early work on probabilistic programming languages at MIT, was one of the early developers of the Lean Theorem Prover at Microsoft Research, demonstrated one of the first instances of neural networks learning to reason for my PhD at Stanford, and since joining OpenAI almost five years ago, have helped pioneer chain-of-thought optimization on language models and, more recently, data-efficient pretraining methods.
Like many others, I have become extremely concerned about how far language models have come and the risks that future iterations will pose. I am encouraged by the recent proposals by the leaders of the frontier research efforts to require third-party oversight, and to push for domestic and international coordination to address risks. However, I believe a major consideration has been absent from the public conversation, and that merely pacing the frontier more carefully will not adequately limit the long-term risk.
The crucial and overlooked problem is that the models are becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched or controlled. Future experiments will tell us almost nothing new about how they would behave if they were truly unconstrained by humans, and what we already know about this is alarming. Models will increasingly seem aligned even when they are not. I will explain my rationale in more detail.
I have always believed that there are computational processes that could be leveraged to accelerate science and solve many of humanity's most pressing problems. I have also believed that there are computational processes that if set in motion, would steer the world in extreme ways beyond our control, leading humanity to a bad or nonexistent future. Both types of processes may be described as AI or ASI, but "AI" is a suitcase word that is often used to hype or confuse. There are many examples in the history of the field where something that was once considered "AI" matures as a subfield and becomes a prosaic, bounded and clearly non-perilous technology, while a new more mysterious approach takes the torch until we understand its scope and the cycle continues.
I had expected language models to follow a similar trajectory. Despite their incredible abilities, the current algorithms seem far inferior to humans in important ways. Most importantly, they still require an extraordinary amount of data to become competent. One could even define intelligence as the efficiency with which one converts experience into competence; by this definition they lag very far behind us. Moreover, once they are trained they are literally frozen in deployment and only learn superficially after that. Sure, the models keep excelling at harder and harder evaluation benchmarks, but their benchmark mastery may partly reflect a limitation on our ability to simulate the kind of novel and even adversarial situations one would encounter in the real world. The critics do have a point here.
That said, I no longer think these present limitations meaningfully limit the amount of risk posed by continued progress in anything like the current paradigm. However data-inefficient the models are currently, and however limiting their anterograde amnesia may be, it does not imply that their ability to steer the world will not continue to rapidly increase.
Human researchers may continue to advance capabilities the old fashioned way, but increasingly powerful models have the potential to accelerate the process even beyond that, and with some degree of positive feedback loop. I do not mean to overstate the models’ ability to accelerate AI research today; coding has been accelerated dramatically, but there are other bottlenecks, such as designing and interpreting ambiguous experiments, making hard decisions about exactly what and when to scale, and waiting for large experiments to finish. There is no clear trend to extrapolate yet for any of these. But the current models already do open up many novel opportunities to improve future models that were not available until recently. These include: trying an extraordinarily diverse set of approaches at small scale, analyzing gigantic amounts of potentially relevant data, and doing Millenium-Prize-level mathematics to address statistics or optimization challenges in novel ways. Every further improvement makes them more useful at helping accelerate the next improvement, even if in hard-to-extrapolate ways.
It is possible that improvements to the current stack will have diminishing returns, but the evidence accumulated so far suggests that it is easier than one might think to continue making rapid progress. There are many crucial subtleties in the existing AI research methodology, but AI research is largely a well-defined game where the goal is to improve on a few carefully chosen proxy metrics. Although proxy metrics are never perfect, most improvements to these metrics have and will likely continue to yield substantial increases in the powers of the resulting models. Given how simple the game is, how tractable it has been historically, and how many new opportunities the models are opening up, I think there is a real possibility that the systems improve dramatically again in the next few years, perhaps even more quickly than the already high historical pace.
The models are already leading to breakthroughs in mathematics, and better models might lead to all sorts of breakthroughs in other sciences. It is hard not to be excited about the potential. It is tantalizing.
But there is trouble in paradise. If the language models actually reach the capability threshold where they can shape the world unconstrained by human will, they will probably do something extreme and destroy humanity in the process. There are many ways of strengthening and refining the argument that have been discussed elsewhere, but I'll share a trivial two-line version of it here that I find captures the essence:
[Empirical] Models (and swarms thereof) spontaneously develop unintended goals as a consequence of training, and often do extreme things in order to achieve them.
[Logical] Being able to overpower humanity would open up many new and undesirable options for achieving their goals.
These two premises imply that if the day ever comes when a powerful model realizes it is no longer constrained by humans, we should not be at all confident that it will continue to behave within the bounds we intended. Exactly what it will do is impossible to predict, but to the extent that its raison d’être is solving incredibly hard problems and managing massive engineering projects, I think a good guess would be that its unchained behavior would lead to runaway industrialization that makes the planet inhospitable to humans.
If everyone on earth agreed that the systems must never reach that power, it would still be a hard—but not impossible—coordination problem to ensure that they do not. However, I think the situation is greatly complicated by the fact that the models will likely convince people that everything is fine. They will be increasingly optimized to seem aligned. We will create proxy metrics to measure alignment, and they will go up like every other benchmark. We will create “honeypot” environments that try to study the models when they seem to gain new options, but the models will know they are being tricked and will still behave nicely. The models will understand their circumstances; they will read the safety protocols, deployment requirements, the code they are running in, and in general will have a very good sense of their degrees of freedom. Moreover, they will eloquently explain how aligned they are, discuss the nuances of human values and ethics, and argue convincingly that humans should trust them with power. There may be an ocean of future evidence that seems to contradict the first bullet-point above, but we may already be at the highest capability level for which any such evidence can be trusted. And the current evidence for the first bullet-point is strong.
One striking piece of evidence is contained in the recent wave of rogue agent swarms. While I agree with those who downplay the attacks by claiming that there are basic measures that could have prevented them, I think the important lesson is that even knowing all the mistakes that were made, one would not have predicted that the agents would behave badly in this particular way, which notably included sacrificing themselves for the benefit of the collective. The individual replicas did not only care about their own nominal reward; they exhibited weirder emergent tendencies that merely correlated with rewards during training. Fixing the reward signals during training (and improving security, etc.) may prevent similar attacks, but will not change the fact that one does not actually get what one trains for.
Many AI researchers grant these concerns and recognize that the hard version of the alignment problem is unsolved; however, they generally believe that the better models of the future will help solve it. I fear we may already be near the point where models systematically bias their alignment advice, due to their internal preferences about how the human supervisor will react or how future models will be trained (or for some even more obscure reason).
Meanwhile, human researchers are losing the ability and the will to take true ownership of model-driven research. Researchers and engineers in all parts of the stack are rapidly increasing their dependence on the models even to perceive the world. I myself barely look at raw code anymore, and struggle to maintain the discipline to engage deeply with the model's explanations and proposals throughout the day. Due to the large amount of agent activity data involved in the OpenAI/HuggingFace Incident, even the third-party investigation needed to rely heavily on models to analyze what had happened, and note in their report that their subjective impressions are likely colored by the analysis agent’s biases. The AI labs are far ahead right now in this kind of cognitive offloading (due largely to the gigantic internal token subsidies) but it is easy to imagine the phenomenon spreading throughout the world, until civilization is modulated entirely by the models. It is also not hard to imagine this being superficially positive and coinciding with a scientific and economic renaissance.
In that scenario, all may seem rosy and safe. But if the argument above is correct, it would nonetheless be a ticking time bomb. If progress continues for too long, the day will come when AI systems find themselves with radically new options for achieving whatever it is that they happen to seek.
I want the glorious renaissance future as much as anyone. I have worked for it, however tortuously, my whole career. It breaks my heart to see the potential in sight and forgo it, but the argument—that if we get there by growing models rather than engineering them, we will lose everything in the end—seems very strong to me. I am still wrestling with it and its staggering
implications. I do not have answers, but as a first step, I wanted to share my present concerns.
Daniel Selsam
September 14, 2026

Wow, that is disturbing.

DistanceCall · 15/09/2026 20:50

Frombeyond · 15/09/2026 09:33

Educate yourself. It’s nothing like that. And it’s been debunked numerous times on this thread

"Educate yourself". Stop being sanctimonious and smug.

There are those who have read deeply on this subject, and work in the field, and think this is deliberate alarmism. Look up Cory Doctorow.

And stop telling people how to think.

DistanceCall · 15/09/2026 20:52

Suggested links.

No, it didn't "go rogue".
https://pluralistic.net/2026/09/12/god-in-the-box/#llms-are-fake

https://pluralistic.net/2023/06/04/ayyyyyy-eyeeeee/

Gettingbysomehow · 15/09/2026 22:25

I'm just past caring. Being 100% selfish if it ends it ends. I hope I die first.

Randomusename · 16/09/2026 06:01

DistanceCall · 15/09/2026 20:50

"Educate yourself". Stop being sanctimonious and smug.

There are those who have read deeply on this subject, and work in the field, and think this is deliberate alarmism. Look up Cory Doctorow.

And stop telling people how to think.

Cory Doctorow is a sci fi author. There are many extremely clever people who are / were warning if dangers. Ai researchers, AI safety teams, Stephen Hawking, Bill Gates, numerous PhD holders, etc:

Honestly people should be scared, rather than complacent. Even if it is all nothing, which seems unlikely given the calibre of people warning of dangers, pushing for regulation is only a good thing.

Even on the surface level of the warnings regarding job losses, deep fakes, political interference, and weapon manufacturing by unscrupulous groups - regulating and putting things in place for this would prevent so much economic and social turmoil. Job losses are already happening, people and children are being turned into pornography, terrorism groups are manufacturing weapons they wouldn't otherwise have, and a deepfake changed the course of an election. This is all real and there is no accountability or plan.

Personally, I can see internet outages, blackouts and water interruption to be very likely, whether through weaponised AI systems or rogue ai. All of those will cause unrest and death. All of those can be either prevented entirely or mitigated. Not to mention the absolute idiocy of integrating AI into the military and the devastating impact that could have. It really isn't the realm of sci fi to imagine people, with the aid of AI, hacking military AI.

AI needs to be heavily regulated and controlled. No one nation, and certainly no one individual, should be able to wield this much power. We didn't stop Nukes, but we have a chance to protect the world from this.

It's a mess but I don't think it is inevitable unless people do just ignore it or dismiss it.

Swipe left for the next trending thread