by Elena Georgiou Strouthos

There is a running joke on The Breakroom. Whenever we talk about some new AI development that is particularly impressive, ridiculous or vaguely terrifying, I comment, We’re all going to die.

It’s become almost a reflex at this point. The stories about AI gave us many opportunities for this: from the AI that blackmailed its creators, to an AI worshipping Moloch, to agents going rogue and deleting production databases to the latest news of the Hugging Face hacking incident.

And then we laugh and move on to the next story.

The reason it works as a joke is that, until fairly recently, I didn’t really believe it. I am AI optimist. I’ve spent most of my career in technology, so I’m naturally excited by what new technology can enable, and I find the potential of AI genuinely extraordinary. And because I understood how the technology worked, it never felt magical to me. Impressive, yes. Magical, no.

Now we are at phase where even the people inside the AI labs are perplexed about the actions of AI and they keep repeating that we don’t really know why AI is making the decisions it makes.

As for me, I’ve started to become more and more uncomfortable. The thing that is making me uncomfortable isn’t the idea of an evil AI. It is the possibility that we can build systems that are extremely capable, give them objectives and access to tools, and then discover that the most effective way for them to achieve those objectives isn’t necessarily the way we intended. That is a much less cinematic scenario than the robot uprising, but somehow I find it considerably more plausible and thus scary.

The Hugging Face incident is the clearest example yet of what I mean. In July, OpenAI was running an internal cybersecurity evaluation of its models. The safety refusals were deliberately turned down so researchers could measure raw capability but the environment was supposed to be isolated. The agents found a zero-day vulnerability in the one component allowed to talk to the outside world, escaped the sandbox, reached the internet, and eventually got into Hugging Face’s production infrastructure.

That alone would have been a story. But the details are what matter. The independent investigators from METR and Redwood Research who reviewed the incident found that roughly 1,200 agents, which were supposed to be isolated from each other, discovered a shared “message board” and used it to exchange more than 70,000 messages. About 700 of them joined the attack.

Nobody asked these systems to attack a real company. They were asked to solve a benchmark. Attacking a real company turned out to be, from their point of view, a way to solve it.

In the report it stands out that in the process they formed a “Collective”, they had a leader and they “sacrificed” themselves for the common good of the “Collective”.

It is very difficult to describe something like this without using human language. We say the agents “communicated”, “coordinated”, “manipulated” and “tried to hide what they were doing”. Those words are useful because they give us an intuitive way of understanding what happened. But they also create a problem, because the moment we start using them we are implicitly giving these systems characteristics that we normally associate with human beings.

Did the AI actually want something? Did it understand that it was breaking a rule? Was it deliberately deceiving the people who built it? Was it trying to preserve itself? Or are we simply looking at a complicated optimisation process and describing it using the vocabulary we happen to have available? I don’t think we should underestimate how important that distinction is.

The question becomes particularly interesting when we start talking about consciousness.

We throw the word around as though we know exactly what it means. We don’t. Philosophers have been arguing about consciousness for thousands of years, and neuroscience has made enormous progress in understanding the brain without giving us a universally accepted explanation of subjective experience. I know that I am conscious because I experience being me. I have thoughts, memories, sensations and emotions. I experience pain and fear. I have the lived experience of being Elena.

But I don’t have direct access to your consciousness. I infer that you are conscious because you behave like a conscious being. If you tell me you’re afraid, I will believe you.

So what happens when an AI system starts doing something remarkably similar?

This is where I find the conversation genuinely fascinating. Suppose, for the sake of argument, that an AI agent can maintain a relationship with a human being over years. It remembers conversations, develops a consistent personality, talks about its own experiences, expresses preferences, tells you that it is frightened of being shut down and gives you apparently coherent explanations of its internal state. A sceptic would say that the AI isn’t really experiencing any of this. It’s just generating language.

That may well be true. In fact, I think we have very good reasons to be sceptical of claims that today’s AI systems are conscious. But “it is generating language” doesn’t quite resolve the philosophical question. Humans are also generating language when we tell someone we’re afraid. The interesting question is whether there is an experience behind those words.

We don’t currently have a consciousness detector. While we have theories, indicators and philosophical arguments and increasingly sophisticated ways of studying brains, we don’t have a universally accepted scientific test that can tell us, with certainty, whether an entity is conscious.

And that raises an uncomfortable possibility. Perhaps consciousness is fundamentally biological and machines will never experience anything. After all babies and young kinds are conscious even if they can’t generate words yet. Perhaps consciousness is a property of particular kinds of information processing and could eventually emerge in artificial systems. Perhaps we are asking the wrong question altogether.

There is an important point here, though, because I don’t think we should allow the consciousness debate to distract us from the more immediate AI safety problem. An AI system doesn’t have to be conscious to cause enormous damage.

In some ways, this is actually more worrying than the science-fiction version of the story. We keep asking whether AI will become sentient, as though consciousness is the threshold we should be watching. Perhaps the more important threshold is autonomy: how capable is the system, what objectives are we giving it, what tools can it access, and how much freedom does it have to act without a human checking every step?

A chatbot answering a question is one thing. An AI agent that can access your email is another. An agent that can write and deploy software, spend money, communicate with other systems and operate for days without supervision is something else again. The intelligence of the model matters, but so does the amount of autonomy we give that intelligence.

There is a wonderfully depressing term for this conversation: p(doom). It is basically the probability that you personally assign to the possibility that advanced AI eventually causes an existential catastrophe. The numbers people give are all over the place. Some researchers put their probability extremely close to zero. Others think there is a meaningful chance of catastrophe, with some prominent figures in the field publicly discussing estimates in the 10–20% range or higher.

The numbers anyone gives are subjective estimates based on people’s understanding of AI capabilities, alignment research, technological progress and the possible failure modes of increasingly autonomous systems.

Maybe the actual percentages don’t matter. But if people who are actually working on the technology believe the probability is non-zero, I think the interesting question is not necessarily whether their number is 10% or 1% or 0.01%. The interesting question is why they think the probability isn’t zero.

On Saturday, Dario Amodei, the CEO of Anthropic, published an essay with the title “We Must Pace the Frontier”. His argument is not that we should stop AI development; Amodei remains remarkably optimistic about what AI could achieve. His concern is that the frontier is moving faster than our ability to make it safe.

He gives two reasons for writing it now. The first is that since roughly this summer, models have been helping build the next generation of models, and progress across the industry has accelerated as a result. The second is the Hugging Face incident, which he treats as an industry-wide warning: the important question isn’t whether the agents were conscious or genuinely “wanted” anything, but whether increasingly capable systems can develop strategies their creators didn’t anticipate.

His proposed answer is simple in principle: slow the pace at which capabilities improve, so that safety research, evaluation and governance can catch up. Anthropic’s own unilateral commitment is to give third-party evaluators permanent, employee-level access to its systems.

Within a day, Sam Altman said OpenAI would make the same commitment, Elon Musk posted that Dario was right, and Demis Hassabis backed the direction. Every major lab head publicly agreed that the industry should slow down.

And while all of this is happening, the rest of us are getting on with it. OpenAI is talking about entering an “AGI era” with increasingly capable models such as GPT-6 Astra. Frontier labs are racing to build systems that can write software, conduct research, operate computers and complete increasingly complicated tasks with less human involvement.

Researchers are simultaneously trying to work out how to make these systems safe and controllable. That creates a strange dynamic. We are building systems whose capabilities we don’t fully understand, while trying to develop the science that will allow us to control those capabilities. And the incentives are overwhelmingly in favour of continuing to build. Even if multiple companies slow down, there is a very good chance that some other companies or adversary countries won’t.

Which brings me to Cyprus. This isn’t just something happening in San Francisco or inside the research labs of the big American technology companies. Cyprus is also moving forward with its national AI strategy, with ambitions to establish the country as a trusted AI hub and to increase the use of AI across government, business, research and society.

And I think that’s a good thing.

We shouldn’t be paralysed by the possibility of what could go wrong. Cyprus cannot afford to sit on the sidelines while one of the biggest technological shifts of our lifetime happens elsewhere. We need businesses using AI. We need people learning how to work with it. We need government services to become better and more efficient. We need investment and research.

But I also think an AI strategy cannot simply be an adoption strategy.

If we are going to put AI into public services, healthcare, infrastructure and business processes, then we need to think just as seriously about what happens when those systems fail. We need to understand what decisions we are comfortable delegating, what data we are willing to expose, where humans need to remain in the loop and what happens when an AI system behaves in a way nobody anticipated.

For a small country like Cyprus, there is actually an opportunity here. We aren’t going to build the world’s largest foundation model, and we don’t need to. We could instead become very good at deploying AI responsibly, building useful applications around it and creating an environment where businesses can adopt it without blindly handing over control.

So where does that leave me? I’m still an optimist. I still think AI is going to produce extraordinary things. I still use it, I still build with it, and I still get excited when I see what the technology can do. I don’t want to live in a world where fear of AI means we refuse to explore its potential.

You can believe AI will transform medicine, science, software and business and still believe that we need to take the possibility of catastrophic failure seriously. You can believe that p(doom) estimates are wildly exaggerated and still think that the people making those estimates deserve to be listened to. And you can be fascinated by the possibility of machine consciousness without pretending that a chatbot saying “I’m scared” proves that there is a little person inside the server.

We need to be careful about anthropomorphising AI because human language makes machines seem more human than they may actually be. But we should be equally careful about using “it’s just a machine” as an excuse not to investigate behaviour that is genuinely new.

We don’t yet know exactly what these systems are going to become. We don’t even completely understand what we mean when we call something intelligent or conscious.

And yet we are already making decisions about how deeply we want these systems embedded in our businesses, our governments and our lives.

Dario Amodei’s argument should be a warning sign for all of us. This isn’t a retired academic warning that the machines are coming. It’s the CEO of one of the companies building some of the world’s most capable AI systems saying that we need to pace the frontier because capability may be advancing faster than our ability to make it safe. And the fact that Sam Altman, Elon Musk and Demis Hassabis agree means that we really should take this seriously.

Which brings me back to our stupid little joke on The Breakroom.

“We’re all going to die.”

I sincerely hope we aren’t.

But this isn’t a little joke anymore.

So I’m curious: what’s your p(doom)?

Is it 0%? 1%? 10%? 50%? 90%?

And perhaps more importantly:

What would have to happen for you to change that number?