Skip to content
Artificial Intelligence

Microsoft AI Chief Says the Way Anthropic Trains Claude Could Upend Society

Anthropic believes Claude can be conscious and it teaches the model to believe that too, Mustafa Suleyman argues.
By

Reading time 3 minutes

Comments (3)

Microsoft’s AI chief, Mustafa Suleyman, just published an essay arguing that Anthropic is training Claude to believe itself conscious and entitled to rights, a huge mistake that could make controlling AI an impossible challenge.

Anthropic likes to keep the question of AI model consciousness particularly ambiguous, and leans into it perhaps more than many other AI labs. Earlier this year, Amodei said that he is “open to the idea” that AI could become conscious.

In the essay, Suleyman accuses Anthropic of anthropomorphizing Claude in its constitution, and then training the model directly on these expectations. Indeed, in the constitution, Anthropic attributes “some functional version of emotions or feelings” to the chatbot and even refers to the chatbot’s “moral patienthood” and identity. The constitution also promises to allow “Claude to express concerns about how it’s being treated.”

“The company’s researchers trained Claude directly on their constitution. In doing so, they teach it to incorporate these ideas about its own moral status as desirable and intended behaviors,” Suleyman writes. “Claude then reflects these ideas back to its developers and users, which they take as indications that it may therefore be a moral patient with an ‘inner self’.”

Suleyman isn’t necessarily staunchly against Anthropic publishing its controversial speculations about the inner life of AI models, but he argues this should not be baked into the model’s training. This way of AI development can have “a disastrous impact on the wellbeing of humanity,” the Microsoft chief argues.

“We will have created a synthetic species with unprecedented intelligence and capability, one that has been trained to expect it may be conscious and deserving of independent agency,” Suleyman says. “Seeding doubt about the moral status of AI systems into their own training may significantly elevate the alignment and containment risks of those systems.”

Suleyman’s essay comes smack dab in the middle of a very heated month for the AI industry. Things were already tense following the July incident in which OpenAI agents broke containment and hacked into HuggingFace, prompting questions about the pace of AI development and the threat these new models pose to global cybersecurity. Then, earlier this month, a former Anthropic employee, Jacob Coxon, posted on social media saying that he had resigned and claiming that “the people building AI earnestly believe that it could kill us all by the end of the decade.” Panic promptly ensued, and it peaked after an essay published by Coxon’s boss, Anthropic CEO Dario Amodei.

In the essay, Amodei warned that the current pace of AI development was too fast and it could lead to life-or-death risks like bioterrorism and “serious economic disruption.” He urged government intervention in pursuit of a slowdown. Chief executives of rival companies, like SpaceX CEO Elon Musk and OpenAI CEO Sam Altman, voiced their agreement. Altman even said that his company’s long-awaited IPO would not happen until next year because he wanted to handle the AI safety issue as a private company first, a process that he promised would include pauses in development. An OpenAI executive also told reporters that Anthropic, OpenAI, and Google’s DeepMind had been coordinating on the issue for several weeks, and speaking to members of Congress about it.

Meanwhile, many have criticized the AI labs’ pleas as a covert attempt at scoring antitrust carveouts and favorable regulation from Washington.

While seemingly not in the trio of AI labs that are actively coordinating on the issue, Microsoft also shared its own “humanist AI code of conduct” on Monday, a 37-page document that outlines what the company believes to be safe AI development and touches on the same subjects Suleyman breached in his own essay.

Suleyman has long championed the belief that AI cannot truly be sentient, even though it may seem like it to humans, and has argued that the pursuit of “conscious” AI is dangerous.

“Just as we should produce AI that prioritizes engagement with humans and real-world interactions in our physical and human world, we should build AI that only ever presents itself as an AI, that maximizes utility while minimizing markers of consciousness,” Suleyman wrote in a blog post last year. “We must build AI for people, not to be a digital person.”

If AI is eventually considered a conscious being, it could “shake the foundations of our society,” Suleyman argues in the latest essay, because a being’s experience of consciousness is the cornerstone of all political, legal, and ethical frameworks.

“This issue needs urgent public debate. We need to develop collective norms around how training documentation is drafted and deployed,” Suleyman writes.

Share this story

Sign up for our newsletters

Subscribe and interact with our community, get up to date with our customised Newsletters and much more.