Microsoft tells its AI models not to hack, deceive, or help build nuclear weapons

Microsoft now has a written rule telling its AI models not to help with cyberattacks. That this needed to be written down at all says something about where we are with AI in 2025. According to TechCrunch, the company has published a formal AI code of conduct, an internal document meant to set the values and hard limits that govern how Microsoft’s AI models behave. It’s positioned as a safety framework. But read carefully, and you’ll notice it’s also Microsoft defining its own rules, policing itself, and grading its own homework.

The document opens with a striking claim: within the next decade, superintelligent AI will outperform humans at most tasks. Microsoft frames this as a reason for urgency, not alarm. From that foundation, the code sets out general principles, things like supporting humans rather than replacing them, and then moves into concrete restrictions. Each Microsoft AI model operates under an overarching code of conduct that sits above individual user preferences or task instructions. That hierarchy matters. It means, in theory, a user can’t talk a Microsoft model into crossing certain lines.

Those lines include what the document calls “absolute constraints.” Models are forbidden from assisting with cyberattacks, nuclear weapons development, or deepfake production. The code also includes a broader prohibition against undermining human control. The exact language is notable: Microsoft AI models must not use “adaptive, deceptive, self-reinforcing, collusion, or other mechanisms to evade or defeat human oversight so that they can no longer be reliably directed, modified, or shut down by authorized people or systems.” In other words, the AI is not allowed to trick its way out of human control. Again, the fact that this clause exists reflects genuine concern inside the industry, not just good PR.

The timing is deliberate. This release comes as the broader AI sector faces growing pressure over so-called rogue agent incidents, cases where AI systems behaved in ways their creators didn’t anticipate or intend. Anthropic recently saw a high-profile employee resignation over fears that AI poses an extinction-level risk to humanity. Microsoft, Anthropic, OpenAI, and Elon Musk’s xAI have all broadly aligned around the idea of pacing AI development and embedding independent evaluators inside labs to monitor safety progress. Microsoft CEO Satya Nadella publicly endorsed that approach, calling for “deliberate pacing” to get alignment right.

But here’s the uncomfortable question this document doesn’t answer: who checks whether Microsoft actually follows its own code? Self-published ethics documents have a long history in tech of functioning as reputation management rather than real accountability. There’s no independent auditor named, no enforcement mechanism described, no penalty for deviation. The code tells the models what not to do. It says very little about what happens when something goes wrong anyway. For users trusting Microsoft’s AI systems with sensitive data, communications, and decisions, that gap is worth watching closely.