Microsoft has officially entered the AI cybersecurity race, unveiling its first specialized security model alongside a new agentic platform at an event in San Francisco. The move positions the tech giant as a direct competitor to Anthropic, Google, and OpenAI in the rapidly growing market for AI-driven threat defense.
MAI-Cyber-1-Flash: A Model Built for Vulnerability Hunting
The centerpiece of Microsoft's announcement is MAI-Cyber-1-Flash, a model the company describes as purpose-built to uncover difficult vulnerabilities within complex codebases. The model is designed to power MDASH, Microsoft's dedicated framework for identifying and remediating software vulnerabilities.
Microsoft claims the model outperforms rival offerings on Cyber Gym, an established benchmark the company referred to as the primary standard for AI cybersecurity evaluation. According to the company, MAI-Cyber-1-Flash is both more powerful and more cost-effective than competing models.
Mustafa Suleyman, CEO of Microsoft AI and co-founder of DeepMind, said the model is already being deployed. He stated that MAI-Cyber-1-Flash, combined with GPT 5.4 within the MDASH harness, outperformed Gemini, GPT 5.5 Cyber, GPT 5.6 Sol, and Mythos 5 on the Cyber Gym benchmark.
Perception: An Agentic Security Platform
Alongside the new model, Microsoft introduced Perception, a platform designed to deploy teams of AI agents that assist with and automate various security workflows. The platform is capable of integrating with MDASH and is built to help enterprise defenders counter AI-driven cyberattacks at scale.
Hayete Gallot, Microsoft's vice president for security, emphasized that as hackers increasingly leverage AI in their attacks, defenders need comparable tools. Perception, she explained, enables organizations to defend against AI-powered threats using AI at the scale and speed that attackers already possess.
The platform operates using three distinct agent teams. Red teams generate detailed simulations of potential attacks, offering context about likely threat actors and the vulnerabilities they might exploit. Blue teams focus on detecting and triaging existing bugs, while green teams execute corrective actions to resolve those issues.
