Story Highlights
- Artificial intelligence is producing discoveries and software faster than humans can verify them.
- Experts warn that critical systems could increasingly depend on code people cannot fully understand.
- Formal verification could mathematically check whether software behaves as intended.
- The authors argue the U.S. should make verified software a national priority.
What Happened
Artificial intelligence is changing the balance between discovery and verification. The authors point to mathematician Jacob Tsimerman, who recently helped review an AI-generated mathematical proof after researchers determined that the system had resolved a longstanding geometry problem. The episode illustrates a broader challenge: AI can now generate complex results at a pace that can exceed the ability of human experts to examine them. The authors argue that human confirmation, rather than discovery itself, is increasingly becoming the limiting factor.
The same concern extends beyond mathematics to software that operates hospitals, financial institutions and power grids. The article cites recent examples of AI identifying previously unknown software vulnerabilities and the widespread disruption caused by a faulty software update in July 2024. At the same time, generative AI is allowing developers to create large amounts of software through natural-language prompts, raising concerns about code that works without being fully understood by the people responsible for it.
- AI-generated discoveries are increasing the demand for human verification.
- Critical infrastructure increasingly depends on complex software.
- AI tools can identify serious vulnerabilities in existing systems.
- Rapid code generation can make software harder to fully audit.
Why It Matters
The issue matters because software failures can have consequences far beyond individual computers. Financial systems, hospitals, transportation networks and power infrastructure depend on programs functioning correctly. When critical code is difficult to understand or verify, identifying weaknesses before they cause damage becomes more difficult. The authors argue that relying only on traditional patching and reactive cybersecurity measures may not be sufficient for systems where a failure could affect large numbers of people.
The proposed answer is greater use of formal verification, a process that uses mathematical proofs to establish whether software follows specified requirements. The approach cannot guarantee that an entire system is safe or that humans have written the correct requirements, but it can eliminate certain categories of software errors. The authors therefore see mathematical verification as a way to make increasingly AI-driven technology more dependable without attempting to formally check every line of code.
- Software reliability is becoming a national infrastructure concern.
- Formal methods can automatically check defined software properties.
- Human judgment remains essential when defining system requirements.
- Verification could reduce specific classes of critical software failures.
Political and Public Context
The debate places artificial intelligence alongside questions of national security, infrastructure resilience and technological leadership. The article points to work involving defense, intelligence, academia and industry on formal methods, while also citing a Senate hearing in which Sen. Mark Warner discussed the ability of AI tools to discover vulnerabilities in sensitive systems. These developments suggest that AI security is becoming an issue extending beyond the technology industry into government and national defense.
The authors call for a coordinated national effort rather than leaving verification entirely to individual companies. Their proposal includes public libraries of verified software components, common standards and benchmarks, improved tools for checking software updates, and training that connects mathematics, computer science, engineering and national security. They also argue that sustained investment in mathematics education and research will be necessary because formal verification depends on mathematical foundations.
- AI security increasingly overlaps with national security policy.
- Government agencies could play a larger role in verification standards.
- Public infrastructure for verified software is one proposed solution.
- Mathematics and computer science education could become more strategically important.
What Happens Next
The central question is whether government agencies, technology companies and operators of critical infrastructure will adopt stronger verification practices as AI-generated software expands. The authors recommend that companies responsible for critical systems build auditing and formal guarantees into development rather than relying primarily on fixes after vulnerabilities are discovered.
The broader policy debate will likely focus on where mathematical verification should be required, which systems should receive priority and who should establish common standards. The authors do not argue that every piece of software must be formally verified. Instead, they propose concentrating rigorous verification on the logic of systems where failures could create significant public harm.
- Watch for new standards governing AI-generated software.
- Watch how critical-infrastructure operators approach formal verification.
- Watch for federal investment in verification research and training.
- Watch whether companies shift from reactive patching toward built-in assurance.


