AI News

OpenAI says an internal version of its next major model, Astra, has produced new results on ten long-standing problems in mathematics and theoretical computer science. The company says the work spans geometry, coding theory, group theory, complexity, cryptography and combinatorics, with each argument later formalized in Lean.

The announcement matters less as a claim that AI has independently solved a broad set of famous problems than as a test of whether language models can generate research-level mathematical arguments that experts can inspect and verify. OpenAI said the problems had seen no progress on their main results for at least a decade, and in many cases much longer. The company also acknowledged that the work remains subject to engagement and scrutiny from the relevant mathematical communities.

Ten problems, across several fields

OpenAI’s list includes new upper bounds for high-dimensional sphere packing, reaching down to the Cohn–Elkies threshold, as well as improved bounds for the size of binary and spherical codes at specified minimum distances. These are central questions in high-dimensional geometry and coding theory.

In group theory and operator algebras, the company reports a construction showing that non-sofic groups exist and says it has disproved Connes’s rigidity conjecture. That conjecture concerns whether certain groups are uniquely determined by their von Neumann algebras.

The results also include new lower bounds for arithmetic circuits and formulas computing the permanent. OpenAI gives an arithmetic-formula lower bound of order n^4/log n. In complexity theory, the company reports an exponential parallel repetition theorem for general two-player quantum games, extending a principle better established in classical settings.

Other reported results concern the closest vector problem, a lattice problem linked to post-quantum cryptography; Ehrhart’s volume conjecture on convex bodies with a single interior lattice point; multicolor triangle Ramsey numbers; and compactness and degeneracy questions in extremal graph theory. OpenAI says the latter results resolve Erdős problems 146 and 180, while the Ramsey result resolves Erdős problem 183.

The announcement follows OpenAI’s earlier disclosure of an AI-generated disproof of the Erdős unit-distance conjecture. The company says that result helped prompt subsequent research by external mathematicians on related problems in additive combinatorics, computational geometry and complexity.

How the results were produced

According to OpenAI, the ten results were generated by an internal version of Astra rather than by a publicly available model. The company says the total token usage required to find the solutions would have cost approximately $2,000 at Sol API rates. That is a vendor estimate, not an independently audited measure of the cost of reproducing the research.

Humans then prepared the arguments into manuscripts with the same model, OpenAI said. The model subsequently formalized each argument in a Lean certificate. OpenAI is also releasing a model-generated narration of the system’s thinking process for each solution, although such narration should not automatically be treated as a complete or faithful record of the model’s internal computation.

The Lean formalizations are important because they provide a machine-checkable representation of the proofs. They do not, by themselves, establish that the underlying claims are mathematically important, novel in every detail or correctly interpreted in their broader research context. Those questions still require examination by specialists and comparison with existing literature.

Evidence, attribution and uncertainty

The available evidence for the claims comes from OpenAI’s official announcement. A separate wire-style listing carries the same headline, but the supplied material does not include independent reporting or the full text of a third-party article. As a result, the strongest claims about the model’s performance, cost and research impact are OpenAI-reported.

OpenAI says it takes responsibility for the correctness of the prepared manuscripts and formalized proofs, while assigning the mathematical arguments themselves to the system. That distinction is unusually explicit for an AI research announcement. It reflects a growing dispute over how authorship and credit should be assigned when a model proposes an argument, humans edit it and formal systems check parts of the result.

The company is not presenting the announcement as a substitute for peer review. Instead, it asks mathematicians to place the results in context and develop the ideas through further research. Until that process occurs, the announcement should be read as a report of potentially significant machine-generated arguments, not as settled evidence that Astra has replaced expert mathematical research.

OpenAI also links the work to ChatGPT for Academic Researchers, an initiative it says will provide 100,000 scientists and mathematicians with free access to its best ChatGPT models. That access program is separate from the Astra results, and the announcement does not show that the reported advances can be reproduced using the publicly distributed tools.

What it means for research teams and AI builders

For mathematical researchers, the most practical signal is the combination of generation and formal verification. A model that can suggest a plausible proof but cannot express it in a system such as Lean remains difficult to trust at scale. Formalization creates a narrower path for checking correctness, although it does not eliminate the need for human judgment about definitions, novelty and significance.

For AI builders, the announcement points toward evaluation suites based on open research problems rather than conventional question-answering benchmarks. OpenAI says it has been evaluating models on such problems during development. This approach tests whether a system can sustain multi-step reasoning, identify useful intermediate structures and produce artifacts that other tools can verify.

The economics are also relevant. OpenAI’s estimated $2,000 inference cost for ten results suggests that research-grade model use may be feasible for well-funded teams, but the figure cannot yet be generalized. It excludes the value of expert review, manuscript preparation, formalization work and the infrastructure needed to reproduce or extend the results. Enterprise buyers should therefore distinguish raw model-compute cost from the full cost of a reliable research workflow.

For founders and product teams, the near-term opportunity is likely to be narrow, tool-connected systems rather than general-purpose automated mathematicians. Workflows that combine models with theorem provers, code execution, literature search and specialist review may be more dependable than systems judged only by fluent explanations.

What to watch next

The first signal will be external verification of the ten claims, particularly the reported disproof of Connes’s rigidity conjecture, the non-sofic group construction and the quantum parallel repetition result. Publication, independent formal checks and detailed responses from specialists will show whether the arguments survive scrutiny and how much is genuinely new.

Researchers should also watch whether the Lean certificates are complete, accessible and usable by others, rather than merely accompanying polished manuscripts. Reproduction by teams without access to OpenAI’s internal Astra system would provide a stronger test of the work’s practical value.

A further question is whether similar results emerge from publicly available models at materially lower cost. That would indicate that the capability is spreading across the field rather than remaining an internal advantage tied to one unreleased system.

Creati.ai perspective

OpenAI’s announcement is significant because it presents AI as a generator of candidate mathematical research, while also exposing the limits of that claim. The most credible part of the workflow is not the model’s narration or the breadth of the list; it is the attempt to convert each argument into a machine-checkable Lean certificate and to state openly how responsibility is divided.

The next stage will be determined outside OpenAI. If independent mathematicians validate, simplify and extend the results, the announcement could mark a meaningful shift in research tooling. If the claims prove difficult to reproduce or place in context, it will instead demonstrate why model output, formal verification and human mathematical judgment must remain separate layers of the same process.

Featured

OpenAI reports ten AI-generated advances in mathematics and theoretical computer science

OpenAI says its Astra model generated ten advances on long-standing math and theoretical computer science problems, with Lean formalization for review.