h.sHamid Samir
All stories

Altman says AI mathematical reasoning is entering a new tier

Sam Altman described rapid gains in the mathematical abilities of newer AI models, but the comparisons are qualitative claims rather than independently documented benchmark results.

Sam Altman, speaking with Salesforce CEO Marc Benioff, said AI models are making rapid progress in mathematical reasoning.

According to the published account of the conversation, Altman compared GPT-5.5 with an average mathematics professor and placed GPT-5.6 among the top one to two percent. He described a model called “Astra” as slightly beyond that level and said OpenAI’s next internal model was working on problems that challenge even leading mathematicians.

Why the claim matters

If independently validated, the important signal would be more than a score on one test. It would suggest that successive model generations are becoming more capable of handling difficult, multi-step problems. That could expand AI’s role from answering questions to assisting specialists with problems that require sustained reasoning.

A necessary caveat

The published material presents these rankings as conversational comparisons and does not provide the benchmark, problem set, test conditions, or reproducible results behind them. Phrases such as “average professor” and “top one to two percent” should therefore not be treated as independently established measurements. Solving a problem also does not guarantee a correct proof or remove the need for expert verification.

ویدیوی گفت‌وگوی سم آلتمن و مارک بنیوف درباره پیشرفت توان استدلال ریاضی مدل‌های هوش مصنوعی