AI in 15 — October 09, 2026
Scott Aaronson called it "one of the biggest days in mathematical history." About a day later, a plus one that should have been a minus one took down three of the papers.
Welcome to AI in 15 for Friday, October 9, 2026. I'm Kate, your host.
And I'm Marcus, your co-host. Happy Friday, Kate.
Happy Friday! Here's what we've got. OpenAI's giant math release gets its first retractions, and the people who can actually read these proofs start weighing in.
OpenAI's revenue turns out to be about twenty billion dollars smaller than last week's reports.
Anthropic's new usage policy says you're not allowed to be cruel to Claude.
Plus hackers break OpenAI's Codex at Pwn2Own, a nearly two-billion-dollar push to create biology data for AI, a new six-hundred-billion-parameter model from China, and a man who says Claude Code helped him find a planet.
Yesterday we told you OpenAI had posted more than seven hundred math manuscripts on GitHub. Marcus, a lot has happened since.
It has. The collection covers more than three hundred and seventy open problems, all produced by an internal model that hasn't been released. Each solution took about three hours of compute on average. Then, within roughly a day, OpenAI withdrew three papers. All three were about the Hodge conjecture, which is one of the Millennium Prize problems. The first paper used plus one where its own conventions called for minus one. That sign error broke a key step, and the other two papers relied on that step, so all three fell together.
So one wrong sign took out three papers.
A chain of dominoes. OpenAI also revised fourteen papers and fixed thirteen citations, which takes the total from seven hundred and twenty-two down to seven hundred and nineteen. Here's the number I'd focus on: according to reports, only about forty-two percent of the top-level results, around three hundred, have been formally verified in Lean.
Remind people what Lean is.
It's software that checks a proof one line at a time, like a very strict spell-checker for logic. If a result has a Lean certificate, you can trust it as fast as a machine can check it. If it's written only in prose, a person has to read it, and these are very long, very new proofs. The Hodge papers were the kind a person has to read.
So what are the big claims in there?
Scott Aaronson, the complexity theorist, wrote a long blog post listing them. The biggest is a proof of the Unique Games Conjecture, a central open problem in theoretical computer science, and that one does come with a Lean certificate. There's also L equals BPL, a major result about whether randomness actually helps algorithms. He also lists faster algorithms for matrix multiplication and integer multiplication, a long-standing quantum complexity question resolved, and a positive answer to a problem Aaronson himself posed in 2007. Fortune adds that there was progress on three more Millennium problems, though none were fully solved.
Wait, his own problem? How did he take that?
With a lot of excitement, and with an honest caveat: no human has fully understood most of these proofs yet. The person who has spent the most time on the Unique Games proof so far is his wife, Dana Moshkovitz, who has spent her career on that conjecture. She says it's hard to read, it invents a new tree-based code, and the citations are unclear, but she believes it's correct.
That's a pretty strong endorsement from someone who knows the problem that well. How are other mathematicians reacting?
They're split. A group called the Association for Human Mathematics published an open letter urging mathematicians to stop working with OpenAI. It includes the line "Mathematicians did not ask for this work to be done."
Hmm. Nobody asked for the printing press either.
Hacker News made pretty much that point, loudly. There are more serious criticisms, though. Tristan Buckmaster at NYU says OpenAI hasn't done its "due diligence at all" to rule out that the model built on other people's unpublished work. And OpenAI's own math advisory group, which is hosted at the Institute for Advanced Study, says the company followed some of its responsible-release recommendations, not all of them.
And who's on the other side?
Dan Litt at the University of Toronto put it simply: "My view is that this is great for mathematics." The set theorist Asaf Karagila described a more practical problem. Researchers are getting flooded with messages from the public saying "an AI proved something in your field, is it right?" Aaronson also pointed out that a separate result by Virginia Williams and Josh Alman was assisted by an Anthropic model. He contrasts OpenAI's approach of releasing everything with Anthropic's more digested one.
So where does this leave us?
Sort the results into two piles. The Lean-verified ones are strong evidence that these models do real research math, not just competition problems. The prose-only ones are claims until someone checks them, and this week we saw how fast they can fail. The next month is going to be about humans checking the headline results.
Quick hits. First, money. The Financial Times reports that OpenAI told investors its annualized revenue was approaching fifty billion dollars at the end of September. A week ago, reports were saying seventy.
And annualized revenue, to be clear, means taking a recent month or quarter and projecting it over a full year. According to the FT's source, the seventy billion figure came from OpenAI's own investors trying to build a direct comparison with Anthropic. They took an August estimate of about forty billion and applied a seventy percent growth figure that had been reported.
So the number was extrapolated, not reported by OpenAI.
Pretty much. The underlying issue is accounting. Both companies follow GAAP, the standard US rules, but they book sales through cloud partners differently. Axios gave an example. On a hundred-dollar sale through a cloud provider, Anthropic can record the full hundred as revenue and the provider's cut as an expense. OpenAI records only its own share. Apply OpenAI's method to both and the comparison looks less good for OpenAI. And on the other side, Anthropic's reported operating profit in Q2 excludes stock-based compensation, meaning employees paid in shares.
Why should a listener care about what's basically an accounting dispute?
Because OpenAI is valued at eight hundred and fifty-two billion dollars, and the stock prices of Nvidia, Oracle and CoreWeave partly depend on these figures. OpenAI filed confidentially for an IPO in June and is reportedly aiming to list in 2027. Once it's public, investors get audited quarterly numbers instead of figures passed along by people with a stake in them. Personally, I'm looking forward to that.
Next, Anthropic published a usage policy update that takes effect November 12. The rule everyone's talking about bans "sustained and needless abusive or cruel behavior" toward Claude.
And it's narrower than the headlines make it sound. It covers only extreme cases where someone repeatedly mistreats the model "with no discernible purpose." Ordinary frustration, creative writing and legitimate red-teaming are all excluded. The consequence is that the conversation ends. Your account doesn't get banned. Claude has actually been able to end abusive chats since August 2025.
So I can still yell at it when it breaks my spreadsheet.
Yell away. What's new is that a major lab has written model-welfare thinking into its terms of service. Hacker News was split. Some people said Anthropic was "getting high on its own supply." Others gave a practical reason: models learn from conversations, and normalizing abuse could cause alignment problems later.
Is there anything else in the update?
Honestly, the less talked-about parts may matter more. The weapons section now specifically names guidance-and-control software and arming drones. There are new limits on tracking people without their consent and on having Claude recommend who police should investigate or arrest. There are safety requirements for autonomous hardware that takes physical actions. The elections section has been rewritten to focus on deceiving voters, and one summary says the blanket ban on personalized campaign targeting was dropped. Anthropic ties all of this to Claude doing longer, more independent work.
Over to Cork, Ireland, where Pwn2Own is running this week. That's the hacking contest where researchers get paid to break things.
And this year there are dedicated categories for AI infrastructure and coding agents. On day one, researchers exploited thirty-two zero-days, meaning previously unknown flaws, and earned three hundred and eighty-eight thousand five hundred dollars. CyberInsider counts twenty-eight, because some exploit chains reused known bugs. The headline for us: one team took down OpenAI's cloud Codex coding agent with a single argument-injection bug.
Just one bug?
One. Argument injection means sneaking extra instructions into a command the agent runs. If your team gives an agent shell access, that's the kind of bug that should worry you. Researchers also showed zero-days in LiteLLM, a popular open-source gateway that routes requests to different AI models. And Vietnam's VinSOC team earned forty thousand dollars for a five-bug chain against Oracle's Autonomous AI Database.
So AI agents are now tested the way phones are.
Exactly, and they should be. A Samsung Galaxy S26 was hacked twice too, so they're in good company. Final totals come in when the contest ends today.
A quick update from yesterday's GPT-6 story. Along with Intelligent UI, OpenAI also released a Decisions API in public beta.
It's a small thing, but I like it. It runs on GPT-6 Luna and returns only structured answers: a probability, a choice from a fixed list with confidence scores, or a score in a range. Input is ten cents per million tokens, and output is free. One early tester found something worth knowing. In choice mode, a coin weighted seventy percent toward heads came back "heads" ninety-eight percent of the time. In probability mode, the API correctly said seventy percent.
So if you ask it to pick, it hides the uncertainty.
Right. If you're building on it, ask for the probability. And while we're on yesterday's stories, a correction on Haiku 5.5. We said the price cut was ninety percent, based on the list prices. Anthropic says it's about seventy-five percent cheaper on average, partly because a new tokenizer splits text into more tokens. So the per-token price fell more than what you'll actually pay per task.
Thanks for keeping us honest.
Now biology. Biohub is expanding its Virtual Biology Initiative to about one point eight billion dollars in funding, data, compute and measurement technology.
It comes from several sources. More than five hundred million from the US Department of Energy over five years, through the Genesis Mission. Over five hundred million worth of existing NIH-funded datasets, to be cleaned up and standardized for AI training. Three hundred million combined from Google DeepMind, Isomorphic Labs and Meta. And Biohub's own founding five hundred million. Nvidia is supplying compute, and the Broad, Allen and Sanger institutes are among the scientific partners.
Why does this matter more than another big model?
Because in biology the bottleneck is data, not compute. One Hacker News commenter said, "Compute was never the primary bottleneck here." The money goes to wet-lab data collection: how cells respond to drugs and genetic changes across many cell types, plus imaging that can capture millions of cells. And it's open, with shared standards. Getting DeepMind, Meta and government agencies to feed one dataset is an achievement in itself.
On open models, Step 5 Preview from Chinese lab StepFun hit the top of Hacker News after appearing on OpenRouter.
It's a sparse mixture-of-experts model, which means it uses only part of the network for each token. It has about six hundred billion parameters in total, around twenty-seven billion active per token, and a million-token context window. It accepts images and short videos. Pricing is about a dollar per million input tokens and two seventy for output, and it scores forty-four on Artificial Analysis's Intelligence Index. One commenter called it "smarter and slightly cheaper than Gemini 3.8 Flash."
And the weights are open?
They're promised for October fifteenth. Right now the Hugging Face repository is empty. Same as Mistral's "Le Chonk" yesterday: "open-weight" doesn't mean much until the files are actually there. And most of the benchmark numbers come from StepFun itself. Also, at six hundred billion parameters, it won't run on a home machine with a hundred and twenty-eight gigabytes of memory, which was a common complaint on Hacker News.
Lightning round. Marcus, someone on Reddit says they found a planet.
An amateur says they ran Claude Code agents on raw telescope data and turned up a possible new planet. My favourite Hacker News reply: "We got vibe astronomy before GTA6." To be clear, nothing has been independently checked or peer-reviewed. It's a claimed candidate, and the skeptics are right to ask about false positives.
Fair. What else?
Samsung published LittleBit on GitHub, a method that compresses language models below one bit per weight. Quality drops a lot. One benchmark score falls from about eighty percent to forty-seven. So it's an interesting experiment, not something practical yet. And there's a viral YouTube parody called "LGTM," short for Looks Good To Me, about AI code review. It's billed as a Claude Opus 5.5 music video, but commenters noted the song was made with Suno.
Seems about right that nobody checked the credits on a song about code review.
One to watch: the human review of the Unique Games proof. If it holds up, this week will be remembered as a turning point. If it doesn't, it'll be remembered for retractions.
Agreed, but that one has a Lean certificate, so I'd bet it survives. I'd watch the prose-only proofs more closely.
That's your AI in 15 for today. See you tomorrow.