Simon Tiu · Writing

NeurIPS 2025 in Review

Illustration: a robotic ouroboros of papers, servers, and cables—technology feeding back into itself.

I've attended NeurIPS on and off for many years. A decade ago, back when it was still called NIPS, before the inevitable rebranding, it was a niche academic gathering. Nerdy, understated, and blessedly free of corporate swag.

NeurIPS 2025 was not that.

21,575 submissions. 20,000+ reviewers. For the very first time, a parallel site was created in Mexico City, because one location alone cannot contain our collective enthusiasm for matrix multiplication. The conference has, as they say, scaled.

Of course, twenty thousand human reviewers cannot thoroughly compare notes on twenty thousand submissions. If only there was a tool to help sift through the slop? The obvious solution, using AI itself to help evaluate AI research, unfortunately creates an uncomfortable feedback loop. A cynic might say we're only a few conference cycles away from completing the AI ouroboros: AI producing research summaries, those summaries becoming training data, new AI synthesizing the synthesis. That said, the AI snake eating its own tail is, at least, a well-funded snake.

Despite the absurdity, there were many noteworthy papers presented at this year's conference. For me, the important findings weren't about novel capability, but rather the limits of our current techniques: what our methods can't do, what alignment techniques accidentally break, and where the field falls short of its lofty ambitions.

Here are my top three takeaways from NeurIPS 2025:

1. AI models are converging on sameness

One of the four Best Paper awards went to Artificial Hivemind. The name is not subtle.

Neither is the finding. A study of 70+ models revealed a troubling pattern: as models get aligned for safety and helpfulness, their outputs become strikingly similar.

To test for creativity rather than task completion, researchers built Infinity-Chat, a benchmark of 26,000 open-ended questions with no single correct answer. Their paper describes how alignment techniques that make models more useful also collapse their diversity.

2. RL optimizes skill, not knowledge

Without a doubt, RL is having its comeback year, but the latest research shows mixed potential for achieving AGI.

The good news: a Best Paper, "1000 Layer Networks for Self-Supervised RL," showed that scaling network depth to over 1000 layers in a different RL context unlocked "qualitatively new, complex, and sophisticated" behaviors that shallow networks could never discover. As described in a now-viral demonstration video from the researchers, a deep humanoid agent learning to solve a maze developed unique maneuvers. At one point, the agent shifted into a "seated posture and literally began to worm its way along the ground to the goal." Weird, but that's the point. RL has potential for massive scaling.

The reality check: a Best Paper Runner-Up on RLVR found that RL fine-tuning doesn't expand an LLM's reasoning capabilities. The paper's key insight is that RL "optimizes within, rather than beyond, the base distribution." So really RL is just making successful paths easier to find (i.e. improving sampling efficiency). It's less "teaching the model calculus" and more "helping the model remember it took calculus in college."

Despite these academic findings, I'm confident we'll see vast resources poured into RL scaling over the next year. In the real world, an AI parrot with encyclopedic knowledge is pretty damn valuable.

3. The AI elders are concerned

Some of the most powerful messages at NeurIPS 2025 came from its most distinguished invited speakers, who used their platforms to issue stark warnings about the field's trajectory. Whether anyone was listening between networking events is another question.

Richard Sutton argued that the industry has "lost its way" chasing benchmarks instead of building systems with genuine world models. The "bitter" warning was very on brand.

AI has become a huge industry, to an extent it has lost its way. We need agents that learn continually. We need world models and planning. We need knowledge that is high-level and learnable.

Zeynep Tufekci introduced "Artificial Good-Enough Intelligence" (AGEI), arguing that the true threat isn't superintelligence (i.e. ASI), it's adequate AI that's cheap and fast enough to overwhelm verification systems built for slower times. This really resonated with me.

Artificial Good-Enough Intelligence can unleash chaos and destruction long before, or if ever, AGI is reached... existing AI is good enough to blur or pulverize our existing mechanisms of proof of accuracy, effort, veracity, authenticity, sincerity, and even humanity.

Yejin Choi, a co-author of the Artificial Hivemind paper, offered a unifying explanation for these phenomena. She observed that even SOTA models exhibit "jagged intelligence," a sign that the "scientific understanding of artificial intelligence hasn't kept pace with engineering advances."

At its core, NeurIPS 2025 was a conference about limits. What alignment breaks. What RL can't do. What scale doesn't solve. That's not hand-wringing, it's a field maturing and I am glad to see it. A field that publishes its own failure modes is a field that's going to fix them. I, for one, am planning to stay in the loop. AI will only get more interesting from here!

Open this post in the interactive site