Reporting and analysis on frontier AI research: new architectures, benchmark results, alignment work, and the labs pushing the state of the art forward.

After a decade in the shadow of pure deep learning, symbolic reasoning is being welded back onto neural systems — not as a rival paradigm, but as scaffolding for reliability.

Learned simulators of physical and social environments are becoming the substrate that reasoning, robotics, and planning increasingly depend on — and the labs that own the best world models may own the next decade.

For years diffusion belonged to images, video, and audio. A new wave of research suggests the same denoising machinery may soon reshape how large language models generate text.

For most of the transformer era, dense attention was treated as a load-bearing wall. A new wave of sparse and structured-attention variants is quietly showing that the wall was decorative.

Rumors of diminishing returns from bigger models miss what is actually happening: the axes of scaling are shifting from parameters toward data quality, test-time compute, and multi-agent orchestration.

Every year a new suite of hard reasoning benchmarks lands, and every year a model clears them within months. What that pattern is really telling us about progress.

A million-token context window sounds like unlimited memory. In practice, attention degrades, retrieval fails, and the illusion of memory quietly costs you accuracy.

Mixture-of-experts models were an academic curiosity a decade ago. They now underpin most frontier systems — for reasons that have as much to do with GPUs as with intelligence.