Skip to content
Adventures and Amusings of a Mathematician
Go back

The Edge of Artificial Intelligence Research: A Citation-Grounded Survey from ICML, ICLR, NeurIPS (2023–2026)

There is a version of this post that is just a list of papers. I want to write the version that explains why each paper matters and what the citation graph around it actually looks like — because the field has fragmented into enough subfields that even practitioners working in adjacent areas often miss the load-bearing references.

The frame I’ll use is domain-by-domain, grounded in specific papers from ICML, ICLR, NeurIPS, and the few adjacent venues that have absorbed work that doesn’t fit cleanly into a generalist ML conference (CoRL for robotics, CVPR/SIGGRAPH for vision and graphics, Nature for biology). I’ll call out arxiv-only and industry-blog releases explicitly when they’re load-bearing — increasingly, the most important results in this field never go through peer review at all.

The through-line, before I get into specifics: pretraining is no longer the research frontier. The frontier has moved to test-time compute, generative simulators, embodied grounding, and the architectural primitives (state-space models, flow matching, mixture-of-experts) that make the new regime tractable. Almost every section below is a variant of that story.


1. Natural Language Processing

NLP went through three distinct phases in the 2023–2026 window: (a) the alignment phase, where the question was “how do we make pretrained LLMs follow instructions”; (b) the architecture phase, where the question was “what replaces or augments the transformer”; and (c) the reasoning phase, where the question is “how do we use RL on verifiable rewards to elicit chain-of-thought.” We are deep into (c) as of 2026.

Alignment and preference optimization

The defining paper of the alignment phase was Direct Preference Optimization (Rafailov et al., NeurIPS 2023, Outstanding Paper). DPO collapsed the RLHF pipeline — reward model + PPO — into a single closed-form contrastive loss. Within twelve months, every open-weights model release post-Llama-2 was using DPO or one of its descendants (IPO, KTO, SimPO).

The follow-up wave matters too:

Architecture: state-space models and hybrids

The transformer’s quadratic-attention bottleneck has been the open architectural problem since 2017. The breakthrough was Mamba (Gu and Dao, 2023, arxiv) and Mamba-2 (Dao and Gu, ICML 2024), which made selective state-space models (SSMs) competitive with transformers at scale by introducing input-dependent transitions and a hardware-aware parallel scan.

Mamba alone didn’t displace transformers — but hybrid SSM-attention stacks did, in the sense that nearly every long-context-optimized open model in 2025 used some hybrid:

The open question for 2026 is whether SSMs can match transformer reasoning on tasks that require precise in-context retrieval; the RULER benchmark (Hsieh et al., COLM 2024) is the standard test bed.

Reasoning, RL on verifiable rewards, and test-time compute

This is the live frontier. The first systematic statement that test-time compute can substitute for parameter count was “Scaling LLM Test-Time Compute Optimally Can Be More Effective Than Scaling Model Parameters” (Snell et al., 2024, arxiv).

The reasoning paradigm crystallized publicly with OpenAI’s o1, but the open-research story runs through:

Mixture-of-experts at scale

The architectural workhorse behind frontier-quality open models is sparse MoE. The reference papers:

Mechanistic interpretability

The most interesting applied-interpretability development is the rise of sparse autoencoders (SAEs) as a tool for finding monosemantic features in residual streams. Key references:

Agents

The agents literature is messier — many of the most-cited results live in industry blog posts, not conference proceedings. The peer-reviewed anchors:


2. Speech

Speech research bifurcated cleanly: a discriminative track that more or less solved general-purpose ASR, and a generative track that has now collapsed the entire ASR-LLM-TTS pipeline into single audio-native models.

ASR

Self-supervised speech representation

TTS and zero-shot voice cloning

The defining shift was treating speech generation as language modeling over neural codec tokens.

Speech-native LLMs

The convergence of ASR, dialog, and TTS into single end-to-end models is the most important development of 2024–2025.

Neural audio codecs

The infrastructure layer:


3. Video

Video is where diffusion transformers won, and where “video generation” and “world model” are visibly converging into the same research object.

Backbone: diffusion transformers

The architectural pivot point was DiT (Peebles and Xie, ICCV 2023) — replacing U-Nets with transformers in latent diffusion. Every frontier video model from 2024 onward is a DiT or a near-relative.

Video generation

Video understanding

Flow matching for video

Worth flagging separately: most frontier video models in 2025 are using rectified flow (Liu et al., ICLR 2023) and flow matching (Lipman et al., ICLR 2023) rather than DDPM-style diffusion. The training objective is simpler, the sampling needs fewer steps, and the empirical quality is at least as good.


4. Sound (non-speech audio)

Music and general audio are smaller communities than speech, but they have crystallized around a few load-bearing papers.

Audio understanding and event detection

Sound generation

Music

Source separation


5. What you might have missed

A few high-impact areas that don’t fit cleanly into the above buckets but should be on any 2026 reading list.

Robotics and Vision-Language-Action models (VLAs)

The robotics community has its own venue (CoRL) and deserves a separate post, but the load-bearing papers:

Computational biology

3D generation and reconstruction

Diffusion architecture

The architectural primitives have shifted under most people’s noses:


6. Reinforcement Learning and World Models

This is the single most active area of deep learning research in 2026, in my reading. Three previously-separate threads — LLM reasoning RL, generative video models, and embodied AI — are visibly converging into a single research program around learning a world model that is good enough to plan in.

World models: the hottest area

The conceptual seed for all of this is older: World Models (Ha and Schmidhuber, NeurIPS 2018). The 2024–2025 wave is what happens when the seed paper finally has compute and data behind it.

RL for reasoning (LLMs)

Embodied / robotics RL

Search, self-play, and program synthesis

Open-ended learning

A smaller community but conceptually important:

The unifying thesis

Read across the 2024–2026 RL literature and the same picture appears: RL is back, but only because we finally have generative simulators worth planning in. DreamerV3 made the case that a learned world model is enough for general control. Genie made the case that internet video is enough to train one. R1 made the case that RL on verifiable rewards in a “world model” of language (a base LLM) elicits reasoning. The next two years will be about whether these threads merge into a single architecture, or whether language reasoning and physical reasoning end up requiring genuinely different machinery.


7. Benchmarks and current SOTA scores

A reference table for the load-bearing benchmarks per subdomain, with the best public results as of early 2026. Scores are approximate, drawn from technical reports, leaderboards, and arxiv evaluation tables; they move month-to-month and rely on self-reported numbers in many cases. Treat this as orientation, not a leaderboard. Where a benchmark is saturated I say so — running it on a frontier model is no longer informative.

The short version of who leads where, before the tables:

NLP — reasoning and knowledge

BenchmarkWhat it measuresCurrent SOTANotes
MMLU-Pro57-subject reasoning + knowledge~88–92% (frontier reasoning models)Original MMLU saturated above 92%; MMLU-Pro is the harder successor.
GPQA DiamondGraduate-level science (PhD-blocking questions)~85–90% (o3-class, Claude Opus 4.x, Gemini 2.5 Pro)Human PhD experts ~65%; reasoning models broke past expert level in late 2024.
AIME 2024 / 2025Olympiad-level math95%+ on AIME 2024 for o3-class; ~80–90% on AIME 2025DeepSeek-R1 reported ~79.8% on AIME 2024; o3 reported 96.7%. AIME 2025 still discriminates.
MATH-500Competition / high-school mathSaturated above 96%No longer informative for frontier models.
FrontierMathResearch-level math (Tao et al. designed)~25–32% for o3-classDesigned to stay unsaturated for years.
Humanity’s Last Exam (HLE)Cross-domain expert-blocking~25–35% (top reasoning models)Best new “stays hard” benchmark; most non-reasoning models still under 10%.
ARC-AGI v1Few-shot abstract visual reasoning~87% (o3 high-compute setting)High-compute runs cost $20+ per task; v1 is effectively retired as a frontier target.
ARC-AGI v2Harder ARC successor<20% across the boardThe current frontier puzzle.

NLP — coding and agents

BenchmarkWhat it measuresCurrent SOTANotes
SWE-bench VerifiedReal GitHub issue resolution~70–82% (Claude Sonnet/Opus 4.x, GPT-5, Gemini 2.5 Pro)Was ~12% (Claude 3 Opus) in early 2024. Largest single-benchmark jump in recent history.
SWE-bench MultimodalBug fixes with visual context~50–60%Newer, less saturated.
LiveCodeBenchContest-style coding (time-stratified)~85–90% (top reasoning models)Time-stratification mitigates contamination.
HumanEval / MBPPFunction-completionSaturated (>95%)Useless for frontier comparison.
τ-bench (tau-bench)Multi-turn tool-use in retail/airline domains~70–80%Better proxy for “real agent” work than single-turn benchmarks.
BFCL v3 (Berkeley Function Calling)Function calling correctness~85–90%Standard tool-use benchmark.
WebArena / VisualWebArenaWeb-browsing agents~40–55%Stays hard; the agent frontier.

NLP — long context

BenchmarkWhat it measuresCurrent SOTANotes
RULER (128k)Long-context retrieval and reasoning~88–92% (Gemini 1.5/2.5 Pro, Claude Sonnet 4.x)The credible long-context test bed; needle-in-haystack is too easy.
RULER (1M)Frontier-context regimeGemini family ~80%+; others drop sharplyFew credible 1M-token systems.
LongBench v2Realistic long-document tasks~50–60%Stays hard.
∞BenchMulti-task long context~60–70% topOlder but still cited.

Speech — ASR

BenchmarkWhat it measuresCurrent SOTANotes
LibriSpeech test-cleanEnglish read speech WER~1.4–1.7% WER (Parakeet, Canary, Whisper-v3)Saturated.
LibriSpeech test-otherNoisier English~2.8–3.5% WERNear-saturated.
Common Voice (multilingual)100+ language WERWhisper-v3 baseline; OWSM/Canary close on covered languagesVery high variance across languages.
FLEURS102-language ASR~10–15% avg WER (top models)The standard multilingual coverage benchmark.
AMI / Earnings-22Meeting / accented speech12–18% WERWhere general ASR still struggles.

Speech — TTS and voice

BenchmarkWhat it measuresCurrent SOTANotes
LibriTTS WER (objective)Synthesis intelligibility<2%Saturated for non-streaming.
SECS / SIM-OSpeaker similarity (zero-shot voice cloning)~0.65–0.75 (F5-TTS, NaturalSpeech 3, Voicebox)Some commercial systems claim higher.
DNSMOS / UTMOSPredicted MOS~4.0–4.4Most frontier systems indistinguishable from ground truth on these proxies.
Moshi latency (full-duplex)End-to-end response time~200msProduction-quality target the open community is chasing.

Video — generation

BenchmarkWhat it measuresCurrent SOTANotes
VBench16-dimension video quality (subject/background consistency, motion smoothness, etc.)Sora, Veo 3, Movie Gen, Kling 2 lead closed; Wan 2.1, HunyuanVideo lead openThe de facto standard.
VBench-Long / VBench++Long video and I2VSame leaders; gap narrows on I2VAdds image-conditioned and long-form.
Movie Gen BenchInternal Meta eval (released)Movie Gen self-reported leaderReproducible recipe; useful sanity check.
EvalCrafterMulti-dimension comparisonClosed > open by 5–15%Aggregate score is fragile; use dimension-by-dimension.

Video — understanding

BenchmarkWhat it measuresCurrent SOTANotes
VideoMMELong/short video QA~75–82% (Qwen2.5-VL, InternVL2.5, Gemini)Best general video-understanding leaderboard.
MVBench20-task video understanding~70–78%Approaching saturation.
EgoSchemaLong egocentric video~65–75%Stays hard; designed to require true temporal reasoning.
NExT-QA / Perception TestCausal/temporal reasoning over video~75–85%The classic comprehension benchmarks.

Sound (non-speech audio)

BenchmarkWhat it measuresCurrent SOTANotes
AudioSet mAP527-class audio tagging~50–52% (BEATs and successors)The reference tagging benchmark.
AudioCaps FADText-to-audio quality (lower is better)~1.3–1.8 (AudioLDM 2, Stable Audio 2)Frechet Audio Distance.
MusicCaps FAD-VGGText-to-music quality~3.5–4.5 (MusicGen, Stable Audio 2)Suno/Udio do not publish on this.
MUSDB18 SDRMusic source separation~10–11 dB (Band-Split RNN, HT Demucs)Higher is better; near practical ceiling.
CLAP zero-shot AudioSetText-audio alignment~50%+ mAPThe audio analogue of CLIP zero-shot.

Robotics and VLAs

BenchmarkWhat it measuresCurrent SOTANotes
LIBERO4-suite manipulation (spatial, object, goal, long)~85–95% success (π0, OpenVLA, RDT)The standard simulated VLA benchmark.
SimplerEnvSim-to-real-aligned manipulation evalπ0, π0.5, RT-2-X, Octo leadDesigned so sim numbers correlate with real-robot performance.
CALVINLong-horizon language-conditioned manipulation~80%+ (top VLAs)Saturating.
Open X-Embodiment evalsCross-embodiment generalizationRT-X, OpenVLA, π0 the reference pointsDataset-paper benchmark.
HumanoidBench / Isaac humanoid suitesWhole-body humanoid controlRL + sim-to-real (Berkeley, NVIDIA, Tesla recipes)No single agreed metric yet.

Biology

BenchmarkWhat it measuresCurrent SOTANotes
CASP15 / CASP16 GDT-TSProtein structure predictionAF2/AF3 ~85–90 GDT-TSThe classical structural-biology benchmark.
AF3 PoseBustersProtein–ligand docking~70–80% success (AlphaFold 3)Major step over classical docking.
RFdiffusion success rateDe novo binder design~10–30% wet-lab hit rateActive area; numbers vary by target class.
ProteinGymVariant effect prediction (ESM-class models)ESM-2/ESM3 lead openStandard zero-shot benchmark for PLMs.
Evo / nucleotide LMsDNA modeling at long contextEvo (StripedHyena) the open referenceMillion-token DNA context.

3D reconstruction and generation

BenchmarkWhat it measuresCurrent SOTANotes
Mip-NeRF 360 / DTU PSNRNovel-view synthesis3D Gaussian Splatting baseline; recent variants push +1–2 dBSaturated as a research target.
CO3D / RealEstate10kFeed-forward 3D reconstruction without posesDUSt3R, MASt3R, VGGTThe pose-free regime.
Tanks and Temples / ETH3DMulti-view stereo3DGS-based and recent feed-forward methodsLong-standing reference.
GSO (Google Scanned Objects)Image-to-3D generationTrellis, InstantMesh, recent DiT-3D variantsNo single agreed metric — mixes CLIP score, LPIPS, F-score.

RL and world models

BenchmarkWhat it measuresCurrent SOTANotes
Atari 100kSample-efficient RL (human-normalized score)DreamerV3 ~120%, IRIS ~100%, EfficientZero V2 ~190%World-model methods now beat humans at 100k frames.
DMC (DeepMind Control)Continuous controlDreamerV3, TD-MPC2 dominantSaturated on many tasks.
CrafterOpen-ended survival (procedural)DreamerV3 superhumanReference for general agents.
Minecraft Diamond (from scratch)Long-horizon explorationDreamerV3 first, no prior knowledgeHeadline claim of the Nature paper.
NetHack Learning EnvironmentHard explorationStays hard; no reliable solverThe unsolved bar.
ProcgenGeneralization across procedural levelsStays hardLess tracked in 2025–2026 but still relevant for generalization claims.
Genie 3 interactive evalMinute-long, prompt-controllable simulationGenie 3 (DeepMind)No public quantitative leaderboard yet — assessed qualitatively.

A few honest caveats on this section:

  1. Many of these numbers are self-reported in technical reports and have not been independently reproduced. Where a frontier closed model claims +2 points over the previous SOTA, treat that as a hint, not a settled fact.
  2. Some benchmarks (HumanEval, MMLU, MATH-500) are contaminated by training data overlap. The credible benchmarks now bake in time-stratification (LiveCodeBench), private test sets (FrontierMath, HLE), or held-out construction (ARC-AGI v2).
  3. The most important capabilities don’t have benchmarks yet. There is no good public benchmark for “can this agent maintain coherent context over a multi-day software-engineering task” or “does this world model permit zero-shot transfer to a real robot.” The frontier is moving faster than the eval community can keep up.

What I’d read first

If you have one weekend and want to recompile your mental model of the field, I’d read in this order:

  1. DreamerV3 — for the structural argument that model-based RL works.
  2. Genie (ICML 2024 Best Paper) — for “video model = world model.”
  3. Mamba-2 — for the post-transformer architectural option.
  4. DeepSeek-R1 — for the test-time-compute regime.
  5. DPO — for what alignment looked like before reasoning RL ate it.
  6. DiT — for the architectural primitive behind every frontier image and video model.
  7. AlphaFold 3 — for what mature scientific deep learning looks like.

The bibliography is intentionally biased toward papers with reproducible methods over papers with marketing. The hardest part of staying current in this field in 2026 isn’t finding the frontier — it’s distinguishing the frontier from the press release. Reading the actual papers, with their actual ablation tables, is still the only reliable filter.


Share this post on:

Previous Post
The Edge of Mathematics Research: A Citation-Grounded Survey of Open Problems and Recent Breakthroughs (2019–2026)
Next Post
The Perception–Planning Gap: What's Actually Hard About Visual AI in 2026