Tag: foundation-models
All the articles with the tag "foundation-models".
-
The State of Robotics in 2026: A Citation-Grounded Survey of Research and Companies
A citation-grounded survey of robotics research as it stands in May 2026 — VLA foundation models (RT-2, OpenVLA, π0, π0.5, Helix, Gemini Robotics), imitation-learning architectures (Diffusion Policy, ACT/ALOHA, RDT-1B), cross-embodiment data (Open X-Embodiment, DROID), whole-body humanoid control (HOVER, ASAP, OmniH2O), reward design with LLMs (Eureka, DrEureka), world models for robotics (Cosmos, Genie, 1X), simulators (Isaac Lab, Genesis, RoboCasa), and the company landscape (Figure, Tesla, Boston Dynamics, Apptronik, Agility, 1X, Unitree, Physical Intelligence, Skild AI, XPENG, UBTECH, Fourier). With venue-by-venue references from ICLR, ICML, NeurIPS, CoRL, RSS, ICRA, and Science Robotics.
-
The Edge of Artificial Intelligence Research: A Citation-Grounded Survey from ICML, ICLR, NeurIPS (2023–2026)
A domain-by-domain survey of the artificial intelligence research frontier, grounded in specific papers from ICML, ICLR, NeurIPS, and adjacent venues (CoRL, CVPR, Nature). Covers natural language processing, speech, video, sound, robotics/VLA, biology, 3D generation, diffusion architecture, and the renaissance of reinforcement learning and world models. The through-line: pretraining is no longer the frontier — test-time compute, generative simulators, and embodied grounding are.
-
The Perception–Planning Gap: What's Actually Hard About Visual AI in 2026
A technical survey of where visual perception and planning research actually stands in 2026. Pixel-level perception is largely solved at the representation layer, but perception-for-action — geometry, physics, dexterity, long-horizon control — is not. Reading the recent literature on JEPA, DreamerV3, Sora-as-world-model, RT-2 / OpenVLA / π0, Helix, Gemini Robotics, DUSt3R / MASt3R / VGGT, and the world-model evaluation papers (WorldModelBench, Physion, IntPhys 2), the through-line is the same: we have strong representations and architectural ideas, but the data, evaluation, and physical-grounding infrastructure to validate them is what's missing.