Datamase

Article Overviews

Every "Article Overview" post from Datamase — short summaries of papers, model releases, and technical reports worth knowing about. Search to find a specific one.

65 article overviews
Kimi K3: Open Frontier Intelligence
Kimi K3: Open Frontier Intelligence
Summary of “Kimi K3: Open Frontier Intelligence”, by researchers at Kimi
JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence
JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence
Summary of “JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence”, by researchers at Jingdong Group
GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation
GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation
Summary of “GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation”, by GigaAI, Tsinghua University
Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs
Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs
Summary of “Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs”, by researchers from Yale University and Google Research
Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent
Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent
Summary of “Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent”, by the Agents-A1 Team, Shanghai Artificial Intelligence Laboratory
ResearchStudio-Idea: An Evidence-Grounded Research-Ideation Skill Suite from ML Conference Outcomes
ResearchStudio-Idea: An Evidence-Grounded Research-Ideation Skill Suite from ML Conference Outcomes
Summary of “ResearchStudio-Idea: An Evidence-Grounded Research-Ideation Skill Suite from ML Conference Outcomes”, by Researchers from Microsoft Research and affiliated universities
Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots
Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots
Summary of “Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots”, by Researchers from Microsoft Research and affiliated universities
AI-Trader: Benchmarking Autonomous Agents in Real-Time Financial Markets
AI-Trader: Benchmarking Autonomous Agents in Real-Time Financial Markets
Summary of “AI-Trader: Benchmarking Autonomous Agents in Real-Time Financial Markets”, by Researchers from the University of Hong Kong
Harness Engineering for Self-Improvement
Harness Engineering for Self-Improvement
Summary of “Harness Engineering for Self-Improvement”, by Lilian Weng
Autoresearch: The feedback loop behind self-improving agents
Autoresearch: The feedback loop behind self-improving agents
Summary of “Autoresearch: The feedback loop behind self-improving agents”, an interview of Roland Gavrilescu by Latent Space
The Twilight of the Chatbots
The Twilight of the Chatbots
Summary of “The Twilight of the Chatbots” by Ethan Mollick
The AI All-You-Can-Eat Buffet is Ending
The AI All-You-Can-Eat Buffet is Ending
Summary of “The AI All-You-Can-Eat Buffet is Ending” from an interview of Gary Marcus by Steve Eisman
A New Look at AI's Impact on Jobs
A New Look at AI's Impact on Jobs
Summary of “A New Look at AI's Impact on Jobs” by researchers from Ramp and Revelio Labs
Accurate Decoding of Natural Sentences from Non-Invasive Brain Recordings
Accurate Decoding of Natural Sentences from Non-Invasive Brain Recordings
Summary of “Accurate Decoding of Natural Sentences from Non-Invasive Brain Recordings” by researchers from Meta and affiliated universities
RoboReward: General-Purpose Vision-Language Reward Models for Robotics
RoboReward: General-Purpose Vision-Language Reward Models for Robotics
Summary of “RoboReward: General-Purpose Vision-Language Reward Models for Robotics” by researchers from Stanford and UCBerkeley
Residual Context Diffusion Language Models
Residual Context Diffusion Language Models
Summary of “Residual Context Diffusion Language Models” by researchers from UCBerkeley and Apple
Qwen-AgentWorld: Language World Models for General Agents
Qwen-AgentWorld: Language World Models for General Agents
Summary of “Qwen-AgentWorld: Language World Models for General Agents” by the Qwen Team
PorTAL: Portal Task Adapters for LLMs
PorTAL: Portal Task Adapters for LLMs
Summary of “PorTAL: Portal Task Adapters for LLMs” by Ben Geist, Ramp Labs
Introducing LongCat-2.0
Introducing LongCat-2.0
Summary of “Introducing LongCat-2.0” by Researchers from Meituan
Can Scale Save Us From Plasticity Loss in Large Language Models?
Can Scale Save Us From Plasticity Loss in Large Language Models?
Summary of “Can Scale Save Us From Plasticity Loss in Large Language Models?” by Researchers from Zyphra
LFM2.5-230M: Built to Run Anywhere
LFM2.5-230M: Built to Run Anywhere
Summary of “LFM2.5-230M: Built to Run Anywhere” by Researchers from Liquid AI
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
Summary of “DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence” by Researchers from DeepSeek-AI
Arctic RL: A Unified, Open Source RL Backend for Enterprise Post-Training
Arctic RL: A Unified, Open Source RL Backend for Enterprise Post-Training
Summary of “Arctic RL: A Unified, Open Source RL Backend for Enterprise Post-Training” by Researchers from Snowflake AI Research
Return on Tokens (ROT)
Return on Tokens (ROT)
Summary of "Return on Tokens (ROT)" by Packy McCormick and Markie Wagner
First steps Towards Automated AI Research
First steps Towards Automated AI Research
Summary of "First steps Towards Automated AI Research" by Researchers from Recursive AI
Mistral OCR 4 Technical Report
Mistral OCR 4 Technical Report
Summary of "Mistral OCR 4 Technical Report" by Researchers from Mistral AI
Laguna M.1/XS.2 Technical Report
Laguna M.1/XS.2 Technical Report
Summary of "Laguna M.1/XS.2 Technical Report" by Researchers from Poolside AI
Krea 2 Technical Report
Krea 2 Technical Report
Summary of “Krea 2 Technical Report" by Researchers from Krea AI
Reinforcement learning towards broadly and persistently beneficial models
Reinforcement learning towards broadly and persistently beneficial models
Summary of “Reinforcement learning towards broadly and persistently beneficial models" by Researchers from OpenAI
LifeSciBench: Evaluating Language Models on Realistic, Expert-Level Tasks in the Life Sciences
LifeSciBench: Evaluating Language Models on Realistic, Expert-Level Tasks in the Life Sciences
Summary of “LifeSciBench: Evaluating Language Models on Realistic, Expert-Level Tasks in the Life Sciences" by Researchers from OpenAI and Tacit Labs
AtomMem: Building Simple and Effective Memory System for LLM Agents via Atomic Facts
AtomMem: Building Simple and Effective Memory System for LLM Agents via Atomic Facts
Summary of “AtomMem: Building Simple and Effective Memory System for LLM Agents via Atomic Facts" by Researchers from USTC and Anhui University
How Long Until AI Doesn't Need Humans?
How Long Until AI Doesn't Need Humans?
Summary of “How Long Until AI Doesn't Need Humans?" by Ajeya Cotra and Timothy B. Lee
Is AI ruining our skills? Early results are in — and they’re not good
Is AI ruining our skills? Early results are in — and they’re not good
Summary of “Is AI ruining our skills? Early results are in — and they’re not good" by Mariana Lenharo at Nature
AI Systems Out-Persuade Expert Humans
AI Systems Out-Persuade Expert Humans
Summary of “AI Systems Out-Persuade Expert Humans" by Researchers at Oxford, UKAI, Stanford, and LSE
GDM AI Control Roadmap
GDM AI Control Roadmap
Summary of “GDM AI Control Roadmap" by Researchers at Google DeepMind
VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models
VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models
Summary of “VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models" by Researchers at Sina Weibo Inc.
The Untrainable
The Untrainable
Summary of “The Untrainable" by Sarah Guo
Olmo-Eval: An evaluation workbench for the model development loop
Olmo-Eval: An evaluation workbench for the model development loop
Summary of “Olmo-Eval: An evaluation workbench for the model development loop" by Researchers from Ai2
Agentic coding and persistent returns to expertise
Agentic coding and persistent returns to expertise
Summary of “Agentic coding and persistent returns to expertise" by Researchers from Anthropic
Dialogues with AI Reduce Beliefs in Misinformation but Build No Lasting Discernment Skills
Dialogues with AI Reduce Beliefs in Misinformation but Build No Lasting Discernment Skills
Summary of “Dialogues with AI Reduce Beliefs in Misinformation but Build No Lasting Discernment Skills" by Researchers at MIT
Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation
Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation
Summary of “Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Condition" by Researchers from the Qwen Team
MolmoMotion: Forecasting Point Trajectories in 3D with Languge Instruction
MolmoMotion: Forecasting Point Trajectories in 3D with Languge Instruction
Summary of “MolmoMotion: Forecasting Point Trajectories in 3D with Language Instruction" by Researchers from the Allen Institute for AI, University of Washington, and UNC-Chapel Hill
Loop Engineering
Loop Engineering
Summary of “Loop Engineering" by Addy Osmani
IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse
IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse
Summary of “IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse”, by Researchers from Tsinghua University and Z.ai
AI demands more engineering discipline. Not less
AI demands more engineering discipline. Not less
Summary of “AI demands more engineering discipline. Not less”, by Charity Majors
From AGI to ASI
From AGI to ASI
Summary of “AGI to ASI”, by researchers from Google DeepMind, the University of Waterloo, Australian National University, and University College London
Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding
Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding
Summary of “Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding”, by researchers from Allen Institute for AI and University of Washington
ParseBench: A Document Parsing Benchmark for AI Agents
ParseBench: A Document Parsing Benchmark for AI Agents
Summary of “ParseBench: A Document Parsing Benchmark for AI Agents”, by researchers from LlamaIndex
Large Language Models Hack Rewards, and Society
Large Language Models Hack Rewards, and Society
Summary of “Large Language Models Hack Rewards, and Society”, by researchers from King’s College London, Fudan University, and the Alan Turing Institute
AI Dark Output: The Visible Cost of Invisible Output
AI Dark Output: The Visible Cost of Invisible Output
Summary of “AI Dark Output: The Visible Cost of Invisible Output”, by Malcolm Spittler and Dylan Patel
Democratic Governance of Frontier AI: A Blueprint for a Federal Framework
Democratic Governance of Frontier AI: A Blueprint for a Federal Framework
Summary of "Democratic Governance of Frontier AI: A Blueprint for a Federal Framework" by researchers at OpenAI
CubePart: An Open-Vocabulary Part-Controllable 3D Generator
CubePart: An Open-Vocabulary Part-Controllable 3D Generator
Summary of “CubePart: An Open-Vocabulary Part-Controllable 3D Generator" by researchers at Roblox, Stanford, and Carnegie Mellon University
MiMo-V2.5-Pro-UltraSpeed: Pushing 1T-Parameter Model Generation Speed to 1000 TPS
MiMo-V2.5-Pro-UltraSpeed: Pushing 1T-Parameter Model Generation Speed to 1000 TPS
Summary of “MiMo-V2.5-Pro-UltraSpeed: Pushing 1T-Parameter Model Generation Speed to 1000 TPS”, by researchers from Xiaomi
Introducing North Mini Code: Cohere’s First Model For Developers
Introducing North Mini Code: Cohere’s First Model For Developers
Summary of “Introducing North Mini Code: Cohere’s First Model For Developers”, by the Cohere Code Agents Team
Releasing the MisoTTS
Releasing the MisoTTS
Summary of “Releasing the MisoTTS”, by Aoden Teo & Cassidy Dalva, Miso Labs
Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories
Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories
Summary of “Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories", by researchers from Google Research and Cornell University
Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
Summary of “Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning" by researchers at NVIDIA
When AI Builds Itself
When AI Builds Itself
Summary of “When AI Builds Itself”, by researchers from Anthropic
GPIC: A Giant Permissive Image Corpus for Visual Generation
GPIC: A Giant Permissive Image Corpus for Visual Generation
Summary of “GPIC: A Giant Permissive Image Corpus for Visual Generation”, by researchers from Stanford, Radical Numerics, University of Michigan, and Salesforce Research
Adapting the Interface, Not the Model: Runtime Harness Adaptation for Deterministic LLM Agents
Adapting the Interface, Not the Model: Runtime Harness Adaptation for Deterministic LLM Agents
Summary of “Adapting the Interface, Not the Model: Runtime Harness Adaptation for Deterministic LLM Agents”, by Tianshi Xu, Huifeng Wen, Meng Li
Automated alignment is harder than you think
Automated alignment is harder than you think
Summary of “Automated alignment is harder than you think”, by Aleksandr Bowkis, Marie Davidsen Buhl, Jacob Pfau, Geoffrey Irving
Cosmos 3: Omnimodal World Models for Physical AI
Cosmos 3: Omnimodal World Models for Physical AI
Summary of “Cosmos 3: Omnimodal World Models for Physical AI”, by NVIDIA AI Research
EMERGENCE WORLD: A Laboratory for Evaluating Long-horizon Agent Autonomy
EMERGENCE WORLD: A Laboratory for Evaluating Long-horizon Agent Autonomy
Summary of “EMERGENCE WORLD: A Laboratory for Evaluating Long-horizon Agent Autonomy”, by researchers at Emergence AI
Measuring the AI Economy
Measuring the AI Economy
Summary of “Measuring the AI economy”, by Anton Korinek (PIIE) and Patrick McKelvey (Bank of Canada)
Building a Hill Climbing Machine
Building a Hill Climbing Machine
Summary of "MAI-Thinking-1: Building a Hill-Climbing Machine", by the Microsoft AI Team.
No article overviews match your search.
All posts on Substack → New posts on AI, video, and the tech world.