Today’s frontier reasoning models have settled into a familiar and predictable pattern of improvement.

To crack the next coding benchmark or computer use task, we create an environment for an agent to operate in, define a set of verifiable or non-verifiable rewards and let the agent hill-climb on these metrics. While this has yielded tremendous gains in a subset of white collar work, these agents have severe limitations. They are incapable of discovering new knowledge since they can only answer questions in the distribution of what humans have previously specified. This is a hindrance when trying to unearth new insights or find an intelligent starting point on the road to new breakthroughs.

Solving long horizon tasks and building systems capable of lifelong learning requires something fundamentally different. To solve a problem no human has solved, models must be able to:

  • pose their own questions by learning their own shortcomings and how to overcome them with new lines of inquiry – not merely understanding how to respond to what humans may desire
  • reward progress for interesting and novel results, not just for accuracy and getting full marks on a benchmark
  • collaborate with teams of agents to harness their insights and improve as a collective when stuck, rather than as an individual (inspired by processes from biological and cultural evolution)

Achieving these goals requires significant investment in research that has not been systematically prioritized by mainstream labs. This work spans techniques and infrastructure for unsupervised environment design for general-sum games, multi-agent evolution, deployment time continual improvement, and much more.

It also necessitates a phase shift in what problems and domains we think about solving. Of all the scientific verticals where this agent could operate from biotechnology to nuclear fusion, it makes most sense to start automating AI research for two reasons. Firstly, since AI can code and AI is code, it is the most immediately approachable problem. Secondly, there are significant compounding benefits to building a system that can solve tough problems in AI from data efficiency to credit assignment, so that if a scientific problem requires a new algorithm or system design, the agent will innovate at lightning speed to get there.

A mission this bold commands a special team with a history of diverse, groundbreaking research spanning all of the above topics and more, assembled across oceans from San Francisco to London, to concentrate time and effort on this singular endeavor. A superhuman team, if you will.

We believe Recursive Super Intelligence is that rare assembly of talent, ready to kickstart a new paradigm and make AI models synonymous with originality and knowledge discovery. The founders share some of the most impressive, star-studded CVs of any Amplify portfolio company, and not just because there’s eight of them. They hail from almost all major AI labs and have written some of the most iconic papers of the last decade. And on a personal note, we’ve been fortunate to know several on this team for many years.

Richard Socher, the CEO of RSI, was one of the first visionaries to see the importance of neural networks in solving natural language processing. His work on GloVE embedding models is now one of the most highly cited papers in the history of NLP. His subsequent startup MetaMind, which enabled enterprise access to computer vision and NLP solutions, was acquired by Salesforce, where he not only spent several years as Chief Scientist but became widely credited for inventing prompt engineeringThe Natural Language Decathlon: Multitask Learning as Question AnsweringBryan McCann, Nitish Shirish Keskar, Caiming Xiong, Richard SocherDeep learning has improved performance on many natural language processing (NLP) tasks individually. However, general NLP models cannot emerge within a paradigm that focuses on the particularities of a single metric, dataset, and task. We introduce the Natural Language Decathlon (decaNLP), a challenge that spans ten tasks: question answering, machine translation, summarization, natural language inference, sentiment analysis, semantic role labeling, zero-shot relation extraction, goal-oriented dialogue, semantic parsing, and commonsense pronoun resolution. We cast all tasks as question answering over a context. Furthermore, we present a new Multitask Question Answering Network (MQAN) jointly learns all tasks in decaNLP without any task-specific modules or parameters in the multitask setting. MQAN shows improvements in transfer learning for machine translation and named entity recognition, domain adaptation for sentiment analysis and natural language inference, and zero-shot capabilities for text classification. We demonstrate that the MQAN's multi-pointer-generator decoder is key to this success and performance further improves with an anti-curriculum training strategy. Though designed for decaNLP, MQAN also achieves state of the art results on the WikiSQL semantic parsing task in the single-task setting. We also release code for procuring and processing data, training and evaluating models, and reproducing all experiments for decaNLP.arXiv:1806.08730v1View paper itself.

Josh Tobin, now CTO of RSI, is no stranger to this problem as he was previously CEO of Amplify portfolio company Gantry. Having started his career as one of the first 25 people at OpenAI, he left to build a continuous machine learning improvement platform that helped machine learning engineers understand how their deployed models are performing, ways to improve it from data curation to experimentation, and then actually operationalize those improvements. After a second stint at OpenAI where he led Deep Research, we are excited to see him back on the frontlines of startups and self improving systems.

Jeff Clune’s litany of work in evolutionary algorithms and open-endedness needs no introduction. Before LLMs even existed, he produced seminal research such as POETPaired Open-Ended Trailblazer (POET): Endlessly Generating Increasingly Complex and Diverse Learning Environments and Their SolutionsRui Wang, Joel Lehman, Jeff Clune, Kenneth O. StanleyWhile the history of machine learning so far largely encompasses a series of problems posed by researchers and algorithms that learn their solutions, an important question is whether the problems themselves can be generated by the algorithm at the same time as they are being solved. Such a process would in effect build its own diverse and expanding curricula, and the solutions to problems at various stages would become stepping stones towards solving even more challenging problems later in the process. The Paired Open-Ended Trailblazer (POET) algorithm introduced in this paper does just that: it pairs the generation of environmental challenges and the optimization of agents to solve those challenges. It simultaneously explores many different paths through the space of possible problems and solutions and, critically, allows these stepping-stone solutions to transfer between problems if better, catalyzing innovation. The term open-ended signifies the intriguing potential for algorithms like POET to continue to create novel and increasingly complex capabilities without bound. Our results show that POET produces a diverse range of sophisticated behaviors that solve a wide range of environmental challenges, many of which cannot be solved by direct optimization alone, or even through a direct-path curriculum-building control algorithm introduced to highlight the critical role of open-endedness in solving ambitious challenges. The ability to transfer solutions from one environment to another proves essential to unlocking the full potential of the system as a whole, demonstrating the unpredictable nature of fortuitous stepping stones. We hope that POET will inspire a new push towards open-ended discovery across many domains, where algorithms like POET can blaze a trail through their interesting possible manifestations and solutions.arXiv:1901.01753v3View paper, which highlighted why populations of agents making mutated improvements outperform maximizing an objective. Since then, his team - of whom Jenny Zhang, Shenghran Hu and Cong Lu have excitingly joined RSI - have pioneered work in pairing evolutionary algorithms with LLMs from Darwin Godel MachinesDarwin Godel Machine: Open-Ended Evolution of Self-Improving AgentsJenny Zhang, Shengran Hu, Cong Lu, Robert Lange, Jeff CluneToday's AI systems have human-designed, fixed architectures and cannot autonomously and continuously improve themselves. The advance of AI could itself be automated. If done safely, that would accelerate AI development and allow us to reap its benefits much sooner. Meta-learning can automate the discovery of novel algorithms, but is limited by first-order improvements and the human design of a suitable search space. The Gödel machine proposed a theoretical alternative: a self-improving AI that repeatedly modifies itself in a provably beneficial manner. Unfortunately, proving that most changes are net beneficial is impossible in practice. We introduce the Darwin Gödel Machine (DGM), a self-improving system that iteratively modifies its own code (thereby also improving its ability to modify its own codebase) and empirically validates each change using coding benchmarks. Inspired by Darwinian evolution and open-endedness research, the DGM maintains an archive of generated coding agents. It grows the archive by sampling an agent from it and using a foundation model to create a new, interesting, version of the sampled agent. This open-ended exploration forms a growing tree of diverse, high-quality agents and allows the parallel exploration of many different paths through the search space. Empirically, the DGM automatically improves its coding capabilities (e.g., better code editing tools, long-context window management, peer-review mechanisms), increasing performance on SWE-bench from 20.0% to 50.0%, and on Polyglot from 14.2% to 30.7%. Furthermore, the DGM significantly outperforms baselines without self-improvement or open-ended exploration. All experiments were done with safety precautions (e.g., sandboxing, human oversight). The DGM is a significant step toward self-improving AI, capable of gathering its own stepping stones along paths that unfold into endless innovation.arXiv:2505.22954v3View paper to HyperAgentsHyperagentsJenny Zhang, Bingchen Zhao, Wannan Yang, Jakob Foerster, Jeff Clune, Minqi Jiang + 2 moreSelf-improving AI systems aim to reduce reliance on human engineering by learning to improve their own learning and problem-solving processes. Existing approaches to self-improvement rely on fixed, handcrafted meta-level mechanisms, fundamentally limiting how fast such systems can improve. The Darwin Gödel Machine (DGM) demonstrates open-ended self-improvement in coding by repeatedly generating and evaluating self-modified variants. Because both evaluation and self-modification are coding tasks, gains in coding ability can translate into gains in self-improvement ability. However, this alignment does not generally hold beyond coding domains. We introduce \textbf{hyperagents}, self-referential agents that integrate a task agent (which solves the target task) and a meta agent (which modifies itself and the task agent) into a single editable program. Crucially, the meta-level modification procedure is itself editable, enabling metacognitive self-modification, improving not only the task-solving behavior, but also the mechanism that generates future improvements. We instantiate this framework by extending DGM to create DGM-Hyperagents (DGM-H), eliminating the assumption of domain-specific alignment between task performance and self-modification skill to potentially support self-accelerating progress on any computable task. Across diverse domains, the DGM-H improves performance over time and outperforms baselines without self-improvement or open-ended exploration, as well as prior self-improving systems. Furthermore, the DGM-H improves the process by which it generates new agents (e.g., persistent memory, performance tracking), and these meta-level improvements transfer across domains and accumulate across runs. DGM-Hyperagents offer a glimpse of open-ended AI systems that do not merely search for better solutions, but continually improve their search for how to improve.arXiv:2603.19461v1View paper (one of the first public demonstrations of an AI improving its own code), and inspired similar research across the ecosystem such as DeepMind’s AlphaEvolve. If the future of AI is evolutionary, I can’t think of any better than Jeff to guide us there.

Tim Rocktaeschel's reputation similarly precedes him, having authored fundamental papers in model collaboration like Rainbow Teaming and Debate with LLMs, generative world models with his mind-boggling work on Genie, and even the invention of RAG itself. In particular, his work on autocurricula such as PLR and ACCELEvolving Curricula with Regret-Based Environment DesignJack Parker-Holder, Minqi Jiang, Michael Dennis, Mikayel Samvelyan, Jakob Foerster, Edward Grefenstette + 1 moreIt remains a significant challenge to train generally capable agents with reinforcement learning (RL). A promising avenue for improving the robustness of RL agents is through the use of curricula. One such class of methods frames environment design as a game between a student and a teacher, using regret-based objectives to produce environment instantiations (or levels) at the frontier of the student agent's capabilities. These methods benefit from their generality, with theoretical guarantees at equilibrium, yet they often struggle to find effective levels in challenging design spaces. By contrast, evolutionary approaches seek to incrementally alter environment complexity, resulting in potentially open-ended learning, but often rely on domain-specific heuristics and vast amounts of computational resources. In this paper we propose to harness the power of evolution in a principled, regret-based curriculum. Our approach, which we call Adversarially Compounding Complexity by Editing Levels (ACCEL), seeks to constantly produce levels at the frontier of an agent's capabilities, resulting in curricula that start simple but become increasingly complex. ACCEL maintains the theoretical benefits of prior regret-based methods, while providing significant empirical gains in a diverse set of environments. An interactive version of the paper is available at accelagent.github.io.arXiv:2203.01302v3View paper laid the groundwork for agents that ask their own questions during training in a previous era of RL. His students now lead research teams globally, and we’re ecstatic former members of his UCL lab like Dominik Schmidt have joined the founding team.

Yuandong Tian spent a decade at Meta after completing his PhD at CMU, where he improved the training and inference of LLMs with work on StreamingLMEfficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, Mike LewisDeploying Large Language Models (LLMs) in streaming applications such as multi-round dialogue, where long interactions are expected, is urgently needed but poses two major challenges. Firstly, during the decoding stage, caching previous tokens' Key and Value states (KV) consumes extensive memory. Secondly, popular LLMs cannot generalize to longer texts than the training sequence length. Window attention, where only the most recent KVs are cached, is a natural approach -- but we show that it fails when the text length surpasses the cache size. We observe an interesting phenomenon, namely attention sink, that keeping the KV of initial tokens will largely recover the performance of window attention. In this paper, we first demonstrate that the emergence of attention sink is due to the strong attention scores towards initial tokens as a "sink" even if they are not semantically important. Based on the above analysis, we introduce StreamingLLM, an efficient framework that enables LLMs trained with a finite length attention window to generalize to infinite sequence lengths without any fine-tuning. We show that StreamingLLM can enable Llama-2, MPT, Falcon, and Pythia to perform stable and efficient language modeling with up to 4 million tokens and more. In addition, we discover that adding a placeholder token as a dedicated attention sink during pre-training can further improve streaming deployment. In streaming settings, StreamingLLM outperforms the sliding window recomputation baseline by up to 22.2x speedup. Code and datasets are provided at https://github.com/mit-han-lab/streaming-llm.arXiv:2309.17453v4View paper and GaLOREGaLore: Memory-Efficient LLM Training by Gradient Low-Rank ProjectionJiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang Wang, Anima Anandkumar, Yuandong TianTraining Large Language Models (LLMs) presents significant memory challenges, predominantly due to the growing size of weights and optimizer states. Common memory-reduction approaches, such as low-rank adaptation (LoRA), add a trainable low-rank matrix to the frozen pre-trained weight in each layer, reducing trainable parameters and optimizer states. However, such approaches typically underperform training with full-rank weights in both pre-training and fine-tuning stages since they limit the parameter search to a low-rank subspace and alter the training dynamics, and further, may require full-rank warm start. In this work, we propose Gradient Low-Rank Projection (GaLore), a training strategy that allows full-parameter learning but is more memory-efficient than common low-rank adaptation methods such as LoRA. Our approach reduces memory usage by up to 65.5% in optimizer states while maintaining both efficiency and performance for pre-training on LLaMA 1B and 7B architectures with C4 dataset with up to 19.7B tokens, and on fine-tuning RoBERTa on GLUE tasks. Our 8-bit GaLore further reduces optimizer memory by up to 82.5% and total training memory by 63.3%, compared to a BF16 baseline. Notably, we demonstrate, for the first time, the feasibility of pre-training a 7B model on consumer GPUs with 24GB memory (e.g., NVIDIA RTX 4090) without model parallel, checkpointing, or offloading strategies.arXiv:2403.03507v2View paper, made LLMs reason continuously with CoconutTraining Large Language Models to Reason in a Continuous Latent SpaceShibo Hao, Sainbayar Sukhbaatar, DiJia Su, Xian Li, Zhiting Hu, Jason Weston + 1 moreLarge language models (LLMs) are typically constrained to reason in the language space, where they express the reasoning process through a chain-of-thought (CoT) to solve complex problems. However, the language space may not always be optimal for reasoning. Most word tokens primarily ensure textual coherence and are not essential for reasoning, while some critical tokens require complex planning and pose challenges to LLMs. To explore the potential of reasoning beyond language, we introduce a new paradigm called Coconut (Chain of Continuous Thought). Coconut utilizes the last hidden state of the LLM as a representation of the reasoning state, termed "continuous thought." Instead of decoding this state into words, we feed it back to the model as the next input embedding directly in the continuous space. This latent reasoning paradigm enables an advanced reasoning pattern, where continuous thoughts can encode multiple alternative next steps, allowing the model to perform a breadth-first search (BFS) rather than committing prematurely to a single deterministic path as in CoT. Coconut outperforms CoT on logical reasoning tasks that require substantial search during planning and achieves a better trade-off between accuracy and efficiency.arXiv:2412.06769v4View paper and also did architecture work on attention sinks and positional encodings. He even published an open source version of AlphaGo called OpenGoELF OpenGo: An Analysis and Open Reimplementation of AlphaZeroYuandong Tian, Jerry Ma, Qucheng Gong, Shubho Sengupta, Zhuoyuan Chen, James Pinkerton + 1 moreThe AlphaGo, AlphaGo Zero, and AlphaZero series of algorithms are remarkable demonstrations of deep reinforcement learning's capabilities, achieving superhuman performance in the complex game of Go with progressively increasing autonomy. However, many obstacles remain in the understanding of and usability of these promising approaches by the research community. Toward elucidating unresolved mysteries and facilitating future research, we propose ELF OpenGo, an open-source reimplementation of the AlphaZero algorithm. ELF OpenGo is the first open-source Go AI to convincingly demonstrate superhuman performance with a perfect (20:0) record against global top professionals. We apply ELF OpenGo to conduct extensive ablation studies, and to identify and analyze numerous interesting phenomena in both the model training and in the gameplay inference procedures. Our code, models, selfplay datasets, and auxiliary data are publicly available at https://ai.facebook.com/tools/elf-opengo/.arXiv:1902.04522v5View paper, where he wrote over 90% of the code for a new infra platform called ELF which made the system train on only 2000 GPUs, despite himself being a senior researcher and manager at the time!

Tim Shi has spent close to a decade in the trenches of AI and startups. After dropping out of his PhD at Stanford where he worked on reinforcement learning, he became an early member of OpenAI where he led work on using AI to control a computer - highly prescient research! Subsequently, he left to be the founder and CTO of Cresta, a text and voice based company that both analyzes customer calls and automates customer experience end to end.

Alexei Dosovitskiy is a legendary figure in computer vision having invented the Vision TransformerAn Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner + 6 moreWhile the Transformer architecture has become the de-facto standard for natural language processing tasks, its applications to computer vision remain limited. In vision, attention is either applied in conjunction with convolutional networks, or used to replace certain components of convolutional networks while keeping their overall structure in place. We show that this reliance on CNNs is not necessary and a pure transformer applied directly to sequences of image patches can perform very well on image classification tasks. When pre-trained on large amounts of data and transferred to multiple mid-sized or small image recognition benchmarks (ImageNet, CIFAR-100, VTAB, etc.), Vision Transformer (ViT) attains excellent results compared to state-of-the-art convolutional networks while requiring substantially fewer computational resources to train.arXiv:2010.11929v2View paper, one of the most influential and highly cited papers in the history of the field. More broadly, his work has covered neural radiance fields, simulators for robotics, a host of architectures for vision learning and even applications in modeling RNA.

Caiming Xiong has a long history straddling both product and research, having joined MetaMind as a senior researcher in 2014 and subsequently spent a decade at Salesforce where he led Applied Research. Over this period, he built a practical pipeline to convert NLP, time series analysis and CV research into real products across Commerce, Marketing, Availability and Sales Clouds across the CRM, and also authored a paper that invented prompt engineering back in 2018.

It’s rare to see so many seemingly disparate threads and remarkable researchers come together, but at the same time, nothing could make more sense. This team is uniquely qualified to attack the most pressing problem in artificial intelligence today and have already started hiring a fantastic crop of infra and research talent. They’re on the cusp of unlocking a universe of insights for scientists everywhere, and we couldn’t be more chuffed to partner with these folks, and get a front row seat to the development of recursive super intelligence.