Amplify has a long history of attending NeurIPS (for the papers, of course) and went 5 strong this year for one of the densest and most exciting conferences we’ve been to in a while. So we (Barr, Sarah, Rohan, Mike, and Neiman) decided to write up a few of the themes that we heard a lot about. From advances in data curation to compute scaling laws and (finally) good audio models coming soon, it’s an exciting time to be a Neural Information Processing System.
Takeaway 1: RL is making a comeback
Over the last decade, reinforcement learning has delivered several successes in solving tasks with a clear reward function (e.g., Atari GamesPlaying Atari with Deep Reinforcement LearningWe present the first deep learning model to successfully learn control policies directly from high-dimensional sensory input using reinforcement learning. The model is a convolutional neural network, trained with a variant of Q-learning, whose input is raw pixels and whose output is a value function estimating future rewards. We apply our method to seven Atari 2600 games from the Arcade Learning Environment, with no adjustment of the architecture or learning algorithm. We find that it outperforms all previous approaches on six of the games and surpasses a human expert on three of them.arXiv:1312.5602v1View paper or GoMastering Chess and Shogi by Self-Play with a General Reinforcement Learning AlgorithmThe game of chess is the most widely-studied domain in the history of artificial intelligence. The strongest programs are based on a combination of sophisticated search techniques, domain-specific adaptations, and handcrafted evaluation functions that have been refined by human experts over several decades. In contrast, the AlphaGo Zero program recently achieved superhuman performance in the game of Go, by tabula rasa reinforcement learning from games of self-play. In this paper, we generalise this approach into a single AlphaZero algorithm that can achieve, tabula rasa, superhuman performance in many challenging domains. Starting from random play, and given no domain knowledge except the game rules, AlphaZero achieved within 24 hours a superhuman level of play in the games of chess and shogi (Japanese chess) as well as Go, and convincingly defeated a world-champion program in each case.arXiv:1712.01815v1View paper), but it was never quite at the forefront of GenAI hype cycles. That might be changing, though: in the context of foundation models (language, vision, speech, and/or multimodal), researchers are using RL to teach agents reasoning in environments formally verifying mathematical or programming solutions.
For example, researchers are generating code in Dafny and Rust and using formal verification to provide mathematical guarantees that AI generated code is correctAlphaVerus: Bootstrapping Formally Verified Code Generation through Self-Improving Translation and TreefinementAutomated code generation with large language models has gained significant traction, but there remains no guarantee on the correctness of generated code. We aim to use formal verification to provide mathematical guarantees that the generated code is correct. However, generating formally verified code with LLMs is hindered by the scarcity of training data and the complexity of formal proofs. To tackle this challenge, we introduce AlphaVerus, a self-improving framework that bootstraps formally verified code generation by iteratively translating programs from a higher-resource language and leveraging feedback from a verifier. AlphaVerus operates in three phases: exploration of candidate translations, Treefinement -- a novel tree search algorithm for program refinement using verifier feedback, and filtering misaligned specifications and programs to prevent reward hacking. Through this iterative process, AlphaVerus enables a LLaMA-3.1-70B model to generate verified code without human intervention or model finetuning. AlphaVerus shows an ability to generate formally verified solutions for HumanEval and MBPP, laying the groundwork for truly trustworthy code-generation agents.arXiv:2412.06176v1View paper. Similarly, digital environments like OSWorld make it easier for researchers to test new algorithms such as GFlowNet finetuningAmortizing intractable inference in large language modelsAutoregressive large language models (LLMs) compress knowledge from their training data through next-token conditional distributions. This limits tractable querying of this knowledge to start-to-end autoregressive sampling. However, many tasks of interest -- including sequence continuation, infilling, and other forms of constrained generation -- involve sampling from intractable posterior distributions. We address this limitation by using amortized Bayesian inference to sample from these intractable posteriors. Such amortization is algorithmically achieved by fine-tuning LLMs via diversity-seeking reinforcement learning algorithms: generative flow networks (GFlowNets). We empirically demonstrate that this distribution-matching paradigm of LLM fine-tuning can serve as an effective alternative to maximum-likelihood training and reward-maximizing policy optimization. As an important application, we interpret chain-of-thought reasoning as a latent variable modeling problem and demonstrate that our approach enables data-efficient adaptation of LLMs to tasks that require multi-step rationalization and tool use.arXiv:2310.04363v3View paper in completing basic computer operations.
From a technical standpoint, we chatted with researchers who are revisiting age-old problems of credit assignmentA Survey of Temporal Credit Assignment in Deep Reinforcement LearningThe Credit Assignment Problem (CAP) refers to the longstanding challenge of Reinforcement Learning (RL) agents to associate actions with their long-term consequences. Solving the CAP is a crucial step towards the successful deployment of RL in the real world since most decision problems provide feedback that is noisy, delayed, and with little or no information about the causes. These conditions make it hard to distinguish serendipitous outcomes from those caused by informed decision-making. However, the mathematical nature of credit and the CAP remains poorly understood and defined. In this survey, we review the state of the art of Temporal Credit Assignment (CA) in deep RL. We propose a unifying formalism for credit that enables equitable comparisons of state-of-the-art algorithms and improves our understanding of the trade-offs between the various methods. We cast the CAP as the problem of learning the influence of an action over an outcome from a finite amount of experience. We discuss the challenges posed by delayed effects, transpositions, and a lack of action influence, and analyse how existing methods aim to address them. Finally, we survey the protocols to evaluate a credit assignment method and suggest ways to diagnose the sources of struggle for different methods. Overall, this survey provides an overview of the field for new-entry practitioners and researchers, it offers a coherent perspective for scholars looking to expedite the starting stages of a new study on the CAP, and it suggests potential directions for future research.arXiv:2312.01072v2View paper for achieving sample efficiency. Knowing which reasoning steps to reward on the way to a correct answer can significantly reduce the amount of attempts and consequently the amount of data / compute you need. Furthermore, making models learn tasks in the right order can boost efficiency, an area of research known as environment design. OMNI-EPICOMNI-EPIC: Open-endedness via Models of human Notions of Interestingness with Environments Programmed in CodeOpen-ended and AI-generating algorithms aim to continuously generate and solve increasingly complex tasks indefinitely, offering a promising path toward more general intelligence. To accomplish this grand vision, learning must occur within a vast array of potential tasks. Existing approaches to automatically generating environments are constrained within manually predefined, often narrow distributions of environment, limiting their ability to create any learning environment. To address this limitation, we introduce a novel framework, OMNI-EPIC, that augments previous work in Open-endedness via Models of human Notions of Interestingness (OMNI) with Environments Programmed in Code (EPIC). OMNI-EPIC leverages foundation models to autonomously generate code specifying the next learnable (i.e., not too easy or difficult for the agent's current skill set) and interesting (e.g., worthwhile and novel) tasks. OMNI-EPIC generates both environments (e.g., an obstacle course) and reward functions (e.g., progress through the obstacle course quickly without touching red objects), enabling it, in principle, to create any simulatable learning task. We showcase the explosive creativity of OMNI-EPIC, which continuously innovates to suggest new, interesting learning challenges. We also highlight how OMNI-EPIC can adapt to reinforcement learning agents' learning progress, generating tasks that are of suitable difficulty. Overall, OMNI-EPIC can endlessly create learnable and interesting environments, further propelling the development of self-improving AI systems and AI-Generating Algorithms. Project website with videos: https://dub.sh/omniepicarXiv:2405.15568v3View paper from Jeff Clune’s team showed great results in creating this learnable frontier in the context of code models, winning a Best Paper Award at the Intrinsically Motivated Open-ended Learning Workshop.
Several breakthroughs using RL made agents play games against each other, known as self play, to massively increase the amount of data they’ve seen and consequent learning. Looking ahead, it seems likely we’ll use self-play to improve language agents too. Recently, Noam Brown mentioned he’s hiring for a multi-agent research team at OpenAI. He highlighted how useful self-play techniques were when he worked on Libratus (a champion poker bot), and was excited that making intermediate agents reason against each other gives us another axis for scaling compute during post-training.
Takeaway 2: Data curation makes models faster, better, and smaller
It’s no secret that better data yields better models. After NeurIPS, we are even more excited about new techniques for analyzing how data influences models. Cohere’s Procedural Knowledge in Pre-Training drives LLM ReasoningProcedural Knowledge in Pretraining Drives Reasoning in Large Language ModelsThe capabilities and limitations of Large Language Models have been sketched out in great detail in recent years, providing an intriguing yet conflicting picture. On the one hand, LLMs demonstrate a general ability to solve problems. On the other hand, they show surprising reasoning gaps when compared to humans, casting doubt on the robustness of their generalisation strategies. The sheer volume of data used in the design of LLMs has precluded us from applying the method traditionally used to measure generalisation: train-test set separation. To overcome this, we study what kind of generalisation strategies LLMs employ when performing reasoning tasks by investigating the pretraining data they rely on. For two models of different sizes (7B and 35B) and 2.5B of their pretraining tokens, we identify what documents influence the model outputs for three simple mathematical reasoning tasks and contrast this to the data that are influential for answering factual questions. We find that, while the models rely on mostly distinct sets of data for each factual question, a document often has a similar influence across different reasoning questions within the same task, indicating the presence of procedural knowledge. We further find that the answers to factual questions often show up in the most influential data. However, for reasoning questions the answers usually do not show up as highly influential, nor do the answers to the intermediate reasoning steps. When we characterise the top ranked documents for the reasoning questions qualitatively, we confirm that the influential documents often contain procedural knowledge, like demonstrating how to obtain a solution using formulae or code. Our findings indicate that the approach to reasoning the models use is unlike retrieval, and more like a generalisable strategy that synthesises procedural knowledge from documents doing a similar form of reasoning.arXiv:2411.12580v2View paper used influence functions – a statistical technique that determines how changing a particular input affects a model’s output – to show that while models learn that factual knowledge derives from a single document, generalizable reasoning emerges from a collection of documents. We noticed equal interest at the conference in Aleksander Madry’s datamodelsDatamodels: Predicting Predictions from Training DataWe present a conceptual framework, datamodeling, for analyzing the behavior of a model class in terms of the training data. For any fixed "target" example $x$, training set $S$, and learning algorithm, a datamodel is a parameterized function $2^S \to \mathbb{R}$ that for any subset of $S' \subset S$ -- using only information about which examples of $S$ are contained in $S'$ -- predicts the outcome of training a model on $S'$ and evaluating on $x$. Despite the potential complexity of the underlying process being approximated (e.g., end-to-end training and evaluation of deep neural networks), we show that even simple linear datamodels can successfully predict model outputs. We then demonstrate that datamodels give rise to a variety of applications, such as: accurately predicting the effect of dataset counterfactuals; identifying brittle predictions; finding semantically similar examples; quantifying train-test leakage; and embedding data into a well-behaved and feature-rich representation space. Data for this paper (including pre-computed datamodels as well as raw predictions from four million trained deep neural networks) is available at https://github.com/MadryLab/datamodels-data .arXiv:2202.00622v1View paper. This approach blindly optimizes for which subset of the data leads to the highest performance on a single task, after which we can draw inferences from the subsets about how we should curate data going forward.
During post-training, scalable oversight techniquesMeasuring Progress on Scalable Oversight for Large Language ModelsDeveloping safe and useful general-purpose AI systems will require us to make progress on scalable oversight: the problem of supervising systems that potentially outperform us on most skills relevant to the task at hand. Empirical work on this problem is not straightforward, since we do not yet have systems that broadly exceed our abilities. This paper discusses one of the major ways we think about this problem, with a focus on ways it can be studied empirically. We first present an experimental design centered on tasks for which human specialists succeed but unaided humans and current general AI systems fail. We then present a proof-of-concept experiment meant to demonstrate a key feature of this experimental design and show its viability with two question-answering tasks: MMLU and time-limited QuALITY. On these tasks, we find that human participants who interact with an unreliable large-language-model dialog assistant through chat -- a trivial baseline strategy for scalable oversight -- substantially outperform both the model alone and their own unaided performance. These results are an encouraging sign that scalable oversight will be tractable to study with present models and bolster recent findings that large language models can productively assist humans with difficult tasks.arXiv:2211.03540v2View paper – methods for aligning models to responsibly improve model capabilities while ensuring that model behavior conforms to norms – have proven successful in generating high quality synthetic data, augmenting evaluation, and supporting higher throughput, higher quality annotation. During the conference, we hosted a breakfast for scalable oversight researchers to discuss this topic further.
While previous researchReflexion: Language Agents with Verbal Reinforcement LearningLarge language models (LLMs) have been increasingly used to interact with external environments (e.g., games, compilers, APIs) as goal-driven agents. However, it remains challenging for these language agents to quickly and efficiently learn from trial-and-error as traditional reinforcement learning methods require extensive training samples and expensive model fine-tuning. We propose Reflexion, a novel framework to reinforce language agents not by updating weights, but instead through linguistic feedback. Concretely, Reflexion agents verbally reflect on task feedback signals, then maintain their own reflective text in an episodic memory buffer to induce better decision-making in subsequent trials. Reflexion is flexible enough to incorporate various types (scalar values or free-form language) and sources (external or internally simulated) of feedback signals, and obtains significant improvements over a baseline agent across diverse tasks (sequential decision-making, coding, language reasoning). For example, Reflexion achieves a 91% pass@1 accuracy on the HumanEval coding benchmark, surpassing the previous state-of-the-art GPT-4 that achieves 80%. We also conduct ablation and analysis studies using different feedback signals, feedback incorporation methods, and agent types, and provide insights into how they affect performance.arXiv:2303.11366v4View paper proposed using LLMs to judge or critique model outputs, new approaches address the limitations of these techniques. For example, models that don’t know the answer to a question can watch stronger models debate each other and select the right answerDebating with More Persuasive LLMs Leads to More Truthful AnswersCommon methods for aligning large language models (LLMs) with desired behaviour heavily rely on human-labelled data. However, as models grow increasingly sophisticated, they will surpass human expertise, and the role of human evaluation will evolve into non-experts overseeing experts. In anticipation of this, we ask: can weaker models assess the correctness of stronger models? We investigate this question in an analogous setting, where stronger models (experts) possess the necessary information to answer questions and weaker models (non-experts) lack this information. The method we evaluate is debate, where two LLM experts each argue for a different answer, and a non-expert selects the answer. We find that debate consistently helps both non-expert models and humans answer questions, achieving 76% and 88% accuracy respectively (naive baselines obtain 48% and 60%). Furthermore, optimising expert debaters for persuasiveness in an unsupervised manner improves non-expert ability to identify the truth in debates. Our results provide encouraging empirical evidence for the viability of aligning models with debate in the absence of ground truth.arXiv:2402.06782v4View paper. This both generates new chains of thought that were not present in the pre-training data and provides a signal on whether it’s true or false, increasing the quantity and quality of supervision data. These methods could be especially helpful in domains without any ground truth such as legal or medical reasoning.
Takeaway 3: Emphasis shift from pretraining scaling law progress to inference-time compute scaling law progress
Noam Brown at the Compound AI Systems workshop talk said: “I’ve never heard any serious AI researcher say that AI is hitting a wall.” Much of the discussion at NeurIPS this year was around where exactly AI progress will (and won’t) come from, with a particular focus on scaling laws.
One buzzy moment was in Ilya Sutskever’s talk, when he stated that “pre-training as we know it will end.” Major AI progress has come from scaling data and compute. While compute is growing, we are pre-training on easily available internet data which is limited. In short, the low hanging fruit is gone. Scaling laws, which have historically driven major breakthroughs during pretraining, also inherently yield slower returns due to their logarithmic nature. It is also more costly to train enormous models – Noam posed the question: “would we pay trillions for better AI?” That said, just because pretraining with the low hanging internet fruit may slow down, pretraining is far from dead; there are tons of other data sources we haven’t yet figured out, like under-utilized domain-specific data.
Much of the conference focus shifted from a focus on pretraining scaling laws to gains from inference-time compute. Beyond exciting post training techniques, over the last year we have seen strategies for how spending compute at inference time can predictably improve an agent’s accuracy. In the summer, work from Stanford and BerkeleyAre More LLM Calls All You Need? Towards Scaling Laws of Compound Inference SystemsMany recent state-of-the-art results in language tasks were achieved using compound systems that perform multiple Language Model (LM) calls and aggregate their responses. However, there is little understanding of how the number of LM calls - e.g., when asking the LM to answer each question multiple times and taking a majority vote - affects such a compound system's performance. In this paper, we initiate the study of scaling properties of compound inference systems. We analyze, theoretically and empirically, how the number of LM calls affects the performance of Vote and Filter-Vote, two of the simplest compound system designs, which aggregate LM responses via majority voting, optionally applying LM filters. We find, surprisingly, that across multiple language tasks, the performance of both Vote and Filter-Vote can first increase but then decrease as a function of the number of LM calls. Our theoretical results suggest that this non-monotonicity is due to the diversity of query difficulties within a task: more LM calls lead to higher performance on "easy" queries, but lower performance on "hard" queries, and non-monotone behavior can emerge when a task contains both types of queries. This insight then allows us to compute, from a small number of samples, the number of LM calls that maximizes system performance, and define an analytical scaling model for both systems. Experiments show that our scaling model can accurately predict the performance of Vote and Filter-Vote systems and thus find the optimal number of LM calls to make.arXiv:2403.02419v2View paper started simply scaling the number of LLM calls per query and filtering responses with majority vote. This demonstrated a path towards cost optimal choices for increasing compute, and there are far more dimensions along which one can scale compute at inference time as well as very different results based on the types of questions you’re asking.
At the conference, we saw how architecture search frameworks such as ArchonArchon: An Architecture Search Framework for Inference-Time TechniquesInference-time techniques, such as repeated sampling or iterative revisions, are emerging as powerful ways to enhance large-language models (LLMs) at test time. However, best practices for developing systems that combine these techniques remain underdeveloped due to our limited understanding of the utility of each technique across models and tasks, the interactions between them, and the massive search space for combining them. To address these challenges, we introduce Archon, a modular and automated framework for optimizing the process of selecting and combining inference-time techniques and LLMs. Given a compute budget and a set of available LLMs, Archon explores a large design space to discover optimized configurations tailored to target benchmarks. It can design custom or general-purpose architectures that advance the Pareto frontier of accuracy vs. maximum token budget compared to top-performing baselines. Across instruction-following, reasoning, and coding tasks, we show that Archon can leverage additional inference compute budget to design systems that outperform frontier models such as OpenAI's o1, GPT-4o, and Claude 3.5 Sonnet by an average of 15.1%.arXiv:2409.15254v6View paper help tractably reduce the large design space of ensembling, ranking, fusion, critiquing and verification to a set of cost optimal hyperparameters. Furthermore, work in Large Language MonkeysLarge Language Monkeys: Scaling Inference Compute with Repeated SamplingScaling the amount of compute used to train language models has dramatically improved their capabilities. However, when it comes to inference, we often limit models to making only one attempt at a problem. Here, we explore inference compute as another axis for scaling, using the simple technique of repeatedly sampling candidate solutions from a model. Across multiple tasks and models, we observe that coverage -- the fraction of problems that are solved by any generated sample -- scales with the number of samples over four orders of magnitude. Interestingly, the relationship between coverage and the number of samples is often log-linear and can be modelled with an exponentiated power law, suggesting the existence of inference-time scaling laws. In domains like coding and formal proofs, where answers can be automatically verified, these increases in coverage directly translate into improved performance. When we apply repeated sampling to SWE-bench Lite, the fraction of issues solved with DeepSeek-Coder-V2-Instruct increases from 15.9% with one sample to 56% with 250 samples, outperforming the single-sample state-of-the-art of 43%. In domains without automatic verifiers, we find that common methods for picking from a sample collection (majority voting and reward models) plateau beyond several hundred samples and fail to fully scale with the sample budget.arXiv:2407.21787v3View paper concluded that environments without a clear ground hinder inference time scaling. The authors suggested that auto-formalizing informal problems, e.g. taking a math question in text and rewriting it as verifiable code, can bypass this bottleneck and we’re excited about how more investment into these data pipelines might progress.
Takeaway 4: Good audio models are coming soon
Our second conference breakfast focused on audio models, spanning researchers from real time speech to speech models to music conditioned generation and background noise reduction.
The biggest takeaway was that even if many researchers have switched focus to post training and agents, pretraining marches on in other modalities: we still have significant amounts of unused web audio data! We have not yet figured out how to use most of this data due to privacy concerns, but research in differential privacy may soon unlock this corpus.
Secondly, it is tricky to have long open ended conversations with speech models today because they simply aren’t good enough. Fortunately, improved tokenization and data curation around a now fixed architecture, specifically streaming ASR-TTS (automatic speech recognition – text to speech), should yield performant systems in the next few years.
Thirdly, most audio models today condition on text, but users want to condition on other modalities to generate more diverse outputs. For example, several papers in the Audio Imagination workshop focused on audio model conditioning, and Sony’s Diff-a-RiffDiff-A-Riff: Musical Accompaniment Co-creation via Latent Diffusion ModelsRecent advancements in deep generative models present new opportunities for music production but also pose challenges, such as high computational demands and limited audio quality. Moreover, current systems frequently rely solely on text input and typically focus on producing complete musical pieces, which is incompatible with existing workflows in music production. To address these issues, we introduce "Diff-A-Riff," a Latent Diffusion Model designed to generate high-quality instrumental accompaniments adaptable to any musical context. This model offers control through either audio references, text prompts, or both, and produces 48kHz pseudo-stereo audio while significantly reducing inference time and memory usage. We demonstrate the model's capabilities through objective metrics and subjective listening tests, with extensive examples available on the accompanying website: sonycslparis.github.io/diffariff-companion/arXiv:2406.08384v2View paper lets you take a track that you’re working on, add an audio clip for a theme you want to incorporate e.g. a hummed melody, and output the musical accompaniment! This opens up fascinating new questions in human computer interaction, as we can continuously supply new pieces of audio to nudge the model in a new direction. Beyond these technical advancements, it was encouraging to learn about active speech data collection efforts in low resource languages Ai4BharatIndicVoices: Towards building an Inclusive Multilingual Speech Dataset for Indian LanguagesWe present INDICVOICES, a dataset of natural and spontaneous speech containing a total of 7348 hours of read (9%), extempore (74%) and conversational (17%) audio from 16237 speakers covering 145 Indian districts and 22 languages. Of these 7348 hours, 1639 hours have already been transcribed, with a median of 73 hours per language. Through this paper, we share our journey of capturing the cultural, linguistic and demographic diversity of India to create a one-of-its-kind inclusive and representative dataset. More specifically, we share an open-source blueprint for data collection at scale comprising of standardised protocols, centralised tools, a repository of engaging questions, prompts and conversation scenarios spanning multiple domains and topics of interest, quality control mechanisms, comprehensive transcription guidelines and transcription tools. We hope that this open source blueprint will serve as a comprehensive starter kit for data collection efforts in other multilingual regions of the world. Using INDICVOICES, we build IndicASR, the first ASR model to support all the 22 languages listed in the 8th schedule of the Constitution of India. All the data, tools, guidelines, models and other materials developed as a part of this work will be made publicly availablearXiv:2403.01926v1View paper, paving the path to even greater access to audio models.
Takeaway 5: New Architectural Analyses Explaining Poor Reasoning Results
Previously most attempts to understand LLMs have focused on examining model outputs, and jailbreak attempts have used complex prompt optimization strategies to exploit model weaknesses. At NeurIPS, it was refreshing to see the spotlight shone on theoretical limitations of the transformer and reasoning architectures instead.
In Transformers Need GlassesTransformers need glasses! Information over-squashing in language tasksWe study how information propagates in decoder-only Transformers, which are the architectural backbone of most existing frontier large language models (LLMs). We rely on a theoretical signal propagation analysis -- specifically, we analyse the representations of the last token in the final layer of the Transformer, as this is the representation used for next-token prediction. Our analysis reveals a representational collapse phenomenon: we prove that certain distinct sequences of inputs to the Transformer can yield arbitrarily close representations in the final token. This effect is exacerbated by the low-precision floating-point formats frequently used in modern LLMs. As a result, the model is provably unable to respond to these sequences in different ways -- leading to errors in, e.g., tasks involving counting or copying. Further, we show that decoder-only Transformer language models can lose sensitivity to specific tokens in the input, which relates to the well-known phenomenon of over-squashing in graph neural networks. We provide empirical evidence supporting our claims on contemporary LLMs. Our theory also points to simple solutions towards ameliorating these issues.arXiv:2406.04267v2View paper, the authors showed how representation collapse means transformers are unable to count or copy, reducing reasoning capabilities. In addition, the Softmax is Not EnoughSoftmax is not Enough (for Sharp Size Generalisation)A key property of reasoning systems is the ability to make sharp decisions on their input data. For contemporary AI systems, a key carrier of sharp behaviour is the softmax function, with its capability to perform differentiable query-key lookups. It is a common belief that the predictive power of networks leveraging softmax arises from "circuits" which sharply perform certain kinds of computations consistently across many diverse inputs. However, for these circuits to be robust, they would need to generalise well to arbitrary valid inputs. In this paper, we dispel this myth: even for tasks as simple as finding the maximum key, any learned circuitry must disperse as the number of items grows at test time. We attribute this to a fundamental limitation of the softmax function to robustly approximate sharp functions with increasing problem size, prove this phenomenon theoretically, and propose adaptive temperature as an ad-hoc technique for improving the sharpness of softmax at inference time.arXiv:2410.01104v3View paper talk at the System 2 Reasoning at Scale Workshop highlighted that even if generalizable neural circuitry is learned, limitations of the softmax function (one of the most important in modern deep learning/pervasive across most models) means the model cannot generalize out of distribution for problems with a constant number of inputs. For the same reason we’re excited about interpretability techniques, granular understanding of how information propagates in transformers will help scale better, more accurate models.
See you next year
Amplify showed up big for NeurIPS this year, with two thematic breakfasts (audio and scalable oversight), our yearly dinner across research areas, plus a professor / grad student lunch. We’ve been investing in AI for 10+ years, and it never gets old seeing the cutting edge every December. See you next year!



