Amplify Partners leads $7 million initial funding for covariant.ai.
How can we build a robot that learns? Before answering this question, let’s start with a simpler one. How might a robot catch a baseball? Nowadays, the standard approach would model the problem with complex differential equations that account for factors like gravity, air resistance and so on. Successfully catching the ball depends on an incredibly precise plan executed in a deterministic environment, which is why a lot of tasks are too difficult to automate. But this is not how we, humans, catch a baseball. Humans, to our credit, figured out how to live in an imprecise world. In learning how to play catch, we might observe someone else first, try it ourselves and continue calibrating our movements until we figure it out or get bored, whichever comes first.
So to teach a robot to play catch, or any number of tasks, we need to teach them how to learn from experience and from demonstration. We must enable them to act on their observations and adapt until they achieved their goal. This is what the all-star team at covariant.ai is set to accomplish. Covariant.ai is set on building brains for robots to automate what was previously impossible.

The covariant.ai Founding Team(left-to-right): Peter Chen (CEO), Pieter Abbeel (President and Chief Scientist), Rocky Duan (CTO), Tianhao Zhang (Research Scientist)
The covariant.ai team builds on their pioneering contributions to machine learning applied to robotics, including work in Deep Reinforcement Learning, Deep Imitation Learning and Meta-Learning. Pieter and his students have been forces of nature at UC Berkeley and OpenAI. As a PhD (focused on deep learning and mathematics) at UC Berkeley, I was lucky enough to witness their efforts first-hand. Deep Reinforcement Learning addresses the problem of an agent interacting with its environment and altering its behaviour in response to the rewards received. The “Deep” part takes advantage of advancements in deep neural networks to learn compact representations of complex, high dimensional data. This is key because the real world is infinite dimensional and the real world is where robots operate. Deep Imitation learning is a class of methods to enable learning via demonstrationsOne-Shot Imitation LearningImitation learning has been commonly applied to solve different tasks in isolation. This usually requires either careful feature engineering, or a significant number of samples. This is far from what we desire: ideally, robots should be able to learn from very few demonstrations of any given task, and instantly generalize to new situations of the same task, without requiring task-specific engineering. In this paper, we propose a meta-learning framework for achieving such capability, which we call one-shot imitation learning. Specifically, we consider the setting where there is a very large set of tasks, and each task has many instantiations. For example, a task could be to stack all blocks on a table into a single tower, another task could be to place all blocks on a table into two-block towers, etc. In each case, different instances of the task would consist of different sets of blocks with different initial states. At training time, our algorithm is presented with pairs of demonstrations for a subset of all tasks. A neural net is trained that takes as input one demonstration and the current state (which initially is the initial state of the other demonstration of the pair), and outputs an action with the goal that the resulting sequence of states and actions matches as closely as possible with the second demonstration. At test time, a demonstration of a single instance of a new task is presented, and the neural net is expected to perform well on new instances of this new task. The use of soft attention allows the model to generalize to conditions and tasks unseen in the training data. We anticipate that by training this model on a much greater variety of tasks and settings, we will obtain a general system that can turn any demonstrations into robust policies that can accomplish an overwhelming variety of tasks. Videos available at https://bit.ly/nips2017-oneshot .arXiv:1703.07326v3View paper. A direct application of this is to train robots to learn from VR demonstrationsDeep Imitation Learning for Complex Manipulation Tasks from Virtual Reality TeleoperationImitation learning is a powerful paradigm for robot skill acquisition. However, obtaining demonstrations suitable for learning a policy that maps from raw pixels to actions can be challenging. In this paper we describe how consumer-grade Virtual Reality headsets and hand tracking hardware can be used to naturally teleoperate robots to perform complex tasks. We also describe how imitation learning can learn deep neural network policies (mapping from pixels to actions) that can acquire the demonstrated skills. Our experiments showcase the effectiveness of our approach for learning visuomotor skills.arXiv:1710.04615v2View paper to do tasks that cannot be programmed. And finally meta-learning tackles learning to learn, so robots get better at generalizing from a diverse array of tasksRL$^2$: Fast Reinforcement Learning via Slow Reinforcement LearningDeep reinforcement learning (deep RL) has been successful in learning sophisticated behaviors automatically; however, the learning process requires a huge number of trials. In contrast, animals can learn new tasks in just a few trials, benefiting from their prior knowledge about the world. This paper seeks to bridge this gap. Rather than designing a "fast" reinforcement learning algorithm, we propose to represent it as a recurrent neural network (RNN) and learn it from data. In our proposed method, RL$^2$, the algorithm is encoded in the weights of the RNN, which are learned slowly through a general-purpose ("slow") RL algorithm. The RNN receives all information a typical RL algorithm would receive, including observations, actions, rewards, and termination flags; and it retains its state across episodes in a given Markov Decision Process (MDP). The activations of the RNN store the state of the "fast" RL algorithm on the current (previously unseen) MDP. We evaluate RL$^2$ experimentally on both small-scale and large-scale problems. On the small-scale side, we train it to solve randomly generated multi-arm bandit problems and finite MDPs. After RL$^2$ is trained, its performance on new MDPs is close to human-designed algorithms with optimality guarantees. On the large-scale side, we test RL$^2$ on a vision-based navigation task and show that it scales up to high-dimensional problems.arXiv:1611.02779v2View paper.

Tianhao Zhang (Research Scientist) shows how through VR tele-operation human operators can intuitively teach robots new skills.
The entire covariant.ai team is responsible for an impressive amount of robotics research coming out of OpenAI and UC Berkeley. I first met Rocky Duan (CTO of covariant.ai) when giving a talk on my research at OpenAI early last fall. I remembered distinctly that Rocky asked very interesting questions. Other researchers in the field, when speaking about Rocky’s work, rarely forgot to mention that he is a “rock star”. Andrej Karpathy, director of AI at Tesla and one of the founding members of OpenAI, shared with me once that Peter Chen (CEO of covariant.ai) and Rocky Duan’s work is “extremely creative” and they had “lots of insightful discussions” during their time at OpenAI. This is a view I’ve heard echoed in the community.
Given all this, you can imagine my excitement when Pieter Abbeel told me that he and his students Rocky, Peter and Tianhao were going to found a startup to make robotics much smarter and solve problems previous automation found impossible. I was thrilled to be in a position to support them.
teaching robots to learn quickly a diverse array of tasks, we enable robotic applications traditionally limited to large-scale operations and capital commitments; we are embarking on an exciting journey to democratize automation.
The technologies developed covariant.ai will radically impact how companies, ranging from retail to manufacturing, make use of industrial robotics. Our mandate at Amplify is to work with early stage top-of-their-game technical founders solving technical problems. We could not be more excited to partner with this all star team pushing the cutting edge of robotics research to the real world.




