As AI becomes more powerful, its influence on society becomes greater. The significance of AI five years ago was miniscule next to today, and future AIs will be strong enough to control the world. With a powerful AI, it is vital to ensure that its objectives are aligned with human values – not just some reasonable-sounding proxy of human values that breaks down under the optimizing power of a superintelligence. The challenge of ensuring that an AI’s objectives are in sync with its user’s wishes is known as the alignment problem (a.k.a. control problem). So far, the alignment problem has not been solved in a way that we can be confident will scale up to arbitrarily powerful intelligences.
Our mission
CORAL is working to develop a mathematical theory of agents to provide the framework for a rigorous and scalable solution to the alignment problem. Such a solution would incorporate formal, mathematically rigorous proofs that, given some reasonable assumptions, would guarantee alignment for a powerful AI. To this end, CORAL is working on the learning theoretic research agenda, which aims to understand computationally bounded agents through the tools of computational learning theory, control theory, algorithmic information theory and categorical systems theory. A well-founded theory of agents will also help us understand and predict the capabilities and behaviors of agentic AI systems and build more a more rigorous understanding of potential failure modes.
Why focus on agents?
The AIs of today can assist with many tasks, but are nowhere near strong enough to wrest control of humanity’s future. AIs that poses an existential risk to humanity will be masters at learning, planning, and adapting in pursuit of their objectives – in other words, they will be powerful agents.
Having a well-founded theory of agents will help us design and understand agentic AIs. This is true regardless of the underlying architecture of an AI. It doesn’t matter what combination of deep learning, chain-of-thought reasoning, and not-yet-invented algorithms constitutes an agentic AI at a lower level; a mathematical theory of agents will still describe it and help us understand it, much like how thermodynamics applies to any physical machine and information theory applies to any communication system.
Additionally, we need a mathematically rigorous descriptive theory of imperfect agents to make sense of the other end of the alignment problem: the human users. Even formulating the alignment problem – much less solving it – requires a rigorous definition of what it means for two agents to be aligned and what it means for an imperfect agent to have preferences.
