Sunayana Rane (Princeton University)
Abstract
Technical AI safety researchers, regulators, and policymakers have a common goal: to protect us from the harms arising from increasingly ubiquitous AI systems in the real world. For technical researchers, this often entails meticulously studying the behavior of existing AI systems and building better, safer ones. For regulators, policymakers, and courts, it involves effectively regulating and adjudicating this space of breakneck technological development. And yet a paradoxical challenge makes this nearly impossible: humans cannot instinctively predict harmful AI behavior, and the parties who could, with appropriate time and resources, help foresee harmful AI behavior are the ones currently least incentivized to do so. Much of the danger in AI behavior is that it is often dissimilar from human behavior in unexpected ways. Drawing together empirical evidence from AI interpretability and cognitive science, this work explains why humans should not be expected to foresee harmful AI actions. Instead, the foreseeability of AI harms largely depends on the due diligence of the AI's creators – due diligence they are currently disincentivized from undertaking. Fortunately, the legal concept of foreseeability also gives us the foundation for the solution: responsibility for preventing or recompensing and AI harm should center on whether, and to what extent, each actor behaved responsibly with their knowledge of the tendencies of the AI system in question.