Welcome to the Metrics and Models homepage!
Talk Title: Computer Vision Research at Google DeepMind
Abstract: In the last few years, large language models (LLMs) have proven to be wildly successful. However, training models with robust visual perception capabilities remains an open research problem. Compared to language, visual data presents challenges with huge dimensionality and how to represent it in a feature space. In this talk, I will be presenting a survey of some of the latest published research this year in computer vision from Google DeepMind. This survey will focus on foundational research presented at conferences such as CVPR or ECCV. Papers this year explored a variety of subtopics within vision; however, some common themes among them emerge. First, they describe the challenges and opportunities of training models on large, diverse video datasets, especially when so much of this data is unlabelled. Second, they reveal the fundamental importance of learning good feature representations from these videos. I will be discussing results in dynamic 4D reconstruction, visual forecasting, representation learning, egocentric video, and 4D human perception.
Bio: Jacob Walker is a research scientist at Google DeepMind. His interests include general Computer Vision, Visual Forecasting, Model-Based Reinforcement Learning, and Representation Learning. He received his PhD from the Robotics Institute at Carnegie Mellon University.