After My Third PhD Year: Reflections on My Research Direction
I recently completed the third year of my PhD and am currently interning in the Bay Area. This transition has given me an opportunity to reflect on my research journey over the past few years—and to thank a mentor who offered me invaluable guidance and support early in my PhD.
Over the past year, I have gone through a period of exploration and adjustment in my research direction. Early in my PhD, I began working on a project centered on physical priors and machine learning. Unfortunately, I was unable to carry it forward. Although the project was eventually paused, the guidance I received—and our discussions about physics and machine learning—have continued to shape my thinking.
At the beginning of my PhD, I strongly believed in the value of physical priors for AI systems. Meanwhile, machine learning, computer vision, and robot learning have been profoundly influenced by large-scale foundation models emerging from industry, especially generative video models. Many technical approaches are gradually converging toward a unified paradigm centered on generative models and spanning pre-training, mid-training, and post-training.
Robot learning is undergoing a similar transformation. New paradigms centered on large-scale robot data and generative models are developing rapidly. Recent systems such as Rhoda AI's Direct Video-Action Model (DVA), DreamZero, and π₀.₅, along with our recent work VERA—which post-trains video models on robot videos—are beginning to demonstrate generalization across environments, tasks, and even robotic platforms.
My recent experience in industry has shown me the potential of video models, once trained at scale, to serve as backbones for robot learning.
After these explorations, I still believe that the ways physics characterizes time, motion, and energy will ultimately need to be incorporated into modern deep learning in some form. Compared with the beginning of my PhD, however, I now have a more concrete understanding of this problem—and a greater appreciation of how difficult it is to connect the two in a genuinely effective way.
I am grateful to everyone working together to explore the future of AI.
博士三年级之后:关于研究方向的一些回顾
最近,我结束了博士三年级,目前正在湾区实习。借这个机会,我想回顾过去几年的研究经历,也感谢一位在我博士初期给予过我许多帮助与指导的老师。
过去一年,我在研究方向上经历了不少探索与调整。博士初期,我曾尝试开展一个围绕物理先验与机器学习的项目。遗憾的是,我当时没能将它继续推进。虽然项目最终暂停,但当时获得的指导,以及我们围绕物理学与机器学习展开的讨论,始终影响着我之后的思考。
博士初期,我非常相信物理先验在 AI 系统中的价值。与此同时,过去几年,机器学习、计算机视觉和机器人学习也受到工业界大规模基础模型——尤其是生成式视频模型——的深刻影响。许多技术路线正逐渐汇聚到一种以生成模型为核心,并贯穿 pre-training、mid-training 和 post-training 的统一范式。
类似的转变也正在机器人学习领域发生。以大规模机器人数据和生成模型为核心的新范式正在快速发展。近期 Rhoda AI 的 Direct Video-Action Model (DVA)、DreamZero、π₀.₅,以及我们最近的工作 VERA——通过在机器人视频上对视频模型进行 post-training——都开始展现出跨场景、跨任务,甚至跨机器人平台的泛化能力。
最近的工业界经历,让我看到了视频模型经过规模化训练后,作为机器人学习 backbone 所展现出的潜力。
经历这些探索之后,我仍然相信,物理学对于时间、运动和能量的刻画,最终需要以某种方式融入现代深度学习体系。只是与博士初期相比,我现在对这个问题有了更具体的认识,也更加明白,要找到真正有效的结合方式并不容易。
感谢所有正在共同探索 AI 未来的人。