Boyang Zhong钟伯扬

I am an M.Sc. student in Robotics, Cognition, and Intelligence at the Technical University of Munich and a research intern at Agile Robots in Munich. I study 3D object understanding and generative world models for robot perception and interaction.

My current work studies interactive world models for articulated object understanding and manipulation with Xi Wang and Prof. Daniel Cremers, and contact-aware video world models for dexterous manipulation.

I am seeking Ph.D. opportunities starting in 2027 in spatial intelligence, interactive world models, and robot learning for manipulation.

Previously, I collaborated with Prof. Slobodan Ilic, Prof. Benjamin Busam, and Prof. Alois Knoll on 3D object perception and robot manipulation. I earned my B.Eng. from Tongji University.

News

🔬 Started a research internship at Agile Robots, working on world models for robotics and dexterous manipulation.

🎉 VBVR, a large-scale video reasoning suite, was accepted to ICML 2026.

More news

🎉 FUNCanon was accepted to ICRA 2026.

Research Interests

  1. Past

    Built foundations in visual localization, 3D object assets, and functional canonicalization for robotic manipulation.

  2. Now

    Studying open-world object representations and contact-aware video world models, including articulated interaction in explorable scenes.

  3. Next

    Connecting object-centric perception and generative world models to controllable physical interaction.

Embodied AI & Robotic Manipulation Object Canonicalization & Shape Representation Generative Worlds & Articulated Interaction Open-World 3D Perception & Object Pose Estimation

Selected Publications

* equal contribution, † corresponding author

PANY paper teaser: arbitrary references, 6D poses, and object-centric geometry
Open-World 3D Perception & Object Pose Estimation

Pose Anything Anywhere: Model-free Object Poses from Arbitrary References

PANY estimates the 6D pose of unseen objects from one or sparse arbitrary RGB or RGB-D references. Its multi-view geometry backbone and pose-graph registration aggregate unposed assist views to improve robustness under occlusion and wide viewpoint changes.

FUNCanon paper teaser: functional alignment, action primitives, and sim-to-real manipulation
Embodied AI & Robotic Manipulation

FUNCanon: Learning Pose-Aware Action Primitives via Functional Object Canonicalization for Generalizable Robotic Manipulation

FUNCanon decomposes long-horizon manipulation into actor–verb–object action chunks. Affordance-guided functional canonicalization aligns objects into shared functional frames, allowing pose-aware action primitives to transfer across object categories and tasks.

VBVR paper teaser: video reasoning tasks and evaluation results
Generative Worlds & Articulated Interaction

A Very Big Video Reasoning Suite

VBVR introduces a large-scale video reasoning suite with 200 curated tasks and over one million clips. Its rule-based, human-aligned benchmark provides verifiable evaluation of temporal, spatial, and causal reasoning in video models.

Experiences

Agile Robots logo
Agile WRD logo
Generative Worlds & Articulated Interaction

World Models for Robotics & Dexterous Manipulation

Research Intern · WRD Team, Agile Robots SE
  • Developing contact-aware video generation and world action models for contact-rich robotic manipulation.
  • Investigating 3D motion and contact representations for physically consistent interaction prediction.

Mentors

Mahdi Mustapha Hamad Tech Lead, Robot Learning Applications · Agile Robots SE
TUM Computer Vision Group banner
Technical University of Munich wordmark
Generative Worlds & Articulated Interaction

Interactive World Models for Understanding and Manipulating Articulated Objects

Master’s Thesis · TUM Computer Vision Group
  • Investigating how scene-level world models can support controllable interaction with articulated objects through object-level articulation reasoning and interaction video generation.
  • Exploring structured 3D articulation and motion guidance; evaluating motion quality, appearance preservation, and scene consistency.

Mentors

Xi Wang Junior Research Group Leader · TUM Computer Vision & AI
Daniel Cremers Professor & Chair · TUM Computer Vision & AI
TUM CAMP logo
Open-World 3D Perception & Object Pose Estimation

Open-World Object Pose & Canonical Representation

Research Collaboration · TUM CAMP
  • Contributed to Pose Anything Anywhere (PANY) during its 2026 submission phase, advancing open-world, model-free 6D pose research.
  • Worked on OV-NOCs in parallel through March 2026, focusing on 3D assets and the data annotation pipeline for canonicalized object representations.

Mentors

Junwen Huang PhD Candidate · TUM CAMP
Slobodan Ilic Adjunct Professor · TUM CAMP; Senior Research Scientist · Siemens
Benjamin Busam Professor & Chair · TUM Photogrammetry and Remote Sensing
TUM Info6 Knoll group mark
Agile Robots logo
Embodied AI & Robotic Manipulation

Functional Canonicalization for Robot Manipulation

Research Project · TUM Info6 / Agile Robots SE
  • Curated category-level 3D assets and pose-canonicalized models for FUNCanon.
  • Organized a functional taxonomy and built affordance visualizations for qualitative analysis.

Mentors

Zhenshan Bing Professor · TUM Info6
Alois Christian Knoll Professor & Chair · TUM Info6
Jianwei Zhang Professor & Head · University of Hamburg TAMS
TRESP Lab logo TRESP Lab
LMU Munich logo
Generative Worlds & Articulated Interaction

Enhancing Video Captioning via Reinforcement Learning

Guided Research · Tresp Lab, LMU Munich
  • Designed a composite reward for temporal alignment, linguistic fidelity, and brevity in highlight-aware video captioning.
  • Compared UVCOM/Lighthouse temporal predictors, clause-level aggregation, and GRPO versus DAPO training on QVHighlights.

Mentors

Ruotong Liao PhD Researcher · TRESP Lab, LMU Munich

Earlier Experience

Production Metrology & Quality Management

Mar – Jun 2022

Research Intern · WZL, RWTH Aachen University

Developed laser spot localization and supported multi-sensor axis-error detection.

Intelligent Driving

Jul – Sep 2021

Intern · SAIC VOLKSWAGEN

Built ORB-SLAM-based indoor perception and Ackermann lane-following; our team placed second in the AutoPro Smart Driving Competition.

Education

Technical University of Munich wordmark

M.Sc. Robotics, Cognition, Intelligence

Technical University of Munich

M.Sc. Robotics, Cognition, Intelligence

Munich, Germany · Master’s thesis: Interactive World Models for Understanding and Manipulating Articulated Objects (ongoing).

Tongji University wordmark

B.Eng. Mechatronics

Tongji University

B.Eng. Mechatronics

Bachelor thesis completed at WZL, RWTH Aachen University.

Selected Awards

Germany

DAAD Contact Scholarship for Foreign Students

Scholarship
Tongji University

Scholarship for Outstanding Social Engagement, Tongji University

First Class & Merit

Projects

OV-NOCs canonicalization pipeline for 3D models, category-level pose datasets, object-centric datasets, and scene datasets
Object Canonicalization & Shape Representation

OV-NOCs

Universal pose conventions and shape representations across open-vocabulary object categories, backed by a scalable canonicalization and annotation pipeline.

The project aims to bring uncategorized 3D models, category-level pose datasets, object-centric datasets, and scene datasets into a consistent canonical coordinate space. I worked on 3D asset preparation and the annotation pipeline for the large-scale canonicalized dataset at TUM CAMP.

Mentors

Junwen Huang PhD Candidate · TUM CAMP
Slobodan Ilic Adjunct Professor · TUM CAMP; Senior Research Scientist · Siemens
Benjamin Busam Professor & Chair · TUM Photogrammetry and Remote Sensing
Training pipeline from my guided research report: video and prompt, candidate captions, temporal and linguistic rewards, and GRPO or DAPO updates
Generative Worlds & Articulated Interaction

Enhancing Video Captioning via Reinforcement Learning

Highlight-aware video captioning · Guided Research at TRESP Lab, LMU Munich

Developed a GRPO training framework that rewards video captions for temporal alignment, linguistic fidelity, and brevity. The temporal reward combines moment retrieval and highlight detection signals from UVCOM/Lighthouse-style predictors; the report compares clause-level reward aggregation and DAPO on QVHighlights, with preliminary TVSum experiments.

Mentors

Ruotong Liao PhD Researcher · TRESP Lab, LMU Munich
Conceptual split view of a SLAM point-cloud map and an indoor drone trajectory plan
Open-World 3D Perception & Object Pose Estimation

Autonomous Drone Navigation with Visual-Inertial SLAM

Mobile Robotics Praktikum · Smart Robotics Lab (SRL), TUM

At TUM’s Smart Robotics Lab, I developed visual-inertial SLAM and trajectory-planning components in C++, then led ROS 1 sim-to-real deployment through control tuning, hardware integration, and mission testing. Our team was the only one to transfer autonomous planning from simulation to real flights; planner tuning and VIO re-initialization raised mission success from 44% to 62% and reduced crashes by 30%.

Mentors

Stefan Leutenegger Professor · TUM Smart Robotics Lab