Scanning
Robotic ultrasound scanning
VLMs can connect live ultrasound observations with probe state and scanning goals, allowing a robotic system to adapt its actions as visual evidence changes.
PhD student · RobUSt, CAMP · TUM
Building the VLM brain for medical robots.
I develop vision-language models as the intelligent brain of medical robots. My research asks how VLMs can build genuine understanding from multimodal observations and use that understanding as the foundation for planning adaptive robot actions.
I apply these ideas to medical ultrasound, from robotic scanning to ultrasound-guided interventions and surgery. An interactive AI avatar embodies this intelligence: it communicates with clinicians and patients while connecting the VLM's understanding and plans to robot execution.
I am a PhD student at RobUSt (Robotics and Ultrasound Team), Technical University of Munich, supervised by Prof. Nassir Navab and mentored by Dr. Yuan Bi. I completed my M.Sc. in Computer Science at TUM with distinction.
Research direction
The VLM sits between multimodal observations and robotic skills. It first builds a grounded understanding of what it observes; that understanding then becomes the basis for planning what the robot should do next.
Multimodal input
Intelligent VLM brain
Form a multimodal state from perception, anatomy, context, and evidence
Use that state to select robot skills, act, and replan from feedback
Good planning begins with genuine understanding.
Robot skills
Cross-cutting foundation
My earlier work on clinician-aligned evaluation informs a central question in this research: does the VLM genuinely understand what it observes, and can that understanding support effective planning in medical settings?
Medical ultrasound
Ultrasound makes understanding and planning inseparable: what the robot sees depends on how it moves, and every new observation can change the next action.
Scanning
VLMs can connect live ultrasound observations with probe state and scanning goals, allowing a robotic system to adapt its actions as visual evidence changes.
Surgery
In dynamic surgical settings, VLMs can combine imaging, context, and robot capabilities to support spatial reasoning, planning, and human coordination.
Recent news
Beyond Scalar Scores accepted to EMNLP 2026 Findings.
Joined the RobUSt team at CAMP, TUM, to work on intelligent VLMs for medical robotics.
Developing agentic metrics for medical generated report evaluation with clinical alignment, in collaboration with Jean-Philippe Corbeil at Microsoft Healthcare.
ReEvalMed accepted to EMNLP 2025 Main.
Reviewer for IEEE Transactions on Medical Imaging and volunteer at EMNLP 2025.
Selected work
My work on clinical alignment, model attention, medical representations, and planning shapes how I approach reliable VLM understanding and action today.
A clinician-validated meta-evaluation benchmark for radiology report metrics, revealing clinical misinterpretations and score-inflating practices in widely used metrics.
Post-trained Qwen-8B and MedGemma-4B with SFT/DPO + LoRA, outperforming 32B medical LLMs on clinical alignment while exposing evaluator bias and robustness failures.
Studies how to evaluate long-horizon surgical planning in safety-critical settings, separating visual grounding failures from planning failures and examining how structural knowledge helps constrain model behavior.
Training-free focus directions in key/query activations steer LLM attention toward task-relevant context, mitigating distraction across multiple LLM families.
Derived representations from UK Biobank k-space using masked autoencoders, enabling robust prediction from undersampled, human-unreadable frequency-domain data.
A multimodal model combining CT radiomics and clinical features for hepatic steatosis risk prediction, with metabolomics used to interpret imaging differences.
A weakly supervised parallel CNN architecture with spatial and cross-network attention mechanisms for fine-grained visual recognition.
A learning-based framework for sparse-scan longitudinal brain MRI registration and brain-state forecasting across aging trajectories.
Service
Contact
I am always happy to have a coffee chat about medical AI, robotics, or possible collaborations.