LLM planning for VLA
Thinking about how LLM agents can plan, code, and act as the decision layer for vision-language-action systems.
PhD student, Technical University of Munich
I work on reliable clinical evaluation, LLM-based planning for VLA and agentic coding, and multimodal medical AI. My research asks how better evaluation can shape better models, especially when clinical judgment and real-world constraints matter.
I am a PhD student at RobUSt (Robotics and Ultrasound Team), Technical University of Munich, supervised by Prof. Nassir Navab and mentored by Dr. Yuan Bi. I completed my M.Sc. in Computer Science at TUM with distinction.
Current focus
Thinking about how LLM agents can plan, code, and act as the decision layer for vision-language-action systems.
Building clinician-aligned evaluation signals that can expose model failures and guide better medical generation models.
Learning useful representations from report, WSI, CT/CTA, MRI, and raw medical signals that humans cannot easily inspect.
Recent news
Joined CAMP at TUM and started my PhD.
Developing agentic metrics for medical generated report evaluation with clinical alignment, in collaboration with Jean-Philippe Corbeil at Microsoft Healthcare.
ReEvalMed accepted to EMNLP 2025 Main.
Beyond Scalar Scores released for LLM-based clinical significance evaluation.
Reviewer for IEEE Transactions on Medical Imaging and volunteer at EMNLP 2025.
Research taste
I see evaluation as more than a leaderboard. When metrics reflect clinical judgment, they become useful training signals that pull models toward safer and more meaningful behavior.
I care about opening the black box: understanding where models attend, why they fail, and how structural knowledge can make their decisions easier to audit.
From cardiac MRI k-space to CT radiomics, I like using AI to mine raw medical data for signals that are present but hard for humans to directly interpret.
Selected work
A clinician-validated meta-evaluation benchmark for radiology report metrics, revealing clinical misinterpretations and score-inflating practices in widely used metrics.
Post-trained Qwen-8B and MedGemma-4B with SFT/DPO + LoRA, outperforming 32B medical LLMs on clinical alignment while exposing evaluator bias and robustness failures.
Studies how to evaluate long-horizon surgical planning in safety-critical settings, separating visual grounding failures from planning failures and examining how structural knowledge helps constrain model behavior.
Training-free focus directions in key/query activations steer LLM attention toward task-relevant context, mitigating distraction across multiple LLM families.
Derived representations from UK Biobank k-space using masked autoencoders, enabling robust prediction from undersampled, human-unreadable frequency-domain data.
A multimodal model combining CT radiomics and clinical features for hepatic steatosis risk prediction, with metabolomics used to interpret imaging differences.
A weakly supervised parallel CNN architecture with spatial and cross-network attention mechanisms for fine-grained visual recognition.
A learning-based framework for sparse-scan longitudinal brain MRI registration and brain-state forecasting across aging trajectories.
Service
Contact
I am happy to have coffe chat.