I am an integrated M.S.–Ph.D. student in the DAVIAN Lab at KAIST AI, advised by Professor Jaegul Choo. My research centers on Vision-Language-Action (VLA) models for robotics.
This focus grew out of a single thread running through my earlier work. I spent several years on multi-modal grounding — localizing sound sources in visual scenes and understanding video — learning to align perception across modalities in space and time, and on language grounding through LLM safety and information extraction. VLA is where these threads converge: grounding both perception and language to drive action in the real world.
Research Interests
2026
2025
2024
Integrated M.S.–Ph.D. in Artificial Intelligence
KAIST AI · DAVIAN Lab
Advisor: Prof. Jaegul Choo
Research focus: Vision-Language-Action, multimodal grounding, language grounding & safety
B.S. in Computer Science & Engineering
Kyung Hee University · Visual AI Lab
Advisor: Prof. Jung Uk Kim
Research focus: Multi-modal learning, Video Understanding, Sound Source Localization
AI Engineer
LETSUR (AI Startup)
Development of AI-based services utilizing LLMs and RAG.
AI Research Intern
ETRI (Electronics and Telecommunications Research Institute)
Research on Video Moment Retrieval and Highlight Detection.
Undergraduate Researcher
Kyung Hee University · Visual AI Lab
Supervisor: Jung Uk Kim.
3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance
[project]Object-aware Sound Source Localization via Audio-Visual Scene Understanding
[poster]Learning to Visually Localize Sound Sources from Mixtures without Prior Source Knowledge
[poster]Audio-Visual Spatial Integration and Recursive Attention for Robust Sound Source Localization
[poster]"My Experience of CVPR Acceptance as an Undergraduate and Advice for Graduate School"
[slides]