Multimodal LLM interaction for home robots
Under Prof. Wei Lou, I lead a project combining voice, gesture, and visual context. I fine-tuned Qwen2.5-VL-7B with DoRA on RefCOCO for visual grounding.
I am a third-year Computer Science undergraduate at PolyU. I work on embodied intelligence, 3D scene graphs, and multimodal systems that connect language with structured 3D environments.
I am currently studying how embodied agents can turn 3D perception into structured scene graphs and use those relations for grounded reasoning.
Under Prof. Wei Lou, I lead a project combining voice, gesture, and visual context. I fine-tuned Qwen2.5-VL-7B with DoRA on RefCOCO for visual grounding.
I quantified a 31-class gap in eBird output coverage and used bootstrap, permutation, and TOST tests to find no reliable gain from hard per-class selection.
I built SurveyLens-1k with 1,000 human-written surveys across 10 disciplines and designed its PDF parsing, filtering, and citation-normalization pipeline.

I implemented the plan-to-CrewAI runtime mapper and a four-axis evaluation harness, then validated Phase 1 GRPO reward training with Qwen3-8B and LoRA.

Two 2026 bronze-medal finishes across bioacoustics and game simulation.
I placed 263rd of 4,094 teams in the BirdCLEF+ 2026 acoustic species identification competition.

Our team placed 659th of 6,807 teams in the 2026 competition and received a Kaggle bronze medal.

I am a third-year Computer Science undergraduate at The Hong Kong Polytechnic University.
My current research is centered on embodied intelligence and 3D scene graphs. I am interested in how agents can organize objects, relations, and language into a useful model of a 3D environment.
Alongside research, I maintain several learning and research tools. Their live links are collected at home.shc66.com.

If you are working on embodied intelligence, 3D scene graphs, or multimodal 3D reasoning, feel free to email me.
haochen8.shi@connect.polyu.hk