Multimodal LLM interaction for home robots
Under Prof. Wei Lou, I lead a project combining voice, gesture, and visual context. I fine-tuned Qwen2.5-VL-7B with DoRA on RefCOCO for visual grounding.
Computer Science undergraduate at PolyU, working on vision-language grounding, LLM multi-agent systems, and benchmarks that make model claims easier to trust.
How do models understand multimodal input, how do agents form executable plans, and how can we evaluate both reliably?
Under Prof. Wei Lou, I lead a project combining voice, gesture, and visual context. I fine-tuned Qwen2.5-VL-7B with DoRA on RefCOCO for visual grounding.
I quantified a 31-class gap in eBird output coverage and used bootstrap, permutation, and TOST tests to find no reliable gain from hard per-class selection.

I built SurveyLens-1k with 1,000 human-written surveys across 10 disciplines and designed its PDF parsing, filtering, and citation-normalization pipeline.

I implemented the plan-to-CrewAI runtime mapper and a four-axis evaluation harness, then validated Phase 1 GRPO reward training with Qwen3-8B and LoRA.

I define the workflow and its boundaries first, then decide where a model is genuinely useful.
An AI mock-interview platform built from Flask, Express, Vue, LiveTalking, and a self-hosted Qwen3 model. I worked on authentication, data modules, VLM video feedback, TTS, and lip-sync.
Qwen3, Azure TTS, Wav2Lip, WebRTC, Vue, Flask, DockerGRE and TOEFL vocabulary practice with FSRS repetition, 12 adaptive question types, and three-tier answer judging.
Next.js 16, Cloudflare D1, PWA 02Upload a slide deck and get a coherent page-by-page explanation that carries context forward.
Next.js, FastAPI, Claude, Qwen-VL 03Deterministic DCF and ratio models run first, while the LLM writes only from the resulting numbers.
Python, Pandas, financial APIs 04An offline-first journal that turns prose plans into calendar events with deterministic reconciliation.
React Native, Expo, Supabase 05My personal paper radar for ACL, NeurIPS, ICLR, and CVPR, with keyword and embedding filters.
Python, arXiv API 06Minimax with alpha-beta pruning and a live D3 game tree streamed over WebSocket.
Java, Spring Boot, React, D3I am a second-year Computer Science undergraduate at The Hong Kong Polytechnic University. My interests sit between vision-language grounding, LLM multi-agent systems, and benchmark design.
I care about more than reaching a higher score. I want to know whether the data is clean, the evaluation is fair, and the reward actually represents the intended objective.
Outside research, I like turning ideas into tools for learning, paper discovery, financial analysis, and personal productivity.

I am seeking a research internship in vision-language grounding, multimodal models, or LLM agents. Email is the simplest way to start a conversation.
haochen8.shi@connect.polyu.hk