Hello, I am Haochen Shi

I help models see, plan, and prove it.

Computer Science undergraduate at PolyU, working on vision-language grounding, LLM multi-agent systems, and benchmarks that make model claims easier to trust.

CLEF 2026Accepted working note, sole author
2 papersUnder review at ACL ARR
PolyU URISFunded project, student PI
263 / 4,094BirdCLEF+ 2026, Kaggle Bronze
3.81GPA, Dean's List
Research notebook

Questions I am investigating

How do models understand multimodal input, how do agents form executable plans, and how can we evaluate both reliably?

PolyU URIS fundedStudent PI, Sep 2025 to present

Multimodal LLM interaction for home robots

Under Prof. Wei Lou, I lead a project combining voice, gesture, and visual context. I fine-tuned Qwen2.5-VL-7B with DoRA on RefCOCO for visual grounding.

acc@0.5: 78.8% → 91.2%mIoU: 80.51.62 seconds per query
Accepted CLEF 2026 working noteSole author

Auditing Perch V2 in BirdCLEF+ 2026

I quantified a 31-class gap in eBird output coverage and used bootstrap, permutation, and TOST tests to find no reliable gain from hard per-class selection.

Rank 263 of 4,094234 scored classes31 unmapped classes
BirdCLEF+ 2026 Kaggle Bronze certificate, rank 263 of 4,094.
Under review at ACL ARRCo-author

SurveyLens: A discipline-aware benchmark for survey generation

I built SurveyLens-1k with 1,000 human-written surveys across 10 disciplines and designed its PDF parsing, filtering, and citation-normalization pipeline.

1,000 surveys10 disciplinesValid samples: 62% → 94%
Dataset construction, structured representation, and dual-lens evaluation.
Under review at ACL ARRCo-author

PersonalPlan: Multi-agent planning for programming learning

I implemented the plan-to-CrewAI runtime mapper and a four-axis evaluation harness, then validated Phase 1 GRPO reward training with Qwen3-8B and LoRA.

3,043 instances gatedTier-1 reward: 0.41 → 0.90156 training steps
Supervised fine-tuning, reward-adaptive GRPO, and multi-agent execution.
A little context

Curious about the score behind the score.

I am a second-year Computer Science undergraduate at The Hong Kong Polytechnic University. My interests sit between vision-language grounding, LLM multi-agent systems, and benchmark design.

I care about more than reaching a higher score. I want to know whether the data is clean, the evaluation is fair, and the reward actually represents the intended objective.

Outside research, I like turning ideas into tools for learning, paper discovery, financial analysis, and personal productivity.

The Hong Kong Polytechnic UniversityB.Sc. in Computer Science, 2024 to 2028, GPA 3.81, Dean's List
Toolkit

Tools I work with

LanguagesPython, TypeScript, Java, C++, Go, SQL
AI and MLPyTorch, Hugging Face, Qwen-VL, Qwen3, LoRA, DoRA, GRPO, CrewAI
FrontendReact, Next.js, Vue.js, React Native, Tailwind CSS
BackendFastAPI, Flask, Spring Boot, Node.js, Docker
DataPostgreSQL, MySQL, Redis, Supabase, ChromaDB
Say hello

Have a research problem worth exploring?

I am seeking a research internship in vision-language grounding, multimodal models, or LLM agents. Email is the simplest way to start a conversation.