/ work

Work

From MLU inference engines, scheduling, and cache systems to Triton/BangC kernels, real-time VLMs, and physical simulation—problems, ownership, and outcomes.

/ internships

Cambricon Technologies · internship

AI Infra Intern, Parallel Computing & Inference

Develop and maintain the vLLM_MLU and SGLang_MLU inference engines for Cambricon MLUs, spanning model enablement, scheduling and cache systems, kernel optimization, accuracy testing, and debugging.

Work and outcomes
  • Enabled and optimized Qwen3.5-and-later and KimiK models, including Prefix Caching, Chunked Prefill, and All-to-All; improved scheduling, refactored cache management, and supported migration from PyTorch to the native MLU backend.
  • Used Torch Profiler to inspect scheduling paths and performance traces, and GenCase to capture kernel shapes; fixed regressions, repeated output, and accuracy issues across long-context, multimodal, and agent workloads.
  • Enabled Qwen3.5 on SGLang_MLU through GDN linear-attention routing, scheduling, and Triton kernels, then integrated and validated MTP speculative decoding to reduce end-to-end generation latency.
  • Evaluated Beam Search for multi-candidate generation and implemented a baseline; developed BangC kernels and fixed issues in TransformerEngine training infrastructure.

vLLM_MLU · SGLang_MLU · Triton · BangC · Qwen3.5 · MTP

China Mobile Migu Video · internship

AI Algorithm Intern | World Cup AI Commentary VLM

Built a VLM project framework for World Cup live-broadcast scenarios, using Qwen3-VL-8B with vLLM streaming inference to support real-time video-stream understanding.

Work and outcomes
  • Integrated YOLO for player detection and developed a tactics-recognition framework for real-time tactical parsing.
  • Built standardized multimodal preprocessing and annotation pipelines; one related CCF-A paper is under submission.

Qwen3-VL-8B · vLLM · VLM · YOLO · Data Annotation

/ projects

Tencent · Interactive Entertainment Group · project

FacialRig with Simulation and Neural Networks

For digital-human facial driving, combined physical simulation, anatomy, and neural networks to build a workflow from implicit-network training and material-space reconstruction to FEM simulation.

Work and outcomes
  • Supported rapid fitting from a single image or 3D scan, and validated expression transfer, animation retargeting, and collision avoidance.
  • Improved realism and controllability of facial animation generation.

FacialRig · FEM · Implicit Network · Physics Simulation · Neural Networks

/ core skills

Inference Engines

vLLM · SGLang · MTP · Prefix Caching

Kernels & Hardware

Triton · CUDA · BangC · MLU

Models & Workloads

Qwen3.5 · KimiK · VLM · Agent

/ education

Fudan University

M.S. Candidate, Biomedical Engineering (AI x Biomedical Engineering)

Published one SCI Q1 journal article as first author and filed three invention patents as first inventor.

Fudan University

B.S., Theoretical and Applied Mechanics

Coursework included mathematical physics methods, ordinary differential equations, computational methods and software, and engineering mathematics.