Cambricon Technologies
AI Infra Intern, Parallel Computing & Inference
AI Infra · LLM Inference & Systems Optimization
Fudan University · Shanghai / 复旦大学
Currently developing and maintaining LLM inference engines for Cambricon MLUs.
My work spans vLLM and SGLang model enablement, scheduling and cache systems, plus CUDA / Triton / BangC kernel optimization.
vLLM_MLU · SGLang_MLU
Scheduling · Cache · MTP
Triton · CUDA · BangC
AI Infra Intern, Parallel Computing & Inference
AI Algorithm Intern | World Cup AI Commentary VLM
FacialRig with Simulation and Neural Networks