cv
Basics
| Name | Yao Fu |
| Label | Deep Learning Engineer at NVIDIA |
| yaof@nvidia.com | |
| Url | https://future-xy.github.io |
| Summary | AI systems researcher focused on generative AI infrastructure, LLM inference, and distributed systems. Ph.D. in Computer Science from the University of Edinburgh, with research experience at NVIDIA, Genmo AI, Microsoft Research Asia, and Tencent, and work published at OSDI, NeurIPS, and JMLR. MLCommons Rising Star, 2024. |
Education
Publications
-
2025 MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems
NeurIPS 2025 Datasets and Benchmarks Track
Projects
- 2022.04 - 2026.01
ServerlessLLM
Locality-Enhanced Serverless Inference for Large Language Models
- Developed loading-optimized checkpoint format achieving 4X faster loading than Safetensors
- Designed live-migration mechanism delivering 2X performance improvement over serverless scheduling
- Built locality-aware server allocation reducing start-up latency by 1.86X
- Achieved 10-200X speedup compared to SOTA systems (Ray Serve, KServe)
- 700+ GitHub stars
- 2022.04 - 2025.03
MoE-Infinity
Activation-Aware Expert Offloading for Efficient MoE Serving
- Created Python binding compatible with HuggingFace Transformers
- Built OpenAI-compatible API server for high-throughput serving
- Optimized tensor prefetching and caching for Mixture of Experts models
- Delivered 9X performance improvement over DeepSpeed Infinity
- 300+ GitHub stars
- 2022.01 - 2023.05
Machine Learning Systems: Design and Implementation
Textbook; authored chapters on Deep Learning Recommendation Systems and Federated Learning
- Chinese Edition: ISBN 978-7-30-263007-4, openmlsys.github.io
- English Edition: openmlsys.github.io/html-en
- 4000+ GitHub stars (openmlsys/openmlsys-zh)
-
Open-MoE-LLM-Leaderboard
A new benchmark to assess performance, cost, and quality of sparse MoE models and systems
- Benchmark
- MoE
- Sparse Models
Languages
| Chinese | |
| Native speaker |
| English | |
| Fluent |