Abstract:

The large-scale deployment of AI applications are bringing increasingly complex inference workloads to cloud datacenters. Supporting these workloads at scale challenges existing cloud infrastructure and calls for end-to-end optimizations tailored to the characteristics of AI workloads. In this talk, I will present our recent work toward scalable AI infrastructure across three layers: runtime optimization, cluster management, and application orchestration. Together, these efforts aim to build an end-to-end infrastructure stack that serves modern AI applications more efficiently and at lower cost.

 

Biography:

Minchen Yu is an Assistant Professor at the School of Data Science, The Chinese University of Hong Kong, Shenzhen. He received his Ph.D. from Hong Kong University of Science and Technology. His research interests cover cloud computing and distributed systems, with a recent focus on AI infrastructure and machine learning systems. His research has been published at various prestigious venues, such as NSDI, USENIX ATC, EuroSys, MLSys, ICML, INFOCOM, SoCC, TON, TACO, etc. His work has been applied in leading cloud platforms including Alibaba Cloud. He received the Best Paper Runner-Up Award at IEEE ICDCS 2021.