
Optimizing LLM Inference: Disaggregated Serving, PD Protocol, & KV Pinning
At PyTorch Conference North America 2026, Nicolò Lucchesi, Research Engineer at Mistral AI and vLLM maintainer, will present joint work with Amazon Web Services (AWS) and Red Hat on how disaggregated serving in vLLM has evolved to support the latest generation of hybrid models.
Join us in San Jose on October 20-21: https://hubs.la/Q04v4SL60
Join us in San Jose on October 20-21: https://hubs.la/Q04v4SL60
PyTorch
Welcome to the official PyTorch YouTube Channel. Learn about the latest PyTorch tutorials, new, and more.
PyTorch is an open source machine learning framework that is used by both researchers and developers to build, train, and deploy ML systems that so...
torch.compile and Diffusers: A Hands-On Guide to Peak Performance - Sayak Paul, Hugging Face
PyTorch
torch compile and Diffusers - A Hands On Guide to Peak Performance - PyTorch Compiler Series
PyTorch
verl: An Open Source Large Scale LLM RL Framework for Agentic Tasks - Yuxuan Tong, Bytedance
PyTorch
Lightning Talk: Sparsifying Vision Transformers with Minimal Accuracy Loss - Jesse Cai, Meta
PyTorch