
Sponsored Session: Lightning Talk: Optimizing Model Inference with PyTorch 2.0 - Devansh Ghatak
Sponsored Session: Lightning Talk: Optimizing Model Inference with PyTorch 2.0 - Devansh Ghatak, Simplismart
This session will explore how to maximize inference performance in PyTorch 2.0 by combining dynamic compilation and CUDA graph capture techniques. We will cover practical strategies including Quantization, Ahead-of-Time (AOT) compilation, and the use of custom fused operators all of which are essential tools for achieving low-latency, production-grade deployments.
This session will explore how to maximize inference performance in PyTorch 2.0 by combining dynamic compilation and CUDA graph capture techniques. We will cover practical strategies including Quantization, Ahead-of-Time (AOT) compilation, and the use of custom fused operators all of which are essential tools for achieving low-latency, production-grade deployments.
PyTorch
Welcome to the official PyTorch YouTube Channel. Learn about the latest PyTorch tutorials, new, and more.
PyTorch is an open source machine learning framework that is used by both researchers and developers to build, train, and deploy ML systems that so...
torch.compile and Diffusers: A Hands-On Guide to Peak Performance - Sayak Paul, Hugging Face
PyTorch
torch compile and Diffusers - A Hands On Guide to Peak Performance - PyTorch Compiler Series
PyTorch
verl: An Open Source Large Scale LLM RL Framework for Agentic Tasks - Yuxuan Tong, Bytedance
PyTorch
Lightning Talk: Sparsifying Vision Transformers with Minimal Accuracy Loss - Jesse Cai, Meta
PyTorch