Beyond GPUs: Networking and Systems Problems in Distributed LLM Serving
Raj Joshi
Red Hat and Harvard University
Pravein Govindan Kannan
IBM Research
Simone Ferlin
Red Hat and Karlstad University
Alex Brooks
Red Hat
LLM serving is rapidly becoming as ubiquitous as traditional web serving. Under the hood, large-scale distributed LLM inference depends on components that touch topics already familiar to the SIGCOMM community: request routing and scheduling, topology-aware placement, storage/cache management, networking and transport, tracing/observability, etc.; many of which can be studied without requiring GPU hardware. This interactive non-paper session will introduce the distributed LLM inference stack in a structured way (prefill vs. decode, disaggregation, batched vs. online serving), then use a live interactive board to drill down into individual components and brainstorm how classic networking/systems concepts and ideas apply. We will also give a brief tour of llm-d, a vendor-neutral open-source framework for distributed LLM serving, and close with practitioner perspectives on current industry pain points. Attendees will leave with a clear mental model of the distributed LLM serving stack, a practical open-source on-ramp via llm-d, and concrete directions for impactful contributions.
- Date and Time: Wednesday, Aug 19, 11:45, Track 1