The 3rd Workshop on Networks for AI Computing (NAIC)
The workshop will take place at Room 403.
08:00 - 09:00 | Registration |
09:00 - 09:15 | Welcoming Remarks |
09:15 - 10:10 | Keynote 1 Speaker: Behnaz Arzani (Microsoft) Title: A Networking Researcher’s Encounters with the AI-Systems Flywheel |
10:15 - 10:45 | Morning Break |
10:45 - 11:35 | Session 1: Efficient and Distributed AI Inference Chair: Alan Liu (UMD) GPU-Centric Stateless LLM Serving With GIGANETS Xiaoliang Wang (Nanjing University); Zhenwei Pi, Rui Zhang (Tensorfer); Camtu Nguyen (Nanjing University) ARK: Avoiding Routing Collisions for KV Cache Transfer in Disaggregated LLM Inference Hung-Chun Lin, Ting-Wei Hsu, Chung-En Ho, Ahmed Saeed (Georgia Institute of Technology) Towards Network-Efficient Cross-Regional Inference via Learned Activation Compression Regan McDonald, Marilyn Rego (University of Michigan); Ertza Warraich (Purdue University); Annus Zulfiqar, Muhammad Shahbaz (University of Michigan) Budget-Adaptive Routing: Skipping the Weak When the Strong Answers Anyway Wei Geng (Technical University of Munich); Nitinder Mohan (TU Delft); Jörg Ott (Technical University of Munich) |
11:35 - 12:15 | Session 2: Observability and Emerging Workloads Chair: Maria Apostolaki (Princeton) RDMATracer: A scalable eBPF-based framework for tracing RDMA syscalls Prankur Gupta, Miao Xu, Maxim Samoylov, Prashanth Kannan, Rajiv Krishnamurthy (Meta); Theophilus A. Benson (Carnegie Mellon University) TrainSketch: Collision-Protected Switch Telemetry for Distributed LLM Training Flows Ao Li (Tsinghua University); Zhuochen Fan (Pengcheng Laboratory); Kaicheng Yang (Peking University); Zeyu Luan (Pengcheng Laboratory); Yong Jiang (Tsinghua University); Kejun Li, Qing Li (Pengcheng Laboratory) Towards Traffic Modelling of Multi-Agent Systems: The Role of Coordination Topology Davide Lamagna, Albert Cabellos (UPC, BarcelonaTech); Alberto Rodriguez-Natal (Cisco); Gábor Rétvári (Budapest University of Technology and Economics); Berta Serracanta (UPC, BarcelonaTech) |
12:15 - 13:15 | Lunch Break |
13:15 - 14:10 | Keynote 2 Speaker: Vyas Sekar (CMU) Title: Design Principles for Systems-meets-AI Abstract: We are entering an era where our ability to develop, deploy, and evolve AI-based systems and systems-for-AI has never been better. At the same time, our capacity to observe, manage, and secure these systems has never been worse. Rather than blindly drink the "AI kool-aid," we ask whether there is a more principled way to harness the power of AI in networked systems design. In this talk, we discuss high-level design requirements and design principles to inform our future efforts in AI-for-systems and systems-for-AI. Bio: Vyas Sekar is the Tan Family Chair Professor at Carnegie Mellon University, co-founder at Rockfish Data, and Chief Scientist at Conviva. He received his B.Tech from IIT Madras (President's Gold Medal) and his Ph.D. from CMU. His work has received several honors, including an NSF CAREER Award, the SIGCOMM Rising Star Award, best paper awards at SIGCOMM/CoNEXT/Multimedia, the NSA Science of Security Prize, and SIGCOMM Test of Time Award. |
14:15 - 14:45 | Session 3: AI-Aware Data and Media Pipelines Chair: Ran Ben Basat (UCL) GLENFINNAN: SmartNIC-Accelerated Data Processing for Efficient Vision AI Pipelines Mike Wong (Princeton University); Ulysses Butler (New York University); Emma Farkash (Princeton University); Praveen Tammana (Indian Institute of Technology Hyderabad); Anirudh Sivaraman (New York University); Ravi Netravali (Princeton University) Revisiting Video Streaming for Real-Time Conversational Generative AI Services Surendra Pathak, Mohamed Mohsin Shaik (George Mason University); Matteo Varvello (Nokia); Bo Han (George Mason University) |
14:45 - 15:15 | Afternoon Break |
15:15 - 16:05 | Session 4: Training Networks: From Fabrics to Global Deployment Chair: Danyang Zhuo (Duke) Dragonfly-Ultra: A Scalable, Low-Cost Network Architecture for High-Performance AI Clusters Rui Zhuang (China Mobile Research Institute); Hui Yuan (Huawei Technologies Co., Ltd.); Junye Zhang, Kefei Liu (China Mobile Research Institute); Jie Li (Huawei Technologies Co., Ltd.); Weiqiang Cheng (China Mobile Research Institute); Yuting Wu, Xinyuan Cao, Xiangyu Chen, Zixuan Guan (Huawei Technologies Co., Ltd.); Shengnan Yue (China Mobile Research Institute); Ruixue Wang (Research Institute of China Mobile Communications Co., Ltd Beijing China); Tong Yang (Peking University); Xiaolong Zheng (Huawei Technologies Co., Ltd.) Ampel: Scheduling at the Network Cut in ML Training Valerio Torsiello, Ayush Mishra, Sushovan Das, Lukas Röllin, Tommaso Bonato, Torsten Hoefler, Laurent Vanbever (ETH Zurich) CCSwitch: A Scalable Data Plane for Non-Blocking In-Network Collective Communication Sumukh Pinge (University of California, San Diego); Hardik Soni, Bob Lantz, Khaled Diab, Lianjie Cao (HPE Labs); Tajana Rosing (University of California, San Diego); Puneet Sharma (HPE Labs) Toward WAN-Aware LLM Training Across Heterogeneous, Geo-Distributed Sites (extended abstract) Ziyue Luo, Jiaxuan Cai, Cedric Le Denmat, Srijith Nair, Fatemeh Nourzad, Rohith Krishnan Sudha, Qinhang Wu (The Ohio State University); Jifan Zhang (University of Wisconsin Madison); Sungjae Lee (The Ohio State University); Zhe Li (Rochester Institute of Technology); Peiwen Qiu (The Ohio State University); Siddharth Shah (Carnegie Mellon University); Rishabh Sharma (University of Wisconsin Madison); Sundararajan Srinivasan (The University of Texas at Austin); Yinglun Xia, Xue Zheng (The Ohio State University); Zidong Liu (ComboCurve); Bicheng Ying (Google); Kaushik Chowdhury (The University of Texas Austin); Gauri Joshi (Carnegie Mellon University); Yingbin Liang (The Ohio State University); Robert Nowak (University of Wisconsin Madison); Srinivasan Parthasarathy (The Ohio State University); Saurav Prakash, Balaraman Ravindran (Indian Institute Of Technology-Madras); Sanjay Shakkottai (The University of Texas at Austin); Ness B. Shroff (The Ohio State University); Haibo Yang (Rochester Institute of Technology); Aylin Yener, Jia Liu (The Ohio State University) |
16:05 - 16:45 | Session 5: Adaptive, Reliable AI Networks BEACON: Benchmarking Adaptability of Hardware-Offloaded Congestion Control Algorithms Meet Dadhania, Ranjitha K, Saptarshi Samanta, Sneha Aravind, Hrushikesh J S (Indian Institute of Technology Hyderabad); Abed Mohammad Kamaluddin, Satananda Burla, Hemant Singh (Marvell Technology Inc.); Praveen Tammana (Indian Institute of Technology Hyderabad) GenCC: Heterogeneous Network Congestion Control using LLMs Neta Rozen-Schiff (TU Berlin); Liron Schiff (Akamai); Stefan Schmid (TU Berlin) Delphinus: Ultra-Fast Link Failure Detection and Recovery for AI Data Center Networks Junye Zhang (China Mobile Research Institute); Jie Li, Zhigang Ji, Hui Yuan (Huawei Technologies Co., Ltd.); Kefei Liu, Rui Zhuang, Ruixue Wang, Weiqiang Cheng (China Mobile Research Institute); Zixuan Guan, Xiaolong Zheng (Huawei Technologies Co., Ltd.) |
16:45 - 17:45 | Panel Discussion: Packets, Code, or Prompts: What Are We Building Next? Moderator: Alan Liu (UMD) Panelists:
|
Generative AI is transforming many aspects of modern society with content ranging from text and image to videos. The Large Language Models (LLMs) and other Artificial Intelligence (AI)/Machine Learning (ML) that enable these generative AI capabilities are placing an unprecedented amount of pressure on modern data centers with anecdotal evidence suggesting that the largest model can take months to train. To support these models, modern distributed training clusters contain tens of thousands of GPUs/TPUs, with many expecting the scale to further increase significantly.
More fundamentally, training these large models introduces network communication patterns that require sophisticated and novel topology, routing, and synchronization. As the adoption and use of such models grows, the data generated and the data required to train and make inferences will place emphasis on the design of novel network primitives. The scale, workload, and performance requirements require us to reconsider every layer of the network stacks and scrutinize the solution from a holistic perspective. The recent industry initiative, Ultra Ethernet Consortium (UEC), is actively working on Ethernet-based network optimizations for AI and HPC workloads. The Open Compute Project (OCP) is geared more toward infrastructure support for AI computing. The standard organizations (e.g., IETF) are also seeking opportunities in networking for AI computing. We believe the networking research community should take a bolder position and bring cutting-edge innovations in this front as well.
The workshop aims to bring together researchers and experts from academia and industry to share the latest research, trends, and challenges in cloud and data center networks for AI computing. We expect it to enrich our understanding of the AI workloads, communication patterns, and their impacts on networks, and help the community to identify future research directions. We encourage lively debate on issues like convergence vs disaggregation, front-end vs back-end, smart edges vs programmable core, and the need for new interconnection, new topology, new transport, and new routing algorithms and protocols.
Topics of interest include, but are not limited to:
- Technologies for RDMA and Ethernet efficiency, performance, security, and extensibility
- Load balancing for distributed learning
- Lossless and loss-tolerant network design
- Host and network integration and coordination
- New transport protocols and congestion control for AI training
- Programmable congestion control
- New network architecture and topologies for AI and HPC
- Offloading in SmartNIC/DPU, host hardware, switch
- Scale-out and scale-up network convergence
- Programmable networks for AI workload
- In-network computing techniques and protocols for distributed training and MPI support
- Application-aware networking for AI training and inference
- Collective communication optimization
- Networking for cross-DC learning
- Network optimization for inference
- Convergence of computing, storage, and networking
- Automated and intelligent AI DCN OAM
- LLM for DCN OAM
- Fault prediction, detection, and root cause analysis
- New measurement and telemetry metrics and methods
- Green data center for energy efficiency
- Traffic characterization for AI workload
- Network simulation and benchmarking for AI workloads
- Networking support for Agentic AI
We invite researchers and practitioners to submit original research papers, including position papers on disruptive ideas, and early-stage work with a potential for full papers in the future.
Reviewing will be double-blind. Authors must make a good faith effort to anonymize their submissions. Papers must not include author names and affiliations, and avoid implicitly disclosing the authors’ identity (e.g., via self-citation, funding acknowledgments).
We accept two types of submissions:
- Regular research papers of up to 6 pages, excluding references and appendices. Submissions must be original, unpublished work, and not under consideration at another conference or journal. Authors of accepted submissions are expected to present their work at the workshop. Accepted submissions will be included in the workshop proceedings.
- Extended abstracts of up to 2 pages, excluding references, in the same format as the regular papers. Submissions are about early-stage works and position papers that are still in progress, for authors to showcase their preliminary ideas to get early-stage feedback at the workshop. The authors are expected to present their work in the form of a lighting talk and/or poster during the workshop. Authors of accepted submissions will have the option to opt out from including the submissions in the workshop proceedings.
Please submit your paper via https://naic26.hotcrp.com/
Submissions must be printable PDF files. When creating your submission, you must use the sigconf proceedings template (two-column format, 10-pt font size) available on the official ACM site. LaTeX submissions should use the acmart.cls template (sigconf option), with the 10-pt font.
The NAIC workshop will feature a best paper award.
| Submission deadline | May 20, 2026 (AoE) |
|---|---|
| Acceptance notification | June 12, 2026 (AoE) |
| Camera-ready deadline | June 20, 2026 |
| Workshop date | August 17, 2026 |
| Organizers | Institution |
|---|---|
| Alan Liu | Maryland |
| Maria Apostolaki | Princeton |
| Danyang Zhuo | Duke |
| Name | Institution |
|---|---|
| Haseeb Ashfaq | NYU |
| Ran Ben Basat | UCL |
| Qizhe Cai | UVa |
| Xiaoqi Chen | Purdue |
| Kuntai Du | TensorMesh |
| Soudeh Ghorbani | JHU |
| Prankur Gupta | Meta |
| Marios Kogias | Imperial |
| Ming Liu | Wisconsin |
| Harsha Madhyastha | USC |
| Amedeo Sapio | NVIDIA |
| Wenfei Wu | PKU |
| Weitao Wang | |
| Jiarong Xing | Rice |
| Qiao Xiang | Xiamen U. |
| Annus Zulfiqar | UMich |
| Hong Zhang | Waterloo |
| Junxue Zhang | USTC |
| Yang Zhou | UC Davis |
| Name | Institution |
|---|---|
| Theophilus A. Benson | CMU |
| Torsten Hoefler | ETH Zurich |
| TV Lakshman | Nokia Bell Labs |
| Haoyu Song | Futurewei |
| Ying Zhang | Meta |
| Zhi-li Zhang | Minnesota |