1st Workshop on Memory-Semantic Networking for AI-Scale Systems (MemNetAI)
The workshop will take place at Room 503.
1:15 – 1:25 PM | Opening & Workshop Introduction — Satananda Burla |
1:25 – 2:05 PM | Keynote: Intelligent Scale-Up Domains for Machine Learning — Rachee Singh, Cornell University |
2:05 – 2:45 PM | Paper Session 1 — Session Chair: Muhammad Shahbaz (University of Michigan) Replacing NVMe Staging in LLM Inference with a High-Bandwidth CXL Memory Expander with an On-Device DMA Controller Veerasenareddy Burru, Pradeep Kumar Nalla, Alok Prasad Enabling Hardware-Software Co-Design for Large-Scale AI Systems via Fine-Grained GPU Memory Tracing Palak Mishra, Rajat Bhardwaj, Yash Verma, Ramanjeet Singh, Rinku Shah |
2:45 – 3:15 PM | Coffee Break |
3:15 – 3:50 PM | Paper Session 2 — Session Chair: Muhammad Shahbaz (University of Michigan) Beyond Monoliths: Enabling Flexible and Composable AI Systems via Memory Disaggregation Divya Kiran Kadiyala, Lianjie Cao, Jinsun Yoo, Puneet Sharma, Samantika Sury, Alexandros Daglis Memory as a First-Class Resource in AI-Factory Simulation Ulf R. Hanebutte, Nikhil Kumar John Stephen |
3:50 – 4:40 PM | Panel: Is Memory the New Network? Rethinking AI Infrastructure Shrijeet Mukherjee, NVIDIA · Rachit Agrawal, Cornell University · Prasun Kapoor, Marvell Technology · Rip Sohan, AMD |
4:40 – 4:45 PM | Closing Remarks — Abed Md Kamaluddin |
Intelligent Scale-Up Domains for Machine Learning
Rachee Singh — Cornell University
Rachee Singh is an Assistant Professor of Computer Science at Cornell University, where she leads the SysPhotonics group. Her group develops systems and algorithms for efficient communication over photonic interconnects, with applications to distributed machine learning and large-scale cloud workloads. She is also an Amazon Scholar with the AWS SageMaker HyperPod team, which develops large-scale ML infrastructure. Her research received the Outstanding Paper Award at OFC 2026 and faculty research awards from Google, Jane Street, LinkedIn, Amazon, and Cisco. Her doctoral work was recognized with the Google PhD Fellowship in Systems and Networking, the ACM SIGCOMM Doctoral Dissertation Award (runner-up), and the Outstanding Doctoral Dissertation Award from UMass Amherst.
Is Memory the New Network? Rethinking AI Infrastructure
As AI systems continue to scale, the defining challenges of AI infrastructure are increasingly moving beyond compute. This panel brings together perspectives from industry and academia to examine how evolving memory, networking, and system architectures will shape the next generation of AI infrastructure.
The discussion will explore whether memory is becoming a defining architectural element for AI systems, how new system abstractions and interconnects may emerge, and how simulation and cross-layer design can help guide the infrastructure of tomorrow.
Moderator: Abed Md Kamaluddin — Marvell Technology
Panelists
Shrijeet Mukherjee — NVIDIA. Shrijeet Mukherjee is VP of Engineering in NVIDIA's GPU Architecture team, focusing on connectivity and heterogeneous computing architectures. Previously, he was Co-Founder and CTO of Enfabrica, where he led the development of technologies for composable computer architectures. His career spans large-scale NUMA systems, GPUs, SmartNICs, DPUs, high-performance networking, and open networking. He serves on the Linux NetDev Society Board and holds more than 70 patents.
Rachit Agarwal — Cornell University. Rachit Agarwal is an Associate Professor of Computer Science at Cornell University, where he works on systems, networking, and algorithms. His research spans high-performance distributed systems, networking, and resource-efficient computing, with broad contributions to the design and analysis of large-scale computing infrastructure.
Prasun Kapoor — Marvell Technology. Prasun Kapoor is AVP, MBE Solutions at Marvell, where he leads the Solutions and Architecture team developing system-level solutions across enterprise networking, cloud, and data-center environments. His work includes data-center-scale simulation, intelligent analytics, autonomous network management, and memory and storage solutions aligned with real-world workloads.
Rip Sohan — AMD. Rip Sohan is an Engineering Fellow at AMD, where he leads end-to-end system architecture for high-performance enterprise and AI NICs. His work centers on hardware–software co-design across the entire networking stack—spanning NIC micro-architecture, firmware, kernel networking, and user-space acceleration. He specializes in RDMA, multipath transport, and emerging Ultra Ethernet Consortium (UEC) ecosystems.
The rapid scaling of foundation model training and AI inference is reshaping data center architectures. AI systems are increasingly limited not just by compute or packet-network bandwidth, but by memory capacity, bandwidth, and tail-latency sensitivity—making memory a primary bottleneck.
Modern AI servers span multiple communication tiers: accelerator-scale interconnects (e.g., NVLink-class fabrics) provide high-bandwidth intra-node communication, while cluster-wide coordination relies on RDMA over InfiniBand or Ethernet. Emerging technologies such as silicon photonics and CXL introduce switched, memory-semantic fabrics that enable rack-scale memory pooling and load/store communication across hosts. While server CPUs support CXL-based memory expansion, coherent memory-fabric integration across heterogeneous AI accelerators remains limited. This gap complicates scalable scale-up memory architectures and motivates research into interoperable memory-fabric designs.
This evolution creates a new architectural tier between backplane interconnects and data center networks. Memory traffic now traverses switched infrastructure and experiences network-like effects—congestion, contention, fairness, and failures—making memory performance dependent on scheduling and congestion control.
AI workloads amplify these challenges: training induces bursty synchronization and memory amplification; inference requires distributed KV cache capacity under strict tail latency constraints; and mixture-of-experts models create skewed, dynamic access patterns. As accelerator, memory, storage, and packet-network tiers increasingly interact, performance becomes tightly coupled across layers. Decisions in one tier—such as congestion control, flow scheduling, or memory placement—can cascade into latency amplification, throughput degradation, or instability in another.
Yet the networking community lacks a unified abstraction for memory-semantic fabrics, principled congestion models for load/store traffic over switched infrastructure, and comprehensive simulation and benchmarking methodologies tailored to AI-scale systems.
MemNetAI aims to define the networking principles of memory-semantic fabrics in AI-scale environments and to build a research community around this emerging frontier.
Topics of interest include, but are not limited to:
A. Scale-Up and Cross-Tier Fabric Interaction (Primary Focus)
- Interaction between accelerator-scale fabrics and rack-scale memory-semantic fabrics
- Congestion propagation and feedback across scale-up, memory, and network tiers
- Congestion control and fairness for load/store traffic
- Tail-latency amplification across fabric boundaries
- Coherence and consistency across heterogeneous fabric domains
B. Memory-Semantic Fabric Design
- Switched load/store fabrics and rack-scale memory pooling
- Credit-based flow control and congestion management
- Multi-tier memory fabrics extending beyond a single rack
- Failure domains, resilience, and recovery in pooled memory systems
- Telemetry and observability for memory-semantic fabrics
- Optical and photonic extensions of memory fabrics
- Compute-enabled memory fabrics and near-memory compute
C. AI-Driven Memory and Tiered Architectures
- Distributed KV-cache architectures
- Memory amplification and burst dynamics in large-scale training
- Elastic and disaggregated memory provisioning
- Tiered memory systems, including interaction with disaggregated storage tiers in AI-scale systems
D. Scheduling and Cross-Layer Control
- Memory placement and migration policies
- Coupling between cluster schedulers and fabric congestion
- Cross-layer coordination across compute, memory, storage, and network tiers
E. Simulation, Metrics, and Benchmark Standardization
- Load/store-aware simulation frameworks
- Cross-layer modeling of scale-up and memory fabrics
- Trace-driven AI workload generation
- Metrics for congestion sensitivity, fairness, and iteration-time impact
- Benchmark standardization efforts, including workload suites and reproducible evaluation methodologies for memory-semantic networking
- Pure DRAM device design
- GPU microarchitecture without fabric implications
- AI model optimization without networking relevance
- Storage-only disaggregation without memory semantics
We invite researchers and practitioners to submit original research papers, including position papers on disruptive ideas and early-stage work with potential to become full papers in the future.
Reviewing will be double-blind. Authors must make a good-faith effort to anonymize their submissions. Papers must not include author names and affiliations, and avoid implicitly disclosing the authors’ identity (e.g., via self-citation, funding acknowledgments).
We accept two types of submissions:
Regular research papers of up to 6 pages, excluding references and appendices. Submissions must be original, unpublished work, and not under consideration at another conference or journal. Authors of accepted submissions are expected to present their work at the workshop. Accepted submissions will be included in the workshop proceedings.
Extended abstracts of up to 2 pages, excluding references, in the same format as the regular papers. Submissions are for early-stage works and position papers that are still in progress, allowing authors to showcase their preliminary ideas and receive early-stage feedback at the workshop. The authors are expected to present their work as a lightning talk and/or a poster during the workshop. Authors of accepted submissions will have the option to opt out of including the submissions in the workshop proceedings.
Please submit your paper via https://memnetai26.hotcrp.com
Submissions must be printable PDF files. When creating your submission, you must use the _sigconf _proceedings template (two-column format, 10-pt font size) available on the official ACM site. LaTeX submissions should use the acmart.cls template (sigconf option), with the 10-pt font.
| Submission deadline | May 21, 2026 |
|---|---|
| Acceptance notification | June 6, 2026 |
| Camera-ready deadline | June 20, 2026 |
| Workshop date | August 17, 2026 |
| Organizers | Institution |
|---|---|
| Rinku Shah | IIIT Delhi |
| Praveen Tammana | IIT Hyderabad |
| Abed Mohammad Kamaluddin | Marvell Technology |
| Satananda Burla | Marvell Technology |
| Technical Program Committee | Institution |
| Abed Mohammad Kamaluddin | Marvell, India |
| Abhijit Das | IIT Hyderabad |
| Arnab Kumar Paul | BITS Goa |
| Divyanshu Saxena | UT Austin |
| Eric Spada | Broadcom |
| Jinsun Yoo | Georgia Institute of Technology |
| Michal Kalderon | Marvell, Israel |
| Mythili Vutukuru | IIT Bombay |
| Nathan Tallent | PNNL |
| Praveen Tammana | IIT Hyderabad |
| Priyanka Naik | IBM Research, India |
| Purushottam Kulkarni (Puru) | IIT Bombay |
| Ravichandra Mynidi | Marvell, India |
| Rinku Shah | IIIT Delhi |
| Satananda Burla | Marvell, USA |
| Sathya Peri | IIT Hyderabad |
| Senad Durakovic | Marvell, USA |
| Ulf Hanebutte | Marvell, USA |
| Annus Zulfiqar | University of Michigan, USA |
| Muhammad Shahbaz | University of Michigan, USA |
| Rip Sohan | AMD, USA |
| Prankur Gupta | Meta, USA |
| Gwangsun Kim | POSTECH, South Korea |
| Rahul Bothra | UIUC & Google |