SCAI Research Symposium

Sustainable Compute & AI Infrastructure (SCAI) Research Symposium

Building sustainable AI infrastructure for a resource-constrained world

A symposium imagining AI infrastructure as a flexible and responsible participant in the broader public-infrastructure ecosystem, designed to respect the resource constraints of existing infrastructure and serve the public good.

SCAI Sustainable Compute and AI logo

Building on the NSF CoDec Expedition

SCAI grows from the UMass-led NSF CoDec Expedition in Computational Decarbonization, an interdisciplinary collaboration spanning six universities and bringing together expertise in theory, AI, systems, energy systems, the built environment, and economics. The symposium extends this foundation toward the coupled AI compute-and-energy infrastructure stack.

Visit NSF CoDec Back to top

Invited speakers/panelists

Erin Baker

Erin Baker

Panelist

Distinguished Professor of Mechanical and Industrial Engineering and Faculty Director of the Energy Transition Institute, UMass Amherst

Biography

Erin Baker is a Distinguished Professor of Mechanical and Industrial Engineering at UMass Amherst and Faculty Director of the Energy Transition Institute. Her research applies operations research and economics to decision-making under uncertainty in energy and the environment, with an emphasis on energy justice and publicly funded energy-technology research and development in the face of climate change.

Her modeling work addresses energy policy and planning across geographic and temporal scales and uses multiple parallel models to derive robust insights. She also studies the sustainability of electricity grids in New England and developing countries, as well as the environmental costs and benefits of offshore wind energy.

Mosharaf Chowdhury

Mosharaf Chowdhury

SpeakerPanelist

Associate Professor of Computer Science and Engineering, University of Michigan

Talk

Toward Energy-Optimal AI

Generative AI adoption and its energy consumption are skyrocketing. Training a single frontier model can consume tens of GWh, and inference is rapidly outpacing training in aggregate energy demand. This surge is inflating operational costs, and power delivery is now the gating factor for bringing new GPU capacity online.

In this talk, I will introduce the ML.ENERGY Initiative, our effort to understand and curtail AI's runaway energy demands on three fronts. First, understanding where energy goes: I will present tools to precisely measure AI energy consumption and findings from benchmarking open-weight models across hardware and serving configurations via the ML.ENERGY Leaderboard. Second, optimizing energy use: I will describe how identifying computations on and off the critical path in distributed training enables precise GPU frequency control, saving energy on non-critical work without slowing down training. Third, exposing tradeoffs: I will present how co-optimizing static and dynamic energy through better kernel scheduling reveals the Pareto frontier between energy and performance, enabling practitioners to make informed deployment decisions under diverse constraints. All tools and systems are open-sourced through the Zeus project. I will conclude with open problems in building energy-optimal AI systems.

Biography

Mosharaf Chowdhury is an Associate Professor of Computer Science and Engineering at the University of Michigan, Ann Arbor, where he leads the SymbioticLab. His research focuses on making AI/ML workloads more efficient, with a particular emphasis on reducing their energy consumption through the ML Energy Initiative. Major open-source projects from his team include Infiniswap, the first scalable memory disaggregation solution; FedScale, a planetary-scale AI/ML platform; TPP, a tiered memory manager integrated into the Linux kernel (v5.18+); and Zeus, the first energy-optimal generative AI stack. Previously, Mosharaf invented the concept of coflows and was one of the original creators of Apache Spark. He has received numerous individual honors, including fellowships and paper awards from NSDI, OSDI, ATC, and MICRO.

Ayse K. Coskun

Ayse K. Coskun

SpeakerPanelist

Professor of Electrical and Computer Engineering and Systems Engineering and Director of CISE, Boston University; Chief Scientist, Emerald AI

Talk

AI Data Centers as Flexible Grid Assets

This talk explores how the rapid growth of AI is transforming data centers into major power consumers—and how, with the right technologies, data centers can become valuable grid resources instead. The talk will discuss emerging approaches to making AI workloads "power-flexible" and enabling data centers to dynamically adjust consumption in response to grid conditions while meeting performance constraints. Drawing from real-world deployments, the talk highlights how grid-interactive data centers can accelerate interconnection and support a more resilient and affordable power grid.

Biography

Professor Ayse Coskun is the Chief Scientist at Emerald AI, a startup focused on enabling power flexibility in data centers. She is also a full professor in the Electrical and Computer Engineering Department at Boston University, where she leads the Center for Information and Systems Engineering and serves as Associate Dean for Research at the College of Engineering (currently on leave from BU).

Prof. Coskun is widely recognized as a leading academic authority on orchestrating data center power demand in response to power grid needs. Her broader research applies AI and machine learning to optimize cloud and high-performance computing systems. She has received many honors for her contributions, including the Ernest Kuh Award for energy-efficient system-level design and an IBM Faculty Award for applying AI-based methods in DevSecOps. Earlier in her career, Prof. Coskun worked in industry at Sun Microsystems (now Oracle). She recently served as Deputy Editor-in-Chief of IEEE Transactions on Computer-Aided Design and holds a PhD in Computer Science and Engineering from the University of California, San Diego.

John Goodhue

John Goodhue

Speaker

Executive Director, Massachusetts Green High Performance Computing Center

Talk

Massachusetts Green High Performance Computing Center

A talk introducing MGHPCC facilities, capabilities, and opportunities for research collaboration

Biography

John Goodhue is Executive Director of the Massachusetts Green High Performance Computing Center, an energy-efficient data center and consortium serving more than 20,000 researchers, students, and educators at Boston University, Harvard University, MIT, Northeastern University, the University of Massachusetts, and other institutions across the Northeast. His work focuses on regional and national collaboration around research-computing infrastructure, broadening access to advanced computing for researchers at small and mid-sized institutions, and developing a diverse next generation of computing professionals.

His research-infrastructure leadership includes the Northeast Storage Exchange, Open Storage Network, Northeast Big Data Hub, Eastern Regional Network, and Northeast Cyberteam. He also brings 30 years of industry experience in networking and high-performance computing, including technology leadership, engineering management, and general management roles at Cisco Systems and BBN Technologies, along with work on the management teams of several Boston-area startups. He holds a B.S. in Computer Science and Engineering from MIT.

Fabio Grimaldi

Fabio Grimaldi

SpeakerPanelist

Senior Sustainability Scientist, Amazon Web Services

Talk

The Sustainability Stack for AI at Scale: An Industry Perspective

AI infrastructure presents one of the single biggest opportunities of our time to align large-scale computing with sustainability goals — but realizing that opportunity requires scientific rigor, operational integration, and cross-industry collaboration that match the complexity of the underlying systems. From embodied carbon in accelerator chips to operational energy across globally distributed fleets, quantifying and reducing the environmental footprint of AI compute demands work across materials science, systems engineering, power systems, and policy. This talk describes how AWS approaches this challenge across interconnected layers. Measurement: building lifecycle assessment methodologies for AI hardware and operations, addressing data availability gaps, and working with service teams to quantify emissions from training and inference at the most granular level. Decarbonization of hardware and operations: making carbon metrics visible to service teams, integrating with financial and planning systems, co-owning reduction goals with infrastructure owners, and embedding sustainability into early-stage hardware and service design. Customer engagement: providing emissions reporting through public-facing tools and supporting responsible use of cloud and AI services. Cross-industry research and standardization: publishing methodology openly, contributing to Product Category Rule development, providing technical input to public policy teams on regulatory frameworks, and collaborating with partners across the full value chain — from manufacturers to academia to initiatives like SCAI — to advance the science of sustainable computing.

Biography

Fabio Grimaldi is a Senior Sustainability Scientist at AWS, where he leads the development of methodologies and large-scale models to track and improve the sustainability performance of AWS cloud services. He drives the adoption of sustainability metrics across AWS services, translating data into actionable insights that help AWS and its customers meet their sustainability goals. Fabio joined Amazon in 2022. He holds a PhD in Chemical Engineering from the University College London.

Laura Haas

Laura Haas

Panelist

Professor, Manning College of Information and Computer Sciences, UMass Amherst

Biography

Laura Haas is a Professor in the Manning College of Information and Computer Sciences at UMass Amherst. She joined UMass in 2017 after a distinguished career at IBM, where she was an IBM Fellow and held leadership roles including Director of the Accelerated Discovery Lab, Director of Computer Science at the Almaden Research Center, and head of IBM Research’s worldwide exploratory science program.

Her foundational contributions to database systems include the Starburst query processor, which became the basis for DB2 LUW; Garlic, an early system for integrating heterogeneous data sources; and Clio, the first semi-automatic tool for heterogeneous schema mapping. As the first permanent dean of Manning CICS, she led substantial growth in the college, expanded faculty and student diversity, oversaw the design and construction of a new academic building, and helped raise more than $100 million from public, industry, and philanthropic sources. Haas is an ACM Fellow, a member of the National Academy of Engineering and IBM Academy of Technology, and a Fellow of the American Academy of Arts and Sciences. She earned her Ph.D. in Computer Science from the University of Texas at Austin and her A.B. in Computer Science from Harvard University.

Deepak Rajagopal

Deepak Rajagopal

SpeakerPanelist

Professor and Co-Chair, Environmental Science and Engineering (D.Env.) Program, UCLA Institute of the Environment and Sustainability

Talk

The Challenges of Measuring AI's Environmental Footprint: An Industrial Ecology Perspective

We can increasingly measure what AI systems consume directly, but it remains difficult to determine their lifecycle environmental impacts. For instance, operational energy and water use can often be measured with relatively direct activity data and established emissions factors. Yet for major AI infrastructure providers, much of the reported greenhouse-gas footprint appears to lie upstream—in semiconductor fabrication, server manufacturing, construction, and other capital goods or downstream (end-of-life treatment) —where data and estimates tend to be incomplete and uncertain. Beyond the own lifecycle of AI systems, their may arise major impacts, both positive and negative, on the broader economy and society that tended to be even harder to isolate and quantify, The result is a measurement asymmetry: some of the impacts that are easiest to quantify might not necessarily be the largest and most uncertain.

Industrial ecology, life-cycle assessment, and energy economics provide mature conceptual frameworks for addressing such issues. They distinguish attributional from consequential effects, embodied from operational impacts, and direct from indirect and economy-wide rebound. They also provide a disciplined way to ask a question often neglected in AI sustainability debates: how far should the system boundary expand before additional precision ceases to justify its cost?

From a public policy standpoint, California’s SB 253 is a first-of-a-kind policy going into effect in 2027 which requires large firms doing business in California to disclose Scope 1, 2, and eventually Scope 3 GHG emissions under standardized rules and phased third-party assurance. Such regulations necessitate consistency, transparency, and accountability in measuring and reporting lifecycle emissions of AI. But disclosure alone will not eliminate uncertainty where underlying supply-chain data remain modeled or incomplete. The central challenge, therefore, is not simply more measurement but distinguish what is directly observed from what is estimated, make uncertainty explicit, and design reporting systems that improve decision-relevant accuracy without creating a false impression of precision.

Biography

Deepak Rajagopal is a Professor in the UCLA Institute of the Environment and Sustainability and Dept. of Urban Planning in the UCLA Luskin School of Public Affairs. His fields of research include Industrial Ecology and Life cycle assessment, applied economic analysis of energy and environmental policies. He is also a faculty Scientist in the Energy Analysis Division at the Lawrence Berkeley National Laboratory. He has a Ph.D. in Energy and Resources from UC Berkeley, MS degrees in Ag. and Resource Economics (UC Berkeley), and Mechanical Engineering (U. of Maryland, College Park) and B.Tech. in Mechanical Engineering (Indian Institute. of Technology, Madras). He has been a post-doctoral researcher at the Energy Biosciences Institute, UC Berkeley and also worked as a Structural Engineer at United Technologies Research Center, E.Hartford, Connecticut.

Shaolei Ren

Shaolei Ren

SpeakerPanelist

Professor of Electrical and Computer Engineering, University of California, Riverside

Talk

Powering AI in a Thirsty World

The rapid growth of artificial intelligence (AI) is driving the construction of gigawatt-scale data centers, placing increasing demands on both power grids and water infrastructure. Yet power and water are tightly coupled: water-intensive cooling can compete for local water resources and create risks to data center resilience, while waterless cooling can increase electricity demand and further stress the local grid. This tradeoff is particularly significant on the hottest days of the year, when both power and water systems are under stress. Despite these interdependencies, power and water are often planned and managed separately, overlooking important opportunities for coordination.

This talk explores how responsible AI infrastructure design and operations can address coupled power-water challenges. I will discuss water-aware computing and cooling, coordination with power systems, and approaches for managing resource tradeoffs while strengthening infrastructure resilience and reducing impacts on surrounding communities.

Biography

Shaolei Ren is a Professor of Electrical and Computer Engineering at the University of California, Riverside. His research broadly focuses on developing modeling frameworks, algorithms, and empirical methodologies to address challenges at the intersection of AI, computing systems, and communities. He is a recipient of the U.S. National Science Foundation CAREER Award (2015) and several paper awards, including at ACM e-Energy (2024, 2016) and IEEE ICC (2016). He received his Ph.D. degree from the University of California, Los Angeles.

Jeremy Rice

Jeremy Rice

SpeakerPanelist

Mechanical Systems Lead, Verrus

Talk

Direct and Indirect Energy Resources Enabling the Data Center as a Grid Asset

As data centers become increasingly integrated into the modern power grid, the demand for operational flexibility has never been greater. This talk explores the multi-faceted power and energy impacts of a flexible data center, specifically focusing on the complex interactions between direct energy resources, such as battery energy storage systems (BESS) and dynamic IT loads and indirect energy resources, such as water usage and flexible temperature interfaces. By examining these variables in concert, we can better understand how data centers can become grid assets, while maintaining the required availability of the IT workloads and constraining the use of the indirect energy resources.

Biography

Jeremy Rice, Ph.D., is a seasoned engineering leader with over two decades of experience spanning the "chip to chiller" stack. Currently, he serves as the Mechanical Systems Lead at Verrus LLC.

Prior to Verrus, Jeremy held a significant tenure at Google within their data center organization, where he focused on asset utilization, system simplification, and acted as a technical liaison between the IT and data center teams. He also brings extensive experience in IT-side hardware, having advanced the state of the art in both air and liquid cooling technologies.

Jeremy holds a Bachelor of Science and a Ph.D. in Mechanical Engineering from the University of Connecticut.

Yuanrui Sang

Yuanrui Sang

Speaker

Assistant Professor of Electrical and Computer Engineering, UMass Amherst

Talk

Flexible Data Centers Scheduling: Economic, Environmental, and Transmission Congestion Impacts

Simultaneously considering optimization of operating cost, greenhouse gas, and toxic emissions, this talk discusses a tri-objective, multi-period, power system-constrained framework to schedule flexible data center load. The framework Models data center power consumption as the sum of latency-critical and best-effort loads and considers the temporal flexibility of best-effort workload. The framework was implemented on standard power system test systems with data centers, and pareto fronts were obtained from the solutions. Trade-offs between different objectives are analyzed, and the impacts on electricity prices and system congestions were discussed.

Biography

Yuanrui Sang is an assistant professor in the Department of Electrical and Computer Engineering at the University of Massachusetts Amherst. Before joining UMass in 2024, she was an assistant professor at The University of Texas at El Paso, and she received her Ph.D. in electrical and computer engineering from The University of Utah in 2019. Her research interests include power system operation and planning, grid-enhancing technologies, and the integration of flexible load, such as data centers and electric vehicles, in power systems.

Prashant Shenoy

Prashant Shenoy

Speaker

Distinguished Professor of Computer Science and Director of the NSF CoDec Expedition, UMass Amherst

Talk

Data Centers, AI Workloads, and Efficiency: A Systems Perspective

The exponential growth of cloud computing has been a defining trend of our time, fueled by rapidly growing demands from online and data-intensive workloads. Despite the end of Denard scaling, the cloud's energy demand grew more slowly than expected over the past decade due to the aggressive implementation of energy-efficiency optimizations. However, the rise of AI workloads, which are often more resource-intensive than traditional cloud workloads, has led to rapid growth in data centers with power-hungry accelerators such as GPUs and TPUs, leading to a resurgence in the cloud's energy consumption and a strain on our electric grids.

In this talk, I will provide a systems perspective on the challenges and opportunities in enhancing the efficiency and sustainability of cloud platforms in the face of rising AI demand. I will discuss how resource management techniques such as workload shifting can enhance the efficiency of cloud platforms by exploiting the spatio-temporal variability in grid demand, energy availability, and electricity prices. I will present initial directions in making current computing systems grid-friendly and discuss approaches for navigating performance, efficiency, and cost tradeoffs that arise in their operations. I will end with several open research challenges that the research community needs to tackle to make AI-driven cloud platforms grid-friendly and ensure their continued growth.

Biography

Prashant Shenoy is currently a Distinguished Professor in the College of Information and Computer Sciences at the University of Massachusetts Amherst. He received the B.Tech degree in Computer Science and Engineering from the Indian Institute of Technology, Bombay and the M.S and Ph.D degrees in Computer Science from the University of Texas, Austin. His research interests lie in distributed systems and networking, with a recent emphasis on cloud and sustainable computing. He has been the recipient of several best paper awards at leading conferences, including two ACM Test of Time Awards. He is a fellow of the ACM, IEEE, AAAS, and AAIA.

Ramesh Sitaraman

Ramesh Sitaraman

Panelist

Distinguished Professor, Manning College of Information and Computer Sciences, UMass Amherst; Chief Consulting Scientist, Akamai Technologies

Biography

Ramesh Sitaraman is a Distinguished Professor in the Manning College of Information and Computer Sciences at UMass Amherst and Chief Consulting Scientist at Akamai Technologies. His research spans Internet-scale distributed systems, including algorithms, architectures, performance, energy efficiency, and user behavior. During his time in industry, he helped create the world’s first major content delivery network and pioneered distributed systems that deliver web content, video, applications, and online services to billions of users.

He is the founding director of UMass Amherst’s interdisciplinary Informatics undergraduate program and a Fellow of both ACM and IEEE. His honors include the inaugural ACM SIGCOMM Networking Systems Award for his work on the Akamai CDN, an Excellence in DASH Award for adaptive-bitrate algorithms used in commercial video streaming, the UMass Amherst Distinguished Teaching Award, and NSF Research Initiation and CAREER awards. He earned his Ph.D. in Computer Science from Princeton University and his B.Tech. from the Indian Institute of Technology Madras.

Karin Strauss

Karin Strauss

SpeakerPanelist

Innovation Strategist and Senior Principal Research Manager, Microsoft Research

Talk

AI Needs a Dose of Its Own Cure to Cut the Carbon. Let’s Do It!

As we ride the Cambrian explosion of AI, gains in the efficiency of resource use have become ever more important. They make the technology more accessible, enabling more models, features, products, and applications, increasing the value of AI. But as this community has pointed out, efficiency could backfire as a climate strategy: making AI more efficient could spur so much additional use that total consumption and absolute emissions might keep climbing. So if efficiency’s shadow twin, the availability of low carbon supply, is neglected, increasing value may come with rising environmental cost. After celebrating the progress this community has made on using resources efficiently, on carbon-aware computing, and on measuring embodied carbon, I will turn to increasing that low carbon supply of electricity and materials to build on, and I will share some of the work we are doing in this space. AI, so often seen as adding pressure against reaching net zero, can instead be a positive force to achieve it. Together, AI that makes resource use more efficient and AI that expands low carbon supply can drive a virtuous cycle, and this community can, and I will argue should, participate in both.

Biography

Karin Strauss is a Senior Principal Research Manager and Innovation Strategist at Microsoft Research and an Affiliate Professor in the Paul G. Allen School of Computer Science & Engineering at the University of Washington. Her work spans computer systems, synthetic biology and environmental sustainability, with research ranging from machine learning hardware and emerging memory technologies to biologically inspired computing. She is best known for pioneering DNA data storage systems, a project that received broad industry and media recognition. More recently, she has focused on making AI and IT infrastructure more sustainable.

Adam Wierman

Adam Wierman

SpeakerPanelist

Carl F Braun Professor of Computing and Mathematical Sciences, Caltech

Talk

Asset or Burden: Navigating the Community Impact of Data Centers

AI-driven data center growth imposes measurable externalities on local communities: noise, backup-generator emissions, water withdrawal, power quality degradation, and the potential for rising retail electricity prices. This talk characterizes what we can currently measure across each channel, then examines approaches for mitigating the community impacts and even providing benefits for community infrastructure using engineering, algorithmic, and policy levers such as siting, workload flexibility, energy and water storage, power-aware cooling, and policy levels.

Biography

Adam Wierman is the Carl F Braun Professor in the Department of Computing and Mathematical Sciences at Caltech. He received his Ph.D., M.Sc., and B.Sc. in Computer Science from Carnegie Mellon University. Adam’s research strives to make the networked systems that govern our world sustainable and resilient. He is best known for his work spearheading the design of algorithms for sustainable and community-centric data centers, including pioneering work on net-zero data centers, data center demand response, geographical load balancing, the public health impact of data centers, and the water usage of data centers. His work has seen significant industry adoption (e.g. through the startup Verrus). Additionally, he is well known for his work on heavy tails, including co-authoring a book on “The Fundamentals of Heavy Tails.” He is an ACM Fellow and an IEEE Fellow. He has received the ACM Sigmetrics Rising Star award, the ACM Sigmetrics Test of Time award, the IEEE INFOCOM Test of Time award, the IEEE Communications Society William R. Bennett Prize, the Caltech IDEA Advocate award, the Caltech GSC Excellence in Mentoring award, multiple teaching awards, and is a co-author of papers that have received “best paper” awards at conferences across computer science, energy systems, and operations research.

Le Xie

Le Xie

Panelist

Gordon McKay Professor of Electrical Engineering and Faculty Director of the Power and AI Initiative, Harvard University

Biography

Le Xie is the Gordon McKay Professor of Electrical Engineering at the Harvard John A. Paulson School of Engineering and Applied Sciences and Faculty Director of the Power and AI Initiative at Harvard SEAS. Before joining Harvard, he served on the faculty of Texas A&M University from 2010 to 2024. He earned his B.E. in Electrical Engineering from Tsinghua University, S.M. in Engineering Sciences from Harvard, and Ph.D. in Electrical and Computer Engineering from Carnegie Mellon University. His industry experience includes work at ISO New England and Edison Mission Energy Marketing and Trading.

His research interests include modeling and control in data-rich large-scale systems, the grid integration of clean-energy resources, and electricity markets. He is an IEEE Fellow and IEEE Power & Energy Society Distinguished Lecturer, and is the lead author of Data Science and Applications for Modern Power Systems.

Juncheng Yang

Juncheng Yang

Speaker

Assistant Professor of Computer Science, Harvard University

Talk

Rethinking Storage for Sustainable AI: From Models to Generated Data

The rapid growth of AI is creating a new storage sustainability challenge. Modern AI systems produce and retain enormous amounts of data—from billions of model checkpoints and fine-tuned variants to an ever-growing volume of AI-generated content. Yet today’s storage systems largely treat these objects as conventional byte streams, ignoring the rich structure and semantics introduced by AI workloads.

In this talk, I will present our recent work on rethinking storage systems for AI data. I will first discuss ZipLLM, which exploits relationships among models and combines model-aware compression with deduplication to substantially reduce the footprint of large model repositories. I will then present TensorDex, which pushes this idea further by treating tensors, rather than files or models, as first-class storage objects and exploiting relationships among tensors across an entire model ecosystem. Finally, I will discuss LatentStore, which revisits a more fundamental question for AI-generated data: do we need to store the generated object at all? By storing compact model-native representations and reconstructing data on demand, LatentStore trades increasingly inexpensive computation for reductions in long-term storage.

Together, these systems illustrate a broader opportunity: rather than applying traditional storage techniques directly to rapidly growing AI data, we can redesign the storage stack around the structure, semantics, and regenerability of AI workloads. I will conclude with a broader vision for sustainable AI storage, where computation and storage are jointly optimized to reduce the growing resource and environmental footprint of AI.

Biography

Juncheng Yang is an Assistant Professor in Harvard John A. Paulson School of Engineering and Applied Sciences. His research interests broadly cover the efficiency, performance, reliability, and sustainability of large-scale data systems and machine learning systems.

Juncheng's works have received best paper awards or honorable mention at VLDB'26, VALUETOOLS'24, NSDI'24, NSDI'21, SOSP'21, and SYSTOR'16. Juncheng was a Facebook Fellow, recognized as a Rising Star in machine learning and systems, and a Google Cloud Research Innovator. His dissertation on designing efficient and scalable cache management systems received the CMU SCS Doctoral Dissertation Award and the ACM SIGOPS Dennis M. Ritchie Doctoral Dissertation Award.

His works have been widely adopted. For example, S3-FIFO and SIEVE are adopted for production at hundreds of companies with more than 60 open-source libraries and packages in 18 programming languages. Moreover, his group maintains libCacheSim, the most popular cache simulation library, and freeinference, a free LLM inference service.

Minlan Yu

Minlan Yu

SpeakerPanelist

Gordon McKay Professor of Computer Science, Harvard University

Talk

Resilient AI Infrastructure: From the GPU to the Grid

Modern AI systems run on massive, costly infrastructure that must operate under three kinds of change: workload changes, as agentic request rates, output lengths, and tool-calling times vary continuously; infrastructure interruptions, as GPU and network failures, preemptions, and maintenance repeatedly halt training at scale; and power changes, as available grid power fluctuates. In this talk, I will present our recent work on resilient AI infrastructure that adapts to all three. For workload changes, we use internal LLM signals to schedule inference requests, from standalone inferences to agentic workflows, cutting latency and improving efficiency. For infrastructure interruptions, we introduce TrainMover, a resilient LLM training runtime which leverages elastic and standby machines to handle interruptions with minimal downtime. For power changes, we started the Harvard Power and AI Initiative to rethink the coordination between the power grid and AI infrastructure—making AI workloads flexible and making grid planning matching such flexibility.

Biography

Minlan Yu is a Gordon McKay professor at the Harvard School of Engineering and Applied Science. She’s the assistant director of the SRC/DARPA JUMP 2.0 ACE Center for Evolvable Computing, and the co-director for the Harvard power and AI initiative. She received her B.A. in computer science and mathematics from Peking University and her M.A. and PhD in computer science from Princeton University. She received the ACM-W rising star award, NSF CAREER award, and ACM SIGCOMM doctoral dissertation award. She served as PC co-chair for SIGCOMM, NSDI, HotNets, and several other conferences and workshops.

Golbon Zakeri

Golbon Zakeri

Panelist

Professor of Mechanical and Industrial Engineering, UMass Amherst

Biography

Golbon Zakeri is a Professor of Operations Research in the Department of Mechanical and Industrial Engineering at UMass Amherst. Her research develops analytics, economic models, and optimization methods for decision-making under uncertainty, with particular emphasis on electricity markets and power systems. She uses mathematical modeling to study policies and system designs that support efficient, reliable, resilient, and equitable energy procurement.

Before joining UMass Amherst, Zakeri was a faculty member at the University of Auckland, where she directed the Electric Power Optimization Centre, served as Deputy Director of the University of Auckland Energy Centre, and was President of the Operations Research Society of New Zealand from 2013 to 2017. Her prior experience also includes research at Argonne National Laboratory. She serves as an Area Editor for Energy and Environment at Operations Research, an editor of the INFORMS-Springer book series, and an associate editor for Computational Management Science. She earned her Ph.D. in Mathematics and Computer Science from the University of Wisconsin–Madison.

Participant and talk information reflects confirmations received to date and will be updated as additional details become available.

Back to top

The shared intellectual framework

One research agenda, three coupled layers

We envision AI infrastructure as a coupled cyber-physical system spanning three interdependent layers. In this vision, models and workloads expose flexibility, computing systems translate that flexibility into reliable operational decisions, and energy systems and public infrastructure shape the physical, economic, and societal conditions under which those decisions are made. This vision motivates a shared research agenda across three coupled layers: AI and workloads, computing systems, and energy and infrastructure.

AI and Workloads

How models and workloads can expose and use flexibility while preserving useful service and research outcomes.

  • Flexible training and inference through checkpointing, pause and resume, batching, model selection, and precision scaling
  • Trade-offs among accuracy, latency, throughput, availability, progress, cost, and resource consumption
  • Efficient model techniques, including compression, sparsity, quantization, and workload-aware optimization
  • Measurement and lifecycle analysis of energy, carbon, water, and material footprints

Computing Systems

How computing platforms can convert workload flexibility and changing resource conditions into reliable runtime action.

  • Scheduling, placement, provisioning, and admission control across accelerators, clusters, clouds, and edge platforms
  • Coordination across heterogeneous and geographically distributed computing resources
  • Telemetry, forecasting, and feedback control across compute, network, storage, cooling, thermal, and power conditions
  • Reliability, recovery, and controlled degradation that preserve application-level service objectives

Energy and Infrastructure

How large, dynamic AI loads interact with electric grids and the wider infrastructure systems that communities depend on.

  • AI load modeling and forecasting, including effects on transmission, distribution, power quality, and stability
  • Datacenter siting and interconnection, resource adequacy, capacity expansion, and coordinated infrastructure planning
  • Grid-responsive control, demand response, ramp management, curtailment and recovery, ancillary services, and on-site resources
  • Markets, tariffs, reliability, affordability, water, land, governance, and impacts on the public good
Coupling is the research problem.

The agenda runs in both directions: infrastructure constraints must inform model and systems design, while workload capabilities and service objectives must inform datacenter, grid, and public-infrastructure planning and control.

AI and machine learning Computer and distributed systems Power systems and control Public infrastructure, planning, and policy
Back to top
Draft · Subject to change

Symposium agenda

Two days of invited talks, panels, research highlights, poster presentations, and structured conversation across AI systems, data-center infrastructure, power systems, environmental impacts, and public priorities. Times, session titles, and participation may change as the program is finalized.

Agenda view

Day 1 — Thursday, September 17

Registration, breakfast, and informal networking
Welcome remarks

Mike Malone, Laura Vandenberg, and Sanjay Raman · Chair: Mohammad Hajiesmaili

Full session details
Mike MaloneVice Chancellor for Research and Engagement, UMass AmherstWelcome remarks
Laura VandenbergAssociate Vice Chancellor and Vice Provost for Research and Engagement; Professor of Environmental Health Sciences, UMass AmherstWelcome remarks
Sanjay RamanDaniel J. Riccio Jr. Dean of Engineering; Professor of Electrical and Computer Engineering, UMass AmherstWelcome remarks
Mohammad HajiesmailiAssociate Professor, Manning College of Information and Computer Sciences, UMass AmherstSession chair
Expeditions Keynotes

Adam Wierman and Prashant Shenoy · Chair: Ramesh Sitaraman

Full session details
Adam Wierman Carl F Braun Professor of Computing and Mathematical Sciences, Caltech Talk: Asset or Burden: Navigating the Community Impact of Data Centers

AI-driven data center growth imposes measurable externalities on local communities: noise, backup-generator emissions, water withdrawal, power quality degradation, and the potential for rising retail electricity prices. This talk characterizes what we can currently measure across each channel, then examines approaches for mitigating the community impacts and even providing benefits for community infrastructure using engineering, algorithmic, and policy levers such as siting, workload flexibility, energy and water storage, power-aware cooling, and policy levels.

Biography: Adam Wierman is the Carl F Braun Professor in the Department of Computing and Mathematical Sciences at Caltech. He received his Ph.D., M.Sc., and B.Sc. in Computer Science from Carnegie Mellon University. Adam’s research strives to make the networked systems that govern our world sustainable and resilient. He is best known for his work spearheading the design of algorithms for sustainable and community-centric data centers, including pioneering work on net-zero data centers, data center demand response, geographical load balancing, the public health impact of data centers, and the water usage of data centers. His work has seen significant industry adoption (e.g. through the startup Verrus). Additionally, he is well known for his work on heavy tails, including co-authoring a book on “The Fundamentals of Heavy Tails.” He is an ACM Fellow and an IEEE Fellow. He has received the ACM Sigmetrics Rising Star award, the ACM Sigmetrics Test of Time award, the IEEE INFOCOM Test of Time award, the IEEE Communications Society William R. Bennett Prize, the Caltech IDEA Advocate award, the Caltech GSC Excellence in Mentoring award, multiple teaching awards, and is a co-author of papers that have received “best paper” awards at conferences across computer science, energy systems, and operations research.

Prashant ShenoyDistinguished Professor of Computer Science and Director of the NSF CoDec Expedition, UMass AmherstTalk: Data Centers, AI Workloads, and Efficiency: A Systems Perspective
Ramesh SitaramanDistinguished Professor, Manning College of Information and Computer Sciences, UMass Amherst; Chief Consulting Scientist, Akamai TechnologiesSession chair
Break
Industry Session I

Karin Strauss and Fabio Grimaldi · Chair: David Irwin

Full session details
Karin StraussInnovation Strategist and Senior Principal Research Manager, Microsoft ResearchTalk: AI Needs a Dose of Its Own Cure to Cut the Carbon. Let’s Do It!
Fabio GrimaldiSenior Sustainability Scientist, Amazon Web ServicesTalk: The Sustainability Stack for AI at Scale: An Industry Perspective
David IrwinProfessor and Associate Department Head of Electrical and Computer Engineering; Adjunct Professor of Computer Science, UMass AmherstSession chair
Lunch
Afternoon welcome remarks

Brian Levine

Full session details
Brian LevineAssociate Dean of Research & Engagement; Distinguished Professor, Manning College of Information and Computer Sciences, UMass AmherstWelcome remarks
Panel I: What Should Academia Solve for the Future of AI Infrastructure?

Karin Strauss, Fabio Grimaldi, Adam Wierman, and Ramesh Sitaraman · Moderator: Laura Haas

Full session details
Karin StraussInnovation Strategist and Senior Principal Research Manager, Microsoft ResearchPanelist
Fabio GrimaldiSenior Sustainability Scientist, Amazon Web ServicesPanelist
Adam WiermanCarl F Braun Professor of Computing and Mathematical Sciences, CaltechPanelist
Ramesh SitaramanDistinguished Professor, Manning College of Information and Computer Sciences, UMass Amherst; Chief Consulting Scientist, Akamai TechnologiesPanelist
Laura HaasProfessor, Manning College of Information and Computer Sciences, UMass AmherstModerator
Faculty and emerging-researcher highlights

Faculty talks: Juncheng Yang and Yuanrui Sang · Job-market talks: Walid Abdelrahman Hanafy, Can Hankendi, Adam Lechowicz, Qingsong Liu, and Christopher Yeh · Chair: Mohammad Hajiesmaili

Full session details
Juncheng YangAssistant Professor of Computer Science, Harvard UniversityTalk: Rethinking Storage for Sustainable AI: From Models to Generated Data
Yuanrui SangAssistant Professor of Electrical and Computer Engineering, UMass AmherstTalk: Flexible Data Centers Scheduling: Economic, Environmental, and Transmission Congestion Impacts
Job-market talks
Walid Abdelrahman Hanafy UMass Amherst Talk: Flex: Grid-Responsive Provisioning and Scheduling for Elastic Cloud Clusters

The talk will explain the workload and temporal coupling inherent in carbon-aware resource provisioning and scheduling for data centers, and why effective management must account for (i) the cluster’s current and anticipated demand and its elasticity, (ii) exogenous grid signals and their variability, and (iii) the trade-off between delaying work and the potential savings enabled by waiting.

To address these challenges, I proposed Flex, a grid-responsive resource manager that jointly provisions cluster capacity and schedules elastic batch jobs. Flex addresses this coupling by computing optimal provisioning and scheduling decisions over recent historical conditions and reusing those decisions at runtime. I show that this approach provides an effective and practical method for grid-responsive management of elastic batch workloads

Can Hankendi Boston University Talk: PALS: Power-Aware LLM Serving for Grid-Interactive AI

Large-scale LLM inference is becoming a significant and increasingly dynamic data-center load, yet today’s serving systems largely treat GPU power as a fixed hardware constraint. This talk presents PALS, a power-aware LLM serving framework that makes GPU power a first-class runtime control knob. At runtime, PALS selects GPU power caps, batch sizes, and tensor-parallel configurations based on profiled power–performance tradeoffs, adapting the serving configuration as the available power budget changes. Implemented within vLLM, PALS shows how inference workloads can respond to changing power constraints while maintaining application-level performance and QoS. I will discuss results across dense and Mixture-of-Experts models and show how application-aware power management can connect LLM serving objectives with data-center and grid-level power requirements. More broadly, PALS illustrates how AI workloads can expose controllable flexibility rather than behaving as fixed electrical loads.

Adam Lechowicz University of Massachusetts Amherst Talk: Unlocking System Control Benefits of AI using Theoretical Modeling

AI and machine learning can improve decision-making in complex systems, but their unreliability remains a major barrier in settings where feasibility and worst-case guarantees matter. This lightning talk presents a perspective on how theoretical modeling can help bridge that gap. Focusing on online decision-making under uncertainty, I describe a “robust algorithm learning” approach: first analytically characterize a certificate set of algorithms that provably satisfy a robustness guarantee, such as a competitive-ratio bound; then use data-driven learning to optimize performance within that "safe" search space. This combines classical theoretical tools for identifying structural guarantees with modern learning methods that adapt algorithms to real problem instances. I illustrate the idea through our SIGMETRICS 2026 work on online smoothed demand management, where the approach is instantiated and evaluated. More broadly, the talk argues that theory can make AI-driven control schemes practical in systems such as data centers and power grids by constraining learning without giving up its performance benefits.

Qingsong Liu Caltech & UMass Amherst Talk: Decisions That Reshape the System: Closed-Loop Learning and Resource Allocation for Stateful AI Infrastructure

Modern computing systems must learn and allocate resources under uncertainty while meeting operational constraints. Yet their decisions often have persistent effects: configuration changes take time to settle, admitted workloads occupy capacity and shape future feedback, and repeated use can alter resource performance. These effects violate standard assumptions that feedback is immediate, resources are consumed only once, or system dynamics are fixed. My research develops algorithmic foundations with provable guarantees for such stateful online decision-making and translates them into closed-loop control for AI infrastructure. I will highlight three themes—convergence-aware learning, reusable capacity management, and deterioration-aware allocation—and show how they motivate a future agenda in capacity management for heterogeneous AI clusters, multi-timescale LLM serving, and agentic infrastructure.

Christopher Yeh Harvard University Talk: Online conformal risk control for energy applications

Integrating AI into modern energy systems requires ensuring safety, even under distribution shift. Online conformal risk control presents a promising approach to achieve long-run online safety guarantees including under distribution shift, but typically without accounting for decision costs. In this work, we demonstrate that the trade-off between decision costs and long-run risk control is naturally formulated as an instance of constrained online convex optimization (COCO) with long-term constraints: the safety loss defines the per-round constraint, while the decision loss defines the per-round objective. Building upon results from the COCO literature, we derive the first sublinear static regret guarantees for online conformal prediction, including in settings where the safety constraint functions are either convex or monotone. We demonstrate the utility of our approach on battery storage arbitrage settings.

Mohammad HajiesmailiAssociate Professor, Manning College of Information and Computer Sciences, UMass AmherstSession chair
Break and refreshments
Technical Session I: Responsible AI Infrastructure

Shaolei Ren and Deepak Rajagopal · Chair: Prashant Shenoy

Full session details
Shaolei RenProfessor of Electrical and Computer Engineering, University of California, RiversideTalk: Powering AI in a Thirsty World
Deepak RajagopalProfessor and Co-Chair, Environmental Science and Engineering Program, UCLA Institute of the Environment and SustainabilityTalk: The Challenges of Measuring AI’s Environmental Footprint: An Industrial Ecology Perspective
Prashant ShenoyDistinguished Professor of Computer Science and Director of the NSF CoDec Expedition, UMass AmherstSession chair
Panel II: Can AI Infrastructure Scale Responsibly? Impacts on the Grid, Water, and Communities

Le Xie, Shaolei Ren, Deepak Rajagopal, and Erin Baker · Moderator: Golbon Zakeri

Full session details
Le XieGordon McKay Professor of Electrical Engineering and Faculty Director of the Power and AI Initiative, Harvard UniversityPanelist
Shaolei RenProfessor of Electrical and Computer Engineering, University of California, RiversidePanelist
Deepak RajagopalProfessor and Co-Chair, Environmental Science and Engineering Program, UCLA Institute of the Environment and SustainabilityPanelist
Erin BakerDistinguished Professor of Mechanical and Industrial Engineering and Faculty Director of the Energy Transition Institute, UMass AmherstPanelist
Golbon ZakeriProfessor of Mechanical and Industrial Engineering, UMass AmherstModerator
Poster session
Dinner and networking

Day 2 — Friday, September 18

Light breakfast and arrival
Welcome remarks

James Allan and Caitlyn Butler

Full session details
James AllanSenior Associate Dean of Operations; Distinguished Professor, Manning College of Information and Computer Sciences, UMass AmherstWelcome remarks
Caitlyn ButlerAssociate Dean for Research and Graduate Affairs, Riccio College of Engineering; Professor of Civil and Environmental Engineering, UMass AmherstWelcome remarks
Industry Session II: Flexible Data Centers

Ayse K. Coskun, Jeremy Rice, and John Goodhue · Chair: Prashant Shenoy

Full session details
Ayse K. CoskunProfessor of Electrical and Computer Engineering and Systems Engineering; Director of the Center for Information and Systems Engineering, Boston University; Chief Scientist, Emerald AITalk: AI Data Centers as Flexible Grid Assets
Jeremy RiceMechanical Systems Lead, VerrusTalk: Direct and Indirect Energy Resources Enabling the Data Center as a Grid Asset
John GoodhueExecutive Director, Massachusetts Green High Performance Computing CenterTalk: Massachusetts Green High Performance Computing Center
Prashant ShenoyDistinguished Professor of Computer Science and Director of the NSF CoDec Expedition, UMass AmherstSession chair
Break
Technical Session II: Frontiers of AI Systems and Networking

Mosharaf Chowdhury and Minlan Yu · Chair: Ramesh Sitaraman

Full session details
Mosharaf ChowdhuryAssociate Professor of Computer Science and Engineering, University of MichiganTalk: Toward Energy-Optimal AI
Minlan YuGordon McKay Professor of Computer Science, Harvard UniversityTalk: Resilient AI Infrastructure: From the GPU to the Grid
Ramesh SitaramanDistinguished Professor, Manning College of Information and Computer Sciences, UMass Amherst; Chief Consulting Scientist, Akamai TechnologiesSession chair
Lunch and structured networking
Panel III: How Flexible Can AI Infrastructure Really Be?

Ayse K. Coskun, Jeremy Rice, Mosharaf Chowdhury, and Minlan Yu · Moderator: David Irwin

Full session details
Ayse K. CoskunProfessor of Electrical and Computer Engineering and Systems Engineering; Director of the Center for Information and Systems Engineering, Boston University; Chief Scientist, Emerald AIPanelist
Jeremy RiceMechanical Systems Lead, VerrusPanelist
Mosharaf ChowdhuryAssociate Professor of Computer Science and Engineering, University of MichiganPanelist
Minlan YuGordon McKay Professor of Computer Science, Harvard UniversityPanelist
David IrwinProfessor and Associate Department Head of Electrical and Computer Engineering; Adjunct Professor of Computer Science, UMass AmherstModerator
Additional speakers

Speakers TBD · Talks TBD

Closing remarks

Mohammad Hajiesmaili, Prashant Shenoy, Ramesh Sitaraman, and David Irwin

Full session details
Mohammad HajiesmailiAssociate Professor, Manning College of Information and Computer Sciences, UMass AmherstClosing remarks
Prashant ShenoyDistinguished Professor of Computer Science and Director of the NSF CoDec Expedition, UMass AmherstClosing remarks
Ramesh SitaramanDistinguished Professor, Manning College of Information and Computer Sciences, UMass Amherst; Chief Consulting Scientist, Akamai TechnologiesClosing remarks
David IrwinProfessor and Associate Department Head of Electrical and Computer Engineering; Adjunct Professor of Computer Science, UMass AmherstClosing remarks
Back to top

Participate in the symposium

Call for Posters

We invite poster submissions spanning AI, computing systems, sustainability, and public infrastructure.

Published or accepted work, work in progress, new research directions, systems, datasets, testbeds, demonstrations, and interdisciplinary research are welcome. The session is designed to foster exchange across research and practitioner communities.

Shared scope. Poster submissions may focus deeply on any one layer of the coupled research agenda or connect multiple layers. Cross-layer work that makes the relationship between AI, computing systems, and power or public infrastructure explicit is particularly encouraged.

Review the research agenda

Submission format

  • A title, author list, affiliations, and contact information
  • A brief abstract of up to 500 words describing the problem, motivation, approach, contribution, and current status of the work
  • The most relevant layer or layers of the coupled research agenda
  • For recently published or accepted work, the full citation and a link to the publication
  • For work in progress, a brief description of preliminary findings, anticipated contributions, or questions on which feedback would be valuable

The poster session will be non-archival and is intended to encourage feedback, exchange, and collaboration. Selected authors will be invited to present their posters in person at the SCAI Research Symposium.

Back to top

Contribute to the conversation

Call for Lightning Talks

New results, open questions, and emerging directions.

The SCAI Research Symposium invites submissions for focused, 5-minute talks that introduce research or important challenges to an interdisciplinary community and spark further conversation.

Shared scope. Talks may engage one layer of the coupled research agenda or connect several. Cross-layer questions, constraints, capabilities, and opportunities are especially welcome.

Review the research agenda

Who should submit

Senior graduate students and postdoctoral researchers on the job market, early-career faculty, and mid-career faculty working across AI, computing systems, power systems and control, or other public-infrastructure domains.

Published, accepted, ongoing, and preliminary work are all welcome; no paper is required.

What works well

  • A recent or ongoing result
  • An emerging direction, open problem, or provocative question
  • A system, dataset, testbed, measurement effort, or deployment lesson
  • A cross-layer technical or societal challenge

Keep the presentation focused rather than compressing a conventional conference talk.

Submission format

  • Talk title, speaker name, and affiliation
  • A brief description of the proposed talk, up to 150 words
  • The most relevant layer or layers of the coupled research agenda
  • A one-sentence answer to: What should the audience remember or discuss after your talk?

The session is non-archival. Participation in the SCAI Research Symposium is by invitation.

Back to top

Invitation-only symposium

The symposium will bring together invited researchers, infrastructure practitioners, utilities, policymakers, UMass campus leadership, state officials, and partners interested in the future of reliable, grid-aware AI systems. If you are interested in attending, please email Mohammad Hajiesmaili at hajiesmaili@cs.umass.edu.

View draft agenda Back to top