EnerTune Baselines

Team project · Research/Academic

Third author on EnerTune (SOSP '26), as an undergraduate research assistant in Yadwadkar's SysML Lab at UT Austin. I rebuilt four GPU-sharing serving systems (GPULets, Usher, FGD, ParvaGPU) from their papers and ported each onto our 16-A100 cluster, which is what made a fair comparison possible and enabled EnerTune's 1.4-2.3x lower energy result. My baseline profiling also fed its 7.3x average cut in profiling time.

gpuletsusherfgdparvagpurebuilt +ported16 × A1001.4–2.3× lower energy7.3× less profiling time (avg)
Schematic of the workflow, not a plot. The numbers under it are the paper's team results: energy against state-of-the-art baselines, profiling time against prior systems' own profiling methods.

Case study

Problem

GPU-sharing serving systems pack models onto shared GPUs to raise utilization. EnerTune argues that optimizing for utilization alone can raise energy use, and showing that fairly meant running the prior systems on the same cluster as EnerTune.

My part

Undergraduate research assistant and third author. I rebuilt the four comparison systems and debugged and ported each research prototype onto the 16-A100 cluster, then profiled the baselines. EnerTune itself, its analytical performance and power models and its energy-aware placement, is the paper's joint work.

Approach

The four systems cover different packing policies: Usher packs aggressively with MPS and replicates models, FGD packs greedily to minimize fragmentation, GPULets caps sharing at two models per GPU, and ParvaGPU runs MPS inside MIG slices. Each had to run on EnerTune's testbed and serve identical workloads, so the energy comparison is like for like.

Evidence

A team result, reported in the paper: EnerTune meets performance SLOs while using 1.4-2.3x less energy than state-of-the-art baselines, and its analytical model cuts profiling time 7.3x on average, up to 17.3x, versus the profiling prior systems rely on (brute-force search for ParvaGPU, kernel-level profiling for Usher and GPULets). The ports are public in the baselines repository; EnerTune's code is in the UT-SysML repository.

Limits

The baseline numbers come from our ports running on EnerTune's testbed and workloads, not from each system's original deployment.

Skills

  • GPU Scheduling
  • ML Inference Serving
  • Multi-Tenant GPUs
  • Systems Research
  • Python