MLLM-Energy

Team project · Research/Academic

Graduate course research (ECE 382V) with a public paper and project site. We profiled InternVL3-8B serving on an NVIDIA A100 with vLLM and Nsight Compute, splitting every request into its four inference phases. At batch size 1, decode takes 88.7% of the phase-attributed runtime while sustaining only 14% SM and 23% DRAM throughput; output-heavy throughput and GPU power both saturate at about a third of the GPU's SMs; and going from batch size 1 to 16 is a 12.9x energy-per-request lever on output-heavy work but only 1.9x on input-heavy work. The measurement groundwork for energy-first serving.

non-decodedecode: 88.7% of time100%14% SM throughput
Schematic, not plotted data. Bar widths follow the phase shares at batch size 1: decode 88.7%, and the hatched 11.3% is all other phases together (vision encoder, MLP connector, prefill). Of the utilization line, only decode's 14% mean SM throughput is measured.

Case study

Problem

Serving frameworks run every phase of a multimodal request (vision encoder, MLP connector, prefill, decode) under one fixed GPU configuration, even though the phases load the hardware very differently. We set out to measure how much of the GPU each phase actually uses, and whether the idle part could be turned into energy savings.

Approach

A four-person course team. We served InternVL3-8B through a patched vLLM on one A100, marked each phase with NVTX ranges, and captured per-kernel SM and DRAM throughput with Nsight Compute. We then swept the active SM count up to all 108 across batch sizes 1-16 for an input-heavy and an output-heavy workload, recording latency, throughput, power and energy per request: 1,080 configurations in all.

Evidence

From the paper: at batch size 1, decode is 88.7% of phase-attributed runtime and the other 11.3% is every remaining phase together (vision encoder 9.5%, prefill 1.8%, MLP connector under 1%). Output-heavy throughput levels off near 36 of 108 SMs at every batch size, and GPU power flattens at the same point: about 230 W for input-heavy work and output-heavy batches of 4 or more, 203-210 W for output-heavy at batch size 1, all under the 250 W TDP.

Limits

One model on one GPU, and NVTX phase times are CPU-observed windows rather than GPU-busy time. Masking SMs on a single job never beat the full GPU's energy per request; potential savings from frequency throttling or co-locating work remain untested.

Skills

  • CUDA
  • vLLM
  • Nsight Compute
  • GPU Profiling
  • LLM Inference
  • Energy Efficiency