MLLM-Energy
Graduate course research (ECE 382V) with a public paper and project site. We profiled InternVL3-8B serving on an NVIDIA A100 with vLLM and Nsight Compute, splitting every request into its four inference phases. At batch size 1, decode takes 88.7% of the phase-attributed runtime while sustaining only 14% SM and 23% DRAM throughput; output-heavy throughput and GPU power both saturate at about a third of the GPU's SMs; and going from batch size 1 to 16 is a 12.9x energy-per-request lever on output-heavy work but only 1.9x on input-heavy work. The measurement groundwork for energy-first serving.
Case study
Problem
Serving frameworks run every phase of a multimodal request (vision encoder, MLP connector, prefill, decode) under one fixed GPU configuration, even though the phases load the hardware very differently. We set out to measure how much of the GPU each phase actually uses, and whether the idle part could be turned into energy savings.
Approach
A four-person course team. We served InternVL3-8B through a patched vLLM on one A100, marked each phase with NVTX ranges, and captured per-kernel SM and DRAM throughput with Nsight Compute. We then swept the active SM count up to all 108 across batch sizes 1-16 for an input-heavy and an output-heavy workload, recording latency, throughput, power and energy per request: 1,080 configurations in all.
Evidence
From the paper: at batch size 1, decode is 88.7% of phase-attributed runtime and the other 11.3% is every remaining phase together (vision encoder 9.5%, prefill 1.8%, MLP connector under 1%). Output-heavy throughput levels off near 36 of 108 SMs at every batch size, and GPU power flattens at the same point: about 230 W for input-heavy work and output-heavy batches of 4 or more, 203-210 W for output-heavy at batch size 1, all under the 250 W TDP.
Limits
One model on one GPU, and NVTX phase times are CPU-observed windows rather than GPU-busy time. Masking SMs on a single job never beat the full GPU's energy per request; potential savings from frequency throttling or co-locating work remain untested.
Skills
- CUDA
- vLLM
- Nsight Compute
- GPU Profiling
- LLM Inference
- Energy Efficiency