The CPUs were doing all the work
Jellyfin was transcoding on CPU. It worked, but two Xeons chewing through a 1080p stream to send someone a 540p one is a waste of a machine that's also doing everything else. I put a Tesla P4 in the server to take that job.
The P4 is a nice card for this — single slot, no power connector, passively cooled, and it was cheap because nobody wants Pascal for AI work anymore. For video encoding it's still fine.
4.25x realtime
Confirmed in a real playback session, not just a settings page. A 1080p60 clip transcoding down to 540p ran at 263 fps, about 4.25x realtime, with both the decode and the encode on the GPU.
The thing worth knowing is that decode and encode are separate dedicated blocks on the card, not the general purpose cores. That matters because Immich runs its photo tagging on the same GPU and the two don't fight over anything except memory.
Some options exist but do nothing
ffmpeg lists encoders the P4's silicon doesn't actually have. Turning them on doesn't error — it silently falls back to CPU, which is the exact thing I installed the card to avoid. These stay off:
- AV1 encoding — needs an RTX 40 series card.
- AV1 decoding — needs RTX 30 or newer.
- 10-bit and 12-bit HEVC range extensions — also RTX 30 or newer.
What it does handle: H.264 both ways, HEVC both ways including 10-bit, and VP9 decode. That covers essentially everything in my library.