‹ Term

GPU rendering benchmark

Term’s terminal rendering moved to an EGL/GLES3 instanced renderer — here is the method, the data and the captures, measured on real hardware.

Written · Updated

4.3ms
GPU frame time · 235 fps
3.4×
vs. the CPU path
0.35ms
GL draw / frame
Results

Two measurements, one corpus, both through the built-in Log frame stats developer switch. The CPU column is the 2026-07 baseline on a Pura 70, kept because the CPU path no longer exists to re-measure; the GPU column is a fresh 2026-08 run on an HBN-AL80. Different phones, so read the speed-up as the order of the win rather than a figure to two decimal places.

Workload CPU path GPU path Speed-up
Full-screen scroll (firehose) 14.6 ms · 68 fps 4.26 ms · 235 fps 3.4×
Slow scroll (~5 rows/frame) 12.8 ms · 76 fps 5.05 ms · 198 fps 2.5×
logFrameStats · CPU column: Pura 70, 2026-07 · GPU column: HBN-AL80, 56×39 grid, 2026-08 · coloured ANSI-SGR corpus

GPU frame time is almost independent of how much scrolled — every frame is a full redraw, and that is cheap for a GPU. The CPU path’s sensitivity to the number of scrolled rows simply disappears.

GPU frame breakdown

About 1400 instanced quads per frame. Glyphs are rasterised into a texture atlas once, and every later frame only re-emits mesh instances.

Stage Average What it does
build 1.82 ms CPU side: walk the grid → instance quads + rasterise new glyphs into the atlas + upload
draw 0.35 ms GL instanced draw
swap 2.08 ms eglSwapBuffers (present)
Frame total 4.26 ms ≈ 235 fps
Re-measured 2026-08-29 on an HBN-AL80 · 56×39 grid at 1260×2312 · 1 440 frames of saturated full-screen scroll

Two 14 MB memory shuffles are gone: no more scroll-shifting the shadow bitmap (~6 ms) and no more copy into the dmabuf (~5.6 ms) — that ~11 ms is the entire saving. The GPU rasterises in place, straight onto the EGL surface.

Captures

The corpus is coloured ANSI-SGR mixed with Chinese, Korean and ASCII, written to the session’s TTY by the benchmark script — no need to touch the phone.

Full-screen firehose scroll: a full screen of coloured ANSI plus Chinese/Korean/ASCII scrolling continuously
P1 · Full-screen scroll — continuous scrolling, every frame a total redraw.
Idle: partial presentation frames driven only by the blinking cursor
P2 · Idle — only the cursor blinks; partial presentation.
Single-row updates: typing at 8 Hz with the rows above preserved pixel for pixel
P3 · Single-row update — typing at 8 Hz, rows above preserved pixel for pixel.
Benchmark script

The script automates all three phases (firehose scroll / idle blink / single-row update) by writing to the TTY of a session on the remote host — zero interaction on the phone.

gpubench.sh
#!/bin/bash
T="${1:-/dev/ttys000}"

# corpus: coloured ANSI-SGR + Chinese/Korean/ASCII, 400 rows
corpus() {
  i=0
  while [ $i -lt 400 ]; do
    printf "\033[3%dm%04d 端末描画性能測定 中文渲染基准 가나다라마 \033[1mBOLD\033[0m ascii-abcdefghijklmnop %05d\n" \
      $((i % 8)) $i $((i * 37))
    i=$((i + 1))
  done
}
C="$(corpus)"

# P1 firehose scroll, 30s (full frames, total damage)
printf "\n\033[36m=== GPUBENCH P1: firehose scroll 30s ===\033[0m\n" > "$T"
end=$((SECONDS + 30))
while [ $SECONDS -lt $end ]; do printf "%s\n" "$C" > "$T"; done

# P2 idle, 60s (blink only → partial presentation frames)
printf "\n\033[36m=== GPUBENCH P2: idle (blink only) 60s ===\033[0m\n" > "$T"
sleep 60

# P3 single-row update, 20s (~8 Hz, typing-shaped partial frames)
printf "\n\033[36m=== GPUBENCH P3: single-row updates 20s ===\033[0m\n" > "$T"
end=$((SECONDS + 20))
n=0
while [ $SECONDS -lt $end ]; do
  printf "\rtyping simulation %06d" $n > "$T"
  n=$((n + 1))
  sleep 0.12
done
printf "\n\033[36m=== GPUBENCH done ===\033[0m\n" > "$T"
How to reproduce
  1. Open a session to the target host; copy gpubench.sh there and find the session’s TTY with who -u.
  2. Run ./gpubench.sh /dev/ttysNNN (the session’s tty).
  3. Collect: hdc -t <target> shell hilog | grep "connrender: gpu" — one stats window every 60 frames.

Two things decide whether a run means anything. Pace the feed. A plain cat is not a sustained load — 200 000 lines / 28 MB was consumed in under a second here, and every window after that is an idle cursor blink at 2 fps; read those as “the renderer only manages 2 fps” and you have measured nothing. Offer a known rate instead and hold it. Watch the clocks. The same screen costs ~11 ms/frame when painted twice a second and ~4.3 ms under sustained load — 2.5×, purely DVFS. That 11 ms is the honest number for the first frame after a pause; the 4.3 ms is the honest number under load. Never compare across the two. Also keep the soft keyboard down: it halves the grid to 56×19 and flatters everything.