Mali GPU with LiteRT-LM

Share
Mali GPU with LiteRT-LM

Running Gemma 4 E2B on a Mali GPU with LiteRT-LM

This guide should work on any armbian minimal/console (trixie) system with a working Mali-g610 GPU.

DO NOT USE DESKTOP

Verified on: Orange Pi 5 Max (Rockchip RK3588, Arm Mali-G610 MC4), Armbian/Debian 13 (trixie), aarch64, 8 GB RAM.

Status summary: Both CPU (XNNPACK) and GPU (WebGPU/Dawn → Vulkan → panvk → panthor) work on the mainline panthor kernel — numbers below. The GPU path needed one line of local patching in
Mesa's panvk driver (raise maxImageDimension3D 512 → 2048 to meet the WebGPU minimums that LiteRT's Dawn library enforces; recipe in §5, rationale in §7.1). The model is multimodal (Text + Vision + Audio): the vision encoder runs on the same patched panvk via--vision-backend gpu (§5 "VL path" results). The vendor Rockchip libmali route and OpenCL remain unavailable on mainline panthor (§7.2, §7.3).

Stack (as shipped):

Gemma 4 E2B (.litertlm, int4)  →  LiteRT-LM CLI  →  LiteRT GPU accelerator (WebGPU/ML Drift)

Dawn (WebGPU impl) → Vulkan → panvk → panthor (kernel) → Mali-G610

1. Prerequisites

  • ARM (aarch64) Linux with a Mali GPU and at least 8 GB RAM.
  • Armbian/Debian 13 (trixie)
  • The panthor (or panfrost for older Mali) kernel driver active:
ls /dev/dri/renderD128          # render node must exist
lsmod | grep panthor            # or panfrost / lima

Full ARMv8.2-A dotprod support is required — LiteRT's ARM64 binaries are built with it and will SIGILL otherwise:lscpu | grep -w asimddp (both A55 and A76 cores list it here).

2. Install Mesa + Vulkan for Mali (needs root)

sudo apt update
# Prefer the newest Mesa (backports on Debian) for the best panvk coverage of Valhall CSF GPUs:
sudo apt install -y -t trixie-backports mesa-vulkan-drivers libvulkan1 vulkan-tools
# Optional (see §7 — currently yields NO usable OpenCL device on panthor):
sudo apt install -y -t trixie-backports mesa-opencl-icd clinfo

Verify the GPU is visible to Vulkan:

vulkaninfo --summary | grep -iE "deviceName|driverName"
# Expect: deviceName = Mali-G610 MC4   driverName = panvk

3. Install the LiteRT-LM CLI (user space, no root)

curl -LsSf https://astral.sh/uv/install.sh | sh
export PATH="$HOME/.local/bin:$PATH"
uv tool install litert-lm        # current: v0.17.x; incl. CPU (XNNPACK/YNNPACK) + GPU (WebGPU/Vulkan) backends

4. Get the model (Gemma 4 E2B, int4 .litertlm, 2.58 GB)

mkdir -p ~/litert-lm-bench && cd ~/litert-lm-bench
curl -L -o gemma-4-E2B-it.litertlm \
  https://huggingface.co/litert-community/gemma-4-E2B-it-litert-lm/resolve/main/gemma-4-E2B-it.litertlm
# or have the CLI fetch it for you:
#   litert-lm benchmark --from-huggingface-repo litert-community/gemma-4-E2B-it-litert-lm gemma-4-E2B-it.litertlm

Other ready models: google/gemma-3n-E2B-it-litert-lm (gemma-3n-E2B-it-int4.litertlm),litert-community/gemma-4-E4B-it-litert-lm, ... (see litert-lm list / HF).

5. Benchmark

Lock the CPU governor to performance first (fair, repeatable numbers):

sudo sh -c 'for c in /sys/devices/system/cpu/cpu[0-7]/cpufreq/scaling_governor; do echo performance > $c; done'

CPU baseline (XNNPACK, 8 threads):

litert-lm benchmark ~/litert-lm-bench/gemma-4-E2B-it.litertlm \
  -p 1024 -d 256 --backend cpu --cpu-thread-count 8 --cache disk --runs 2

GPU (patched panvk, §7.1). One-time local Mesa build — stays entirely in userland, the system Mesa is untouched:

# 1) Mesa source + one-line patch (26.1.2)
cd ~/litert-lm-bench
curl -L -o mesa.tar.xz https://archive.mesa3d.org/mesa-26.1.2.tar.xz
tar -xf mesa.tar.xz && cd mesa-26.1.2
python3 - <<'PY'
p = 'src/panfrost/vulkan/panvk_vX_physical_device.c'
s = open(p).read()
s = s.replace('.maxImageDimension3D = PAN_ARCH <= 10 ? (1 << 9) : (1 << 14),',
              '.maxImageDimension3D = (1 << 11), /* 2048 — meets WebGPU min */')
open(p, 'w').write(s)
PY

# 2) build deps (the LLVM/CLC chain is required — panvk bakes CLC-compiled SPIR-V for libpan)
sudo apt install -y --no-install-recommends meson ninja-build pkg-config gcc g++ python3-mako \
  python3-yaml python3-ply libdrm-dev llvm-19-dev libllvmspirvlib-19-dev spirv-tools \
  libclang-19-dev libclang-cpp19-dev

# 3) panvk-only build (~10–25 min on the 5 Max)
meson setup build-panvk -Dvulkan-drivers=panfrost -Dgallium-drivers= -Dbuildtype=release \
  -Dllvm=enabled -Dmesa-clc=auto -Dplatforms= -Degl=disabled -Dgbm=disabled -Dglx=disabled \
  -Dopengl=false -Dglvnd=disabled -Dtools= -Dbuild-tests=false
ninja -C build-panvk -j6

# 4) benchmark with the patched driver, process-scoped. First put the WebGPU prebuilts
#    (libLiteRtTopKWebGpuSampler.so + libwebgpu_dawn.so etc., litert-lm repo tag v0.17.1,
#    prebuilt/linux_arm64) in ~/litert-lm-bench/prebuilt_v0171/ — see Troubleshooting.
export BUILD=~/litert-lm-bench/mesa-26.1.2/build-panvk
export VK_ICD_FILENAMES="$BUILD/src/panfrost/vulkan/panfrost_devenv_icd.aarch64.json"
export LD_LIBRARY_PATH="$BUILD/src/panfrost/vulkan:$HOME/litert-lm-bench/prebuilt_v0171"
export PATH="$HOME/.local/bin:$PATH"            # uv-installed CLI
litert-lm benchmark ~/litert-lm-bench/gemma-4-E2B-it.litertlm \
  -p 1024 -d 256 --backend gpu --cache disk --runs 2

First GPU run compiles the WebGPU/Vulkan shaders (one-time) and uploads GPU-rearranged weights (~0.8 GB weight cache); --cache disk persists both next to the model, so later loads start fast — the compile cost is NOT repaid on every run.

Results on this Orange Pi 5 Max (Mali-G610 MC4)

Backend Prefill (tk/s) Decode (tk/s) TTFT (s)
CPU (XNNPACK, 8 threads) 123.6 12.3 8.4
GPU (WebGPU/Dawn→Vulkan, patched panvk §5/§7.1) 374.0 8.7 2.9

Google's on-device-class references for Gemma 4 E2B: S26 Ultra GPU 3808/52 tk/s, Raspberry Pi 5 CPU 133/7.6 tk/s. On this board the GPU wins the prefill race by ~3.0× (374 vs 124 tk/s) and TTFT by ~2.9× (2.9 vs 8.4 s), while decode stays on the CPU's side (8.7 vs 12.3 tk/s) — as expected, decode is the bottleneck on Mali-class hardware; the Mali GPU is a prefill/TTFT accelerator here, not a decode accelerator.

GPU clock note (measured, no guesswork): the Mali-G610 is an integrated GPU sharing the board's DDR. The devfreq governor (simple_ondemand, stock) boosts it to 1 GHz under load and parks it at 200 MHz idle — confirmed by sampling cur_freq at 400 ms during the run (233/263 samples at 1000 MHz, temps 41→56 °C, no thermal trip). Leave the governor alone: forcingperformance/min/max via sysfs was counterproductive (idle readbacks that look like the clock collapsed, and it destabilized perfectly good runs). The numbers above are at the stock governor's real 1 GHz boost.

Results: Gemma 4 E4B (4B) — text GPU via the web flavor, vision GPU straight from stock

E4B (the 4B sibling) ships as three files in the HF repo: the stock gemma-4-E4B-it.litertlm (text+vision+audio, 3.66 GB) plus an -gpu and a -web (text-only, 2.97 GB) variant. On this 8 GB board the stock file cannot run on the GPU — two structural walls, both measured:

panthor job watchdog: a fixed one-shot pipeline dispatch (dmesg: job timeout ... seqno=144, same job on every E4B attempt) exceeds the driver's compiled-in 1 s job timeout even at the real 1 GHz boost (E2B's equivalent is seqno=77 and fits). The GPU is reset, Dawn reports VK_ERROR_DEVICE_LOST, generation aborts.

8 GB RAM ceiling: GPU buffers (pinned, unrescalable shmem) peak near 4–5 GB on top of the 3.66 GB model; every run ended in a global OOM-kill (EXIT=137) during iteration 2 even with a 2048-token KV cap + ringbuffers + disk cache.

The -web flavor dodges both — its finer op layout splits the killer dispatch under the 1 s bar and shrinks the GPU working set (peak 4.3 GB observed, holds 1 GHz the whole run):

Gemma 4 E4B lane Flavor Prefill (tk/s) Decode (tk/s) TTFT (s)
CPU (XNNPACK, 8 threads) stock 52.6 5.58 19.6
GPU (patched panvk) web (text-only) 75.1 4.54 13.9

(reproducible to the decimal across --runs 2; init 6.2 s; note --cache disk does not persist for the web flavor — expect a recompile each run. --speculative-decoding true gave no gain:4.37 vs 4.54 tok/s decode (this build ships no draft model).) Reference: Raspberry Pi 5 (16 GB, CPU) 51/3.2/20.5 — this board matches or beats it on every column. Same shape as E2B: GPU wins prefill +1.4× and TTFT 19.6→13.9 s, but decode stays CPU-favored (4.5 vs 5.6 tk/s).

E4B vision (VL) on GPU works — cpu text + gpu vision

The stock E4B's vision encoder is a separate 1477-op subgraph that fits under the panthor watchdog and inside 8 GB even though the full stock model's text lane does not. Drive it via the Engine API (this repo's vl_bench.py --model ...), LLM lane = CPU, vision encoder = GPU (patched panvk): the test image (resized to 912×672 → 2394 patches, near E4B's max_num_patches 2520) is encoded on the Mali.

E4B VL (LLM lane = cpu) vision cpu vision gpu (patched panvk)
TTFT (s) 12.3–12.5 9.6
Decode (tok/s) 7.33–7.37 7.28–7.39
Reply (same image) "coyote walking on a dirt path" identical

Text-only CPU control: TTFT ≈ 2.0 s (7-token prompt, decode ~7.8 tok/s). Isolated vision-encode cost: ~10.4 s CPU vs ~7.6 s GPU — the Mali cuts E4B vision latency by ~2.9 s/image (~27%). No watchdog hits, no OOM (peaks well below ceiling; vision weights are a fraction of a GB). Same conclusion as E2B: for vision work keep the LLM on CPU and let the GPU run the encoder.

# 1. SET ENVIRONMENT FOR GPU VISION DELEGATE
  export LITERT_VISION_DELEGATE=gpu

  # 2. RUN VL HARNESS FOR STOCK GEMMA-4-E4B
~/.local/share/uv/tools/litert-lm/bin/python vl_bench.py \
    --text cpu \
    --vision gpu \
    --image \
    --iters 3 \
    --model gemma-4-E4B-it.litertlm
# Ensure directory exists and download model
mkdir -p ~/litert-lm-bench
curl -L -o ~/litert-lm-bench/gemma-4-E4B-it-web.litertlm \
  https://huggingface.co/litert-community/gemma-4-E4B-it-litert-lm/resolve/main/gemma-4-E4B-it-web.litertlm

# Execute benchmark with standard LiteRT-LM flags
litert-lm benchmark ~/litert-lm-bench/gemma-4-E4B-it-web.litertlm \
  --backend=gpu \
  -p 1024 \
  -d 256 \
  --runs 2

Results: VL (vision) path

The benchmark sub-command only benches the text path. To time the vision encoder, drive the same engine via the Python API (vl_bench.py in this directory) — it loads the model, streams a real image (test_multi.jpg, httpbin.org/image/jpeg) + text through vision_backend=cpu|gpu and prints TTFT (vision encode + prefill + first token), per-chunk decode rate, and the reply. One sentence prompt, 13-token answer, ~290-token total prefill (280 vision + text), generations deterministic (temperature=0). Medians over 3+ runs after compile:

Combo (text / vision) VL TTFT (s) Decode (tok/s) Reply matches CPU?
cpu / cpu 8.93–9.40 17.4 baseline
cpu / gpu (patched panvk) 6.16–6.62 17.2 yes, identical
gpu / gpu (patched panvk) 5.62–6.16 9.8 yes, identical

Text-only controls (7-token prompt): CPU TTFT ≈ 0.8 s, GPU TTFT ≈ 0.57 s.

Isolating the vision cost (VL TTFT − text-only TTFT for the same text backend): ~8.1 s on a CPU vision encoder vs ~5.4 s GPU — the Mali GPU cuts vision-encoder latency by ~2.7–3 s per image (~30%), and the full-GPU combo is ~37% faster end-to-end (5.6 vs 8.9 s). The patched panvk covers the model's separate VISION_ENCODER subgraph too (no extra Dawn limits tripped).

Caveat: for short generations (< ~20 tokens) GPU decode is slower than CPU (9.8 vs 17.4 tok/s) — per-iteration GPU sync overhead — so for chatty/VL replies the CPU text lane with GPU vision is often the best mix (6.2 s TTFT, CPU-fast decode).

Benchmarking the VL path yourself

# 1. FETCH SAMPLE TEST IMAGE
  curl -sL -o ~/litert-lm-bench/test_multi.jpg https://httpbin.org/image/jpeg

  # 2. SET PYTHON INTERPRETER PATH
  VL_PY=~/.local/share/uv/tools/litert-lm/bin/python

  # 3. BASELINE: TEXT (CPU) + VISION (CPU)
  $VL_PY ~/litert-lm-bench/vl_bench.py --text cpu --vision cpu --image --iters 3

  # 4. HYBRID: TEXT (CPU) + VISION (GPU) — EXPORT PANVK MESA ENVS FIRST
  export BUILD=~/litert-lm-bench/mesa-26.1.2/build-panvk
  export VK_ICD_FILENAMES="$BUILD/src/panfrost/vulkan/panfrost_devenv_icd.aarch64.json"
  export LD_LIBRARY_PATH="$BUILD/src/panfrost/vulkan:$HOME/litert-lm-bench/prebuilt_v0171"

  $VL_PY ~/litert-lm-bench/vl_bench.py --text cpu --vision gpu --image --iters 3

  # 5. FULL ACCELERATION: TEXT (GPU) + VISION (GPU)
  $VL_PY ~/litert-lm-bench/vl_bench.py --text gpu --vision gpu --image --iters 3

  # 6. TEXT-ONLY CONTROLS (ISOLATE VISION-ENCODER COST)
  $VL_PY ~/litert-lm-bench/vl_bench.py --text cpu --iters 3
  $VL_PY ~/litert-lm-bench/vl_bench.py --text gpu --iters 3

Prints per-iteration TTFT / decode tok/s / reply. Run cases sequentially — parallel GPU runs contend for the Mali GPU and inflate the numbers (observed 2× TTFT when two ran at once).

6. Inference

# 1. DIRECT CLI EVALUATION 
  litert-lm run ~/litert-lm-bench/gemma-4-E2B-it.litertlm \
    --cache disk --prompt "What is the capital of France?"

  # 2. INTERACTIVE REPL MODE
  litert-lm run ~/litert-lm-bench/gemma-4-E2B-it.litertlm

  # 3. OPENAI-COMPATIBLE API SERVER
  litert-lm serve ~/litert-lm-bench/gemma-4-E2B-it.litertlm

Vision (and audio) input uses run with attachments — one per --attachment, placed before the first user text (images and audio can be mixed):

litert-lm run ~/litert-lm-bench/gemma-4-E2B-it.litertlm \
  --attachment ~/litert-lm-bench/test_multi.jpg \
  --prompt "What is in this image? Answer in one sentence." \
  --vision-backend cpu          
  # or gpu (patched panvk, §5) — ~2.7–3 s faster vision encode

--vision-backend/--audio-backend pick the encoder lane independently of --backend (which chooses the LLM lane). Like --backend gpu, --vision-backend gpu needs the §5 env (VK_ICD_FILENAMES + LD_LIBRARY_PATH) exported for that run.

Notes:

--cache disk persists compiled artifacts next to the model — the first GPU load compiles shaders, later loads start instantly. The load time is NOT paid on every run. Observed caches: <model>_*_mldrift_program_cache.bin (compiled kernels) and<model>_*_mldrift_weight_cache.bin (GPU-rearranged weights, ~0.8 GB). The vision encoder's kernels live in the same cache files — first GPU --vision-backend gpu run compiles, later
runs reuse.

GPU (patched panvk) is ~2.8× faster at prefill and ~2.7× on TTFT, but ~30% slower at decode — pick the lane by workload (§7.4).

Vision: --vision-backend gpu cuts vision-encode latency by ~2.7–3 s per image; CPU text + GPU vision is the best blend for short/chatty replies (§5 VL results).

Speculative decoding (--speculative-decoding true) can lift decode on CPU+GPU for rewrite/summarize/coding style prompts (Gemma 4 E2B supports it).

7. GPU deep dive: the blocker and how it was unblocked here

Everything below was reproduced with LiteRT-LM v0.17.1 (CLI + litert-lm-api), Mesa 26.1.2 (trixie-backports), kernel 7.2.4-edge-rockchip64 (mainline panthor).

7.1 The panvk limit blocker — RESOLVED with a one-line local patch

LiteRT's WebGPU accelerator uses Dawn, which enforces the WebGPU spec minimums against the Vulkan driver and refuses to proceed when they are unmet. Stock panvk reports:

maxImageDimension3D = 512      # WebGPU spec requires ≥ 2048

Dawn logs exactly this (PhysicalDeviceVk.cpp:794: "Insufficient Vulkan limits for maxTextureDimension3D ... must be at least 2048"), and stock panvk then dies with a null-pointer dispatch (SIGSEGV, pc=0x0) inside the WebGPU path on the first real GPU execution. The root cause is the driver limit, not packaging: the crash is identical even with the WebGPU prebuilts (libLiteRtTopKWebGpuSampler.so + friends) fetched from the litert-lm repo prebuilt/linux_arm64 @ v0.17.1 on LD_LIBRARY_PATH.

The fix (validated on this board): panvk caps 3D textures at 512 for Valhall
(PAN_ARCH <= 10), but the Mali-G610 hardware handles 2048³ — raising the advertised limit to 2048 makes Dawn accept the adapter:

-.maxImageDimension3D = PAN_ARCH <= 10 ? (1 << 9) : (1 << 14),
+.maxImageDimension3D = (1 << 11),       /* 2048 — meets WebGPU min (was 512 on arch <= 10) */

Why it's safe: Dawn only validates the advertised limit against its spec minimum, and the value also caps future image allocations — so if a kernel ever genuinely requested a 2048³ 3D texture, panvk would fail cleanly at allocation instead of corrupting anything. LiteRT/ML-Drift's LLM and vision-encoder kernels allocate buffers and 2D textures only, so real GPU behaviour is unchanged; Dawn simply stops rejecting the adapter, the SIGSEGV disappears, and the full benchmark runs (§5). Confirmed via vulkaninfo on the patched build: maxImageDimension3D = 2048. All other WebGPU minimums (per-stage descriptors, workgroup sizes, buffer ranges) were already comfortably met.

The build is process-scoped (VK_ICD_FILENAMES + LD_LIBRARY_PATH per run), the system Mesa is never touched, and reverting is unsetting two variables. Unless/until Mesa raises this limit upstream for Valhall, keep this build around for LiteRT-LM GPU runs.

7.2 The vendor route (Rockchip libmali) — needs a different kernel

Rockchip's proprietary blob would satisfy Dawn (it exposes Vulkan 1.3 with full limits), but the blob's userspace talks to Rockchip's proprietary kbase kernel driver (/dev/mali) and cannot attach to the mainline panthor driver:

"libmali-valhall-g610-g13p0-gbm"      → loads, exports NO Vulkan ICD (0 vk_* symbols)
"libmali-valhall-g610-g24p0-gbm"
  libMaliVulkan.so.1 (api 1.3.276)    → loads, "No mali devices found" then SIGSEGV (no /dev/mali)

Installable packages exist (tsukumijima/libmali-rockchip releases, e.g.
v1.9-1-20260312-bd33ee2), but they all presume the vendor kernel.

Unblock: flash a vendor-kernel image (kernel with kbase/mali.ko, e.g. Armbian images with the Rockchip BSP kernel), then install
libmali-valhall-g610-g13p0-gbm (or g24p0) and point Dawn/Vulkan at it
(VK_DRIVER_FILES/LD_LIBRARY_PATH). That same blob also provides OpenCL 3.0 (libMaliOpenCL.so + /etc/OpenCL/vendors/mali.icd), which would additionally satisfy the "would OpenCL be faster?" curiosity — for running LLMs, both WebGPU/Vulkan and OpenCL on the same GPU land in the same class of throughput; the win is the mature compiler's fast startup, not raw speed.

7.3 OpenCL on the current kernel — not available

Mesa's rusticl on panthor currently exposes zero devices (clinfo -l shows only the rusticl platform, no device). So there is no OpenCL device at all on the mainline stack, and LiteRT-LM has no Linux OpenCL accelerator anyway.

7.4 Practical advice for this board today

  • Chat / streaming (decode-bound): CPU is the better lane — 12.3 vs 8.7 tk/s decode.
  • Long prompts / RAG / document Q&A (prefill-bound): GPU pays off — 345 vs 124 tk/s prefill,
    3.1 vs 8.4 s TTFT.

Images / VL: --vision-backend gpu cuts vision-encode latency by ~2.7–3 s (~30%). Shortest VL TTFT is text+vision both on GPU (5.6 vs 8.9 s all-CPU), but keep the LLM on CPU when replies are short and chatty (6.2 s TTFT plus CPU-fast decode).

  • --speculative-decoding true can lift decode further on both lanes (Gemma 4 E2B supports it).

Serving an OpenAI-compatible endpoint (litert-lm serve)

litert-lm serve exposes an OpenAI API on 0.0.0.0:9379, so LAN clients (LiteCode, opencode, aider, Continue, scripts) use the board like any OpenAI endpoint. Import a model once (litert-lm import ./gemma-4-E2B-it.litertlm, registry at ~/.litert-lm/models/), then litert-lm serve. Per-model settings come from ~/.litert-lm/config.json (default + models.<id>), so clients stay plain-OpenAI:

{
  "default": { "backend": "cpu", "cpu_thread_count": 8, "cache": "disk" },
  "models": { "gemma-4-E2B-it.litertlm": { "speculative_decoding": true, "max_num_tokens": 4096 } }
}

Endpoints: GET /v1/models and POST /v1/chat/completions (streaming + non-streaming). OpenAI tools / tool_choice are bridged to the model's function-calling template — verified to return valid tool_calls JSON (non-streaming and streamed) and to complete the full agent cycle (assistant tool_call → tool result → final answer). /v1/embeddings exists but is not tied to LLM models.

Even though litert-lm describe reports Supports Function Call: NO for Gemma 4, the serve layer still bridges tools into the model template (empirically verified); that flag only means the CLI's interactive run has no native FC template.

RAM budget (E2B, CPU): max_num_tokens preallocates KV/ringbuffer arenas up front.Measured process RSS: 4096 → ~3.1 GB (4.9 GB free — recommended), 8192 → ~7 GB (1.1 GB free — works but leaves no headroom). Keep 4096 and let clients budget. Config keys (0.15+): backend, vision_backend, audio_backend, cpu_thread_count, cache,
max_num_tokens, speculative_decoding, thinking, thinking_budget,
gpu_decode_steps_per_sync, sampling. CLI flags override config. For a persistent border, run under systemd.

Verified client on this board

LiteCode (razvanneculai/litecode) — works end-to-end. Its Planner
({"synthesis","tasks"} strict JSON) and Executor (raw file content, no fences) prompts fit the 4096 budget; single-request task runs completed correctly and files linter-clean.
litecode.json:

{
  "provider": { "baseURL": "http://<yourIP>:9379/v1", "apiKey": "", "model": "gemma-4-E2B-it.litertlm" },
  "tokenLimit": 4096, "reservedOutputTokens": 1500, "systemPromptBudget": 1000, "maxParallelExecutors": 1
}

Troubleshooting

FATAL ERROR: This binary was compiled with dotprod enabled... → CPU lacks FEAT_DotProd, not supported by the ARM64 prebuilt wheels. RK3588 has it.

GPU benchmark errors but CPU works → with stock Mesa this is the §7.1 Dawn limit crash (maxImageDimension3D = 512); with the patched build, check vulkaninfo --summary again — panvk must list the Mali device and report maxImageDimension3D = 2048. ExportVK_ICD_FILENAMES+LD_LIBRARY_PATH for every run that should use the patched driver.

Could not load shared library libLiteRtTopKWebGpuSampler.so → the CLI wheel does not bundle the WebGPU samplers; fetch them from the litert-lm repo prebuilt/linux_arm64/ at the matching version tag (e.g. v0.17.1) and export their directory on LD_LIBRARY_PATH.

First GPU run is slow (one-time kernel compilation) — normal; use --cache disk so later loads skip it. Caches persist next to the model.

Bigger models (e.g. stock gemma-4-E4B-it.litertlm, 3.66 GB) fail on the GPU even with the patched driver: panthor ... job timeout in dmesg (fixed 1 s driver watchdog, not tunable; the model has a one-shot dispatch that exceeds it even at the real 1 GHz boost) theVK_ERROR_DEVICE_LOST, and/or a global OOM-kill (EXIT=137) when pinned GPU buffers + the model outgrow 8 GB. E2B fits, E4B-stock doesn't. Use the text-only -web flavor (gemma-4-E4B-it-web.litertlm, 2.97 GB) for E4B-on-GPU — its finer kernels stay under the watchdog and its GPU set peaks ~4.3 GB (§5 E4B section).

Join the conversation on the Armbian Forum and let us know how it runs on your hardware!

Read more