Mali GPU with LiteRT-LM
Running Gemma 4 E2B on a Mali GPU with LiteRT-LM
This guide should work on any armbian minimal/console (trixie) system with a working Mali-g610 GPU.
DO NOT USE DESKTOP
Verified on: Orange Pi 5 Max (Rockchip RK3588, Arm Mali-G610 MC4), Armbian/Debian 13 (trixie), aarch64, 8 GB RAM.
Status summary: Both CPU (XNNPACK) and GPU (WebGPU/Dawn → Vulkan → panvk → panthor) work on the mainline panthor kernel — numbers below. The GPU path needed one line of local patching in
Mesa's panvk driver (raise maxImageDimension3D 512 → 2048 to meet the WebGPU minimums that LiteRT's Dawn library enforces; recipe in §5, rationale in §7.1). The model is multimodal (Text + Vision + Audio): the vision encoder runs on the same patched panvk via--vision-backend gpu (§5 "VL path" results). The vendor Rockchip libmali route and OpenCL remain unavailable on mainline panthor (§7.2, §7.3).
Stack (as shipped):
Gemma 4 E2B (.litertlm, int4) → LiteRT-LM CLI → LiteRT GPU accelerator (WebGPU/ML Drift)
Dawn (WebGPU impl) → Vulkan → panvk → panthor (kernel) → Mali-G610
1. Prerequisites
- ARM (aarch64) Linux with a Mali GPU and at least 8 GB RAM.
- Armbian/Debian 13 (trixie)
- The
panthor(orpanfrostfor older Mali) kernel driver active:
ls /dev/dri/renderD128 # render node must exist
lsmod | grep panthor # or panfrost / lima
Full ARMv8.2-A dotprod support is required — LiteRT's ARM64 binaries are built with it and will SIGILL otherwise:lscpu | grep -w asimddp (both A55 and A76 cores list it here).
2. Install Mesa + Vulkan for Mali (needs root)
sudo apt update
# Prefer the newest Mesa (backports on Debian) for the best panvk coverage of Valhall CSF GPUs:
sudo apt install -y -t trixie-backports mesa-vulkan-drivers libvulkan1 vulkan-tools
# Optional (see §7 — currently yields NO usable OpenCL device on panthor):
sudo apt install -y -t trixie-backports mesa-opencl-icd clinfo
Verify the GPU is visible to Vulkan:
vulkaninfo --summary | grep -iE "deviceName|driverName"
# Expect: deviceName = Mali-G610 MC4 driverName = panvk
3. Install the LiteRT-LM CLI (user space, no root)
curl -LsSf https://astral.sh/uv/install.sh | sh
export PATH="$HOME/.local/bin:$PATH"
uv tool install litert-lm # current: v0.17.x; incl. CPU (XNNPACK/YNNPACK) + GPU (WebGPU/Vulkan) backends
4. Get the model (Gemma 4 E2B, int4 .litertlm, 2.58 GB)
mkdir -p ~/litert-lm-bench && cd ~/litert-lm-bench
curl -L -o gemma-4-E2B-it.litertlm \
https://huggingface.co/litert-community/gemma-4-E2B-it-litert-lm/resolve/main/gemma-4-E2B-it.litertlm
# or have the CLI fetch it for you:
# litert-lm benchmark --from-huggingface-repo litert-community/gemma-4-E2B-it-litert-lm gemma-4-E2B-it.litertlm
Other ready models: google/gemma-3n-E2B-it-litert-lm (gemma-3n-E2B-it-int4.litertlm),litert-community/gemma-4-E4B-it-litert-lm, ... (see litert-lm list / HF).
5. Benchmark
Lock the CPU governor to performance first (fair, repeatable numbers):
sudo sh -c 'for c in /sys/devices/system/cpu/cpu[0-7]/cpufreq/scaling_governor; do echo performance > $c; done'
CPU baseline (XNNPACK, 8 threads):
litert-lm benchmark ~/litert-lm-bench/gemma-4-E2B-it.litertlm \
-p 1024 -d 256 --backend cpu --cpu-thread-count 8 --cache disk --runs 2
GPU (patched panvk, §7.1). One-time local Mesa build — stays entirely in userland, the system Mesa is untouched:
# 1) Mesa source + one-line patch (26.1.2)
cd ~/litert-lm-bench
curl -L -o mesa.tar.xz https://archive.mesa3d.org/mesa-26.1.2.tar.xz
tar -xf mesa.tar.xz && cd mesa-26.1.2
python3 - <<'PY'
p = 'src/panfrost/vulkan/panvk_vX_physical_device.c'
s = open(p).read()
s = s.replace('.maxImageDimension3D = PAN_ARCH <= 10 ? (1 << 9) : (1 << 14),',
'.maxImageDimension3D = (1 << 11), /* 2048 — meets WebGPU min */')
open(p, 'w').write(s)
PY
# 2) build deps (the LLVM/CLC chain is required — panvk bakes CLC-compiled SPIR-V for libpan)
sudo apt install -y --no-install-recommends meson ninja-build pkg-config gcc g++ python3-mako \
python3-yaml python3-ply libdrm-dev llvm-19-dev libllvmspirvlib-19-dev spirv-tools \
libclang-19-dev libclang-cpp19-dev
# 3) panvk-only build (~10–25 min on the 5 Max)
meson setup build-panvk -Dvulkan-drivers=panfrost -Dgallium-drivers= -Dbuildtype=release \
-Dllvm=enabled -Dmesa-clc=auto -Dplatforms= -Degl=disabled -Dgbm=disabled -Dglx=disabled \
-Dopengl=false -Dglvnd=disabled -Dtools= -Dbuild-tests=false
ninja -C build-panvk -j6
# 4) benchmark with the patched driver, process-scoped. First put the WebGPU prebuilts
# (libLiteRtTopKWebGpuSampler.so + libwebgpu_dawn.so etc., litert-lm repo tag v0.17.1,
# prebuilt/linux_arm64) in ~/litert-lm-bench/prebuilt_v0171/ — see Troubleshooting.
export BUILD=~/litert-lm-bench/mesa-26.1.2/build-panvk
export VK_ICD_FILENAMES="$BUILD/src/panfrost/vulkan/panfrost_devenv_icd.aarch64.json"
export LD_LIBRARY_PATH="$BUILD/src/panfrost/vulkan:$HOME/litert-lm-bench/prebuilt_v0171"
export PATH="$HOME/.local/bin:$PATH" # uv-installed CLI
litert-lm benchmark ~/litert-lm-bench/gemma-4-E2B-it.litertlm \
-p 1024 -d 256 --backend gpu --cache disk --runs 2
First GPU run compiles the WebGPU/Vulkan shaders (one-time) and uploads GPU-rearranged weights (~0.8 GB weight cache); --cache disk persists both next to the model, so later loads start fast — the compile cost is NOT repaid on every run.
Results on this Orange Pi 5 Max (Mali-G610 MC4)
| Backend | Prefill (tk/s) | Decode (tk/s) | TTFT (s) |
|---|---|---|---|
| CPU (XNNPACK, 8 threads) | 123.6 | 12.3 | 8.4 |
| GPU (WebGPU/Dawn→Vulkan, patched panvk §5/§7.1) | 374.0 | 8.7 | 2.9 |
Google's on-device-class references for Gemma 4 E2B: S26 Ultra GPU 3808/52 tk/s, Raspberry Pi 5 CPU 133/7.6 tk/s. On this board the GPU wins the prefill race by ~3.0× (374 vs 124 tk/s) and TTFT by ~2.9× (2.9 vs 8.4 s), while decode stays on the CPU's side (8.7 vs 12.3 tk/s) — as expected, decode is the bottleneck on Mali-class hardware; the Mali GPU is a prefill/TTFT accelerator here, not a decode accelerator.
GPU clock note (measured, no guesswork): the Mali-G610 is an integrated GPU sharing the board's DDR. The devfreq governor (simple_ondemand, stock) boosts it to 1 GHz under load and parks it at 200 MHz idle — confirmed by sampling cur_freq at 400 ms during the run (233/263 samples at 1000 MHz, temps 41→56 °C, no thermal trip). Leave the governor alone: forcingperformance/min/max via sysfs was counterproductive (idle readbacks that look like the clock collapsed, and it destabilized perfectly good runs). The numbers above are at the stock governor's real 1 GHz boost.
Results: Gemma 4 E4B (4B) — text GPU via the web flavor, vision GPU straight from stock
E4B (the 4B sibling) ships as three files in the HF repo: the stock gemma-4-E4B-it.litertlm (text+vision+audio, 3.66 GB) plus an -gpu and a -web (text-only, 2.97 GB) variant. On this 8 GB board the stock file cannot run on the GPU — two structural walls, both measured:
panthor job watchdog: a fixed one-shot pipeline dispatch (dmesg: job timeout ... seqno=144, same job on every E4B attempt) exceeds the driver's compiled-in 1 s job timeout even at the real 1 GHz boost (E2B's equivalent is seqno=77 and fits). The GPU is reset, Dawn reports VK_ERROR_DEVICE_LOST, generation aborts.
8 GB RAM ceiling: GPU buffers (pinned, unrescalable shmem) peak near 4–5 GB on top of the 3.66 GB model; every run ended in a global OOM-kill (EXIT=137) during iteration 2 even with a 2048-token KV cap + ringbuffers + disk cache.
The -web flavor dodges both — its finer op layout splits the killer dispatch under the 1 s bar and shrinks the GPU working set (peak 4.3 GB observed, holds 1 GHz the whole run):
| Gemma 4 E4B lane | Flavor | Prefill (tk/s) | Decode (tk/s) | TTFT (s) |
|---|---|---|---|---|
| CPU (XNNPACK, 8 threads) | stock | 52.6 | 5.58 | 19.6 |
| GPU (patched panvk) | web (text-only) | 75.1 | 4.54 | 13.9 |
(reproducible to the decimal across --runs 2; init 6.2 s; note --cache disk does not persist for the web flavor — expect a recompile each run. --speculative-decoding true gave no gain:4.37 vs 4.54 tok/s decode (this build ships no draft model).) Reference: Raspberry Pi 5 (16 GB, CPU) 51/3.2/20.5 — this board matches or beats it on every column. Same shape as E2B: GPU wins prefill +1.4× and TTFT 19.6→13.9 s, but decode stays CPU-favored (4.5 vs 5.6 tk/s).
E4B vision (VL) on GPU works — cpu text + gpu vision
The stock E4B's vision encoder is a separate 1477-op subgraph that fits under the panthor watchdog and inside 8 GB even though the full stock model's text lane does not. Drive it via the Engine API (this repo's vl_bench.py --model ...), LLM lane = CPU, vision encoder = GPU (patched panvk): the test image (resized to 912×672 → 2394 patches, near E4B's max_num_patches 2520) is encoded on the Mali.
| E4B VL (LLM lane = cpu) | vision cpu | vision gpu (patched panvk) |
|---|---|---|
| TTFT (s) | 12.3–12.5 | 9.6 |
| Decode (tok/s) | 7.33–7.37 | 7.28–7.39 |
| Reply (same image) | "coyote walking on a dirt path" | identical |
Text-only CPU control: TTFT ≈ 2.0 s (7-token prompt, decode ~7.8 tok/s). Isolated vision-encode cost: ~10.4 s CPU vs ~7.6 s GPU — the Mali cuts E4B vision latency by ~2.9 s/image (~27%). No watchdog hits, no OOM (peaks well below ceiling; vision weights are a fraction of a GB). Same conclusion as E2B: for vision work keep the LLM on CPU and let the GPU run the encoder.
# 1. SET ENVIRONMENT FOR GPU VISION DELEGATE
export LITERT_VISION_DELEGATE=gpu
# 2. RUN VL HARNESS FOR STOCK GEMMA-4-E4B
~/.local/share/uv/tools/litert-lm/bin/python vl_bench.py \
--text cpu \
--vision gpu \
--image \
--iters 3 \
--model gemma-4-E4B-it.litertlm# Ensure directory exists and download model
mkdir -p ~/litert-lm-bench
curl -L -o ~/litert-lm-bench/gemma-4-E4B-it-web.litertlm \
https://huggingface.co/litert-community/gemma-4-E4B-it-litert-lm/resolve/main/gemma-4-E4B-it-web.litertlm
# Execute benchmark with standard LiteRT-LM flags
litert-lm benchmark ~/litert-lm-bench/gemma-4-E4B-it-web.litertlm \
--backend=gpu \
-p 1024 \
-d 256 \
--runs 2Results: VL (vision) path
The benchmark sub-command only benches the text path. To time the vision encoder, drive the same engine via the Python API (vl_bench.py in this directory) — it loads the model, streams a real image (test_multi.jpg, httpbin.org/image/jpeg) + text through vision_backend=cpu|gpu and prints TTFT (vision encode + prefill + first token), per-chunk decode rate, and the reply. One sentence prompt, 13-token answer, ~290-token total prefill (280 vision + text), generations deterministic (temperature=0). Medians over 3+ runs after compile:
| Combo (text / vision) | VL TTFT (s) | Decode (tok/s) | Reply matches CPU? |
|---|---|---|---|
| cpu / cpu | 8.93–9.40 | 17.4 | baseline |
| cpu / gpu (patched panvk) | 6.16–6.62 | 17.2 | yes, identical |
| gpu / gpu (patched panvk) | 5.62–6.16 | 9.8 | yes, identical |
Text-only controls (7-token prompt): CPU TTFT ≈ 0.8 s, GPU TTFT ≈ 0.57 s.
Isolating the vision cost (VL TTFT − text-only TTFT for the same text backend): ~8.1 s on a CPU vision encoder vs ~5.4 s GPU — the Mali GPU cuts vision-encoder latency by ~2.7–3 s per image (~30%), and the full-GPU combo is ~37% faster end-to-end (5.6 vs 8.9 s). The patched panvk covers the model's separate VISION_ENCODER subgraph too (no extra Dawn limits tripped).
Caveat: for short generations (< ~20 tokens) GPU decode is slower than CPU (9.8 vs 17.4 tok/s) — per-iteration GPU sync overhead — so for chatty/VL replies the CPU text lane with GPU vision is often the best mix (6.2 s TTFT, CPU-fast decode).
Benchmarking the VL path yourself
# 1. FETCH SAMPLE TEST IMAGE
curl -sL -o ~/litert-lm-bench/test_multi.jpg https://httpbin.org/image/jpeg
# 2. SET PYTHON INTERPRETER PATH
VL_PY=~/.local/share/uv/tools/litert-lm/bin/python
# 3. BASELINE: TEXT (CPU) + VISION (CPU)
$VL_PY ~/litert-lm-bench/vl_bench.py --text cpu --vision cpu --image --iters 3
# 4. HYBRID: TEXT (CPU) + VISION (GPU) — EXPORT PANVK MESA ENVS FIRST
export BUILD=~/litert-lm-bench/mesa-26.1.2/build-panvk
export VK_ICD_FILENAMES="$BUILD/src/panfrost/vulkan/panfrost_devenv_icd.aarch64.json"
export LD_LIBRARY_PATH="$BUILD/src/panfrost/vulkan:$HOME/litert-lm-bench/prebuilt_v0171"
$VL_PY ~/litert-lm-bench/vl_bench.py --text cpu --vision gpu --image --iters 3
# 5. FULL ACCELERATION: TEXT (GPU) + VISION (GPU)
$VL_PY ~/litert-lm-bench/vl_bench.py --text gpu --vision gpu --image --iters 3
# 6. TEXT-ONLY CONTROLS (ISOLATE VISION-ENCODER COST)
$VL_PY ~/litert-lm-bench/vl_bench.py --text cpu --iters 3
$VL_PY ~/litert-lm-bench/vl_bench.py --text gpu --iters 3
Prints per-iteration TTFT / decode tok/s / reply. Run cases sequentially — parallel GPU runs contend for the Mali GPU and inflate the numbers (observed 2× TTFT when two ran at once).
6. Inference
# 1. DIRECT CLI EVALUATION
litert-lm run ~/litert-lm-bench/gemma-4-E2B-it.litertlm \
--cache disk --prompt "What is the capital of France?"
# 2. INTERACTIVE REPL MODE
litert-lm run ~/litert-lm-bench/gemma-4-E2B-it.litertlm
# 3. OPENAI-COMPATIBLE API SERVER
litert-lm serve ~/litert-lm-bench/gemma-4-E2B-it.litertlm
Vision (and audio) input uses run with attachments — one per --attachment, placed before the first user text (images and audio can be mixed):
litert-lm run ~/litert-lm-bench/gemma-4-E2B-it.litertlm \
--attachment ~/litert-lm-bench/test_multi.jpg \
--prompt "What is in this image? Answer in one sentence." \
--vision-backend cpu
# or gpu (patched panvk, §5) — ~2.7–3 s faster vision encode
--vision-backend/--audio-backend pick the encoder lane independently of --backend (which chooses the LLM lane). Like --backend gpu, --vision-backend gpu needs the §5 env (VK_ICD_FILENAMES + LD_LIBRARY_PATH) exported for that run.
Notes:
--cache disk persists compiled artifacts next to the model — the first GPU load compiles shaders, later loads start instantly. The load time is NOT paid on every run. Observed caches: <model>_*_mldrift_program_cache.bin (compiled kernels) and<model>_*_mldrift_weight_cache.bin (GPU-rearranged weights, ~0.8 GB). The vision encoder's kernels live in the same cache files — first GPU --vision-backend gpu run compiles, later
runs reuse.
GPU (patched panvk) is ~2.8× faster at prefill and ~2.7× on TTFT, but ~30% slower at decode — pick the lane by workload (§7.4).
Vision: --vision-backend gpu cuts vision-encode latency by ~2.7–3 s per image; CPU text + GPU vision is the best blend for short/chatty replies (§5 VL results).
Speculative decoding (--speculative-decoding true) can lift decode on CPU+GPU for rewrite/summarize/coding style prompts (Gemma 4 E2B supports it).
7. GPU deep dive: the blocker and how it was unblocked here
Everything below was reproduced with LiteRT-LM v0.17.1 (CLI + litert-lm-api), Mesa 26.1.2 (trixie-backports), kernel 7.2.4-edge-rockchip64 (mainline panthor).
7.1 The panvk limit blocker — RESOLVED with a one-line local patch
LiteRT's WebGPU accelerator uses Dawn, which enforces the WebGPU spec minimums against the Vulkan driver and refuses to proceed when they are unmet. Stock panvk reports:
maxImageDimension3D = 512 # WebGPU spec requires ≥ 2048
Dawn logs exactly this (PhysicalDeviceVk.cpp:794: "Insufficient Vulkan limits for maxTextureDimension3D ... must be at least 2048"), and stock panvk then dies with a null-pointer dispatch (SIGSEGV, pc=0x0) inside the WebGPU path on the first real GPU execution. The root cause is the driver limit, not packaging: the crash is identical even with the WebGPU prebuilts (libLiteRtTopKWebGpuSampler.so + friends) fetched from the litert-lm repo prebuilt/linux_arm64 @ v0.17.1 on LD_LIBRARY_PATH.
The fix (validated on this board): panvk caps 3D textures at 512 for Valhall
(PAN_ARCH <= 10), but the Mali-G610 hardware handles 2048³ — raising the advertised limit to 2048 makes Dawn accept the adapter:
-.maxImageDimension3D = PAN_ARCH <= 10 ? (1 << 9) : (1 << 14),
+.maxImageDimension3D = (1 << 11), /* 2048 — meets WebGPU min (was 512 on arch <= 10) */
Why it's safe: Dawn only validates the advertised limit against its spec minimum, and the value also caps future image allocations — so if a kernel ever genuinely requested a 2048³ 3D texture, panvk would fail cleanly at allocation instead of corrupting anything. LiteRT/ML-Drift's LLM and vision-encoder kernels allocate buffers and 2D textures only, so real GPU behaviour is unchanged; Dawn simply stops rejecting the adapter, the SIGSEGV disappears, and the full benchmark runs (§5). Confirmed via vulkaninfo on the patched build: maxImageDimension3D = 2048. All other WebGPU minimums (per-stage descriptors, workgroup sizes, buffer ranges) were already comfortably met.
The build is process-scoped (VK_ICD_FILENAMES + LD_LIBRARY_PATH per run), the system Mesa is never touched, and reverting is unsetting two variables. Unless/until Mesa raises this limit upstream for Valhall, keep this build around for LiteRT-LM GPU runs.
7.2 The vendor route (Rockchip libmali) — needs a different kernel
Rockchip's proprietary blob would satisfy Dawn (it exposes Vulkan 1.3 with full limits), but the blob's userspace talks to Rockchip's proprietary kbase kernel driver (/dev/mali) and cannot attach to the mainline panthor driver:
"libmali-valhall-g610-g13p0-gbm" → loads, exports NO Vulkan ICD (0 vk_* symbols)
"libmali-valhall-g610-g24p0-gbm"
libMaliVulkan.so.1 (api 1.3.276) → loads, "No mali devices found" then SIGSEGV (no /dev/mali)
Installable packages exist (tsukumijima/libmali-rockchip releases, e.g.v1.9-1-20260312-bd33ee2), but they all presume the vendor kernel.
Unblock: flash a vendor-kernel image (kernel with kbase/mali.ko, e.g. Armbian images with the Rockchip BSP kernel), then installlibmali-valhall-g610-g13p0-gbm (or g24p0) and point Dawn/Vulkan at it
(VK_DRIVER_FILES/LD_LIBRARY_PATH). That same blob also provides OpenCL 3.0 (libMaliOpenCL.so + /etc/OpenCL/vendors/mali.icd), which would additionally satisfy the "would OpenCL be faster?" curiosity — for running LLMs, both WebGPU/Vulkan and OpenCL on the same GPU land in the same class of throughput; the win is the mature compiler's fast startup, not raw speed.
7.3 OpenCL on the current kernel — not available
Mesa's rusticl on panthor currently exposes zero devices (clinfo -l shows only the rusticl platform, no device). So there is no OpenCL device at all on the mainline stack, and LiteRT-LM has no Linux OpenCL accelerator anyway.
7.4 Practical advice for this board today
- Chat / streaming (decode-bound): CPU is the better lane — 12.3 vs 8.7 tk/s decode.
- Long prompts / RAG / document Q&A (prefill-bound): GPU pays off — 345 vs 124 tk/s prefill,
3.1 vs 8.4 s TTFT.
Images / VL: --vision-backend gpu cuts vision-encode latency by ~2.7–3 s (~30%). Shortest VL TTFT is text+vision both on GPU (5.6 vs 8.9 s all-CPU), but keep the LLM on CPU when replies are short and chatty (6.2 s TTFT plus CPU-fast decode).
--speculative-decoding truecan lift decode further on both lanes (Gemma 4 E2B supports it).
Serving an OpenAI-compatible endpoint (litert-lm serve)
litert-lm serve exposes an OpenAI API on 0.0.0.0:9379, so LAN clients (LiteCode, opencode, aider, Continue, scripts) use the board like any OpenAI endpoint. Import a model once (litert-lm import ./gemma-4-E2B-it.litertlm, registry at ~/.litert-lm/models/), then litert-lm serve. Per-model settings come from ~/.litert-lm/config.json (default + models.<id>), so clients stay plain-OpenAI:
{
"default": { "backend": "cpu", "cpu_thread_count": 8, "cache": "disk" },
"models": { "gemma-4-E2B-it.litertlm": { "speculative_decoding": true, "max_num_tokens": 4096 } }
}
Endpoints: GET /v1/models and POST /v1/chat/completions (streaming + non-streaming). OpenAI tools / tool_choice are bridged to the model's function-calling template — verified to return valid tool_calls JSON (non-streaming and streamed) and to complete the full agent cycle (assistant tool_call → tool result → final answer). /v1/embeddings exists but is not tied to LLM models.
Even though litert-lm describe reports Supports Function Call: NO for Gemma 4, the serve layer still bridges tools into the model template (empirically verified); that flag only means the CLI's interactive run has no native FC template.
RAM budget (E2B, CPU): max_num_tokens preallocates KV/ringbuffer arenas up front.Measured process RSS: 4096 → ~3.1 GB (4.9 GB free — recommended), 8192 → ~7 GB (1.1 GB free — works but leaves no headroom). Keep 4096 and let clients budget. Config keys (0.15+): backend, vision_backend, audio_backend, cpu_thread_count, cache,max_num_tokens, speculative_decoding, thinking, thinking_budget,gpu_decode_steps_per_sync, sampling. CLI flags override config. For a persistent border, run under systemd.
Verified client on this board
LiteCode (razvanneculai/litecode) — works end-to-end. Its Planner
({"synthesis","tasks"} strict JSON) and Executor (raw file content, no fences) prompts fit the 4096 budget; single-request task runs completed correctly and files linter-clean.litecode.json:
{
"provider": { "baseURL": "http://<yourIP>:9379/v1", "apiKey": "", "model": "gemma-4-E2B-it.litertlm" },
"tokenLimit": 4096, "reservedOutputTokens": 1500, "systemPromptBudget": 1000, "maxParallelExecutors": 1
}
Troubleshooting
FATAL ERROR: This binary was compiled with dotprod enabled... → CPU lacks FEAT_DotProd, not supported by the ARM64 prebuilt wheels. RK3588 has it.
GPU benchmark errors but CPU works → with stock Mesa this is the §7.1 Dawn limit crash (maxImageDimension3D = 512); with the patched build, check vulkaninfo --summary again — panvk must list the Mali device and report maxImageDimension3D = 2048. ExportVK_ICD_FILENAMES+LD_LIBRARY_PATH for every run that should use the patched driver.
Could not load shared library libLiteRtTopKWebGpuSampler.so → the CLI wheel does not bundle the WebGPU samplers; fetch them from the litert-lm repo prebuilt/linux_arm64/ at the matching version tag (e.g. v0.17.1) and export their directory on LD_LIBRARY_PATH.
First GPU run is slow (one-time kernel compilation) — normal; use --cache disk so later loads skip it. Caches persist next to the model.
Bigger models (e.g. stock gemma-4-E4B-it.litertlm, 3.66 GB) fail on the GPU even with the patched driver: panthor ... job timeout in dmesg (fixed 1 s driver watchdog, not tunable; the model has a one-shot dispatch that exceeds it even at the real 1 GHz boost) theVK_ERROR_DEVICE_LOST, and/or a global OOM-kill (EXIT=137) when pinned GPU buffers + the model outgrow 8 GB. E2B fits, E4B-stock doesn't. Use the text-only -web flavor (gemma-4-E4B-it-web.litertlm, 2.97 GB) for E4B-on-GPU — its finer kernels stay under the watchdog and its GPU set peaks ~4.3 GB (§5 E4B section).
Join the conversation on the Armbian Forum and let us know how it runs on your hardware!