feat(serve): set APEX I-Compact as default, harden benchmark workflow

Serving: - make serve now launches Claude-distilled APEX 35B-A3B (16GB) with 2 parallel slots and 256K context as the daily driver - add serve-custom for ad-hoc model testing - add flush-gpu to reclaim unified memory after stuck runs Benchmarks: - default Vulkan-only backends (ROCm trails at long context) - add --backends filter to run-baseline.sh - fix backend filter substring bug (grep -qFx for exact line match) - fix model filter regex metacharacter bug (grep -qiF for literal) - respect --tg in long-context tests instead of hardcoded n=32 ROCm bump to 7.2.1 (kernel 6.18.4+ patch); keep 7.2 as optional. Catalog: - add mudler APEX I-Compact (Claude-distilled 35B, 17GB) - add 0xSero REAP-40 (pruned 122B-A10B, 46GB) - update download instructions: hf download (huggingface-cli is gone)
2026-04-13 01:11:46 +02:00
parent 474d94a07e
commit 15bb6a8ed9
7 changed files with 36 additions and 12 deletions
--- a/scripts/benchmark/setup.sh
+++ b/scripts/benchmark/setup.sh
@@ -15,8 +15,8 @@ log_header "Benchmark Setup"
 # ── 1. Check toolbox containers ──────────────────────────
 log_info "Checking toolbox containers..."

-REQUIRED_TOOLBOXES=("llama-vulkan-radv" "llama-rocm-7.2")
-OPTIONAL_TOOLBOXES=("llama-rocm-6.4.4" "llama-vulkan-amdvlk")
+REQUIRED_TOOLBOXES=("llama-vulkan-radv" "llama-rocm-7.2.1")
+OPTIONAL_TOOLBOXES=("llama-rocm-7.2" "llama-rocm-6.4.4" "llama-vulkan-amdvlk")

 existing=$(detect_toolbox_names 2>/dev/null || true)
 missing=()