strix-halo-optimizations

Files

Felipe Cardoso dd403a907c feat(serve): add optimized llama-server launcher with n-gram speculation

Add `make serve` and `make serve-ngram` for launching llama-server with
baked-in optimal settings (Vulkan RADV, q4_0 KV cache, flash attention,
no-mmap, full GPU offload). N-gram speculative decoding gives 1.1-1.4x
tg speedup on repetitive content without upstream PR dependencies.
Update Phase 5 status: MTP is months away (4 unmerged PRs, no MoE
support), draft-model speculation stalled on ROCm buffer crashes.

2026-03-30 21:12:30 +02:00

agentic-benchmarks.md

feat: add Qwen3.5 model catalog and agentic evaluation framework

2026-03-26 00:20:23 +01:00

architecture.md

fix(docs): address review findings — accuracy, consistency, completeness

2026-03-25 21:44:16 +01:00

benchmarking.md

fix(docs): address review findings — accuracy, consistency, completeness

2026-03-25 21:44:16 +01:00

bios-vram-guide.md

docs: add README, CLAUDE.md, AGENTS.md, and full docs/ suite