Sparse-attention and KV-cache experiments for Qwen3.5-4B on an Intel Vulkan GPU. Reproducible measurements, including unsuccessful experiments.
Sparse-attention and KV-cache experiments for Qwen3.5-4B on an Intel Vulkan GPU. Reproducible measurements, including unsuccessful experiments.
Specification
- Language
- C++
- Created
- 2026-09-10
- Repository
- leonardoeloy / moba-qwen3.5-4b-gguf
Release notes
32K scored-prefill: 12.43 minutes dense versus 8.20 minutes sparse, with +0.74% sampled perplexity and 64/64 continuation agreement. One timed pass per mode on one passage. Short-context results are essentially tied. Cache and sharing experiments are documented separately.