Qwen 3.8 Flash-Next on local hardware: Strata, SGLang, llama.cpp

Sum Codes
Sum Codes
Sep 30, 2026  #qwenchat #Qwen38 #LocalAI

Qwen3.8-Flash-Next changes the local inference equation: a huge model with only a small fraction of its parameters active for each token.
This video explores its mixture-of-experts design, Gated DeltaNet and attention, Qwen Sparse Attention, HyperConnections, predictive lookup embeddings, and multi-token prediction—and why supporting the architecture is only the beginning.

We compare the approaches of llama.cpp, SGLang, EXL3, Trellis and Strata, from broadly portable compatibility to specialized execution and intelligent use of GPU memory, CPU compute, system RAM and SSD storage.

The central question: if only a small portion of an enormous model is needed at any moment, what is the smartest way to use the hardware we already have?

#qwenchat #Qwen38 #LocalAI #LlamaCpp #SGLang #Strata #MixtureOfExperts #AIInference

00:00 A different inference problem
00:13 Total parameters versus active parameters
00:32 Gated DeltaNet, attention and the hybrid architecture
00:55 PLE: a massive lookup system
01:18 MTP: propose, verify, accept
01:40 Why inference support is difficult
02:08 llama.cpp and the Qwen4Exp milestone
02:32 SGLang and day-zero support
02:52 Compatibility versus optimization
03:26 EXL3 and Trellis: specialized execution
04:03 Strata: GPU, CPU, RAM and SSD
04:33 Concurrent CPU and GPU computation
04:57 Combining PLE and multi-token prediction
05:19 Comparing the approaches
06:03 How specialized ideas reach mainstream engines
06:14 Why Flash-Next changes the hardware equation
06:45 Making smarter use of the hardware we have