vramarcade

arcade / model

qwen3.8-flash-next

Qwen3.8-Flash-Next (UD-Q3_K_XL, Qwen4 preview, 176.9B-A6B)

Qwen4-architecture preview, running on an unmerged llama.cpp PR. The only entry here that is not fully GPU-resident: experts in RAM, n-gram table paged off disk. Cold load is ~15 minutes and it thinks at xhigh by default, so expect it to be the slowest column on the grid by some distance.

show

Everything it built 4 result(s)