N
Hacker Next
new
past
show
ask
show
jobs
submit
login
▲
I Rented a 96 GB GPU and Took Uncensored Qwen3.8 From 44 to 125 tok/s
(
aseemshrey.com
)
6 points by
LuD1161
12 hours ago
|
1 comment
add comment
Rendered at 16:48:33 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
karmakaze 43 minutes ago
[-]
I did a similar thing running Q6_K model and Q8_0 DFlash2 (draft=7) quants:
DavidAU/Qwen3.8-27B-Cold-Fusion-GAIN-V1.1-NM-DAU-NEO-MAX-MTP-GGUF incoai/Qwen3.8-27B-DFlash2-GGUF
using llama.cpp PR/commit
https://github.com/ggml-org/llama.cpp/pull/27342
on an AMD R9700 (32GB)