Compute/benchmark/2026-10-05

Qwen Flash Next 125B Model Runs on Mac Mini and Strix Halo with SSD Streaming

A user on r/LocalLLaMA successfully ran the Qwen Flash Next 125B model on a 64GB Mac Mini and a Strix Halo mini PC, utilizing SSD streaming and speculative decoding. Performance metrics were reported, including decode speeds of 44-59 tok/s with speculative decoding and ~1,400 tok/s prefill on the Strix Halo, and a 27% GPU wait time on experts during decode on the Mac Mini. The engine used for this setup is reported to be open.

3 articles from 1 outlet covered this story. The underlying claim is sourced from a benchmark.

What do all outlets agree on?

1 outlet covered “Qwen Flash Next 125B Model Runs on Mac Mini and Strix Halo with SSD Streaming”. All of them report the following:

  • Qwen Flash Next (125B) model was run
  • Model ran on consumer hardware (Mac Mini, Strix Halo mini PC)
  • Utilized SSD streaming
  • Involved Mixture of Experts (MoE) architecture
  • Performance metrics were reported
  • Engine is open

Which outlets covered this?

All 3 articles found on this story, grouped by the stance of the piece. Every link goes to the original publisher.

What related stories are there?

Which companies does this involve?

Get the week in AI in one email

What happened, which outlets reported it, and where their coverage differed. One issue a week.

The first issue hasn’t gone out yet. Subscribe and it’s the one you’ll get.

We’ll send the digest and nothing else. One-click unsubscribe. Privacy.