Agentic AI · on-device · all Blackwell

Agentic AI is moving on-device.
sparkinfer makes every Blackwell GPU fast for it.

The inference runtime for the age of personal agents — reproducible, source-verified MoE/LLM decode across consumer & edge Blackwell: from RTX Spark on your desk to DGX Spark in the workstation.

NVIDIA RTX Spark — personal AI computer for the age of agents

Target GPUs

NVIDIA Blackwell, on-device — consumer sm_120 + edge sm_121 (not datacenter sm_100)
NVIDIA RTX Spark

RTX Spark GB10sm_121

Personal AI PC · flagship target

"A new class of processor for the age of personal agents." Run agents locally and privately on a Windows PC.

128 GB unified · 1 PFLOP FP4 · 120B-param LLMs
NVIDIA × Microsoft ↗
NVIDIA DGX Spark

DGX Sparksm_121

AI workstation

Desk-side Blackwell GB10 workstation — the same agentic inference stack, scaled for builders and teams.

Blackwell GB10 · unified memory
DGX Spark ↗
NVIDIA GeForce RTX 5090

RTX 5090sm_120

Consumer Blackwell · current dev GPU

Desktop flagship — big-model local inference on a 32 GB card. Where the frontier is measured today.

32 GB GDDR7 · ~1.79 TB/s
RTX 5090 ↗
NVIDIA RTX PRO 6000 Blackwell

RTX PRO 6000sm_120

Workstation flagship

96 GB Blackwell — a full MoE plus all experts and a large KV cache stay resident, no paging.

96 GB GDDR7 · ~1.79 TB/s
RTX PRO ↗

Qwen3.6-35B-A3B SOTA · optimization journey

Qwen3.5 · prefill optimization journey

Qwen3-MoE 30B-A3B · optimization journey

128-tok decode · 0.6 → 485 tok/s
long-context · 71 → 393 tok/s

sparkinfer vs llama.cpp

same GPU · same GGUF · 128-token decode

Qwen3-MoE 30B-A3B · per-context decode

Hugging Face Qwen/Qwen3-30B-A3B-GGUF
sparkinfer vs llama.cpp · 128 / 512 / 4k / 16k / 32k decode

Qwen3.6-35B-A3B SOTA · decode

Hugging Face unsloth/Qwen3.6-35B-A3B-GGUF
sparkinfer vs llama.cpp · 128 / 512 / 4k / 16k / 32k decode

Qwen3.6-35B-A3B SOTA · prefill

Hugging Face unsloth/Qwen3.6-35B-A3B-GGUF
sparkinfer vs llama.cpp · 128 / 512 / 4k / 16k / 32k prefill

Qwen3.5 (Qwythos) · per-context

Hugging Face empero-ai/Qwythos-9B-Claude-Mythos-5-1M-GGUF
sparkinfer vs llama.cpp · 128 / 4k / 32k / 64k / 128k decode

Auto-eval labels

deterministic — from verified speedup over same-box main

Evaluated PRs

bot labels + comments · never auto-merges · split by scored metric

Decode

tok/s · strongest decode context
PRarealabeltok/svs frontierproof

Prefill

pp tok/s · strongest prefill context
PRarealabelpp tok/svs frontierproof