Explore Model Recipes

Browse ready-to-run inference recipes for NVIDIA DGX Spark systems.

Runtime
Registry
Min Nodes
Parallelism?
Quant
Params
1.7B
397B
Ctx?
32K
256K
Showing 38 of 38 recipes

Qwen/Qwen3-1.7B-GGUF

llama-cpprpc

qwen3-1.7b-llama-cpp

Qwen3 1.7B (Q8_0 GGUF) -- small test model via llama.cpp

1.7Bq8_0TP 11 nodes

Qwen/Qwen3-1.7B

sglangtorch

qwen3-1.7b-sglang

Qwen3 1.7B -- small test model, solo or cluster (SGLang)

1.7Bbfloat16TP 11+ nodes

Qwen/Qwen3-1.7B

vllmtorch

qwen3-1.7b-vllm

Qwen3 1.7B -- small test model, solo or cluster

1.7Bbfloat16TP 11+ nodes

Qwen/Qwen3-Coder-Next-FP8

sglangtorch

qwen3-coder-next-fp8-sglang

Qwen3 Coder Next (upstream FP8 quant) -- cluster only

80Bfp8TP 22+ nodes195K ctx

Qwen/Qwen3.5-0.8B

sglangtorch

qwen3.5-0.8b-bf16-sglang

Qwen3.5 0.8B (upstream, no quant) SGLang

bfloat16TP 11+ nodes256K ctxspec: nextn

Qwen/Qwen3.5-122B-A10B-FP8

sglangtorch

qwen3.5-122b-a10b-fp8-sglang

Qwen3.5 122B A10B (upstream, FP8 quant) SGLang

fp8TP 22+ nodes256K ctxspec: nextn

unsloth/Qwen3.5-122B-A10B-GGUF

llama-cpprpc

qwen3.5-122b-gguf-q4km-llama-cpp

Qwen3.5 122B (GGUF, unsloth quant, Q4_K_M)

122Bq4_k_mTP 11+ nodes256K ctx

Qwen/Qwen3.5-27B-FP8

sglangtorch

qwen3.5-27b-fp8-sglang

Qwen3.5 27B (upstream, FP8 quant) SGLang

fp8TP 11+ nodes256K ctxspec: nextn

Qwen/Qwen3.5-2B

sglangtorch

qwen3.5-2b-bf16-sglang

Qwen3.5 2B (upstream, no quant) SGLang

bfloat16TP 11+ nodes256K ctxspec: nextn

Qwen/Qwen3.5-35B-A3B

sglangtorch

qwen3.5-35b-a3b-bf16-sglang

Qwen3.5 35B A3B (upstream, no quant) SGLang

bfloat16TP 11+ nodes256K ctxspec: nextn

Qwen/Qwen3.5-35B-A3B-FP8

sglangtorch

qwen3.5-35b-a3b-fp8-sglang

Qwen3.5 35B A3B (upstream, FP8 quant) SGLang

fp8TP 11+ nodes256K ctxspec: nextn

unsloth/Qwen3.5-397B-A17B-GGUF

llama-cpprpc

qwen3.5-397b-gguf-q3km-llama-cpp

Qwen3.5 397B (GGUF, unsloth quant, Q3_K_M)

397Bq3_k_mTP 12+ nodes256K ctx

unsloth/Qwen3.5-397B-A17B-GGUF

llama-cpprpc

qwen3.5-397b-gguf-q6k-llama-cpp

Qwen3.5 397B (GGUF, unsloth quant, Q6_K)

397Bq6_kTP 14+ nodes256K ctx

Qwen/Qwen3.5-4B

sglangtorch

qwen3.5-4b-bf16-sglang

Qwen3.5 4B (upstream, no quant) SGLang

bfloat16TP 11+ nodes256K ctxspec: nextn

Qwen/Qwen3.5-9B

sglangtorch

qwen3.5-9b-bf16-sglang

Qwen3.5 9B (upstream, no quant) SGLang

bfloat16TP 11+ nodes256K ctxspec: nextn

cyankiwi/MiniMax-M2.7-AWQ-4bit

vllmtorch

minimax-m2.7-awq4-vllm

MiniMax-M2.7-AWQ — AWQ quantized MiniMax M2.7

awq4TP 22+ nodes192K ctx

Intel/Qwen3-Coder-Next-int4-AutoRound

vllmtorch

qwen3-coder-next-int4-autoround-vllm

Qwen3-Coder-Next-int4-Autoround

int4TP 11 nodes256K ctx

Qwen/Qwen3-VL-Embedding-8B

vllmtorch

qwen3-vl-embedding-8b-vllm

Qwen3-VL-Embedding-8B — vision-language embedding model for multimodal retrieval

bfloat16TP 11 nodes32K ctx

Qwen/Qwen3-VL-Reranker-8B

vllmtorch

qwen3-vl-reranker-8b-vllm

Qwen3-VL-Reranker-8B — vision-language reranker model for multimodal retrieval

bfloat16TP 11 nodes32K ctx

Qwen/Qwen3.6-27B-FP8

vllmtorch

qwen3.6-27b-fp8-mtp-vllm

Qwen3.6-27B-FP8 — native FP8 format

fp8TP 11+ nodes256K ctxspec: mtp

Qwen/Qwen3.6-27B-FP8

vllmtorch

qwen3.6-27b-fp8-vllm

Qwen3.6-27B-FP8 — native FP8 format

fp8TP 11+ nodes256K ctx

Qwen/Qwen3.6-35B-A3B-FP8

vllmtorch

qwen3.6-35b-a3b-fp8-mtp-vllm

Qwen3.6-35B-A3B-FP8 — native FP8 format

TP 11+ nodes256K ctxspec: mtp

Qwen/Qwen3.6-35B-A3B-FP8

vllmtorch

qwen3.6-35b-a3b-fp8-vllm

Qwen3.6-35B-A3B-FP8 — native FP8 format

TP 11+ nodes256K ctx

google/gemma-4-26B-A4B-it

vllmray

gemma4-26b-a4b

vLLM serving Gemma4-26B-A4B

bfloat16TP 21+ nodes256K ctx

cyankiwi/GLM-4.7-Flash-AWQ-4bit

vllmtorch

glm-4.7-flash-awq

vLLM serving cyankiwi/GLM-4.7-Flash-AWQ-4bit with speed optimization patch

awq4TP 11+ nodes198K ctx

QuantTrio/MiniMax-M2-AWQ

vllmray

minimax-m2-awq

vLLM serving MiniMax-M2-AWQ with Ray distributed backend

awq4TP 22+ nodes125K ctx

cyankiwi/MiniMax-M2.5-AWQ-4bit

vllmray

minimax-m2.5-awq

vLLM serving MiniMax-M2.5-AWQ with Ray distributed backend

awq4TP 22+ nodes125K ctx

nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4

vllmtorch

nemotron-3-nano-nvfp4

vLLM serving Nemotron-3-Nano-NVFP4 on a SINGLE NODE ONLY!

nvfp4TP 11 nodes256K ctx

nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4

vllmray

nemotron-3-super-nvfp4

vLLM serving Nemotron-3-Super-120B using CUTLASS kernels

nvfp4TP 21+ nodes256K ctx

openai/gpt-oss-120b

vllmtorch

openai-gpt-oss-120b

vLLM serving openai/gpt-oss-120b with MXFP4 quantization and FlashInfer

mxfp4TP 11 nodes

Qwen/Qwen3-Coder-Next-FP8

vllmray

qwen3-coder-next-fp8

vLLM serving Qwen3-Coder-Next-FP8

fp8TP 21+ nodes128K ctx

Intel/Qwen3-Coder-Next-int4-AutoRound

vllmtorch

qwen3-coder-next-int4-autoround

Qwen3-Coder-Next-int4-Autoround

int4TP 11 nodes256K ctx

Qwen/Qwen3.5-122B-A10B-FP8

vllmray

qwen3.5-122b-fp8

vLLM serving Qwen3.5-122B-FP8

fp8TP 22+ nodes256K ctx

Intel/Qwen3.5-122B-A10B-int4-AutoRound

vllmray

qwen3.5-122b-int4-autoround

vLLM serving Qwen3.5-122B-INT4-Autoround

int4TP 21+ nodes256K ctx

Qwen/Qwen3.5-35B-A3B-FP8

vllmray

qwen3.5-35b-a3b-fp8

vLLM serving Qwen3.5-35B-A3B-FP8

fp8TP 21+ nodes256K ctx

Intel/Qwen3.5-397B-A17B-int4-AutoRound

vllmray

qwen3.5-397b-int4-autoround

EXPERIMENTAL recipe for Qwen3.5-397B-INT4-Autoround (please refer to README for details! Use with `--no-ray` parameter!)

int4TP 22+ nodes256K ctx

MiniMaxAI/MiniMax-M2.5

vllmray

minimax-m2.5

vLLM serving MiniMax-M2.5 with Ray distributed backend

fp8TP 42+ nodes125K ctx

Qwen/Qwen3.5-397B-A17B-FP8

vllmray

qwen3.5-397b-a17B-fp8

vLLM serving Qwen3.5-397B-A17B-FP8

fp8TP 41+ nodes256K ctx