NVIDIA DGX SPARK · CHINA COMMUNITY

Your desk is now
an AI supercomputing hub

FusionXpark packs the NVIDIA GB10 Grace Blackwell superchip, unleashing 1 PFLOPS of AI compute in a compact chassis — keeping 200B-parameter models running locally.

Official NVIDIA DGX Spark product image: desktop supercomputer and laptop forming a local AI workstation
FUSIONXPARK Grace Blackwell, now on your desk.
AI PERFORMANCE 1 PFLOPS FP4
UNIFIED MEMORY 128 GB LPDDR5x
SUPERCHIP GB10 Grace Blackwell
DGX SPARK GRACE BLACKWELL 1PANEL MAXKB LOCAL AI

DESKTOP SUPERCOMPUTING

Small footprint.
Serious compute.

From model loading and context inference to multi-node scaling, DGX Spark compresses data-center-class AI architecture into a desktop form — personal AI infrastructure, not just another workstation.

See what it can do
01

AI PERFORMANCE

1 PetaFLOPS

High-density AI performance for FP4 inference workloads — right on your desk.

02

UNIFIED MEMORY

128 GB

Larger models, longer contexts, and multimodal inputs run stably in a unified memory architecture.

03

MODEL SCALE

200B

Supports 200B-parameter-class models, with the DeepSeek, Llama, Gemma, and Qwen ecosystems.

04

SCALE OUT

ConnectX

Dual-node interconnect leaves headroom for labs, team prototypes, and small clusters.

FROM MODEL TO ACTION

Keep AI local —
and put it to real work.

01 · LOCAL MODEL

Local LLM Workstation

Run private models, RAG, agents, and multimodal prototypes — sensitive data never leaves the office.

02 · LIVE DEMO

A Tangible AI Demo Environment

Show real workflows like translation, review, and knowledge-base Q&A — turning compute specs into experience.

03 · COMMUNITY

Developer Fan Community

Share images, tutorials, and local deployment experience around DGX Spark, 1Panel, and MaxKB.

DUAL-NODE LOCAL CLUSTER

Two units stacked —
flagship models, fully local.

Two DGX Spark units interconnect over ConnectX with RoCE, forming a 2 PFLOPS / 256 GB unified-memory local compute block. Community-verified: both DeepSeek-V4-Flash and Qwen3.8-Flash-Next serve stably on this dual-node setup — data never leaves the room.

Two DGX Spark units stacked and interconnected via network cables, forming a dual-node local AI cluster
2× DGX SPARK ConnectX interconnect · stack it, cluster it

DEEPSEEK-V4-FLASH-0731

Full-context concurrency
2 streams × 1M tokens (KV pool limit)
Everyday team load
6 streams × 300K tokens, comfortably resident
High-throughput tier
16 streams × 200K tokens · ~315 tok/s

QWEN3.8-FLASH-NEXT

Full multimodal
1 stream × 900K tokens (with vision) · 64 tok/s
Team concurrency
2–4 streams × 900K tokens · ~115 tok/s
Sustained speed
~47 tok/s · 70 tok/s peak
DeepSeek logo TWO-NODE · vLLM TP=2

DeepSeek-V4-Flash-0731

The official NVFP4 sparse-MLA path with DSpark speculative decoding (MTP×5) squeezes a 1M-token context into two desktop units. Long documents, repo-scale RAG, and multi-stream chat all hold up.

Context limit
1,048,576 tokens (1M)
Single-stream decode
~62–83 tok/s (128K context)
6-stream total throughput
~160–191 tok/s
High-throughput (200K · 16 streams)
~315 tok/s
KV cache pool
~2.49M tokens (~18 GiB per node)
Qwen logo TWO-NODE · SGLang TP2

Qwen3.8-Flash-Next

A full-multimodal flagship with NVFP4 expert weights + FP8 n-gram drafts, SGLang dual-node tensor parallelism, and CUDA Graph acceleration. 900K context including vision — multimodal workflows stay local.

Context limit
900K tokens (with vision input)
Single stream
64 tok/s (70 peak / 47 typical)
2–4 streams
~115 tok/s
Resident VRAM per node
~63–76 GB
Weight format
NVFP4 experts + FP8 n-gram (135 GB)

* Performance figures come from community dual-node testing (the MiaAI-Lab project repos and the NVIDIA developer forums); real-world results vary with quantization, concurrency, and context length. Individuals and small teams can get near-data-center-class LLM serving from a chassis-sized local cluster.

POWERED BY DGX-CN.COM

One compute box.
Two real workflows.

Preloaded with the 1Panel ops panel and the MaxKB knowledge base — hardware, models, and business scenarios in one local solution.

T TRANSLATION BOX

Translation Box

A localized multilingual expert for document translation, video subtitles, and live meeting interpretation. Content is processed locally, with industry glossary support.

R CONTRACT REVIEW BOX

Contract Review Box

A 24/7 AI legal specialist. Automatically benchmarks against historical contracts via RAG, flags risky clauses, and suggests revisions.

FREQUENTLY ASKED

About DGX Spark —
key questions, answered.

Is dgx-cn.com an official NVIDIA website?

No. dgx-cn.com is a Chinese fan site for DGX Spark, sharing hardware information, local deployment experience, and community application solutions with developers in China.

Which local LLMs can DGX Spark run?

It targets mainstream open-model ecosystems such as DeepSeek, Qwen, Meta Llama, and Google Gemma. The runnable scale depends on model precision, quantization, and context length.

What scenarios is DGX Spark good for?

Local LLM inference, RAG knowledge bases, agent development, multimodal prototypes, and AI demos that demand data privacy and low latency.

How do I get the AI Box images and deployment tutorials?

Scan the WeChat QR code at the bottom of the page and add the note “粉丝” (fan) to get the Translation Box and Contract Review Box images plus deployment discussions.

DGX SPARK CHINA COMMUNITY

Let's turn desktop compute
into real productivity.

Get pricing and lead times, AI Box images and deployment tutorials, and join the DGX Spark developer community in China.

  • Latest pricing & lead times
  • Tech discussion & image sharing
  • MaxKB scenario demos
WeChat inquiry QR code for the DGX Spark fan community
WECHAT Scan to chat · mention “购买” (purchase)