Newsletter
Join the Community
Subscribe to our newsletter for the latest news and updates
Source: LMSYS primary technical report on Ling-3.0-flash batch-1 decoding.
Ling-3.0 Flash DSpark has an attention-grabbing number: 1,120 output tokens per second from one request. That is fast enough to sound like a new baseline for AI inference. It is not. The result comes from a carefully optimized, four-Blackwell, synthetic benchmark—and that narrowness is exactly why the engineering work is useful.
Ling-3.0 Flash DSpark is a draft model and serving path built to accelerate single-request generation for Ling-3.0-flash. In an August 21, 2026 primary technical report, the Ant Ling and RadixArk SGLang teams say their tuned NEXTN path moved from 288 to 606 output tokens per second, while DSpark reached 1,120 tokens per second and 0.78 ms mean time per output token on four NVIDIA Blackwell GPUs. Those are vendor-and-partner-reported measurements, not an independent universal speed rating. The controlled run used one concurrent request, greedy decoding, a synthetic random workload, fixed 8,192-token inputs, and 1,024-token outputs. NVIDIA's CUDA guide supports one mechanism in the optimization stack—overlapping dependent kernels—but does not validate the model benchmark. The practical takeaway is narrower: inference teams with comparable Blackwell hardware can reproduce the published setup; everyone else should treat 1,120 tok/s as a test target, not an expected user experience.
This is a news analysis of public technical evidence, not a hands-on performance review. The headline measurements come from the teams that built the model and serving path. The article treats the LMSYS post and Ant Ling release as primary evidence for their exact setup, NVIDIA's CUDA guide as external support for the PDL mechanism, and IT Home only as independent confirmation that Ling-3.0 flash-family checkpoints are publicly available. No independent source reproduced the four-Blackwell benchmark, so every speed comparison below stays tied to the authors' hardware, command, workload, and date.
The teams did not get the headline number from one magic kernel. They describe a sequence of changes to the NEXTN speculative-decoding path, followed by a switch to a DSpark draft model.
The cleanest comparison is tuned NEXTN versus DSpark. The report says both used the same machine, command, fixed random workload, and 1,000 requests. That makes the 1.53 ms to 0.78 ms mean-TPOT change more useful than comparing isolated peak numbers from different runs.
The baseline-to-tuned NEXTN path is still important. It shows how much latency can hide outside the model's obvious matrix multiplications: a CPU waiting on device values, gaps between GPU graphs, small kernels that cannot overlap, and choices that push work onto the critical path.
Speculative decoding uses a smaller draft process to propose several tokens, then asks the target model to verify them together. If the target accepts a longer block, the server commits more tokens per expensive verification step. DSpark changes the drafting strategy so the team reports an average accepted length of 9.95 tokens on its synthetic workload.
Before DSpark, the teams attacked the serving loop itself. Their report says a blocking device-to-host read forced the CPU to wait for GPU progress each step. Removing that wait let the host prepare and enqueue later work while the GPU was still busy.
They also used Programmatic Dependent Launch, or PDL, to start dependent kernels with less dead space between them. NVIDIA documents that PDL can overlap a secondary kernel with unfinished work from a primary kernel on compute-capability 9.0 or newer hardware. NVIDIA also warns that this creates an opportunity for overlap, not a guarantee. Kernel shape and remaining work still decide whether it helps.
The reusable lesson is not “install DSpark and expect 1,120 tok/s.” It is profile the whole decode step before buying more hardware. A host synchronization point can erase the overlap you thought CUDA graphs provided. A synthetic batch-1 benchmark can also hide the behavior that matters in production: mixed prompt sizes, stochastic sampling, multiple users, request scheduling, memory pressure, and tool-calling pauses.
The report includes launch commands and a fixed benchmark recipe. That is unusually helpful. An inference team can run the published setup first, verify that its hardware and software reproduce the same order of magnitude, and then change one workload dimension at a time.
For a normal AI user, tokens per second describe only the generation phase. They do not measure time to first token, tool latency, network time, answer quality, or whether an agent chooses the right action. A model that streams extremely fast can still feel slow if it spends seconds preparing a prompt or waiting on external tools.
The result still matters because it shows how serving software can change the experience without retraining the full target model. Faster single-request decode could make local coding assistants, interactive agents, and long answers feel more immediate—if the gain survives realistic hardware and workloads.

Source: LMSYS primary technical report. The chart reports the authors' four-B200, TP4, BF16, concurrency-one setup and has not been independently reproduced.
The chart supports three bounded statements: tuned NEXTN beat the reported baseline, DSpark beat tuned NEXTN in the controlled 1,000-request comparison, and the reported best configuration reached 1,120 output tok/s with 0.78 ms mean TPOT.
It does not establish a universal Ling-3.0-flash speed. Every headline run used four NVIDIA Blackwell GPUs, tensor parallelism across four devices, BF16, greedy decoding, concurrency one, and a synthetic random workload with fixed 8,192-token inputs and 1,024-token outputs. The authors explicitly say accept length depends on prompt and output distribution.
Independent evidence is narrower. IT Home reported that Ling-3.0 flash-family Base checkpoints are publicly available for further training and research, but it did not rerun this serving benchmark. The bounded X and Reddit check also found adjacent hardware reports, not a comparable four-Blackwell reproduction. The chart is credible primary evidence for the team's setup, not independent proof of what your server will deliver.
This sequence separates three questions that marketing charts often blur: can the benchmark be reproduced, does the optimization transfer, and does the product experience improve?
The public report is detailed about its own setup, but several production questions remain open:
Those gaps do not invalidate the work. They define the next test. This is a strong engineering report with a reproducible starting point, not a neutral declaration that one model now generates at 1,120 tok/s everywhere.
Bottom line: DSpark is worth reproducing if single-request latency is a real bottleneck in your SGLang deployment. For everyone else, the report is a valuable optimization playbook, not a speed promise.
Is Ling-3.0 Flash DSpark available now?
Yes. Ant Ling announced the Ling-3.0-flash-dspark release on August 21, 2026, and the primary technical report includes an SGLang launch recipe. Availability does not guarantee that every SGLang version, accelerator, or quantization path supports the same setup.
Does Ling-3.0 Flash really generate 1,120 tokens per second?
The Ant Ling and SGLang teams report 1,120 output tok/s in a specific synthetic run on four Blackwell GPUs at concurrency one. No qualifying independent reproduction was found. Treat the number as a reproducible hypothesis for comparable systems, not a general speed rating.
What is TPOT, and why is it different from throughput?
TPOT means time per output token during generation. The SGLang report notes that its TPOT excludes time to first token, while output throughput divides total generated tokens by total benchmark wall time. They answer related but different performance questions.
Will DSpark make a local AI app faster on consumer hardware?
Possibly, but this report does not prove it. Consumer GPUs and unified-memory systems have different bandwidth, memory, precision, and software constraints. Reproduce a supported configuration on your hardware and measure finished-task latency before drawing a conclusion.
Should a production inference team switch to DSpark now?
Only after a controlled comparison on real traffic shapes. Start with the published benchmark, then test your prompt distribution, concurrency, sampling, tail latency, quality, memory, and failure recovery. Keep the current serving path available until the new path passes those checks.
Discover practical AI developer tools at AIToolHunt.