DGX Spark · GB10 / SM121

MiniMax H3 on DGX Spark

896×512 · 124 frames · 3 prompts · 4 DiT weight formats × 2 attention backends

固定 workload

任务Text-to-video + audio
视频896×512,124 帧,24 FPS,约 5.17 秒
采样20 steps,Euler / simple,CFG 1.0
输入Ocean whale / Chef close-up / Rainy city car
Seeds42 / 31415 / 27182
ConditionerQwen3-VL-32B NVFP4 AWQ weights,全部固定
Warmup每组合 1 个 2-step 请求,排除计时
正式结果8 组合 × 3 prompts = 24 个视频
E2E 从提交请求到视频与音频输出完成,包括 text encode、denoise、VAE decode 与保存。

变量与来源

变量配置来源
DiT weightsBF16(官方原版)Comfy-Org / MiniMax-H3
DiT weightsINT8 ConvRotComfy-Org / MiniMax-H3
DiT weightsFP8 scaledComfy-Org / MiniMax-H3
DiT weightsNVFP4lilcheaty / MiniMax-H3-NVFP4
AttentionBF16 attention--
AttentionSage2++(QK INT8 / PV FP8)SageAttention 2.2
INT8 / FP8 / NVFP4 均指 DiT weights;Sage2++ 不等于“FP8 attention”,本实验也没有 NVFP4 attention。
01 · Precision map

权重精度、运行时 dtype 与 Attention 精度分开看

精度色签:BF16FP16FP8NVFP4INT8FP32

Qwen3-VL conditioner

WeightsNVFP4-AWQ
Runtime dtypeFP16

所有 8 个组合固定不变。

MiniMax-H3 DiT

WeightsBF16 / INT8 / FP8 / NVFP4
Runtime tensorsBF16

低比特 checkpoint 使用各自的量化 linear path;不是整条 DiT 都以该低比特计算。

Attention

Dense SDPABF16
Sage2++QK INT8 · PV FP8

本实验没有 NVFP4 attention。

VAE decode

Video VAEFP16
Audio VAEFP32

视频与音频 decode 均包含在 E2E 中。

Cold-start 统一内存峰值

每组合重建容器,采样间隔 0.25 秒;同一 Ocean whale / seed 42 / 20-step 请求。System peak 包括服务与模型启动到请求完成的主机总内存压力。

DiT weightsAttentionCheckpoint filesSystem peakContainer cgroup peakvs BF16 baseline
BF16BF16 attention81.752 GiB
113.928 GiB
85.714 GiBbaseline
BF16Sage2++81.752 GiB
113.895 GiB
85.869 GiB−0.033 GiB
INT8 ConvRotBF16 attention39.554 GiB
90.512 GiB
43.170 GiB−23.415 GiB · 20.55%
INT8 ConvRotSage2++39.554 GiB
90.540 GiB
43.305 GiB−23.388 GiB · 20.53%
FP8 scaledBF16 attention39.542 GiB
91.054 GiB
43.157 GiB−22.874 GiB · 20.08%
FP8 scaledSage2++39.542 GiB
90.521 GiB
43.249 GiB−23.406 GiB · 20.54%
NVFP4BF16 attention31.692 GiB
75.916 GiB
36.927 GiB−38.012 GiB · 33.36%
NVFP4Sage2++31.692 GiB
76.015 GiB
36.405 GiB−37.913 GiB · 33.28%
DGX Spark 使用统一内存。GB10 的 nvidia-smi 不提供独立 framebuffer memory;这里的 System peak = MemTotal − min(MemAvailable)。Container cgroup peak 仅作交叉检查,因为 GPU/UVM 驱动分配不会全部计入容器 cgroup。Checkpoint files 是 DiT + 固定 Qwen + 两套 VAE 文件之和,不等于运行峰值。
表格与视频卡片中的 FP8 / NVFP4 默认指 DiT weights;只有明确写 Sage2++ 时,才表示 QK INT8 / PV FP8 的 Attention 路径。
02 · Benchmark · mean of 3 prompts

NVFP4 weights + Sage2++ 平均 200.38 秒

DiT weightsAttentionMean E2ESage 加速vs BF16 baselineMean PSNR
BF16BF16 attention340.488s1.000×
BF16Sage2++306.451s1.111×1.111×25.970 dB
INT8 ConvRotBF16 attention333.178s1.022×25.962 dB
INT8 ConvRotSage2++296.009s1.126×1.150×24.512 dB
FP8 scaledBF16 attention268.389s1.269×22.050 dB
FP8 scaledSage2++232.719s1.153×1.463×22.166 dB
NVFP4BF16 attention232.360s1.465×17.767 dB
NVFP4Sage2++200.383s1.160×1.699×18.289 dB
1.699×最快组合相对官方 BF16 weights + BF16 attention baseline
11–16%Sage2++ 在四套 weights 上的 matched E2E 延迟降低

PSNR:每个 prompt/seed 逐帧对官方 BF16 weights + BF16 attention 输出计算,再对三条视频的 PSNR(dB) 取算术平均。它反映轨迹相似度,不是感知质量分数。

03 · Attention latency profile

Attention 占 DiT forward 的 24.1%

50 次 main attention 调用的 CUDA-event 归因

Attention 24.1%
其余 DiT 75.9%
21.106s完整 DiT forward
5.087sMain attention
16.018sAttention API 外
Numerator 包括 selected attention API 内的 layout/API overhead,不包括 QKV 与 output projections。

为什么 kernel 加速落到 E2E 后变小

实测 matched 50-step E2E
1.080×
由 denoise saving 反推的 integrated attention
1.538×
若 integrated attention 真能保持 2.8×
约 1.158× E2E
这张 profile 来自同一 GB10、同一 896×512/124-frame workload 的 独立 SGLang mixed-FP8 full-resident path,不是第 1 页 community benchmark 的直接计时分解。
04 · Case 1 / 3 · seed 42

Case 1 · Ocean whale

BF16 weights · BF16 attn342.481s · PSNR ∞
INT8 weights · BF16 attn330.472s · 36.192dB
FP8 weights · BF16 attn266.382s · 28.505dB
NVFP4 weights · BF16 attn232.339s · 21.757dB
BF16 weights · Sage2++304.443s · 34.425dB
INT8 weights · Sage2++294.671s · 33.514dB
FP8 weights · Sage2++232.557s · 28.631dB
NVFP4 weights · Sage2++200.390s · 23.461dB
前四项:BF16 attention后四项:Sage2++(QK INT8 / PV FP8)可单独播放,或用右上角按钮同步播放
05 · Case 2 / 3 · seed 31415

Case 2 · Chef close-up

BF16 weights · BF16 attn338.492s · PSNR ∞
INT8 weights · BF16 attn334.517s · 20.693dB
FP8 weights · BF16 attn268.387s · 19.395dB
NVFP4 weights · BF16 attn232.366s · 15.579dB
BF16 weights · Sage2++304.467s · 20.705dB
INT8 weights · Sage2++296.684s · 19.587dB
FP8 weights · Sage2++234.919s · 19.517dB
NVFP4 weights · Sage2++200.335s · 15.274dB
前四项:BF16 attention后四项:Sage2++(QK INT8 / PV FP8)可单独播放,或用右上角按钮同步播放
06 · Case 3 / 3 · seed 27182

Case 3 · Rainy city car

BF16 weights · BF16 attn340.491s · PSNR ∞
INT8 weights · BF16 attn334.546s · 21.001dB
FP8 weights · BF16 attn270.399s · 18.250dB
NVFP4 weights · BF16 attn232.376s · 15.966dB
BF16 weights · Sage2++310.441s · 22.782dB
INT8 weights · Sage2++296.671s · 20.435dB
FP8 weights · Sage2++230.681s · 18.350dB
NVFP4 weights · Sage2++200.425s · 16.131dB
前四项:BF16 attention后四项:Sage2++(QK INT8 / PV FP8)可单独播放,或用右上角按钮同步播放
07 · Conclusion

效率最优不等于像素轨迹最接近

效率优先:NVFP4 DiT weights + Sage2++

三 prompt mean;相对官方 BF16 baseline 降低 41.15% 延迟。

200.38s
更保守:FP8 scaled DiT weights + Sage2++

相对 baseline 1.463×;mean PSNR 22.166 dB。

232.72s
只换 attention:BF16 weights + Sage2++

保留官方 weights;E2E 从 340.49s 降到 306.45s。

1.111×

如何解读

  • NVFP4 + Sage2++ 是 3 个 case 上一致最快的组合,单 case E2E 为 200.335–200.425 秒。
  • NVFP4 的 mean PSNR 更低,说明相对 BF16 reference 的像素轨迹差异更大;不等价于视频语义或感知质量必然更差。
  • 三个 case 的完整视频已分别放在上方三个区块,应结合播放结果而不是只看 PSNR。
  • Attention profile 显示 75.9% 的 DiT 时间在 attention API 外,因此优化 QKV/output/MLP 与编译路径同样重要。
质量结论的当前范围:3 prompts × 1 seed。更强结论仍需更多 prompts/seeds、人工偏好与声明明确的感知指标。