test_gpu_scripts

History

zulifeng 375d439abb feat: 新增 H20 支持、优化算力测试精度并修复多项稳定性问题

- gpu_specs: 新增 H20/H20-3e (中国合规版 H200) 规格定义，并修复
  GPU 名称匹配顺序，避免 "H200" 被 "H20" 子串误匹配
- benchmark(compute): 引入 L2 cache 规避的 matrix pool 轮换 +
  可选 torch.compile(max-autotune)，FP8 增加 _scaled_mm 探测，
  显著提升 FP16/BF16/FP8 实测吞吐准确性
- benchmark(memory): nvbandwidth 增加 --disableAffinity 规避
  fabricmanager NVML 不兼容；全 0 结果时自动回退到 PyTorch；
  D2D 平均值排除对角线零值
- nccl: 各通信操作 (AllReduce/AllToAll/Broadcast 等) 使用独立
  带宽阈值比例，避免 AllToAll 误报 WARN
- rdma: 仅按 link_layer=InfiniBand 过滤端口，无 IB 硬件或全 DOWN
  时直接 SKIP 而非报错
- stress: 计算矩阵尺寸封顶 4096，并改为先并发派发再统一同步，
  修复 8 卡串行执行导致 duration 严重超时的问题
- report: 兼容 RDMA SKIP 状态与 PyTorch 回退场景的 Memory 判定，
  避免回退结果被误判为 FAIL
- config: 新增 benchmark.compute.use_compile 开关

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

2026-05-12 21:41:46 +08:00

__init__.py

init: project scaffolding with README, config, and requirements

2026-04-25 17:23:27 +08:00

benchmark.py

feat: 新增 H20 支持、优化算力测试精度并修复多项稳定性问题

2026-05-12 21:41:46 +08:00

gpu_info.py

fix: resolve stress OOM, D2D efficiency calculation, NCCL execution failures

2026-05-07 18:09:22 +08:00

gpu_specs.py

feat: 新增 H20 支持、优化算力测试精度并修复多项稳定性问题