NVIDIA Developer Blog· Tanya Lenz·· 10 天前AI 评分36
NVIDIA Dynamo-Triton 集成 TensorRT 多设备推理,支持单网络跨多 GPU 运行
Simplifying Model Serving Across Multiple GPUs with NVIDIA TensorRT Multi-Device Integration in NVIDIA Dynamo-Triton
AI 导读
NVIDIA 在 Dynamo-Triton 中集成 TensorRT 多设备推理,让单个 TensorRT 网络借助 NCCL 分布式集合通信跨多块 GPU 执行,同时保留 TensorRT 的推理优化。该能力自 TensorRT 11.0 起获得完整支持,用于应对生成式 AI 超出单 GPU 的算力与显存需求。
来源:NVIDIA Developer Blog · developer.nvidia.com