跳到正文
原文
Google DeepMind·· 2026-06-11精选AI 评分71

Google DeepMind 发布 DiffusionGemma:文本生成最高提速 4 倍

DiffusionGemma: 4x faster text generation

AI 导读

Google DeepMind 发布实验性开源模型 DiffusionGemma,采用文本扩散方式并行生成整块文本,在专用 GPU 上最高实现 4 倍推理提速,单张 NVIDIA H100 上超过 1000 tokens/秒,RTX 5090 上超过 700 tokens/秒。

推荐理由

官方给出 26B MoE 文本扩散模型的速度数据、硬件门槛与质量取舍,便于判断本地低并发场景是否值得尝试。

来源:Google DeepMind · deepmind.google