多码网
返回 AI
AIlibraryPython

vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

77,621 stars15,919 forks4,367 issuesApache-2.0更新于 2026/4/21
amdblackwellcudadeepseekdeepseek-v3gptgpt-ossinferencekimillamallmllm-servingmodel-servingmoeopenaipytorchqwenqwen3tputransformer
暂无详细内容

同类推荐