BentoML (Bento) is an open-source framework and managed cloud platform for deploying and scaling AI/ML model inference in production.
BentoML is an open-source model serving framework founded in 2019 by Chaoyu Yang, headquartered in San Francisco. It helps ML teams package and deploy models as production-ready inference APIs, supporting popular open models (Llama, DeepSeek, Flux, Qwen) as well as custom models across different ML frameworks.
In February 2026, BentoML was acquired by Modular; the product now operates under the 'Bento' brand while continuing to build on the original open-source BentoML project, which remains available on GitHub and is used by more than 10,000 organizations.
Bento supports multiple serving patterns — real-time interactive inference, asynchronous tasks, batch inference, and multi-step workflow orchestration — with performance tuning for latency, throughput, and cost.
Operational features include intelligent auto-scaling with cold-start acceleration and scale-to-zero, version control with rollbacks, canary/shadow/A/B testing for deployments, and observability with LLM-specific monitoring.
The core BentoML framework is free and open source, and can be self-hosted on-premises, on Kubernetes, or across multiple clouds. The managed Bento Cloud offering is billed on a usage basis, metered by the second of active GPU compute (deployments scaled to zero incur no charge), with new accounts receiving an initial free credit allotment to test deployments.
The core BentoML framework is free and open source. The managed Bento Cloud platform is billed on a usage basis for GPU compute, with free starter credits for new accounts.
BentoML was acquired by Modular in February 2026 and now operates under the 'Bento' brand as part of Modular.
Yes, the open-source framework can be self-hosted on-premises, on Kubernetes, or across multiple cloud providers.
It supports popular open models like Llama, DeepSeek, Flux, and Qwen, as well as custom models across various ML frameworks.