BentoML Review, Pricing & Features

BentoML (Bento) is an open-source framework and managed cloud platform for deploying and scaling AI/ML model inference in production.

Category
AI Infrastructure & MLOps
Pricing
Free (open source) + usage-based cloud, from Free (open source); Bento Cloud usage-based with free starter credits
Verified
Not yet
Last updated
July 18, 2026
Founded
2019
Headquarters
San Francisco, California, USA
APIOpen SourceAIFreemiumSelf-Hosted

Overview

BentoML is an open-source model serving framework founded in 2019 by Chaoyu Yang, headquartered in San Francisco. It helps ML teams package and deploy models as production-ready inference APIs, supporting popular open models (Llama, DeepSeek, Flux, Qwen) as well as custom models across different ML frameworks.

In February 2026, BentoML was acquired by Modular; the product now operates under the 'Bento' brand while continuing to build on the original open-source BentoML project, which remains available on GitHub and is used by more than 10,000 organizations.

Key Features

Bento supports multiple serving patterns — real-time interactive inference, asynchronous tasks, batch inference, and multi-step workflow orchestration — with performance tuning for latency, throughput, and cost.

Operational features include intelligent auto-scaling with cold-start acceleration and scale-to-zero, version control with rollbacks, canary/shadow/A/B testing for deployments, and observability with LLM-specific monitoring.

Pricing

The core BentoML framework is free and open source, and can be self-hosted on-premises, on Kubernetes, or across multiple clouds. The managed Bento Cloud offering is billed on a usage basis, metered by the second of active GPU compute (deployments scaled to zero incur no charge), with new accounts receiving an initial free credit allotment to test deployments.

Key Features

Pros & Cons

Pros

  • Core framework is free and open source
  • Framework-agnostic — supports many popular open and custom models
  • Flexible deployment: self-hosted or managed cloud
  • Usage-based cloud billing avoids paying for idle compute (scale-to-zero)

Cons

  • Bento Cloud GPU pricing is usage-based and can be harder to predict than flat subscription pricing
  • Recent acquisition by Modular (Feb 2026) means the product and roadmap are in transition
  • Self-hosting requires Kubernetes/infrastructure expertise to get full value

Pricing

Frequently Asked Questions

Is BentoML free?

The core BentoML framework is free and open source. The managed Bento Cloud platform is billed on a usage basis for GPU compute, with free starter credits for new accounts.

Who owns BentoML now?

BentoML was acquired by Modular in February 2026 and now operates under the 'Bento' brand as part of Modular.

Can BentoML be self-hosted?

Yes, the open-source framework can be self-hosted on-premises, on Kubernetes, or across multiple cloud providers.

What models does BentoML support?

It supports popular open models like Llama, DeepSeek, Flux, and Qwen, as well as custom models across various ML frameworks.

Related Tools