Compare the best AI infrastructure and MLOps tools for training, deploying, and monitoring ML models.
38 tools
Tools in AI Infrastructure & MLOps
MLflow — MLflow is an open-source platform for managing the machine learning and generative AI lifecycle, covering experiment tracking, model packaging, a model registry, evaluation, and deployment.
Kubeflow — Kubeflow is an open-source machine learning toolkit that runs training, tuning, pipeline orchestration, and model serving natively on Kubernetes.
DVC — DVC (Data Version Control) is a free, open-source, Git-compatible tool for versioning datasets, machine learning models, and reproducible ML pipelines.
ClearML — ClearML is an open-source MLOps and AI infrastructure platform for experiment tracking, data versioning, orchestration, and model deployment, self-hosted or managed.
Feast — Feast is an open source feature store that lets machine learning and AI teams define, store, and serve features consistently for both model training and real-time inference.
BentoML — BentoML (now Bento) is an open-source AI model serving framework and managed cloud platform for deploying and scaling machine learning inference in production.
Weights & Biases — AI developer platform for experiment tracking, model evaluation, and MLOps used by machine learning and GenAI teams to build, monitor, and govern models.
Comet — Comet is Perplexity AI's agentic, Chromium-based web browser with a built-in AI assistant that can answer questions, summarize pages, and complete multi-step tasks.
Evidently — Evidently AI is an open source ML and LLM observability and evaluation framework, paired with a hosted Cloud platform, used to test, monitor, and evaluate AI models and data pipelines with over 100 built-in metrics.
Arize Phoenix — Arize Phoenix is a free, open-source AI observability and evaluation platform for tracing, evaluating, and debugging LLM applications and AI agents.
Langfuse — Langfuse is an open-source LLM engineering platform for tracing, evaluating, and managing prompts in production AI applications, now owned by ClickHouse.
Helicone — Helicone is an open-source LLM observability platform and AI gateway that logs, monitors and helps optimize the cost and performance of applications built on OpenAI, Anthropic and other LLM providers.
OpenLLMetry — OpenLLMetry is a free, open-source SDK built on OpenTelemetry that adds observability and tracing to LLM applications with a single line of code, created by Traceloop and now backed by ServiceNow.
PromptLayer — LLM engineering platform for managing, versioning, evaluating, and monitoring prompts and AI agents in production.
LangSmith — LangSmith is LangChain's agent engineering platform for tracing, evaluating, and deploying LLM applications and AI agents in development and production.
vLLM — vLLM is an open-source, high-throughput inference and serving engine for large language models, built around the PagedAttention memory management technique.
LiteLLM — Open-source LLM gateway and Python SDK providing a unified, OpenAI-compatible API across 100-plus LLM providers.
OpenRouter — OpenRouter is a unified API gateway that lets developers access 400-plus large language models from providers like OpenAI, Anthropic, and Google through a single API with pay-as-you-go, usage-based pricing.
Together AI — Together AI is an AI-native cloud platform providing inference APIs, GPU clusters, and fine-tuning for open-source and frontier AI models, used by developers and enterprises building AI applications.
Fireworks AI — Fireworks AI is a generative AI inference platform that lets developers and enterprises run, fine-tune, and deploy open-weight and custom large language models at high speed through a single OpenAI-compatible API.
Replicate — Replicate is a cloud platform for running and fine-tuning open-source machine learning models via a simple API, billed per second of compute used.
Modal — Modal is a serverless cloud platform that lets developers run Python-native AI, ML and data workloads, including GPU jobs, with per-second billing and no infrastructure management.
RunPod — A GPU cloud platform for AI and machine learning workloads, offering on-demand Pods, autoscaling serverless inference, and multi-node clusters billed by the second.
Baseten — Baseten is a managed AI model inference platform for deploying, scaling, and serving open-source and custom machine learning models in production.
Anyscale — Anyscale is a managed platform built on the open-source Ray framework for running distributed AI training, data processing, and inference workloads at scale.
Ray — Open-source distributed computing framework for scaling Python and AI/ML workloads from a laptop to a large GPU cluster
Label Studio — Label Studio is an open-source, multi-modal data labeling and annotation platform for computer vision, NLP, audio, and LLM evaluation, built by HumanSignal.
Argilla — Argilla is an open-source data annotation and curation platform for building high-quality AI and LLM training datasets, now part of Hugging Face.
CVAT — CVAT is a free, open-source, web-based tool for annotating images and video for computer vision and machine learning, with an optional managed SaaS and enterprise self-hosted edition.
Labelbox — Labelbox is a data engine and training-data platform that helps AI teams label, manage, and evaluate the data used to build and improve machine learning models.
SuperAnnotate — SuperAnnotate is an AI data annotation and AI DataOps platform that helps teams label, curate, and manage image, video, text, and audio data to train and fine-tune machine learning and LLM models.
Roboflow — Roboflow is an end-to-end computer vision platform for annotating images, training models, and deploying them to the cloud or edge devices.
Ultralytics — Ultralytics is the company behind the YOLO family of open-source computer vision models, offering object detection, segmentation, and pose estimation plus a hosted platform for training and deploying them.
Agenta — Agenta is an open-source LLMOps platform for prompt management, evaluation, and observability of LLM applications.
GoModel — GoModel is an open-source, MIT-licensed AI gateway written in Go that gives applications one OpenAI- and Anthropic-compatible endpoint across 22+ model providers with caching and cost controls.
Khoj — Open-source, self-hostable AI second brain that lets users chat with their notes and the web, build custom agents and automate research.
LLMKube — Open-source Kubernetes operator for deploying and scaling self-hosted LLM inference across GPU fleets.
Onyx Community Edition — Onyx Community Edition is the free, MIT-licensed, open-source version of Onyx (formerly Danswer), an AI enterprise search and assistant platform that connects to 40+ knowledge sources for chat, RAG, and autonomous agents.