OpenAI still offers the widest single-vendor model range and the largest ecosystem around it. The weights are closed, so there is no self-hosting path if that matters to you. The entries below separate the providers that host their own models from the ones that serve other people's, because the two carry very different lock-in.
8 tools reviewedLast reviewed Ranked by us, not by votes
What counts as LLM API providers
LLM API providers sell programmatic access to large language models running on someone else's infrastructure. You send a prompt over HTTP and get generated text back, paying for usage rather than buying and operating GPUs yourself. The category covers three shapes: first-party labs that serve the models they build, cloud platforms that resell many vendors' models behind one contract and one bill, and inference specialists that host open-weight models for speed or price.
How we judged them
Entries were judged on model quality and range, latency, the shape of pricing and how many separate meters it involves, rate limit and quota behaviour, data handling (in particular whether inputs are used to improve the vendor's models), and enterprise terms such as compliance coverage, deployment options and data residency. Every entry was checked against the vendor's own site during research, and nothing read second-hand is marked verified. Prices are described by shape only, because published rates change often enough that a specific figure would be stale before the page aged. Ordering reflects breadth of fit for a general buyer, so specialists rank lower because they serve narrower needs, not because they perform worse at what they do.
All 8 LLM API providers in this guide, in ranked order.
platform.openai.com · Usage-based per token, metered separately for input, output and cached input, with a cheaper asynchronous batch tier, an intermediate flex tier, and a higher-priced fast tier; some tools billed per call, per minute, or per session.
Best for Teams that want the widest single-vendor model range and the largest surrounding ecosystem of SDKs, tutorials, and third-party integrations.
OpenAI sells API access to its own GPT model family alongside speech, image, and embedding models.
Strengths
Broad lineup spanning text, reasoning, audio transcription, image, video, and embedding models under one account
Its request and response format has become the de facto standard that many competitors deliberately copy, which lowers switching costs in both directions
Documented cost controls: asynchronous batch processing at roughly half price, cached input at a large discount, and an intermediate flex tier
Usage tiers that raise rate limits automatically as account spend grows, rather than requiring a sales conversation
Managed fine-tuning with a discount available in exchange for data sharing
Where it falls short
Weights are closed, so there is no self-hosting option
Older model snapshots are retired on a published deprecation schedule, so integrations need periodic migration work
Billing spans many separate meters (input tokens, output tokens, cached input, cache writes, per-call tool use, file storage, per-session containers), which makes forecasting harder than a single token rate
The fast processing tier is billed at a multiple of standard rates
Fine-tuning adds hourly training charges on top of standard inference rates for the deployed model
claude.com · Usage-based per token with separate input and output rates by model tier, an asynchronous batch path at about half the standard rate, and prompt caching billed separately for cache writes and cache reads.
Best for Long-running agent and coding workloads, and buyers who want one model family reachable through several different cloud contracts.
Anthropic sells API access to its Claude model family for text and image input with text output.
Strengths
Tiered lineup running from a top capability model down to a fast low-cost one, so workloads can be routed by difficulty
Large context windows on current models, with the leading tiers documented at one million tokens
The same models are served through Amazon Bedrock, Google Cloud, and Microsoft Foundry, so procurement can route through an existing cloud agreement instead of a new vendor contract
Batch processing at roughly half price plus prompt caching priced separately for writes and reads
Model IDs are pinned snapshots rather than evergreen pointers, so production behaviour does not shift underneath you without an explicit migration
Where it falls short
Closed weights with no self-hosting option
Current models accept text and image input and return text only: no audio or video input, and no image generation
Because every model ID is a pinned snapshot, deprecations force periodic migration work
One model in the current lineup is limited availability and invitation only, with no self-serve sign-up
A tokenizer change in a recent model generation means the same text can produce roughly 30 percent more tokens than on older models, which affects cost comparisons
ai.google.dev · Free tier with capped rate limits, then usage-based per token across standard, batch, flex and priority modes; caching billed per token plus hourly storage; certain tools billed per request or per thousand queries.
Best for Cost-sensitive and high-volume workloads, free-tier prototyping, and teams already committed to Google Cloud.
Google serves its Gemini models through a developer API via AI Studio and through Vertex AI for enterprise deployments.
Strengths
A genuine free tier for prototyping, which is uncommon among frontier labs
Four processing modes (standard, batch, flex, priority) that let cost and latency be traded off per workload
Large context windows and native multimodal coverage including image, video, and audio
Grounding with Google Search and Maps available as a billed built-in tool rather than a separate integration
Two front doors covering different buyers: AI Studio for developers, Vertex AI for enterprise governance
Where it falls short
Google states that free-tier content is used to improve its products, making the free tier unsuitable for confidential data (the paid tier states the opposite)
The split between the Gemini API and Vertex AI means different auth and quota models, adding work when a prototype moves to production
Closed weights, so the Gemini models cannot be self-hosted
Context caching is billed twice over: per cached token and by hourly storage
Numerous non-token meters (per image, per second of video, per thousand grounding queries) complicate cost modelling
aws.amazon.com · Usage-based per token on demand, a discounted asynchronous batch mode, and a commitment-based provisioned throughput option billed by reserved capacity; charges appear on the existing AWS bill.
azure.microsoft.com · Usage-based per token for serverless model endpoints, hourly compute charges for managed deployments, and commitment-based provisioned throughput units; billed through the Azure subscription.
Best for Microsoft-estate organisations that want OpenAI and other vendors' models under Entra identity, Purview data governance, and Defender security.
Microsoft's platform for building and governing AI applications and agents, providing API access to models from many vendors on Azure.
Strengths
Very large catalogue spanning multiple labs including OpenAI, Anthropic, Meta, Google, xAI and Hugging Face, plus Microsoft's own models
Native integration with Entra ID, Purview and Defender gives one governance and identity plane rather than a separate one per vendor
Several hosting paths including serverless endpoints, container apps, and app service integration
Provisioned throughput available for predictable capacity
Built-in tracing and monitoring, plus runtime protections aimed at prompt attacks and sensitive data leakage
Where it falls short
Requires an Azure tenant and Azure-shaped identity, so it does not suit non-Microsoft estates
The product has been renamed more than once (Azure OpenAI Service, then Azure AI Foundry, then Microsoft Foundry), leaving inconsistent naming across documentation, SDKs and older guides
Model availability and quota vary by Azure region
Provisioned throughput units require a capacity commitment rather than pure pay-as-you-go
The catalogue runs to thousands of models and documentation depth is uneven across them
groq.com · Free plan with capped rate limits, then a usage-based paid Developer plan charged per token (and per second for audio models), with enterprise arrangements quoted on request.
together.ai · Usage-based per token, per image, per video or per minute for serverless; hourly per-GPU billing for dedicated instances and clusters with lower rates for longer reservations; fine-tuning billed per million training tokens; storage billed per GiB-month.
Best for Teams standardising on open-weight models that want one vendor covering serverless calls, dedicated endpoints, fine-tuning, and training clusters.
Together AI hosts open-weight models as serverless endpoints and also rents dedicated GPU capacity for inference and training.
Strengths
One catalogue covering chat, vision, image, audio, video, transcription, embeddings, rerank and moderation
Three capacity models (serverless per token, hourly dedicated instances, reserved GPU clusters) so a workload can scale up without changing vendor
Dedicated instances support custom models, not just the public catalogue
Fine-tuning covering supervised training, full fine-tuning, and direct preference optimisation
Published rate card across all three capacity models rather than sales-gated pricing
Where it falls short
The catalogue centres on open-weight models, so closed frontier models are largely unavailable
Dedicated instances and GPU clusters bill by the hour whether or not they are serving traffic
Reserved cluster pricing involves commitments ranging from roughly a week to six months or more
Storage is billed separately per GiB-month on top of compute
Serverless model availability shifts as open-weight releases arrive and older ones are retired
mistral.ai · Usage-based per token for the hosted API, with separate subscription plans, and quoted pricing for enterprise and self-managed deployments.
Best for European buyers with data residency or sovereignty requirements, and teams that want the option to self-host the same model family they call over an API.
A French model developer offering its own models through a hosted API, with part of the lineup released under open licences and deployable on your own infrastructure.
Strengths
Mixes open-weight models (Apache 2.0 and modified MIT) with commercial ones, so some models can run on hardware you control
Self-managed and on-premise deployment offered explicitly, with the stated position that nothing leaves your perimeter
European infrastructure and in-region inference for residency and sovereignty requirements
Specialised models beyond chat, including OCR, audio transcription, code generation, embeddings, and safety classifiers
Hybrid deployment intended to keep one codebase working across cloud and on-premise environments
Where it falls short
Smaller model catalogue than the large US labs, with fewer options at the top capability tier
Licences vary per model (Apache 2.0 on some, modified MIT on others, commercial on the rest), so open weights cannot be assumed across the lineup
Enterprise and self-managed deployment pricing is not published and requires contacting sales
Self-hosting the open-weight models means supplying and operating your own GPU capacity, moving cost and reliability onto your team
Smaller third-party integration ecosystem than the OpenAI-compatible incumbents