Public APIs
All posts

Best OpenAI API Alternatives in 2026, Ranked

Quick answer: Google's Gemini API is the best all-round OpenAI alternative, with an OpenAI-compatible endpoint and free Flash models. DeepSeek is the cheapest serious paid option at $0.15 per million input tokens off-peak. Groq and Cerebras give you free open-weights models fast, and Ollama or vLLM remove the API bill entirely if you can host the model yourself.

Most "OpenAI API alternatives" pages are written by a vendor on the list. This one is not: we run a directory of public APIs, OpenAI included, and we list the competitors next to it. Every price, limit and base URL below was read from the provider's own documentation or API on 30 September 2026.

What makes something a real OpenAI API alternative?

Three things, and most lists only check the first:

  1. Model quality for your task. This moves every few months, so test on your own prompts rather than trusting a leaderboard.
  2. Drop-in compatibility. If the provider speaks the OpenAI Chat Completions format, switching is a base URL and a key. If it does not, you are rewriting your client code.
  3. Where the model can run. A hosted API, an aggregator in front of many hosted APIs, or open weights you can run on your own hardware. Only the last one survives a vendor changing its terms.

What are the best OpenAI API alternatives?

Ranked for a developer who wants to cut cost or dependency on OpenAI without rewriting their application.

  1. Google Gemini API. The strongest general replacement. Its OpenAI compatibility layer lives at generativelanguage.googleapis.com/v1beta/openai/, and the free tier lists every current Flash and Flash-Lite model as free of charge for input and output. On the paid tier, Gemini 3.1 Flash-Lite costs $0.25 per million input tokens and $1.50 per million output tokens.
  2. DeepSeek. The price floor for a capable hosted model. deepseek-flash costs $0.15 per million input tokens off-peak ($0.30 at peak) and $0.60 per million output, with cache hits at a fraction of a cent. It serves both the OpenAI format at api.deepseek.com and an Anthropic-format endpoint, so it slots into almost any client.
  3. Anthropic Claude. The pick when output quality on long, careful work matters more than price. There is an OpenAI SDK compatibility layer at api.anthropic.com/v1/, but Anthropic's own docs call it intended for testing and comparing models rather than long-term production. It ignores strict on tool schemas and does not support prompt caching, so plan to move to the native API if you stay.
  4. Groq. Free, fast inference on open-weights models through api.groq.com/openai/v1. The free plan allows openai/gpt-oss-120b 30 requests per minute and 1,000 per day. Good for prototypes and for latency-sensitive features.
  5. OpenRouter. One key, one OpenAI-compatible endpoint, and 464 models from dozens of providers when we queried its model list today. 16 of those are :free variants, capped at 50 requests a day, or 1,000 a day once you have bought at least $10 of credit. The best choice if you want to switch models by changing one string.
  6. Cerebras. The other free, very fast host for open weights, OpenAI-compatible at api.cerebras.ai/v1. The free trial gives gpt-oss-120b 1 million tokens a day at 5 requests per minute: generous volume, tight concurrency.
  7. Mistral. A European provider with open-weights and proprietary models and a free Experiment plan for trying the API. Read the terms before sending real data: requests on the free plan can be used to train Mistral's models.
  8. Cohere. Strongest on retrieval work: embeddings, reranking and grounded generation. It offers a Compatibility API at api.cohere.ai/compatibility/v1 for the OpenAI SDK, and it is listed in our directory as Cohere.
  9. Ollama. Not a vendor at all. It runs open-weights models on your own machine and exposes OpenAI-compatible /v1/chat/completions, /v1/embeddings and /v1/responses endpoints on localhost:11434. The client needs an API key value, and Ollama ignores it.
  10. vLLM. The self-hosted answer for production traffic: an OpenAI-compatible server built for throughput on your own GPUs. Heavier to operate than Ollama, much faster under load.

Which options are free?

Free means different things here, so it is worth being precise:

| Option | What is free | The catch | |---|---|---| | Gemini API | Current Flash and Flash-Lite models, input and output | Free-tier rate limits, and Google's free-tier data terms | | Groq | gpt-oss-120b: 30 requests a minute, 1,000 a day | Only the models Groq hosts | | Cerebras | 1M tokens a day on gpt-oss-120b | 5 requests a minute | | OpenRouter | 16 :free models | 50 requests a day without credit | | Mistral | Experiment plan | Phone verification; requests may train models | | Ollama, llama.cpp, vLLM | Everything | You provide the hardware |

If your app only needs a small model for classification, extraction or summaries, the free tiers above cover a real side project. For anything customer-facing, budget for paid usage from day one: free-tier limits are per account, and none of them come with an uptime promise.

Can you use the OpenAI SDK with other providers?

Yes. Seven of the hosted providers ranked above publish an OpenAI-compatible base URL, each quoted in this post, and Ollama and vLLM expose one on your own machine. In most codebases the switch is two lines:

from openai import OpenAI

client = OpenAI(
    base_url="https://api.groq.com/openai/v1",  # the provider's compatible endpoint
    api_key="YOUR_PROVIDER_KEY",
)

Compatible is not identical, though. Parameters a provider does not support are usually ignored without an error rather than refused, which is how structured output quietly stops being structured. After switching, test tool calling, JSON output and streaming explicitly, not just a plain chat message.

Which self-hosted servers are actually maintained?

Open weights are only an exit route if the server running them is alive. Read live from GitHub today:

| Project | Stars | Last push | Status | |---|---|---|---| | Ollama | 181,959 | 2026-09-30 | Active | | llama.cpp | 129,969 | 2026-09-30 | Active | | vLLM | 92,990 | 2026-09-30 | Active | | LiteLLM (proxy) | 59,931 | 2026-09-30 | Active | | LocalAI | 49,341 | 2026-09-30 | Active | | SGLang | 36,668 | 2026-09-30 | Active | | Hugging Face TGI | 10,884 | 2026-03-21 | Archived |

The last row is the one to notice. Text Generation Inference was the default recommendation for self-hosting open models for years, still appears on older lists, and its repository is now archived. If you are running it, plan a move to vLLM or SGLang. LiteLLM is in the table because it solves a different problem: a proxy that puts one OpenAI-format endpoint in front of all the providers above.

What our directory shows about AI APIs

Our AI category lists 33 APIs out of the 1,567 public APIs published today, and 30 of those 33 authenticate with a plain API key. That is the other practical reason OpenAI-compatible providers are easy to swap: you are replacing a key in an environment variable, not re-doing an OAuth integration.

The category also shows that "AI API" is a wider shelf than chat models. It includes aggregators like Eden AI and AI/ML API that route to many model vendors, computer vision and moderation APIs, and task APIs for summarising or translating text. If the job is one narrow task, a task API from the text analysis category can be cheaper and more predictable than prompting a general model. And if you are staying put, OpenAI is listed there too.

Which one should you pick?

| If this is you | Pick | |---|---| | Want the closest general replacement with a free tier | Gemini API | | Cost is the whole question | DeepSeek | | Quality on long, careful work matters most | Anthropic Claude, native API | | Prototyping on free open-weights models | Groq or Cerebras | | Want to switch models without switching vendors | OpenRouter | | Data cannot leave your machines | Ollama for development, vLLM for production |

Frequently asked questions

Is there a free alternative to the OpenAI API?

Yes. Google's Gemini API lists its current Flash models as free of charge on the free tier, Groq and Cerebras host open-weights models like gpt-oss-120b for free within rate limits, and OpenRouter offers 16 free models. Running a model locally with Ollama or llama.cpp costs nothing beyond your hardware.

What is the cheapest OpenAI API alternative?

Among capable hosted models, DeepSeek: deepseek-flash costs $0.15 per million input tokens and $0.60 per million output tokens off-peak, and cached input costs a fraction of a cent. Prices move often, so check the provider's pricing page on the day you commit.

Is there an open source alternative to the OpenAI API?

The models are open weights rather than a single open source API, but the servers that run them are open source and speak the OpenAI format. Ollama and llama.cpp suit a laptop or a single server, and vLLM and SGLang suit production GPUs. Avoid starting on Hugging Face TGI, whose repository is now archived.

Can I switch from OpenAI without changing my code?

Mostly. Point the OpenAI SDK at another provider's compatible base URL, swap the API key and change the model name. Then test tool calls, JSON output and streaming, because unsupported parameters are often ignored silently instead of raising an error.

Does the Claude API work with the OpenAI SDK?

Yes, through Anthropic's compatibility layer at api.anthropic.com/v1/. Anthropic describes it as a way to test and compare models rather than a long-term production integration, and features like prompt caching and strict tool schemas need the native API.