LLM API

Mistral API pricing

Mistral API pricing separates standard input and output tokens from cached-input and Batch processing discounts.

Decision summary: Consider Mistral when its model portfolio, open-weight options, deployment choices, and token discounts fit the workload and operating requirements.
Pricing checked 2026-07-15high confidence

Pricing overview

Usage-based per-model token pricing with documented Batch and cached-input discounts; other products and deployment modes use separate terms.

Source-tracked model data

Mistral API model prices in one table

Example cost uses 1,000,000 input tokens and 300,000 output tokens at each model's default tracked rate, with no cache discount.

ModelInput priceCached inputOutput priceContextExample cost
Mistral Medium 3.5
Stable
$1.50 / 1M$0.15 / 1M$7.50 / 1M256K tokens$3.75

Cost boundary: this example compares token charges only. It does not assume equal output quality, latency, retry rates, or output length across models.

What affects cost

  • Selected Mistral model and service
  • Input and output token volume
  • Cached versus uncached repeated context
  • Synchronous versus Batch processing
  • OCR, audio, fine-tuning, or dedicated deployment units

Lower-cost options from the same provider

  • Mistral Small for cost-sensitive tasks after quality testing
  • Batch processing for eligible asynchronous jobs
  • Cached input for repeated prompts and stable shared context

Alternative providers or products

  • OpenAI API
  • Anthropic API
  • Google Gemini API
  • Together AI or GroqCloud for supported open-model inference

Best for

  • teams evaluating Mistral-hosted API models
  • batch workloads that can use asynchronous processing
  • repeated-context workloads eligible for cached-input discounts

Not ideal for

  • buyers treating every Mistral product as one token rate
  • real-time workloads modeled using Batch discounts
  • teams that have not verified the exact active model identifier
FAQ

Mistral API pricing questions

How is Mistral API usage billed?

Text-model usage is generally billed per million input and output tokens, with separate discounts or units for caching, Batch, OCR, audio, fine-tuning, and deployment products.

Can every workload use the 50% Batch discount?

No. Batch is intended for asynchronous processing. Real-time requests should be budgeted using the applicable standard model rate.