LLM API

GroqCloud API pricing

GroqCloud API pricing lists per-model input and output token rates for hosted open models, with spend controls and separate rates for audio services.

Decision summary: Consider GroqCloud when its supported model catalog, token rates, service limits, and spend controls fit the production workload.
Pricing checked 2026-07-15high confidence

Pricing overview

Usage-based model inference priced per million tokens, plus separate per-character or per-audio-hour units for non-text services.

Source-tracked model data

GroqCloud API model prices in one table

Example cost uses 1,000,000 input tokens and 300,000 output tokens at each model's default tracked rate, with no cache discount.

ModelInput priceCached inputOutput priceContextExample cost
GPT-OSS 120B on Groq
Stable
$0.15 / 1M$0.075 / 1M$0.60 / 1M131.1K tokens$0.33
GPT-OSS 20B on Groq
Stable
$0.075 / 1M$0.037 / 1M$0.30 / 1M131.1K tokens$0.165
Llama 4 Scout on Groq
Preview
$0.11 / 1MNot tracked$0.34 / 1M131.1K tokens$0.212
Qwen3-32B on Groq
Preview
$0.29 / 1MNot tracked$0.59 / 1M131.1K tokens$0.467

Cost boundary: this example compares token charges only. It does not assume equal output quality, latency, retry rates, or output length across models.

What affects cost

  • Selected model endpoint
  • Input and output token volume
  • Text, speech, transcription, or other service unit
  • Retries and application-level agent loops
  • Production rate limits and enterprise-only model requirements

Lower-cost options from the same provider

  • Llama 3.1 8B Instant for suitable lower-complexity work
  • GPT OSS 20B when its task quality is sufficient
  • Cap generated output and monitor retries for every endpoint

Alternative providers or products

  • Together AI
  • OpenRouter
  • Direct model-provider APIs
  • Self-hosted inference when operational cost is understood

Best for

  • teams evaluating Groq-hosted open models
  • applications that need explicit per-model token rates
  • buyers that want dashboard usage and spend controls

Not ideal for

  • teams requiring models outside the current Groq catalog
  • budgets that mix text and audio billing units
  • workloads that have not tested endpoint availability and rate limits
FAQ

GroqCloud API pricing questions

Are all GroqCloud models billed at the same token rate?

No. Each listed model has its own input and output rate, while speech and transcription services use different billing units.

Does the lowest token rate identify the right Groq model?

No. Token cost should be evaluated with task quality, output length, retries, context limits, endpoint availability, and rate limits.