LLM API

Together AI API Pricing and Model Token Costs

Together AI bills serverless text models by input, cached input where available, and output tokens. Dedicated endpoints use hardware-time pricing instead, so the two deployment modes need separate cost models.

Decision summary: Consider Together AI when its open-model catalog and serverless, Batch, or dedicated deployment options match the workload and operating model.
Pricing checked 2026-07-15high confidence

Together AI pricing per token

Serverless text rates are listed per 1 million input, cached-input, and output tokens. The table also converts each tracked rate to 1,000 tokens for smaller workload estimates.

Serverless vs dedicated endpoints

Serverless models use token billing. Dedicated endpoints are billed by reserved hardware time while running, so a token-only estimate does not represent dedicated deployment cost.

Together AI Batch pricing

Batch discounts apply only to eligible models and asynchronous workloads. Confirm current eligibility before using a Batch discount in a budget.

Pricing overview

Billing varies by serverless model, cached-input support, Batch eligibility, modality, or dedicated endpoint hardware.

ModelInput priceCached inputOutput priceDecision note
DeepSeek V4 Pro$1.74 / 1M
$0.00174 / 1K
$0.20 / 1M
$0.0002 / 1K
$3.48 / 1M
$0.00348 / 1K
Official Together serverless chat-model rate.
MiniMax M3$0.30 / 1M
$0.0003 / 1K
$0.06 / 1M
$0.00006 / 1K
$1.20 / 1M
$0.0012 / 1K
Official Together serverless chat-model rate.
Kimi K2.7 Code$0.95 / 1M
$0.00095 / 1K
$0.19 / 1M
$0.00019 / 1K
$4.00 / 1M
$0.004 / 1K
Official Together serverless chat-model rate.
GLM-5.2$1.40 / 1M
$0.0014 / 1K
$0.26 / 1M
$0.00026 / 1K
$4.40 / 1M
$0.0044 / 1K
Official Together serverless chat-model rate.
Kimi K2.6$1.20 / 1M
$0.0012 / 1K
$0.20 / 1M
$0.0002 / 1K
$4.50 / 1M
$0.0045 / 1K
Official Together serverless chat-model rate.
gpt-oss-120B$0.15 / 1M
$0.00015 / 1K
Not listed$0.60 / 1M
$0.0006 / 1K
Official Together serverless chat-model rate; no cached-input value is shown in the tracked row.
StackLens assessment

Choose a Together AI model by workload shape

StackLens assessment based on the tracked serverless rates below. Test output quality and latency on your own tasks before routing production traffic.

Lowest tracked token rate

gpt-oss-120B

Worth testing when a low listed input and output rate matters and cached-input pricing is not required.

Cached-input workloads

MiniMax M3 or another row with a listed cache rate

May fit repeated-context workloads when the exact request pattern qualifies for cached-input billing.

Provider portability

Models available from more than one host

Compare the same model across providers when routing flexibility, regional availability, or operational fit matters.

What affects cost

  • Selected serverless model and modality
  • Input, cached input, and output volume
  • Batch eligibility and processing mode
  • Dedicated endpoint hardware and running time
  • Image, video, audio, embedding, or reranking units

Lower-cost options from the same provider

  • gpt-oss-20B or other lower-rate models when task quality is sufficient
  • Cached-input pricing on supported chat models
  • Batch processing for selected asynchronous serverless workloads

Alternative providers or products

  • GroqCloud
  • OpenRouter
  • Direct model-provider APIs
  • Self-hosted inference when hardware and operations are included

Best for

  • teams evaluating hosted open-model inference
  • workloads that can compare serverless and Batch processing
  • buyers considering dedicated endpoints at sustained scale

Not ideal for

  • budgets that mix serverless and dedicated billing units
  • workloads assuming every model supports cached input
  • real-time traffic modeled with Batch discounts
Cost example

Example serverless token cost

One request with 1,000,000 input tokens and 300,000 output tokens, using listed rates and no cache discount. This compares token charges, not model quality.

gpt-oss-120B: $0.15 input + $0.18 output = $0.33

DeepSeek V4 Pro: $1.74 input + $1.044 output = $2.784

Budget review

Together AI budget checks

  • Do not apply a cached-input rate to a model that does not list one.
  • Keep serverless token billing separate from dedicated endpoint hardware time.
  • Use Batch discounts only for eligible asynchronous workloads.
  • Measure output length and retries because both can change total token cost.
FAQ

Together AI API pricing questions

How does Together AI serverless pricing differ from dedicated endpoints?

Serverless text models are generally billed by input and output tokens, while dedicated endpoints are billed by reserved hardware time while the endpoint is running.

Can every Together AI model use cached-input or Batch discounts?

No. Cached-input and Batch support are model- and workflow-specific. Use the current model catalog and Batch documentation before budgeting a discount.

How much does 1 million tokens cost on Together AI?

The price depends on the selected model and whether tokens are input, cached input, or output. StackLens currently tracks serverless input rates from $0.15 to $1.74 per 1 million tokens and output rates from $0.60 to $4.50 for the models in this table.

Which Together AI model has the lowest tracked token price?

Among the rows currently tracked on this page, gpt-oss-120B has the lowest listed input and output rates. That does not make it the best model for every task; test task quality, latency, and output length before choosing it.