Together AI pricing per token
Serverless text rates are listed per 1 million input, cached-input, and output tokens. The table also converts each tracked rate to 1,000 tokens for smaller workload estimates.
Together AI bills serverless text models by input, cached input where available, and output tokens. Dedicated endpoints use hardware-time pricing instead, so the two deployment modes need separate cost models.
Serverless text rates are listed per 1 million input, cached-input, and output tokens. The table also converts each tracked rate to 1,000 tokens for smaller workload estimates.
Serverless models use token billing. Dedicated endpoints are billed by reserved hardware time while running, so a token-only estimate does not represent dedicated deployment cost.
Batch discounts apply only to eligible models and asynchronous workloads. Confirm current eligibility before using a Batch discount in a budget.
Billing varies by serverless model, cached-input support, Batch eligibility, modality, or dedicated endpoint hardware.
| Model | Input price | Cached input | Output price | Decision note |
|---|---|---|---|---|
| DeepSeek V4 Pro | $1.74 / 1M $0.00174 / 1K | $0.20 / 1M $0.0002 / 1K | $3.48 / 1M $0.00348 / 1K | Official Together serverless chat-model rate. |
| MiniMax M3 | $0.30 / 1M $0.0003 / 1K | $0.06 / 1M $0.00006 / 1K | $1.20 / 1M $0.0012 / 1K | Official Together serverless chat-model rate. |
| Kimi K2.7 Code | $0.95 / 1M $0.00095 / 1K | $0.19 / 1M $0.00019 / 1K | $4.00 / 1M $0.004 / 1K | Official Together serverless chat-model rate. |
| GLM-5.2 | $1.40 / 1M $0.0014 / 1K | $0.26 / 1M $0.00026 / 1K | $4.40 / 1M $0.0044 / 1K | Official Together serverless chat-model rate. |
| Kimi K2.6 | $1.20 / 1M $0.0012 / 1K | $0.20 / 1M $0.0002 / 1K | $4.50 / 1M $0.0045 / 1K | Official Together serverless chat-model rate. |
| gpt-oss-120B | $0.15 / 1M $0.00015 / 1K | Not listed | $0.60 / 1M $0.0006 / 1K | Official Together serverless chat-model rate; no cached-input value is shown in the tracked row. |
StackLens assessment based on the tracked serverless rates below. Test output quality and latency on your own tasks before routing production traffic.
gpt-oss-120B
Worth testing when a low listed input and output rate matters and cached-input pricing is not required.
MiniMax M3 or another row with a listed cache rate
May fit repeated-context workloads when the exact request pattern qualifies for cached-input billing.
Models available from more than one host
Compare the same model across providers when routing flexibility, regional availability, or operational fit matters.
One request with 1,000,000 input tokens and 300,000 output tokens, using listed rates and no cache discount. This compares token charges, not model quality.
gpt-oss-120B: $0.15 input + $0.18 output = $0.33
DeepSeek V4 Pro: $1.74 input + $1.044 output = $2.784
Serverless text models are generally billed by input and output tokens, while dedicated endpoints are billed by reserved hardware time while the endpoint is running.
No. Cached-input and Batch support are model- and workflow-specific. Use the current model catalog and Batch documentation before budgeting a discount.
The price depends on the selected model and whether tokens are input, cached input, or output. StackLens currently tracks serverless input rates from $0.15 to $1.74 per 1 million tokens and output rates from $0.60 to $4.50 for the models in this table.
Among the rows currently tracked on this page, gpt-oss-120B has the lowest listed input and output rates. That does not make it the best model for every task; test task quality, latency, and output length before choosing it.