# Unmatched Price Performance

Fast responses, scalable performance, and costs you can plan for.

## Large Language Models

| AI Model | Current Speed(Tokens per Second) | Input Token Price(Per Million Tokens) | Output Token Price(Per Million Tokens) |
| --- | --- | --- | --- |
| AI Model<br>GPT OSS 20B 128k | Current Speed<br>1,000 TPS | Input Token Price(Per Million Tokens)<br>$0.075(13.3M / $1)* | Output Token Price(Per Million Tokens)<br>$0.30(3.33M / $1)* |
| AI Model<br>GPT OSS Safeguard 20B | Current Speed<br>1,000 TPS | Input Token Price(Per Million Tokens)<br>$0.075(13.3M / $1)* | Output Token Price(Per Million Tokens)<br>$0.30(3.33M / $1)* |
| AI Model<br>GPT OSS 120B 128k | Current Speed<br>500 TPS | Input Token Price(Per Million Tokens)<br>$0.15(6.67M / $1)* | Output Token Price(Per Million Tokens)<br>$0.60(1.66M / $1)* |
| AI Model<br>Llama 4 Scout (17Bx16E) 128k | Current Speed<br>594 TPS | Input Token Price(Per Million Tokens)<br>$0.11(9.09M / $1)* | Output Token Price(Per Million Tokens)<br>$0.34(2.94M / $1)* |
| AI Model<br>Qwen3 32B 131k | Current Speed<br>662 TPS | Input Token Price(Per Million Tokens)<br>$0.29(3.44M / $1)* | Output Token Price(Per Million Tokens)<br>$0.59(1.69M / $1)* |
| AI Model<br>Llama 3.3 70B Versatile 128k | Current Speed<br>394 TPS | Input Token Price(Per Million Tokens)<br>$0.59(1.69M / $1)* | Output Token Price(Per Million Tokens)<br>$0.79(1.27M / $1)* |
| AI Model<br>Llama 3.1 8B Instant 128k | Current Speed<br>840 TPS | Input Token Price(Per Million Tokens)<br>$0.05(20M / $1)* | Output Token Price(Per Million Tokens)<br>$0.08(12.5M / $1)* |

*Approximate number of tokens per $

## Large Language Models (Enterprise-only)

| AI Model |
| --- |
| AI Model<br>Minimax M2.5 | 
| AI Model<br>Qwen3-VL 32B |

## Text-to-Speech Models

| AI Model | Characters /s | Price (Per M Characters) |
| --- | --- | --- |
| AI Model<br>Canopy Labs Orpheus English | Characters /s<br>100 | Price<br>$22.00 |
| AI Model<br>Canopy Labs Orpheus Arabic Saudi | Characters /s<br>100 | Price<br>$40.00 |

## Automatic Speech Recognition (ASR) Models

| AI Model | Speed Factor | Price(Per Hour Transcribed) |
| --- | --- | --- |
| AI Model<br>Whisper V3 Large | Speed Factor<br>217x | Price<br>$0.111* |
| AI Model<br>Whisper Large v3 Turbo | Speed Factor<br>228x | Price<br>$0.04* |

*Audio is billed at a minimum of 10s per request.

## Prompt Caching

| Model | Uncached Input Tokens (Per M Tokens) | Cached Input Tokens (Per M Tokens) | Output Tokens (Per M Tokens) |
| --- | --- | --- | --- |
| Model<br>moonshotai/kimi-k2-instruct-0905 | Uncached Input Tokens (Per M Tokens)<br>$1.00 | Cached Input Tokens (Per M Tokens)<br>$0.50 | Output Tokens (Per M Tokens)<br>$3.00 |
| Model<br>openai/gpt-oss-120b | Uncached Input Tokens (Per M Tokens)<br>$0.15 | Cached Input Tokens (Per M Tokens)<br>$0.075 | Output Tokens (Per M Tokens)<br>$0.60 |
| Model<br>openai/gpt-oss-20b | Uncached Input Tokens (Per M Tokens)<br>$0.075 | Cached Input Tokens (Per M Tokens)<br>$0.0375 | Output Tokens (Per M Tokens)<br>$0.30 |

Note: No extra fee for the caching feature itself. The discount only applies when a cache hit occurs.

## Built-In Tools (Compound)

| Tool | Price | Parameter |
| --- | --- | --- |
| Tool<br>Basic Search | Price<br>$5 / 1000 requests | Parameter<br>web_search |
| Tool<br>Advanced Search | Price<br>$8 / 1000 requests | Parameter<br>web_search |
| Tool<br>Visit Website | Price<br>$1 / 1000 requests | Parameter<br>visit_website |
| Tool<br>Code Execution | Price<br>$0.18 / hour | Parameter<br>code_interpreter |
| Tool<br>Browser Automation | Price<br>$0.08 / hour | Parameter<br>browser_automation |

## Built-In Tools (GPT-OSS)

| Tool | Price | Parameter |
| --- | --- | --- |
| Tool<br>Browser Search - Basic Search | Price<br>$5 / 1000 requests | Parameter<br>browser_search - browser.search |
| Tool<br>Browser Search - Visit Website | Price<br>$1 / 1000 requests | Parameter<br>browser_search - browser.open |
| Tool<br>Code Execution - Python | Price<br>$0.18 / hour | Parameter<br>code_interpreter - python |

## About Our Pricing

No Surprise Inference Bills

Other inference providers spike costs without warning. Some hide behind elastic pricing. Groq pricing is linear and predictable, with no hidden costs or idle infrastructure. Every new user is growth, not risk, and you can keep margins secure.

## Compound Systems

Intelligent Tool Selection Across Multiple Models

Compound AI systems are powered by multiple openly-available models already supported in GroqCloud to intelligently and selectively use tools to answer user queries, starting first with web search and code execution. Pricing is passed through to the underlying models and server side tools that are part of the compound AI system.

## Batch API

Process Large-Scale Workloads Asynchronously

Batch processing lets you run thousands of API requests at scale by submitting your workload as an asynchronous batch of requests to Groq with 50% lower cost, no impact to your standard rate limits, and 24-hour to 7 day processing window.

## Build Fast

Seamlessly integrate Groq starting with just a few lines of code.
