Skip to main content
Groq provides free API access to various open-source models with extremely fast inference speeds powered by their LPU (Language Processing Unit) technology.

Overview

Groq offers free access to multiple language models optimized for their custom LPU hardware, delivering industry-leading inference speeds.

Rate Limits

Each model has specific rate limits:

Available Models

Text Generation

Llama 4 Scout

Latest Llama model with 30K tokens/min

Llama 3.3 70B

Powerful 70B parameter model

Groq Compound

Groq’s proprietary model with 70K tokens/min

Qwen 3 32B

Efficient multilingual model

Speech Recognition

  • Whisper Large v3 - High-accuracy speech recognition
  • Whisper Large v3 Turbo - Faster speech recognition

Safety Models

  • Llama Guard 4 12B - Content moderation
  • Llama Prompt Guard - Prompt injection detection

API Usage

Getting Started

1

Create Account

Sign up at console.groq.com
2

Generate API Key

Create an API key from your dashboard
3

Start Building

Use the OpenAI-compatible API for inference

Key Features

  • Fastest inference speeds in the industry
  • OpenAI-compatible API
  • Multiple model options including Llama, Mistral, and Groq models
  • Audio transcription with Whisper
  • Content moderation models
  • Generous free tier limits

Performance Highlights

Ultra-Fast

Industry-leading inference speeds

High Throughput

Up to 70,000 tokens/minute

Low Latency

Millisecond response times

LPU Technology

Custom hardware for LLM inference

Multiple Models

Wide selection of open models

Audio Support

Whisper for speech recognition

Additional Resources

Groq Console

Access the platform

Documentation

API documentation