# tokengratis.id — data lengkap (llms-full.txt) > Direktori free tier & free credits API LLM, di-aggregate otomatis dari sumber komunitas Indonesia. Aggregator, bukan verifier — tidak ada klaim "verified" di mana pun. Terakhir di-update: 2026-07-31. Field yang tidak muncul di sebuah entri berarti sumber datanya memang tidak menyediakan info itu secara terstruktur — bukan ditebak, bukan dikosongkan paksa jadi placeholder. ## Agnes AI Modalitas: text, vision, image, video | Model | Context | Modality | Rate Limit | | --- | --- | --- | --- | | agnes-1.5-flash | 256K | text + vision | 30 RPM | | agnes-2.0-flash | 256K | text + vision | 30 RPM | | agnes-image-2.0-flash | 4K | image | 30 RPM (1K) | | agnes-image-2.1-flash | 4K | image | 30 RPM (1K) | | agnes-video-v2.0 | 4K | video | 2 RPM | Sumber: - freellm.net (https://freellm.net) — synced 2026-07-31 ## Aion Labs URL: https://www.aionlabs.ai Base URL: https://api.aionlabs.ai/v1 Permanent free tier, no credit card required. 15 RPM, 20K tokens/day. Specialized for roleplay and storytelling. Modalitas: text Gratis: 20K token/hari | Model | Context | Max Output | Modality | Rate Limit | | --- | --- | --- | --- | --- | | Aion 2.0 | 128K | 32K | Text (roleplay) | 15 RPM, 20K TPD | | Aion 2.5 | 128K | 32K | Text (roleplay) | 15 RPM, 20K TPD | | Aion 3.0 | 128K | 32K | Text (roleplay, reasoning) | 15 RPM, 20K TPD | | Aion 3.0 Mini | 128K | 32K | Text (roleplay, reasoning) | 15 RPM, 20K TPD | | Aion-RP 1.0 (8B) | 32K | 32K | Text (roleplay) | 15 RPM, 20K TPD | Sumber: - mnfst/awesome-free-llm-apis (https://github.com/mnfst/awesome-free-llm-apis) — synced 2026-07-31 - freellm.net (https://freellm.net) — synced 2026-07-31 ## Cerebras URL: https://cloud.cerebras.ai/ Base URL: https://api.cerebras.ai/v1 Free tier with payment method required. Ultra-fast inference. 1M tokens/day cap. 64K context on free tier. Modalitas: text, image Gratis: 1M token/hari | Model | Context | Max Output | Modality | Rate Limit | | --- | --- | --- | --- | --- | | gemma-4-31b | 131K (65K on free) | 32K (free) / 40K (paid) | Text + Image | 15 RPM, 30K TPM, 1M TPD | | gpt-oss-120b | 131K (65K on free) | 32K (free) / 40K (paid) | Text | 5 RPM, 30K TPM, 1M TPD | | Llama 3.1 70B | 131K | - | text | - | | zai-glm-4.7 (deprecated Aug 2026) | 131K (64K on free) | 40K | Text | 5 RPM, 30K TPM, 1M TPD | Sumber: - mnfst/awesome-free-llm-apis (https://github.com/mnfst/awesome-free-llm-apis) — synced 2026-07-31 - freellm.net (https://freellm.net) — synced 2026-07-31 - cheahjs/free-llm-api-resources (https://github.com/cheahjs/free-llm-api-resources) — synced 2026-07-31 ## Chutes.ai Modalitas: text | Model | Context | Modality | Rate Limit | | --- | --- | --- | --- | | DeepSeek-R1 | 131K | text + reasoning | Community-powered, no hard cap | | Llama 3.1 70B | 131K | text | Community-powered, no hard cap | Sumber: - freellm.net (https://freellm.net) — synced 2026-07-31 ## Cloudflare Workers AI URL: https://dash.cloudflare.com/profile/api-tokens Base URL: https://api.cloudflare.com/client/v4/accounts/{account_id}/ai/run 10,000 Neurons/day free. 50+ models available on free tier. Modalitas: text, code, embeddings Gratis: 10,000 Neurons/hari | Model | Context | Max Output | Modality | Rate Limit | | --- | --- | --- | --- | --- | | @cf/meta/llama-4-scout-17b-16e-instruct | Up to 10M | Shared w/ context | Multimodal | 10K neurons/day (shared) | | @cf/moonshotai/kimi-k2.7-code | 262K | Shared w/ context | Text (code) | 10K neurons/day (shared) | | @cf/google/gemma-4-26b-a4b-it | 256K | Shared w/ context | Text | 10K neurons/day (shared) | | @cf/meta/llama-3.3-70b-instruct-fp8-fast | 131K | Shared w/ context | Text | 10K neurons/day (shared) | | @cf/zhipuai/glm-4.7-flash | 131K | Shared w/ context | Text | 10K neurons/day (shared) | | @cf/mistralai/mistral-small-3.1-24b-instruct | 128K | Shared w/ context | Text | 10K neurons/day (shared) | | @cf/openai/gpt-oss-120b | 128K | Shared w/ context | Text | 10K neurons/day (shared) | | Mistral 7B | 33K | - | text | - | | Qwen 1.5 7B | 33K | - | text | - | | @cf/deepseek-ai/deepseek-r1-distill-qwen-32b | 32K | Shared w/ context | Text (reasoning) | 10K neurons/day (shared) | | @cf/aisingapore/gemma-sea-lion-v4-27b-it | 8K | 128K | - | 10,000 neurons/day | | @cf/baai/bge-base-en-v1.5 | 8K | - | - | - | | @cf/baai/bge-large-en-v1.5 | 8K | - | - | - | | @cf/baai/bge-m3 | 8K | 1K | - | - | | @cf/baai/bge-small-en-v1.5 | 8K | - | - | - | | @cf/google/gemma-2b-it-lora | 8K | - | - | - | | @cf/google/gemma-7b-it-lora | 8K | - | - | - | | @cf/ibm-granite/granite-4.0-h-micro | 8K | 131K | - | 10,000 neurons/day | | @cf/meta-llama/llama-2-7b-chat-hf-lora | 8K | - | - | - | | @cf/meta/llama-3.1-8b-instruct-fp8 | 8K | 32K | - | 10,000 neurons/day | | @cf/meta/llama-3.2-1b-instruct | 8K | 60K | - | 10,000 neurons/day | | @cf/meta/llama-3.2-3b-instruct | 8K | 80K | - | 10,000 neurons/day | | @cf/meta/llama-guard-3-8b | 8K | - | - | 10,000 neurons/day | | @cf/mistral/mistral-7b-instruct-v0.2-lora | 8K | - | - | 10,000 neurons/day | | @cf/moondream/moondream3.1-9B-A2B | 8K | - | - | - | | @cf/moonshotai/kimi-k2.6 | 8K | - | - | 10,000 neurons/day | | @cf/nvidia/nemotron-3-120b-a12b | 8K | 256K | - | 10,000 neurons/day | | @cf/openai/gpt-oss-20b | 8K | 16K | - | 10,000 neurons/day | | @cf/qwen/qwen2.5-coder-32b-instruct | 8K | 32K | - | 10,000 neurons/day | | @cf/qwen/qwen3-30b-a3b-fp8 | 8K | - | - | 10,000 neurons/day | | @cf/qwen/qwq-32b | 8K | 24K | - | - | | @cf/zai-org/glm-5.2 | 8K | 256K | - | 10,000 neurons/day | | @cf/google/embeddinggemma-300m | - | - | embedding | - | | @cf/pfnet/plamo-embedding-1b | - | - | embedding | - | | @cf/qwen/qwen3-embedding-0.6b | - | - | embedding | - | | Gemma 2B Instruct (LoRA) | - | - | - | 10,000 neurons/day | | Gemma 7B Instruct (LoRA) | - | - | - | 10,000 neurons/day | | Llama 2 7B Chat (LoRA) | - | - | - | 10,000 neurons/day | | Llama 3.2 11B Vision Instruct | - | - | - | 10,000 neurons/day | | Llama 3.3 70B Instruct (FP8) | - | - | - | 10,000 neurons/day | | Llama 4 Scout Instruct | - | - | - | 10,000 neurons/day | | Qwen QwQ 32B | - | - | - | 10,000 neurons/day | Sumber: - mnfst/awesome-free-llm-apis (https://github.com/mnfst/awesome-free-llm-apis) — synced 2026-07-31 - freellm.net (https://freellm.net) — synced 2026-07-31 - cheahjs/free-llm-api-resources (https://github.com/cheahjs/free-llm-api-resources) — synced 2026-07-31 - models.dev (https://models.dev) — synced 2026-07-31 ## Cohere URL: https://dashboard.cohere.com/api-keys Base URL: https://api.cohere.com/v2 Free "Trial" API key, no credit card. 1,000 API calls/month. Non-commercial use only. Modalitas: text, image Gratis: 1,000 calls/bln | Model | Context | Max Output | Modality | Rate Limit | | --- | --- | --- | --- | --- | | Command A+ (218B) | 436K | 64K | Text + Image | 20 RPM | | Command A (111B) | 288K | 8K | Text | 20 RPM | | Command A Reasoning | 288K | ~4K | Text (reasoning) | 20 RPM | | Aya Expanse 32B | 128K | ~4K | Text | 20 RPM | | Command A Vision | 128K | ~4K | Text + Image | 20 RPM | | Command R+ | 128K | 4K | Text | 20 RPM | | Command R7B | 128K | 4K | Text | 20 RPM | | Command R7B Arabic | 128K | ~4K | Text | 20 RPM | | Aya Vision 32B | 16K | ~4K | Text + Image | 20 RPM | | Command A Translate | ~9K | ~4K | Text | 20 RPM | Sumber: - mnfst/awesome-free-llm-apis (https://github.com/mnfst/awesome-free-llm-apis) — synced 2026-07-31 - freellm.net (https://freellm.net) — synced 2026-07-31 - cheahjs/free-llm-api-resources (https://github.com/cheahjs/free-llm-api-resources) — synced 2026-07-31 ## GitHub Models URL: https://github.com/marketplace/models Base URL: https://models.github.ai/inference Free prototyping for all GitHub users. 45+ models. Per-request limits (8K in / 4K out). Modalitas: text, vision, image | Model | Context | Max Output | Modality | Rate Limit | | --- | --- | --- | --- | --- | | gpt-4.1 | 1M | 32K | Text | 10 RPM, 50 RPD | | gpt-4.1-mini | 1M | 32K | Text | 15 RPM, 150 RPD | | Llama-4-Scout-17B-16E-Instruct | 512K | ~4K | Text + Vision | 15 RPM, 150 RPD | | AI21 Jamba 1.5 Large | 256K | - | text + reasoning | - | | Llama-4-Maverick-17B-128E-Instruct-FP8 | 256K | ~4K | Text + Vision | 10 RPM, 50 RPD | | gpt-5 | 200K | 32K | Text | 10 RPM, 50 RPD | | o4-mini | 200K | 100K | Text (reasoning) | 10 RPM, 50 RPD | | Llama-3.3-70B-Instruct | 131K | ~4K | Text | 15 RPM, 150 RPD | | Mistral Large (24.11) | 131K | - | text + image + reasoning | - | | Phi-4 | 131K | - | text + reasoning | - | | gpt-4o | 128K | 16K | Text + Vision | 10 RPM, 50 RPD | | Mistral-Small-3.1 | 128K | ~4K | Text + Vision | 15 RPM, 150 RPD | | DeepSeek-R1 | 64K | 8K | Text (reasoning) | 15 RPM, 150 RPD | 35 more models Sumber: - mnfst/awesome-free-llm-apis (https://github.com/mnfst/awesome-free-llm-apis) — synced 2026-07-31 - freellm.net (https://freellm.net) — synced 2026-07-31 ## Glhf.chat Modalitas: text | Model | Context | Modality | Rate Limit | | --- | --- | --- | --- | | Llama 3.1 70B | 131K | text | Unlimited for free models | | Mixtral 8x7B | 33K | text | Unlimited for free models | Sumber: - freellm.net (https://freellm.net) — synced 2026-07-31 ## Google Gemini URL: https://aistudio.google.com/app/apikey Base URL: https://generativelanguage.googleapis.com/v1beta Free tier unavailable in EU/UK/Switzerland. Free-tier prompts may be used by Google to improve products. Modalitas: text, vision, image, audio, video | Model | Context | Max Output | Modality | Rate Limit | | --- | --- | --- | --- | --- | | Gemini 2.5 Flash | 1M | 65K | Text + Image + Audio + Video | 15 RPM, 1,500 RPD | | Gemini 2.5 Flash-Lite | 1M | 65K | Text + Image + Audio + Video | 30 RPM, 1,500 RPD | | Gemini 2.5 Pro | 1M | 65K | Text + Image + Audio + Video | 5 RPM, 50 RPD | | Gemini 3 Flash (Preview) | 1.0M | - | text | Preview limits | | Gemini 3.1 Flash-Lite | 1M | 65K | Text + Image + Audio + Video | 30 RPM, 1,500 RPD | | Gemini 3.5 Flash | 1M | 65K | Text + Image + Audio + Video | 15 RPM, 1,500 RPD | | Gemini 3.5 Flash-Lite | 1M | 65K | Text + Image + Audio + Video | 30 RPM, 1,500 RPD | | Gemini 3.6 Flash | 1M | 65K | Text + Image + Audio + Video | 15 RPM, 1,500 RPD | | Gemini Flash Latest | 1.0M | 64K | vision + audio + reasoning | - | | Gemini Flash-Lite Latest | 1.0M | 64K | vision + audio + reasoning | - | | Gemma 4 26B A4B IT | 262K | - | vision + reasoning | - | | Gemma 4 31B IT | 262K | - | vision + reasoning | - | | gemini-robotics-er-1.6-preview | 131K | 64K | - | 250,000 tokens/minute, 20 requests/day, 5 requests/minute | | Gemini 2.5 Flash TTS | - | - | - | 10,000 tokens/minute, 10 requests/day, 3 requests/minute | | Gemini 3.1 Flash TTS | - | - | - | 10,000 tokens/minute, 10 requests/day, 3 requests/minute | | Gemini Robotics-ER 1.5 | - | - | - | 250,000 tokens/minute, 20 requests/day, 10 requests/minute | | Gemma 3 12B Instruct | - | - | - | 15,000 tokens/minute, 14,400 requests/day, 30 requests/minute | | Gemma 3 1B Instruct | - | - | - | 15,000 tokens/minute, 14,400 requests/day, 30 requests/minute | | Gemma 3 27B Instruct | - | - | - | 15,000 tokens/minute, 14,400 requests/day, 30 requests/minute | | Gemma 3 4B Instruct | - | - | - | 15,000 tokens/minute, 14,400 requests/day, 30 requests/minute | | Gemma 4 26B A4B Instruct | - | - | - | 16,000 tokens/minute, 14,400 requests/day, 30 requests/minute | | Gemma 4 31B Instruct | - | - | - | 16,000 tokens/minute, 14,400 requests/day, 30 requests/minute | Sumber: - mnfst/awesome-free-llm-apis (https://github.com/mnfst/awesome-free-llm-apis) — synced 2026-07-31 - freellm.net (https://freellm.net) — synced 2026-07-31 - cheahjs/free-llm-api-resources (https://github.com/cheahjs/free-llm-api-resources) — synced 2026-07-31 - models.dev (https://models.dev) — synced 2026-07-31 ## Groq URL: https://console.groq.com/keys Base URL: https://api.groq.com/openai/v1 Free tier, no credit card. Ultra-fast LPU inference. Modalitas: text | Model | Context | Max Output | Modality | Rate Limit | | --- | --- | --- | --- | --- | | GPT-OSS Safeguard 20B | 131K | - | text + reasoning | 1,000 requests/day, 8,000 tokens/minute | | groq/compound | 131K | 8K | Text | 30 RPM, 250 RPD | | groq/compound-mini | 131K | 8K | Text | 30 RPM, 250 RPD | | llama-3.1-8b-instant | 131K | 131K | Text | 30 RPM, 14,400 RPD | | llama-3.3-70b-versatile | 131K | 32K | Text | 30 RPM, 1,000 RPD | | Moonshot Kimi K2 | 131K | - | text | - | | Moonshot Kimi K2 0905 | 131K | - | text | - | | openai/gpt-oss-120b | 131K | 65K | Text | 30 RPM, 1,000 RPD | | openai/gpt-oss-20b | 131K | 65K | Text | 30 RPM, 1,000 RPD | | qwen/qwen3.6-27b | 131K | 16K | Text | 30 RPM, 1,000 RPD | | allam-2-7b | 8K | - | - | 7,000 requests/day, 6,000 tokens/minute | | meta-llama/llama-prompt-guard-2-22m | 8K | - | - | - | | meta-llama/llama-prompt-guard-2-86m | 8K | - | - | - | | canopylabs/orpheus-arabic-saudi | 4K | 50K | - | - | | canopylabs/orpheus-v1-english | 4K | 50K | - | - | | Llama 3.1 8B | - | - | - | 14,400 requests/day, 6,000 tokens/minute | | Llama 3.3 70B | - | - | - | 1,000 requests/day, 12,000 tokens/minute | | whisper-large-v3 | - | - | text | 20 RPM, 2,000 RPD | | whisper-large-v3-turbo | - | - | text | 20 RPM, 2,000 RPD | Sumber: - mnfst/awesome-free-llm-apis (https://github.com/mnfst/awesome-free-llm-apis) — synced 2026-07-31 - freellm.net (https://freellm.net) — synced 2026-07-31 - cheahjs/free-llm-api-resources (https://github.com/cheahjs/free-llm-api-resources) — synced 2026-07-31 - models.dev (https://models.dev) — synced 2026-07-31 ## Hugging Face URL: https://huggingface.co/settings/tokens Base URL: https://router.huggingface.co/v1 100K monthly Inference Provider credits for free users. Routes to Fireworks, Together, Hyperbolic, Nebius, Novita, DeepInfra and others. Thousands of models. Modalitas: text Gratis: 100K kredit/bln | Model | Context | Max Output | Modality | Rate Limit | | --- | --- | --- | --- | --- | | Qwen2.5-7B-Instruct | 131K | ~4K | Text | Credit-metered | | Meta-Llama-3.1-8B-Instruct | 128K | ~4K | Text | Credit-metered | | Phi-3.5-mini-instruct | 128K | ~4K | Text | Credit-metered | | Mistral-7B-Instruct-v0.3 | 32K | ~4K | Text | Credit-metered | | Mixtral-8x7B-Instruct-v0.1 | 32K | ~4K | Text | Credit-metered | thousands of community models Sumber: - mnfst/awesome-free-llm-apis (https://github.com/mnfst/awesome-free-llm-apis) — synced 2026-07-29 - freellm.net (https://freellm.net) — synced 2026-07-29 - cheahjs/free-llm-api-resources (https://github.com/cheahjs/free-llm-api-resources) — synced 2026-07-29 ## Kilo Code URL: https://kilo.ai Base URL: https://api.kilo.ai/api/gateway Free models with no credit card required. kilo-auto/free auto-router routes to minimax/minimax-m2.5:free (80%) and stepfun/step-3.5-flash:free (20%). Modalitas: text, code | Model | Context | Max Output | Modality | Rate Limit | | --- | --- | --- | --- | --- | | nvidia/nemotron-3-super-120b-a12b:free | 262K | 32K | Text | ~200 req/hr | | x-ai/grok-code-fast-1:free | 256K | - | Text (code) | ~200 req/hr | | minimax/minimax-m2.5:free | 196K | 8K | Text | ~200 req/hr | | arcee-ai/trinity-large-thinking:free | 131K | - | Text (reasoning) | ~200 req/hr | | bytedance-seed/dola-seed-2.0-pro:free | 131K | - | Text | ~200 req/hr | | openrouter/free | Varies | Varies | Text | ~200 req/hr | Sumber: - mnfst/awesome-free-llm-apis (https://github.com/mnfst/awesome-free-llm-apis) — synced 2026-07-29 - freellm.net (https://freellm.net) — synced 2026-07-29 ## Kilo Gateway URL: https://kilo.ai/docs/gateway OpenAI-compatible gateway routing to various providers. Free models work without an account. | Model | Rate Limit | | --- | --- | | Cohere North Mini Code | 200 requests/hour per IP, shared across all free models | | Kilo Auto Free (Router) | 200 requests/hour per IP, shared across all free models | | Ling 3.0 Flash | 200 requests/hour per IP, shared across all free models | | NVIDIA Nemotron 3 Nano Omni 30B A3B (Reasoning) | 200 requests/hour per IP, shared across all free models | | NVIDIA Nemotron 3 Super 120B A12B | 200 requests/hour per IP, shared across all free models | | NVIDIA Nemotron 3 Ultra 550B A55B | 200 requests/hour per IP, shared across all free models | | NVIDIA Nemotron 3.5 Content Safety | 200 requests/hour per IP, shared across all free models | | OpenRouter Free Models (Router) | 200 requests/hour per IP, shared across all free models | | Poolside Laguna S 2.1 | 200 requests/hour per IP, shared across all free models | | Poolside Laguna XS 2.1 | 200 requests/hour per IP, shared across all free models | | StepFun Step 3.7 Flash | 200 requests/hour per IP, shared across all free models | Sumber: - cheahjs/free-llm-api-resources (https://github.com/cheahjs/free-llm-api-resources) — synced 2026-07-31 ## LLM7.io URL: https://token.llm7.io Base URL: https://api.llm7.io/v1 Zero-friction API gateway. No registration needed for basic access. 30+ models. GDPR-compliant. Modalitas: text, vision, audio, code Gratis: Free, no signup | Model | Context | Max Output | Modality | Rate Limit | | --- | --- | --- | --- | --- | | Gemini 3.1 Flash Lite | 1.0M | 64K | vision + audio + reasoning | - | | Codestral (latest) | 256K | 4K | text | - | | MiniMax-M2.7 | 205K | - | reasoning | - | | GPT OSS 20B | 131K | - | reasoning | - | | mistral-small-3.1-24b | 32K | - | Text | 30 RPM (120 with token) | | deepseek-r1-0528 | - | - | Text (reasoning) | 30 RPM (120 with token) | | deepseek-v3-0324 | - | - | Text | 30 RPM (120 with token) | | gemini-2.5-flash-lite | - | - | Text + Vision | 30 RPM (120 with token) | | gpt-4o-mini | - | - | Text + Vision | 30 RPM (120 with token) | | qwen2.5-coder-32b | - | - | Text (code) | 30 RPM (120 with token) | ~24 more models Sumber: - mnfst/awesome-free-llm-apis (https://github.com/mnfst/awesome-free-llm-apis) — synced 2026-07-31 - freellm.net (https://freellm.net) — synced 2026-07-31 - models.dev (https://models.dev) — synced 2026-07-31 ## Mistral AI URL: https://console.mistral.ai/api-keys Base URL: https://api.mistral.ai/v1 Free "Experiment" plan, no credit card. ~1B tokens/month. Prompts may be used to improve models. Modalitas: text, image, code Gratis: ~1B token/bln | Model | Context | Max Output | Modality | Rate Limit | | --- | --- | --- | --- | --- | | Codestral | 256K | 256K | Code | ~1 RPS, 500K TPM | | Ministral 14B | 256K | 256K | Text | ~1 RPS, 500K TPM | | Ministral 8B | 256K | 256K | Text | ~1 RPS, 500K TPM | | Mistral Large 3 | 256K | 256K | Text | ~1 RPS, 500K TPM | | Mistral Medium 3.5 (128B) | 256K | 256K | Text + Image + Code | ~1 RPS, 500K TPM | | Mistral Small 4 | 256K | 256K | Text + Image + Code | ~1 RPS, 500K TPM | | Ministral 3B | 128K | 128K | Text | ~1 RPS, 500K TPM | | Mistral 7B | 33K | - | text | - | | Mixtral 8x7B | 33K | - | text | - | Sumber: - mnfst/awesome-free-llm-apis (https://github.com/mnfst/awesome-free-llm-apis) — synced 2026-07-31 - freellm.net (https://freellm.net) — synced 2026-07-31 - cheahjs/free-llm-api-resources (https://github.com/cheahjs/free-llm-api-resources) — synced 2026-07-31 ## ModelScope URL: https://modelscope.cn/my/myaccesstoken Base URL: https://api-inference.modelscope.cn/v1 Free API-Inference for registered users. Requires Alibaba Cloud account binding + real-name verification. Modalitas: text, vision | Model | Context | Max Output | Modality | Rate Limit | | --- | --- | --- | --- | --- | | GLM-5.2 | 1.0M | 128K | reasoning | - | | MiniMax-M3 | 512K | - | vision + reasoning | - | | Kimi K2.5 | 262K | 64K | vision + reasoning | - | | MiniMax-M2.5-highspeed | 205K | - | reasoning | - | | GLM-4.7-FlashX | 200K | - | reasoning | - | | GLM-5.1 | 200K | - | reasoning | - | | deepseek-ai/DeepSeek-V3.2 | 8K | - | - | - | | deepseek-ai/DeepSeek-V4-Flash | 8K | - | - | - | | deepseek-ai/DeepSeek-V4-Pro | 8K | - | - | - | | LLM-Research/Llama-4-Maverick-17B-128E-Instruct | 8K | 4K | - | - | | MedAIBase/AntAngelMed | 8K | - | - | - | | meituan-longcat/LongCat-Flash-Lite | 8K | 320K | - | - | | MiniMax/MiniMax-M1-80k | 8K | - | - | - | | mistralai/Mistral-Large-Instruct-2407 | 8K | - | - | - | | MusePublic/Qwen-Image-Edit | 8K | - | - | - | | opencompass/CompassJudger-1-32B-Instruct | 8K | - | - | - | | OpenGVLab/InternVL3_5-241B-A28B | 8K | - | - | - | | PaddlePaddle/ERNIE-4.5-0.3B-PT | 8K | - | - | - | | PaddlePaddle/ERNIE-4.5-21B-A3B-PT | 8K | - | - | - | | PaddlePaddle/ERNIE-4.5-300B-A47B-PT | 8K | - | - | - | | PaddlePaddle/ERNIE-4.5-VL-28B-A3B-PT | 8K | - | - | - | | Qwen/Qwen3-14B | 8K | - | - | - | | Qwen/Qwen3-235B-A22B | 8K | - | - | - | | Qwen/Qwen3-235B-A22B-Instruct-2507 | 8K | - | - | - | | Qwen/Qwen3-235B-A22B-Thinking-2507 | 8K | - | reasoning | - | | Qwen/Qwen3-30B-A3B | 8K | - | - | - | | Qwen/Qwen3-30B-A3B-Thinking-2507 | 8K | - | reasoning | - | | Qwen/Qwen3-32B | 8K | - | - | - | | Qwen/Qwen3-4B | 8K | - | - | - | | Qwen/Qwen3-8B | 8K | - | - | - | | Qwen/Qwen3-Coder-30B-A3B-Instruct | 8K | - | - | - | | Qwen/Qwen3-Next-80B-A3B-Instruct | 8K | - | - | - | | Qwen/Qwen3-Next-80B-A3B-Thinking | 8K | - | reasoning | - | | Qwen/Qwen3-VL-235B-A22B-Instruct | 8K | - | - | - | | Qwen/Qwen3-VL-8B-Instruct | 8K | - | - | - | | Qwen/Qwen3-VL-8B-Thinking | 8K | 32K | reasoning | - | | Qwen/Qwen3.5-122B-A10B | 8K | - | - | - | | Qwen/Qwen3.5-397B-A17B | 8K | - | - | - | | Shanghai_AI_Laboratory/Intern-S1 | 8K | - | - | - | | Shanghai_AI_Laboratory/Intern-S1-mini | 8K | - | - | - | | Shanghai_AI_Laboratory/Intern-S2-Preview | 8K | - | - | - | | stepfun-ai/Step-3.5-Flash | 8K | - | - | - | | stepfun-ai/Step-3.7-Flash | 8K | - | - | - | | Tencent-Hunyuan/Hy3 | 8K | - | - | - | | XGenerationLab/XiYanSQL-QwenCoder-32B-2412 | 8K | - | - | - | | XGenerationLab/XiYanSQL-QwenCoder-32B-2504 | 8K | - | - | - | | Qwen/Qwen3.5-27B | - | - | Text | 2,000 RPD total; <=500 RPD/model (dynamic) | | Qwen/Qwen3.5-35B-A3B | - | - | Text | 2,000 RPD total; <=500 RPD/model (dynamic) | API-Inference-enabled models Sumber: - mnfst/awesome-free-llm-apis (https://github.com/mnfst/awesome-free-llm-apis) — synced 2026-07-31 - freellm.net (https://freellm.net) — synced 2026-07-31 - models.dev (https://models.dev) — synced 2026-07-31 ## NVIDIA NIM URL: https://build.nvidia.com/explore/discover Base URL: https://integrate.api.nvidia.com/v1 Free with NVIDIA Developer Program membership. 100+ models. Rate-limited (no daily token cap). Modalitas: text, vision, image, audio, video, embeddings, reranking | Model | Context | Max Output | Modality | Rate Limit | | --- | --- | --- | --- | --- | | deepseek-ai/deepseek-v4-flash | 1M | ~64K | Text | ~40 RPM | | meta/llama-guard-4-12b | 1.0M | - | text + image | Up to 40 RPM | | minimaxai/minimax-m3 | 1M | ~64K | Text | ~40 RPM | | z-ai/glm-5.2 | 1.0M | - | text + reasoning | Up to 40 RPM | | mistralai/mistral-medium-3.5-128b | 262K | 262K | Text | ~40 RPM | | moonshotai/kimi-k2.6 | 262K | - | text + image + video + reasoning | Up to 40 RPM | | nvidia/nemotron-3-super-120b-a12b | 262K | 262K | Text | ~40 RPM | | nvidia/nemotron-3-ultra-550b-a55b | 262K | 262K | Text | ~40 RPM | | poolside/laguna-xs-2.1 | 262K | - | text + reasoning | Up to 40 RPM | | stepfun-ai/step-3.7-flash | 262K | - | text + image + video + reasoning | Up to 40 RPM | | Inkling | 256K | - | vision + audio + reasoning | - | | Nemotron 3 Nano Omni 30B A3B Reasoning | 256K | 64K | vision + audio + reasoning | - | | 01-ai/yi-large | 131K | 4K | text | Up to 40 RPM | | adept/fuyu-8b | 131K | - | text | Up to 40 RPM | | ai21labs/jamba-1.5-large-instruct | 131K | - | text | Up to 40 RPM | | aisingapore/sea-lion-7b-instruct | 131K | - | text | Up to 40 RPM | | baai/bge-m3 | 131K | 1K | text | Up to 40 RPM | | bigcode/starcoder2-15b | 131K | - | text | Up to 40 RPM | | databricks/dbrx-instruct | 131K | - | text | Up to 40 RPM | | deepseek-ai/deepseek-coder-6.7b-instruct | 131K | - | text | Up to 40 RPM | | google/codegemma-1.1-7b | 131K | - | text | Up to 40 RPM | | google/codegemma-7b | 131K | - | text | Up to 40 RPM | | google/deplot | 131K | - | text | Up to 40 RPM | | google/gemma-2b | 131K | - | text | Up to 40 RPM | | google/recurrentgemma-2b | 131K | - | text | Up to 40 RPM | | ibm/granite-3.0-3b-a800m-instruct | 131K | - | text | Up to 40 RPM | | ibm/granite-3.0-8b-instruct | 131K | - | text | Up to 40 RPM | | ibm/granite-34b-code-instruct | 131K | - | text | Up to 40 RPM | | ibm/granite-8b-code-instruct | 131K | - | text | Up to 40 RPM | | Llama 3.3 Nemotron Super 49B v1 | 131K | - | reasoning | - | | meta/codellama-70b | 131K | - | text | Up to 40 RPM | | meta/llama-3.1-70b-instruct | 131K | - | text | Up to 40 RPM | | meta/llama-3.2-11b-vision-instruct | 131K | - | text + image | Up to 40 RPM | | meta/llama-3.2-3b-instruct | 131K | - | text | Up to 40 RPM | | meta/llama2-70b | 131K | - | text | Up to 40 RPM | | microsoft/kosmos-2 | 131K | - | text | Up to 40 RPM | | microsoft/phi-3-vision-128k-instruct | 131K | - | text | Up to 40 RPM | | microsoft/phi-3.5-moe-instruct | 131K | 4K | text | Up to 40 RPM | | mistralai/codestral-22b-instruct-v0.1 | 131K | - | text | Up to 40 RPM | | mistralai/mistral-7b-instruct-v0.3 | 131K | - | text | Up to 40 RPM | | mistralai/mixtral-8x22b-v0.1 | 131K | - | text | Up to 40 RPM | | nv-mistralai/mistral-nemo-12b-instruct | 131K | 4K | text | Up to 40 RPM | | nvidia/cosmos-reason2-8b | 131K | 16K | text + image + video + reasoning | Up to 40 RPM | | nvidia/llama-3.1-nemotron-51b-instruct | 131K | - | text | Up to 40 RPM | | nvidia/llama-3.1-nemotron-70b-instruct | 131K | 8K | text | Up to 40 RPM | | nvidia/llama-3.3-nemotron-super-49b-v1.5 | 131K | - | text + reasoning | Up to 40 RPM | | nvidia/llama3-chatqa-1.5-70b | 131K | - | text | Up to 40 RPM | | nvidia/mistral-nemo-minitron-8b-8k-instruct | 131K | - | text | Up to 40 RPM | | nvidia/nemotron-4-340b-instruct | 131K | - | text | Up to 40 RPM | | nvidia/nemotron-4-340b-reward | 131K | - | text | Up to 40 RPM | | nvidia/nemotron-nano-3-30b-a3b | 131K | - | text + reasoning | Up to 40 RPM | | nvidia/nemotron-parse | 131K | - | text | Up to 40 RPM | | nvidia/neva-22b | 131K | - | text | Up to 40 RPM | | nvidia/nvclip | 131K | - | text | Up to 40 RPM | | nvidia/riva-translate-4b-instruct | 131K | - | text | Up to 40 RPM | | nvidia/vila | 131K | - | text | Up to 40 RPM | | openai/gpt-oss-120b | 131K | 131K | Text | ~40 RPM | | openai/gpt-oss-20b | 131K | 131K | Text | ~40 RPM | | writer/palmyra-creative-122b | 131K | - | text | Up to 40 RPM | | writer/palmyra-fin-70b-32k | 131K | - | text | Up to 40 RPM | | writer/palmyra-med-70b | 131K | - | text | Up to 40 RPM | | writer/palmyra-med-70b-32k | 131K | - | text | Up to 40 RPM | | zyphra/zamba2-7b-instruct | 131K | - | text | Up to 40 RPM | | deepseek-ai/deepseek-v4-pro | 128K | ~64K | Text | ~40 RPM | | google/gemma-4-31b-it | 128K | 8K | Text | ~40 RPM | | Llama 3.1 Nemotron Safety Guard 8B v3 | 128K | - | text | - | | meta/llama-3.3-70b-instruct | 128K | 4K | Text | ~40 RPM | | mistralai/mistral-large-2-instruct | 128K | 4K | Text | ~40 RPM | | mistralai/mistral-nemotron | 128K | 8K | Text | ~40 RPM | | Nemotron Mini 4B Instruct | 128K | 8K | text | - | | Nemotron Nano 12B v2 VL | 128K | - | vision + reasoning | - | | nvidia/llama-3.1-nemotron-ultra-253b-v1 | 128K | 4K | Text | ~40 RPM | | nvidia/nemotron-3-nano-30b-a3b | 128K | 32K | Text | ~40 RPM | | nvidia/nemotron-3.5-content-safety | 128K | 8K | text + image + reasoning | Up to 40 RPM | | meta/llama-3.2-1b-instruct | 60K | - | text | Up to 40 RPM | | google/diffusiongemma-26b-a4b-it | 8K | 128K | - | - | | meta/llama-3.1-8b-instruct | 8K | 16K | - | - | | meta/llama-3.2-90b-vision-instruct | 8K | - | - | - | | mistralai/mistral-large-3-675b-instruct-2512 | 8K | - | - | - | | nvidia/ising-calibration-1.5-31b | 8K | - | - | - | | nvidia/llama-3.1-nemoguard-8b-content-safety | 8K | - | - | - | | nvidia/llama-3.1-nemoguard-8b-topic-control | 8K | - | - | - | | nvidia/llama-3.1-nemotron-nano-8b-v1 | 8K | 16K | - | - | | nvidia/llama-3.1-nemotron-nano-vl-8b-v1 | 8K | 16K | - | - | | nvidia/nvidia-nemotron-nano-9b-v2 | 8K | - | - | - | | nvidia/riva-translate-4b-instruct-v1.1 | 8K | 4K | - | - | | nvidia/riva-translate-4b-instruct-v2 | 8K | - | - | - | | nvidia/embed-qa-4 | - | - | embedding | Up to 40 RPM | | nvidia/llama-3.2-nemoretriever-1b-vlm-embed-v1 | - | - | embedding + rerank | Up to 40 RPM | | nvidia/llama-3.2-nv-embedqa-1b-v1 | - | - | embedding | Up to 40 RPM | | nvidia/llama-nemotron-embed-1b-v2 | - | - | embedding + text + image | Up to 40 RPM | | nvidia/llama-nemotron-embed-vl-1b-v2 | 32K | 2K | embedding + text + image | Up to 40 RPM | | nvidia/nemoretriever-parse | - | - | rerank | Up to 40 RPM | | nvidia/nemotron-3-embed-1b | - | - | embedding | Up to 40 RPM | | nvidia/nv-embed-v1 | 32K | 2K | embedding + text | Up to 40 RPM | | nvidia/nv-embedcode-7b-v1 | 32K | 2K | embedding + text | Up to 40 RPM | | nvidia/nv-embedqa-e5-v5 | - | - | embedding | Up to 40 RPM | | nvidia/nv-embedqa-mistral-7b-v2 | - | - | embedding | Up to 40 RPM | | snowflake/arctic-embed-l | - | - | embedding | Up to 40 RPM | Sumber: - mnfst/awesome-free-llm-apis (https://github.com/mnfst/awesome-free-llm-apis) — synced 2026-07-31 - freellm.net (https://freellm.net) — synced 2026-07-31 - cheahjs/free-llm-api-resources (https://github.com/cheahjs/free-llm-api-resources) — synced 2026-07-31 - models.dev (https://models.dev) — synced 2026-07-31 ## Ollama Cloud URL: https://ollama.com/settings/keys Base URL: https://api.ollama.com Free tier with qualitative usage limits. 400+ models from Ollama library. Not OpenAI SDK-compatible; uses Ollama API. Modalitas: text, code | Model | Context | Max Output | Modality | Rate Limit | | --- | --- | --- | --- | --- | | kimi-k2:1t-cloud | 262K | Model-dependent | Text | Session/weekly limits (unpublished) | | deepseek-r1:cloud | 128K | Model-dependent | Text (reasoning) | Session/weekly limits (unpublished) | | deepseek-v3.1:671b-cloud | 128K | Model-dependent | Text | Session/weekly limits (unpublished) | | glm-4.6:cloud | 128K | Model-dependent | Text | Session/weekly limits (unpublished) | | gpt-oss:120b-cloud | 128K | Model-dependent | Text | Session/weekly limits (unpublished) | | qwen3-coder:480b-cloud | 128K | Model-dependent | Text (code) | Session/weekly limits (unpublished) | 30 more cloud models Sumber: - mnfst/awesome-free-llm-apis (https://github.com/mnfst/awesome-free-llm-apis) — synced 2026-07-29 ## OpenCode Zen URL: https://opencode.ai/docs/zen/ AI gateway with curated models. Modalitas: text, vision, audio | Model | Context | Max Output | Modality | | --- | --- | --- | --- | | DeepSeek V4 Flash | 1.0M | - | reasoning | | Laguna S 2.1 | 1.0M | - | reasoning | | MiMo-V2.5 | 1.0M | - | vision + audio + reasoning | | Nemotron 3 Ultra 550B A55B | 1.0M | - | reasoning | | North Mini Code | 256K | 64K | reasoning | | ling-3.0-flash-free | 8K | - | - | | big-pickle | 200K | 32K | - | | Nemotron 3 Ultra Free | 1M | 128K | - | Sumber: - freellm.net (https://freellm.net) — synced 2026-07-31 - cheahjs/free-llm-api-resources (https://github.com/cheahjs/free-llm-api-resources) — synced 2026-07-31 - models.dev (https://models.dev) — synced 2026-07-31 ## OpenRouter URL: https://openrouter.ai/settings/keys Base URL: https://openrouter.ai/api/v1 ~22 free models (marked with :free suffix). OpenAI SDK-compatible. Modalitas: text, vision, image, audio, video, code Gratis: ~22 model free | Model | Context | Max Output | Modality | Rate Limit | | --- | --- | --- | --- | --- | | NVIDIA: Nemotron 3 Ultra (free) | 1M | 66K | Text | 200 req/day (free tier) | | Google: Gemma 4 26B A4B (free) | 262K | 33K | Text + Vision + Video | 20 RPM, 50 RPD | | Google: Gemma 4 31B (free) | 262K | 33K | Text + Vision + Video | 20 RPM, 50 RPD | | Ling-3.0-flash (free) | 262K | 33K | Text | 20 RPM, 50 RPD | | NVIDIA: Nemotron 3 Super (free) | 262K | 262K | Text | 20 RPM, 50 RPD | | Poolside: Laguna S 2.1 (free) | 262K | 33K | Text | 20 RPM, 50 RPD | | Poolside: Laguna XS 2.1 (free) | 262K | 33K | Text | 20 RPM, 50 RPD | | Cohere: North Mini Code (free) | 256K | 64K | Text | 20 RPM, 50 RPD | | NVIDIA: Nemotron 3 Nano 30B A3B (free) | 256K | - | Text | 20 RPM, 50 RPD | | NVIDIA: Nemotron 3 Nano Omni (free) | 256K | 66K | Text + Vision + Audio + Video | 200 req/day (free tier) | | OpenAI: gpt-oss-20b (free) | 131K | 33K | Text | 20 RPM, 50 RPD | | NVIDIA: Nemotron 3.5 Content Safety (free) | 128K | 8K | Text + Vision | 200 req/day (free tier) | | NVIDIA: Nemotron Nano 12B 2 VL (free) | 128K | 128K | Text + Vision + Video | 20 RPM, 50 RPD | | NVIDIA: Nemotron Nano 9B V2 (free) | 128K | - | Text | 20 RPM, 50 RPD | Sumber: - openrouter.ai/api/v1/models (https://openrouter.ai/api/v1/models) — synced 2026-07-31 - mnfst/awesome-free-llm-apis (https://github.com/mnfst/awesome-free-llm-apis) — synced 2026-07-31 - freellm.net (https://freellm.net) — synced 2026-07-31 - cheahjs/free-llm-api-resources (https://github.com/cheahjs/free-llm-api-resources) — synced 2026-07-31 ## OVHcloud AI Endpoints URL: https://www.ovhcloud.com/en/public-cloud/ai-endpoints/catalog/ Base URL: https://oai.endpoints.kepler.ai.cloud.ovh.net/v1 Free anonymous tier (no API key, no signup): 2 RPM per IP per model. 20+ open-weight models hosted in EU. OpenAI SDK-compatible. Modalitas: text, vision, code Gratis: 2 RPM/IP | Model | Context | Max Output | Modality | Rate Limit | | --- | --- | --- | --- | --- | | Qwen3-Coder-30B-A3B-Instruct | 262K | ~32K | Text (code) | 2 RPM (anonymous) | | Meta-Llama-3_3-70B-Instruct | 131K | ~4K | Text | 2 RPM (anonymous) | | Qwen3-32B | 131K | ~32K | Text | 2 RPM (anonymous) | | Qwen3.5-397B-A17B | 131K | ~32K | Text | 2 RPM (anonymous) | | Qwen3.5-9B | 131K | ~8K | Text | 2 RPM (anonymous) | | Qwen3.6-27B | 131K | ~32K | Text | 2 RPM (anonymous) | | gpt-oss-120b | 128K | ~32K | Text | 2 RPM (anonymous) | | gpt-oss-20b | 128K | ~8K | Text | 2 RPM (anonymous) | | Mistral-Nemo-Instruct-2407 | 128K | ~4K | Text | 2 RPM (anonymous) | | Mistral-Small-3.2-24B-Instruct | 128K | ~4K | Text | 2 RPM (anonymous) | | Qwen2.5-VL-72B-Instruct | 128K | ~8K | Text + Vision | 2 RPM (anonymous) | | Mistral-7B-Instruct-v0.3 | 32K | ~4K | Text | 2 RPM (anonymous) | Sumber: - mnfst/awesome-free-llm-apis (https://github.com/mnfst/awesome-free-llm-apis) — synced 2026-07-31 - freellm.net (https://freellm.net) — synced 2026-07-31 ## SambaNova URL: https://cloud.sambanova.ai/apis Base URL: https://api.sambanova.ai/v1 Free tier, no credit card. Ultra-fast RDU inference. 20 RPM, 200K tokens/day. Modalitas: text, image, video Gratis: 200K token/hari | Model | Context | Max Output | Modality | Rate Limit | | --- | --- | --- | --- | --- | | DeepSeek-V3.1 | 128K | ~8K | Text | 20 RPM, 20 RPD, 200K TPD | | DeepSeek-V3.2 (Preview) | 128K | ~8K | Text | 20 RPM, 20 RPD, 200K TPD | | gemma-4-31B-it (Preview) | 128K | ~128K | Text + Image + Video | 20 RPM, 20 RPD, 200K TPD | | gpt-oss-120b | 128K | ~128K | Text | 20 RPM, 20 RPD, 200K TPD | | Meta-Llama-3.3-70B-Instruct | 128K | ~3K | Text | 20 RPM, 20 RPD, 200K TPD | | MiniMax-M2.7 | 128K | ~192K | Text | 20 RPM, 20 RPD, 200K TPD | Sumber: - mnfst/awesome-free-llm-apis (https://github.com/mnfst/awesome-free-llm-apis) — synced 2026-07-31 - freellm.net (https://freellm.net) — synced 2026-07-31 ## SiliconFlow URL: https://cloud.siliconflow.cn/account/ak Base URL: https://api.siliconflow.cn/v1 Permanently free models, no credit card required. 200+ paid models also available. Modalitas: text Gratis: Free (permanen) | Model | Context | Max Output | Modality | Rate Limit | | --- | --- | --- | --- | --- | | deepseek-ai/DeepSeek-R1-Distill-Qwen-7B | 131K | Configurable | Text (reasoning) | 30 RPM, 60K TPM | | Qwen/Qwen3-8B | 131K | 131K | Text | 30 RPM, 60K TPM | Sumber: - mnfst/awesome-free-llm-apis (https://github.com/mnfst/awesome-free-llm-apis) — synced 2026-07-31 - freellm.net (https://freellm.net) — synced 2026-07-31 ## Grok (xAI) Modalitas: text | Model | Context | Modality | Rate Limit | | --- | --- | --- | --- | | Grok-2 | 131K | text | $25/month free credits, resets monthly | | Grok-2 Mini | 131K | text | $25/month free credits, resets monthly | Sumber: - freellm.net (https://freellm.net) — synced 2026-07-31 ## Z AI (Zhipu AI) URL: https://open.bigmodel.cn/usercenter/apikeys Base URL: https://open.bigmodel.cn/api/paas/v4 Permanent free models, no credit card required. Modalitas: text, image Gratis: Free (permanen) | Model | Context | Max Output | Modality | Rate Limit | | --- | --- | --- | --- | --- | | GLM-4.7-Flash | 200K | 128K | Text (reasoning) | 1 concurrent request | | GLM-4.5-Air | 131K | - | reasoning | - | | GLM-4.5-Flash | 128K | ~96K | Text (reasoning) | 1 concurrent request | | GLM-4.6V-Flash | 128K | ~4K | Text + Image | 1 concurrent request | Sumber: - mnfst/awesome-free-llm-apis (https://github.com/mnfst/awesome-free-llm-apis) — synced 2026-07-31 - freellm.net (https://freellm.net) — synced 2026-07-31