AI Gateway

Beget AI Gateway

Connect modern language models to your products and business processes through Beget Cloud. One API key, a single balance and payment in rubles — no separate accounts with different providers, foreign payment methods or multiple integrations.

Beget AI Gateway
  • Easy connection
  • Flexible tariffing
  • Full integration

Pricing and models

AI Gateway pricing and models

Choose a model for your task and pay only for actual usage

AI Gateway offers a catalog of models for different scenarios — from text generation and processing to complex analysis and coding.

  • The cost depends on the selected model and the number of processed tokens
  • The price is shown in rubles per 1 million input and output tokens.
  • All costs are charged to your Beget balance — regardless of the model provider.

Available models

All
General-purpose
Fast and affordable
Long context
Complex tasks
Experimental
Coding
Value for money
Agent workflows
Recommended by Beget
OpenAI
GPT-5.6 Sol
OpenAI's previous-generation flagship for complex assistants, coding, and document analysis. The new model in this family is GPT-6.1 Sol.
Max context
1.1M Tokens
Input tokens / 1M
2,49 €
Output tokens / 1M
12,46 €
OpenAI
GPT-5.6 Luna
OpenAI's lightweight model family for classification, short responses, and batch requests. The new model in this family is GPT-6 Luna.
Max context
1.1M Tokens
Input tokens / 1M
0,25 €
Output tokens / 1M
1,50 €
Anthropic
Claude Sonnet 5
Anthropic's workhorse for instructions, coding, analytics, and complex conversations. The new model in this family is Sonnet 5.5.
Max context
1M Tokens
Input tokens / 1M
2,49 €
Output tokens / 1M
12,46 €
Anthropic
Claude Opus 5
Anthropic's model for multi-step reasoning, architectural decisions, and complex agentic workflows. The new model in this family is Opus 5.5.
Max context
1M Tokens
Input tokens / 1M
6,23 €
Output tokens / 1M
31,16 €
Google
Gemini 3.1 Pro Preview
Google's Pro model: analyzing long documents, knowledge bases, and conversations with a large 1M context.
Max context
1M Tokens
Input tokens / 1M
2,49 €
Output tokens / 1M
14,96 €
Google
Gemini 3.1 Flash Lite Preview
Flash Lite: low latency and cost with a 1M context — for classification and high-throughput workloads.
Max context
1M Tokens
Input tokens / 1M
0,31 €
Output tokens / 1M
1,87 €
DeepSeek
DeepSeek V4 Pro
Strong quality at a low cost — a go-to choice for results without a premium budget.
Max context
600K Tokens
Input tokens / 1M
0,84 €
Output tokens / 1M
2,49 €
DeepSeek
DeepSeek V4 Flash
A high-throughput model for simple responses, classification, and large request volumes.
Max context
1M Tokens
Input tokens / 1M
0,28 €
Output tokens / 1M
0,83 €
DeepSeek
DeepSeek R1 0528
A reasoning model for planning, multi-step analysis, and tasks that require a chain of reasoning.
Max context
131.1K Tokens
Input tokens / 1M
0,88 €
Output tokens / 1M
3,12 €
Qwen
Qwen3.7 Max
A general-purpose Max model with a 1M context and strong Russian-language support.
Max context
1M Tokens
Input tokens / 1M
1,84 €
Output tokens / 1M
5,52 €
Qwen
Qwen3 Coder Next
A specialized coding model for code generation, review, explanation, and automated tests.
Max context
256K Tokens
Input tokens / 1M
0,23 €
Output tokens / 1M
1,13 €
Kimi
Kimi K2.6
A model for agent and research workflows: multi-step reasoning, tool calling, and a 262K context.
Max context
262.1K Tokens
Input tokens / 1M
0,59 €
Output tokens / 1M
3,06 €
GLM
GLM 5.2
An open-weight model for agent pipelines and multi-step tasks. Good value for money.
Max context
202.8K Tokens
Input tokens / 1M
1,48 €
Output tokens / 1M
5,49 €
MiniMax
MiniMax M3
An affordable model with a 1M context for high-throughput tasks and long inputs.
Max context
1M Tokens
Input tokens / 1M
0,38 €
Output tokens / 1M
1,50 €
Meta
Llama 4 Maverick
Meta's open-weight model — a familiar stack for experiments and general-purpose tasks.
Max context
1M Tokens
Input tokens / 1M
0,44 €
Output tokens / 1M
1,25 €
xAI
Grok 4.20
A record 2M context — for research and analyzing entire logs and long conversations.
Max context
2M Tokens
Input tokens / 1M
1,56 €
Output tokens / 1M
3,12 €
OpenAI
text-embedding-3-small
Compact OpenAI embeddings: semantic search, RAG, and clustering at a low cost.
Max context
8.2K Tokens
Input tokens / 1M
0,03 €
Output tokens / 1M
0 €
OpenAI
text-embedding-3-large
Higher-dimensional OpenAI embeddings — maximum accuracy for semantic search and RAG over large text corpora.
Max context
8.2K Tokens
Input tokens / 1M
0,17 €
Output tokens / 1M
0 €
Anthropic
Claude Fable 5.1
Anthropic's flagship: the highest Intelligence and Coding Index scores for long agent and reasoning workflows.
Max context
1M Tokens
Input tokens / 1M
12,46 €
Output tokens / 1M
62,32 €
Google
Gemini 3.8 Flash
Google's Flash model: high quality at a low price with a 1M context — fast responses without a major drop in quality.
Max context
1M Tokens
Input tokens / 1M
0,94 €
Output tokens / 1M
4,68 €
Qwen
Qwen3.8 Max
Qwen's current Max model: general-purpose tasks, Russian-language support, and a long 1M context.
Max context
1M Tokens
Input tokens / 1M
2,49 €
Output tokens / 1M
7,48 €
Kimi
Kimi K3
Kimi for agent pipelines and tool calling with a long 1M context.
Max context
1M Tokens
Input tokens / 1M
3,74 €
Output tokens / 1M
18,70 €
GLM
GLM 5.3
GLM for agent workflows and tool calling. A 1M context and a strong Coding Index at a moderate price.
Max context
1M Tokens
Input tokens / 1M
1,75 €
Output tokens / 1M
5,49 €
xAI
Grok 4.6
xAI's general-purpose model for reasoning and coding with a 500,000-token context window. The new model in this family is Grok 4.7.
Max context
500K Tokens
Input tokens / 1M
2,49 €
Output tokens / 1M
7,48 €
OpenAI
GPT-6.1 Sol
OpenAI's next-generation flagship for agentic coding, document analysis, and multi-step workflows.
Max context
1.1M Tokens
Input tokens / 1M
2,49 €
Output tokens / 1M
12,46 €
OpenAI
GPT-6 Astra
OpenAI's model for complex analysis, engineering, research, and long-running agentic tasks.
Max context
1.1M Tokens
Input tokens / 1M
12,46 €
Output tokens / 1M
62,32 €
OpenAI
GPT-6 Luna
OpenAI's lightweight next-generation model for conversations, classification, and low-cost batch requests.
Max context
1.1M Tokens
Input tokens / 1M
0,13 €
Output tokens / 1M
0,63 €
Anthropic
Claude Sonnet 5.5
Anthropic's next-generation workhorse for coding, instructions, business documents, and analytics.
Max context
1M Tokens
Input tokens / 1M
2,49 €
Output tokens / 1M
12,46 €
Anthropic
Claude Opus 5.5
Anthropic's model for complex tasks, in-depth analysis, and long-running agentic workflows.
Max context
1M Tokens
Input tokens / 1M
4,99 €
Output tokens / 1M
24,93 €
xAI
Grok 4.7
xAI's next-generation general-purpose model for reasoning and coding with a 500,000-token context window.
Max context
500K Tokens
Input tokens / 1M
2,49 €
Output tokens / 1M
7,48 €

Where AI Gateway is used

Where AI Gateway is used
  • AI in websites and apps

    AI in websites and apps

    Text generation, smart search and recommendations through a single API

  • Chatbots and AI assistants

    Chatbots and AI assistants

    Assistants for customers and employees: answers and consultations

  • Working with documents

    Working with documents

    Analysis, summarization, comparison and structuring of information

  • Content generation

    Content generation

    Generating and editing texts, translations and program code

  • Process automation

    Process automation

    Reports, requests, data classification and repetitive tasks

How AI Gateway works

AI Gateway lets you connect AI to a product or internal processes through a single API — without separate integrations with providers.

Your application calls Beget AI Gateway, and you choose a model from the catalog. The gateway passes the request to the model and returns the answer, so you can switch models without reworking the integration.

  1. 1

    Create an API key

    Create an API key, set a monthly budget and specify a limit on requests per minute.

  2. 2

    Choose a model

    Choose a suitable model from the catalog and use different models for different tasks without separate integrations.

  3. 3

    Send a request through a single API

    Connect the API to your application or service. The OpenAI-compatible format simplifies integration and switching models.

  4. 4

    Pay for usage through Beget

    Requests are paid from your Beget balance in rubles — without separate payments to providers or foreign services.

One gateway for working with AI

One API for different models

Work with the LLM catalog through a single OpenAI-compatible API. Instead of separate integrations with each provider — one API key and a single way to connect.

Access without separate provider accounts

There is no need to register with each LLM provider, obtain separate keys and maintain multiple integrations. Access to models is provided through Beget AI Gateway.

Quick connection of AI Gateway

Quick connection

Create an API key in the control panel and connect AI to your project. The OpenAI-compatible API format lets you use familiar tools and reduces the integration effort.

  • Quick start
  • Single API
  • Simple integration
  • Easy to switch between models

Freedom to choose a model

Choose a model for a specific task: text generation, data analysis, coding, working with documents, building assistants and other scenarios. If requirements change, you can use another available model through the same API.

Payment in rubles through Beget

Use AI and pay for requests from a single Beget balance. No need to pay different foreign providers yourself or maintain several separate accounts.

Need help getting connected?

Tell us about your task — we will help you get started with AI Gateway and its use in your project.

By clicking 'Submit Request', you confirm your consent to the processing of personal data.

AI Gateway FAQ

Frequently Asked Questions