Quick Integration with the Uncensored AI API
Integrate your application with our open LLM model in seconds using the OpenAI-compatible API. Simple authentication, standard endpoints, and native streaming support for real-time responses without filters.
Authentication and API Key
Para utilizar a api ia sem censura, você precisa de uma chave de API válida. Crie sua conta na página de criação de chave com apenas e-mail e senha. A chave é exibida imediatamente após o registro e pode ser regenerada a qualquer momento, invalidando a anterior. Utilize o cabeçalho Authorization: Bearer em todas as requisições. Não há necessidade de cartão de crédito para a versão trial, que oferece crédito inicial válido por 7 dias.
If your key is incorrect or expired, the API will return a 401 error. Remember that each account has only one active key at a time. Keep your key secure and avoid exposing it in public repositories to ensure your credit is used only for legitimate requests.
First Request
The main endpoint is POST /v1/chat/completions. This is the industry standard, allowing immediate integration with existing clients. Send your prompt and receive text generated by the uncensored model. The model is optimized to not refuse adult or controversial topics, provided they do not involve minors.
Set the base URL to https://api.apiiasemcensura.com/v1. Make sure to include the model ID uncensored in the request body. The JSON structure follows the OpenAI standard, facilitating migration from existing projects without complex code adaptations.
curl https://api.apiiasemcensura.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'This example demonstrates a basic request. Replace YOUR_API_KEY with your actual key. The model will process the text and return the full response, respecting the defined token limit.
Python SDK for Quick Integration
Use the official OpenAI SDK in Python, adjusting only the base URL. This allows you to leverage mature, well-documented libraries for response handling. The openai library is the standard choice for Python developers seeking stability and ease of use.
Initialize the client pointing to the uncensored AI API server. Set the model to uncensored and send the message. The code returns a response object containing the generated text and usage metadata, such as consumed tokens. This facilitates cost and performance monitoring in real time.
from openai import OpenAI
client = OpenAI(base_url="https://api.apiiasemcensura.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)Adjust the temperature and max_tokens parameters to control creativity and response length. Remember that the total context is limited to 100,000 tokens combining prompt and completion.
Node.js SDK for Web Applications
For Node.js applications, integration follows the same logic as Python. Use the openai package and configure the base URL correctly. This ensures compatibility with modern JavaScript and TypeScript ecosystems.
Create an OpenAI client instance with the API key and service base URL. Call the chat.completions.create() method with the desired messages. The uncensored model will respond directly to the provided context, without additional content filters for legal adult text.
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.apiiasemcensura.com/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);Handle errors appropriately. If the balance is insufficient, you will receive a 402 error. If the rate limit is exceeded, a 429 error will indicate that you should wait before sending new requests.
Response Streaming (SSE)
Server-Sent Events (SSE) support allows receiving responses token by token. This improves the user experience by showing text being generated in real time. Ideal for chatbots and interactive user interfaces requiring low perceived latency.
Set stream: true in the request. The server will send JSON data chunks containing parts of the response. Processing them in real time allows you to display the text progressively. This is crucial for applications where response speed is a competitive advantage.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)Remember to manage the connection state and handle potential interruptions. Streaming does not change the token limit or the model used, only how data is delivered to the client.
Limits, Errors, and Context
The API enforces strict limits to ensure stability. The limit is 300 requests per minute per key. The request body cannot exceed 8 MB. The total context allows 100,000 tokens combining prompt and completion. Exceeding these limits will result in specific errors.
Common errors include 401 for invalid key, 402 for insufficient credit, and 429 for rate limit reached. The GET /v1/models endpoint lists available models, returning information about the uncensored model. Use this endpoint to check status and availability.
API specifications
If your tool speaks the OpenAI API, these are the details that matter.
| Parameter | Details |
|---|---|
| API format | OpenAI-compatible: any OpenAI SDK or client works — change the base URL and the key |
| Base URL | https://api.apiiasemcensura.com/v1 |
| Model | uncensored |
| API key | Bearer token in the Authorization header |
| Endpoints | POST /v1/chat/completions · GET /v1/models |
| Other parameters | temperature, top_p, stop, seed and the two penalties are passed through |
| Function calling | Yes — tools, tool_choice; replies carry tool_calls, also when streaming; send results back as role: tool |
| Max output | prompt + completion fit within 100,000 tokens; max_tokens optional, no separate output cap |
| Structured output | JSON object mode via response_format json_object |
| Max context | 100,000 tokens (prompt + completion together) |
| Streaming | Yes — server-sent events; the last chunk carries token usage |
| Parallel requests | 8 requests at the same time per key |
| Response headers | X-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency |
| Max body | 8 MB request body |
| Requests per minute | 300 requests per minute per key |
| Token prices | input $0.25 / 1M tokens, output $1.00 / 1M tokens |
| Free trial | $0.50 of credit valid 7 days, no card needed · Trial key: 2 parallel requests, 60 req/min; full limits (8 and 300) after first top-up |
| Billing | pay as you go from prepaid credit; nothing is charged for failed or refused requests |
| Payment | crypto: USDT on TRON or USDC on Base, $10–$500, any whole sum |
| Subscription | no monthly fee; paid credit does not expire |
| Volume bonus | +5% on $50+, +10% on $100+ |
| Sign-in | Google or e-mail and password |
| Content policy | uncensored for adults; the only hard rule: no sexual content involving minors |
| Keys | one key per account, regenerate any time (the old one stops working) |
When a request fails
Every error is JSON with a type you can switch on. You are never charged for an error.
| HTTP | Type | What to do |
|---|---|---|
400 | bad_request | invalid JSON, empty messages, bad parameter, or prompt + max_tokens over the window — fix and resend |
401 | missing_key · invalid_key · key_revoked | no key, wrong key, or a key replaced by a newer one |
402 | no_credit | out of credit; add credit and retry |
403 | content_blocked | refused by the content policy |
404 | not_found | only /v1/chat/completions and /v1/models exist |
413 | request_too_large | request body larger than 8 MB |
429 | rate_limited · concurrency | over 300/min or 8 parallel — back off and retry |
503 | upstream_busy | model busy — retry in a few seconds |
Frequently asked questions
What does 'uncensored' mean in this context?
It means the model does not refuse to generate adult, fictional, or controversial content, provided it is legal. The only strict limit is sexual content involving minors, which is always blocked. There are no artificial filters for topics common in roleplay or creative text generation.
Can I use the same API key in multiple projects?
Yes, each account has a single key that can be used in as many projects as you like. If you need to revoke access, you can regenerate the key at any time, which invalidates the old key immediately.
How does the credit system work?
Payment is prepaid with no monthly subscription. Purchased credits never expire. You can top up starting from $10, with a 5% bonus for $50 purchases and 10% for $100. The trial account offers $0.50 valid for 7 days with no card required.
Your key is one form away
Create an account, copy the key, and change the base URL. That's all the configuration.