Quickstart: Uncensored Chat API

Get started with the uncensored chat API in minutes. This quickstart guide covers authentication, basic requests, streaming, and key constraints for your first integration.

Get API key

Authentication & Base URL

The API follows the OpenAI-compatible format. Set your base URL to https://api.uncensoredchat.top/v1 and use your API key for authentication. Generate your key on the Get API key page via Google or email. The key is shown immediately and is the only active key for your account; generating a new one replaces the old. Keep your key secure, as it provides direct access to your prepaid credit.

Sending a Chat Completion Request

Send a POST request to /v1/chat/completions with your model ID set to uncensored. The model is an open-weight LLM tuned for unrestricted responses to lawful adult, fictional, or controversial topics. It does not use model aggregation. Below is a standard request using curl.

curl https://api.uncensoredchat.top/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "uncensored",
    "messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
  }'

Responses include token usage in the final chunk. Input tokens are charged at $0.25 per 1M, output at $1.00 per 1M. Errors and refusals do not consume credit.

Python SDK Integration

Use the official OpenAI Python SDK or any compatible client. Configure the base URL and API key to point to our service. The model parameter must be set to uncensored. This approach works for both synchronous and asynchronous calls, allowing you to integrate the uncensored chat API into existing Python workflows without custom HTTP logic.

from openai import OpenAI

client = OpenAI(base_url="https://api.uncensoredchat.top/v1", api_key="YOUR_KEY")

resp = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)

Ensure your SDK version supports the parameters you need, such as temperature and top_p, which are fully supported for controlling response randomness.

Node SDK Usage

Integrate with Node.js using the OpenAI npm package. Set the baseURL to our endpoint and provide your API key. The uncensored llm api behaves identically to standard OpenAI endpoints but serves our dedicated model. This makes it easy to swap in our service into existing Node applications that already use OpenAI-compatible clients.

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.uncensoredchat.top/v1", apiKey: process.env.API_KEY });

const resp = await client.chat.completions.create({
  model: "uncensored",
  messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);

Handle responses as standard JSON objects. The API returns token counts in the last chunk of the response, allowing you to track usage accurately for your prepaid account.

Streaming Responses (SSE)

Enable streaming by setting stream: true in your request. The API returns Server-Sent Events (SSE) for real-time token delivery. This is ideal for chat interfaces where you want to display tokens as they are generated. The final chunk contains the full token usage statistics, which are necessary for accurate billing and monitoring.

stream = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Tell the story in second person."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Streaming does not change the pricing model; you are still charged per token. Ensure your client handles SSE events correctly to avoid missing the final usage data.

Rate Limits, Errors & Constraints

The API enforces strict limits: 300 requests per minute, 8 concurrent requests per key, and an 8 MB request body. Context window is 64,000 tokens total, with a max output of 16,000 tokens (2,048 if unspecified). Common errors include 401 for invalid keys, 402 for insufficient credit, and 429 for rate limits. Credit is prepaid via crypto and never expires. Refusals for minor sexual content are free. Function calling and JSON mode are supported for structured outputs.

Questions and answers

What is the context window size?

The total context window is 64,000 tokens, combining prompt and completion. The maximum output per request is 16,000 tokens, or 2,048 tokens if you do not set the max_tokens parameter.

How do I handle errors like 402 or 429?

A 402 error indicates insufficient prepaid credit; top up via crypto (USDT or USDC) to continue. A 429 error means you have hit the rate limit of 300 requests per minute or 8 concurrent requests. Wait and retry, or reduce your concurrency.

Is the model the same as GPT-4 or Llama?

No. The model ID is <code>uncensored</code>, an open-weight model run on our own servers. It is not GPT, Claude, Gemini, Grok, DeepSeek, Qwen, or Llama. It is tuned for unrestricted responses to lawful adult and controversial topics.

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.

Get API key