Understanding Uncensored Inference
When developers seek an uncensored chat solution, they often encounter complex aggregators that route requests through multiple models or apply hidden moderation layers. This API provides a direct path to a single, open-weight model hosted on dedicated servers. Unlike commercial giants that fine-tune for politeness, this inference engine is tuned to answer questions without refusing lawful adult, fictional, or controversial topics. The model does not distinguish between a casual query and a deep-dive into niche subject matter, making it ideal for creative writing, roleplay, or security research where context matters more than brand recognition.
It is important to note that 'uncensored' does not mean 'unlimited.' The model has a hard content limit: it will refuse requests involving sexual content with minors, but otherwise, it remains permissive. The context window supports 64,000 tokens for the combined prompt and completion, ensuring that long conversations or large documents remain coherent. By avoiding model routing, you get consistent behavior and predictable latency, which is critical for applications that rely on specific tone or style consistency.
API Key Setup and Authentication
Getting started with the uncensored api requires minimal friction. You can sign up using 'Continue with Google' or by creating an account with an email and password. No phone number verification is required, and no credit card is needed for the initial trial. Upon creation, your API key is displayed immediately. Each account is limited to one active key; generating a new one invalidates the previous key, so keep it secure.
The base URL for all requests is https://api.uncensoredchat.top/v1. This URL is compatible with the official OpenAI SDKs and most other OpenAI-compatible clients. You simply pass your API key in the Authorization header as a Bearer token. For example, Authorization: Bearer YOUR_API_KEY. This standard approach means you can use familiar libraries like openai in Python or the OpenAI package in Node.js without writing custom HTTP clients. The model identifier to use in your requests is uncensored.
Constructing the Chat Completion Payload
The core interaction happens via the POST /v1/chat/completions endpoint. You send a JSON payload containing your messages, defined by a role (system, user, or assistant) and content. The model responds with a text completion. Here is a basic example of a request:
curl https://api.uncensoredchat.top/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'The request body includes the model field set to uncensored, an array of messages, and optional parameters like temperature or top_p. The maximum output per request is 16,000 tokens, or 2,048 tokens if you do not specify max_tokens. The request body size is limited to 8 MB, which accommodates substantial context but prevents oversized payloads from clogging the pipeline.
Parameters like seed allow for deterministic outputs if you need reproducible results for testing. The presence_penalty and frequency_penalty help manage how often the model repeats concepts. This flexibility allows you to tune the model's creativity and verbosity to match your application's needs without relying on external routing logic.
Handling Streaming Responses
For applications that benefit from real-time feedback, the API supports Server-Sent Events (SSE) streaming. Instead of waiting for the entire response, you receive tokens as they are generated. This is crucial for chat interfaces where latency matters. In streaming mode, the final chunk of the response includes the token usage statistics, allowing you to track costs accurately.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)When using the OpenAI SDKs, you can enable streaming by setting stream=True in Python or stream: true in Node.js. The SDK handles the parsing of SSE events, yielding partial content chunks. This approach reduces perceived latency for end-users and allows for progressive rendering of text. Remember that streaming does not affect pricing; you are charged for the total tokens generated, regardless of whether the response was streamed or buffered.
Enabling Function Calling and Tools
The API supports function calling, allowing the model to generate structured data that your application can execute. You provide a list of tool definitions in the tools parameter, and the model can choose to call them based on the user's input. This is useful for building assistants that interact with external services, such as fetching weather data or updating a database.
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.uncensoredchat.top/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);When the model decides to use a tool, it returns a response with a tool_calls field containing the function name and arguments. Your application must execute the function and then send the result back to the model in a subsequent message. This loop enables complex, multi-step workflows. The model's ability to understand context allows it to handle multi-turn tool usage, where one function's output informs another. This capability transforms the uncensored chatbot api from a simple text generator into a robust agent framework.
Forcing JSON Output Mode
In many integrations, you need the model to return strictly valid JSON, especially when parsing data for frontend applications or backend services. The API supports a json_object response format. By setting response_format: { "type": "json_object" }, you instruct the model to constrain its output to valid JSON syntax. This reduces the need for post-processing validation and ensures that your application can reliably parse the response.
from openai import OpenAI
client = OpenAI(base_url="https://api.uncensoredchat.top/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)This feature is particularly useful for data extraction tasks, where you want the model to summarize text into structured fields or extract entities. While the model is not guaranteed to produce perfect JSON in all edge cases, this mode significantly increases the probability of valid output. It is a powerful tool for developers who want to combine natural language understanding with structured data processing in a single API call.
Managing Rate Limits and Concurrency
To ensure fair usage and system stability, the API enforces specific rate limits. You are allowed 300 requests per minute per API key. Additionally, you can have up to 8 concurrent requests at the same time. If you exceed these limits, you will receive a 429 Too Many Requests error. It is advisable to implement exponential backoff in your client code to handle these errors gracefully.
The 8 MB request body limit also plays a role in concurrency, as large payloads consume more resources. For high-throughput applications, consider batching requests or using streaming to reduce the load on your client. These limits are generous enough for most small to medium-sized applications but may require optimization for enterprise-scale deployments. Keep an eye on your token usage, as high-volume applications can exhaust prepaid credit quickly.
Billing and Token Calculation
Billing is prepaid and based on real token usage. The pricing is transparent: $0.25 per 1 million input tokens and $1.00 per 1 million output tokens. Errors and refusals are free, so you only pay for successful completions. Credit never expires, and there are no monthly subscriptions or hidden fees. You can top up your balance with crypto, specifically USDT (TRC20) or USDC (Base).
The minimum top-up amount is $10, and the maximum is $500. For larger top-ups, you receive bonus credit: +5% for amounts from $50 and +10% for amounts from $100. This bonus structure rewards higher usage. You can manage your balance and view usage history through the dashboard. Since credit is prepaid, you have full control over your spending. There are no credit cards or PayPal options, keeping the process decentralized and fast.
Privacy and Content Limits
Privacy is a key consideration for many developers. Your account only requires an email address, and no phone number is needed. The prompts you send to the model are not used for training other models, ensuring that your data remains private. This is particularly important for businesses that handle sensitive information or proprietary content. The model's uncensored nature means it does not impose arbitrary content filters, but it does enforce a hard limit on sexual content involving minors.
This content limit is always active and cannot be disabled. If a request clearly involves such content, the model will refuse it. For all other lawful adult, controversial, or niche topics, the model remains permissive. This balance between permissiveness and basic safety makes the API suitable for a wide range of use cases, from creative writing to adult entertainment applications. The lack of data training ensures that your unique prompts remain yours, providing a secure environment for experimentation.