nsfwllms.comDocs
NSFW LLM: Uncensored LLM API Quickstart
Get started with our uncensored LLM API in minutes. This guide covers authentication, chat completions, streaming, and tool calling using standard OpenAI-compatible clients.
- per 1M input tokens
- $0.25
- Output tokens / 1M
- $1.00
- token context
- 100,000
- trial credit
- $0.50
- requests per minute
- 300
Base URL & Authentication
Our API follows the OpenAI standard, making integration straightforward. Use the base URL https://api.nsfwllms.com/v1 for all requests. Authentication is handled via a Bearer token in the Authorization header. You can generate your key on the Get API key page using just an email and password. No phone number or credit card is required to start with the free trial credit.
First Request
Send a standard chat completion request to test connectivity. The model identifier is uncensored. This request demonstrates the basic text-in, text-out flow without any special parameters.
curl https://api.nsfwllms.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'
This returns a standard response object containing the generated text. If you receive a 401, verify your API key. A 402 indicates your prepaid credit balance is insufficient.
Python SDK Integration
Use the official openai Python package to interact with the endpoint. Set the base URL to point to our servers and provide your API key. This allows you to leverage familiar SDK patterns for your uncensored LLM applications.
from openai import OpenAI
client = OpenAI(base_url="https://api.nsfwllms.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)
The client handles serialization and retry logic automatically. You can pass model="uncensored" to ensure you are hitting the correct abliterated model.
Node SDK Usage
For JavaScript or TypeScript projects, initialize the OpenAI client with the custom base URL. This approach works with any OpenAI-compatible client library that supports custom endpoints.
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.nsfwllms.com/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);
Ensure you are using a recent version of the SDK that supports the baseUrl configuration. This ensures compatibility with our streaming and tool calling features.
Streaming Responses (SSE)
Enable streaming by setting stream: true in your request. The API returns Server-Sent Events (SSE) that you can process incrementally. This reduces perceived latency for chat interfaces.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
Each chunk contains partial text deltas. Concatenate these deltas to reconstruct the full response. Streaming works identically to the standard OpenAI protocol, so existing stream parsers should work without modification.
Rate Limits & Constraints
Each API key is limited to 300 requests per minute. The maximum request body size is 8 MB. The context window supports up to 100,000 tokens total for prompt and completion combined. If you exceed the rate limit, you will receive a 429 error. You can regenerate your API key at any time, which immediately revokes the previous key.
Technical reference
Before you integrate, here is exactly what you get with a key.
| Feature | Support |
|---|---|
| API format | OpenAI-compatible: any OpenAI SDK or client works — change the base URL and the key |
| API key | Bearer token in the Authorization header |
| Model ID | uncensored |
| Endpoints | POST /v1/chat/completions · GET /v1/models |
| Base URL | https://api.nsfwllms.com/v1 |
| Sampling parameters | temperature, top_p, stop, seed and the two penalties are passed through |
| Max context | 100,000 tokens (prompt + completion together) |
| Streaming | Yes — server-sent events; the last chunk carries token usage |
| JSON mode | response_format: {"type": "json_object"} |
| Max output | up to 16,000 tokens per request (default 2,048) |
| Tools / tool calls | Supported: tools + tool_choice, tool_calls in the reply (streamed too), tool results as role: tool messages |
| Headers | X-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency |
| Parallel requests | 8 requests at the same time per key |
| Max body | 8 MB request body |
| Rate limit | 300/min per key |
| Billing | prepaid credit, charged by real token usage; errors and refusals are free |
| Trial credit | $0.50 for 7 days, no card |
| Credit expiry | no monthly fee; paid credit does not expire |
| Token prices | input $0.25 / 1M tokens, output $1.00 / 1M tokens |
| Top-up | crypto: USDT on TRON or USDC on Base, $10–$500, any whole sum |
| Volume bonus | +5% from $50, +10% from $100 |
| Content policy | uncensored for adults; the only hard rule: no sexual content involving minors |
| Account | sign in with Google or with e-mail + password |
| Key management | one active key per account; a new key replaces the old one |
When a request fails
Every error is JSON with a type you can switch on. You are never charged for an error.
| HTTP | Type | What to do |
|---|---|---|
400 | bad_request | malformed request or too long for the context window |
401 | missing_key · invalid_key · key_revoked | no key, wrong key, or a key replaced by a newer one |
402 | no_credit | balance is empty — top up, requests resume at once |
403 | content_blocked | sexual content involving minors — refused, not billed |
404 | not_found | unknown endpoint |
413 | request_too_large | body over 8 MB |
429 | rate_limited · concurrency | slow down: rate or parallel limit reached |
503 | upstream_busy | model busy — retry in a few seconds |
Questions and answers
What happens if my prepaid credit runs out?
You will receive a 402 error on subsequent requests. You can top up your account from $10 using crypto (USDT or USDC). Bonus credits are added automatically at higher tiers.
Is the uncensored model GPT-4?
No. The model ID is <code>uncensored</code>. It is an open-weight model tuned for unrestricted text generation, running on our own GPU servers. It is not GPT, Claude, or any other vendor's model.
How do I handle streaming errors?
Check the HTTP status code of the response stream. If the stream ends prematurely or returns an error object, inspect the <code>error</code> field for details like rate limits or invalid parameters.