qwenapis.comDocs
Qwen API: First request in five minutes
Get your first response from our uncensored Qwen API in under five minutes. This quickstart guide covers installation, authentication, and core functionality using standard OpenAI-compatible clients.
Base URL and Authentication
Our API follows the standard OpenAI format, making it easy to switch your existing client configuration. The base URL for all requests is https://api.qwenapis.com/v1. You will need an API key to authenticate your requests. If you haven't created one yet, visit the Get API key page to generate a key using just your email and password. No credit card is required to start with the trial credit.
Each account is limited to one active key at a time. You can regenerate your key at any time, which immediately revokes the previous one. This ensures that if your key is exposed, you can secure your account instantly without waiting for a monthly cycle. Keep your key secure in your environment variables or secret manager.
Make Your First Request
Verify your setup by sending a simple text completion request. The following example uses curl to send a prompt to the uncensored model. Replace YOUR_API_KEY with the key generated in the previous step. The model ID is set to uncensored, which is tuned to answer without content refusals for lawful adult use.
curl https://api.qwenapis.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'If successful, you will receive a JSON response containing the generated text. If you receive a 401 error, check that your key is correct. A 402 error indicates your prepaid credit has been exhausted.
Python SDK Integration
For Python developers, the official OpenAI SDK works directly with our endpoint. Install the library using pip, then configure the client to point to our base URL. This allows you to use familiar methods like chat.completions.create() without changing your application logic significantly. The SDK handles JSON serialization and retry logic for you.
from openai import OpenAI
client = OpenAI(base_url="https://api.qwenapis.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)Remember to set the api_key and base_url correctly. The model name remains uncensored. This approach ensures compatibility with code written for other providers, making it easy to switch or test different endpoints.
Node SDK Usage
Node.js developers can use the @anthropic-ai/sdk or the official openai npm package. Configure the client with our base URL and your API key. The structure of the request is identical to other OpenAI-compatible services. This allows you to integrate uncensored Qwen into your existing Node.js applications with minimal code changes.
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.qwenapis.com/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);Ensure you handle the asynchronous response appropriately. The SDK will return the full completion object, including token usage statistics. This helps you monitor your prepaid credit consumption in real-time.
Enable Streaming
For lower latency and a better user experience, enable streaming in your requests. Streaming sends the response token by token as it is generated, allowing your application to display output progressively. Set the stream parameter to true in your request. The SDK will provide an iterable stream of chunks.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)This is particularly useful for long-form content generation or interactive applications. Each chunk contains partial text, which you can append to your display buffer. The stream ends when the model completes the response or hits the context limit.
Limits, Errors, and Context
Our API has specific limits to ensure fair usage. You are allowed 300 requests per minute per key. The maximum request body size is 8 MB. The context window is 100,000 tokens, covering both the prompt and the completion. If you exceed the rate limit, you will receive a 429 error. If your credit is insufficient, you will receive a 402 error. For invalid keys, you will receive a 401 error. These errors help you manage your application's flow and credit usage effectively.
Technical reference
If your tool speaks the OpenAI API, these are the details that matter.
| Parameter | Details |
|---|---|
| Protocol | OpenAI-compatible: any OpenAI SDK or client works — change the base URL and the key |
| Model ID | uncensored |
| Base URL | https://api.qwenapis.com/v1 |
| Authentication | Bearer token in the Authorization header |
| Endpoints | POST /v1/chat/completions · GET /v1/models |
| SSE streaming | Yes — server-sent events; the last chunk carries token usage |
| Completion length | 16,000 tokens max; 2,048 if max_tokens is not set |
| Sampling parameters | temperature, top_p, stop, seed and the two penalties are passed through |
| Context window | 100,000 tokens (prompt + completion together) |
| Tools / tool calls | Yes — tools, tool_choice; replies carry tool_calls, also when streaming; send results back as role: tool |
| Structured output | JSON object mode via response_format json_object |
| Rate limit | 300/min per key |
| Concurrency | up to 8 in parallel per key |
| Response headers | X-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency |
| Max body | up to 8 MB per request |
| How you pay | prepaid credit, charged by real token usage; errors and refusals are free |
| Top-up | USDT (TRC20) or USDC (Base), any whole amount from $10 to $500 |
| Token prices | input $0.25 / 1M tokens, output $1.00 / 1M tokens |
| Bonus credit | +5% from $50, +10% from $100 |
| Credit expiry | paid credit never expires, no subscription |
| Free trial | $0.50 of credit valid 7 days, no card needed |
| Keys | one active key per account; a new key replaces the old one |
| Content policy | adult content allowed; sexual content involving minors is refused |
| Account | Google or e-mail and password |
Errors and what to do
Every error is JSON with a type you can switch on. You are never charged for an error.
| Code | Type | Meaning |
|---|---|---|
400 | bad_request | invalid JSON, empty messages, bad parameter, or prompt + max_tokens over the window — fix and resend |
401 | missing_key · invalid_key · key_revoked | check the Authorization header or use your current key |
402 | no_credit | out of credit; add credit and retry |
403 | content_blocked | refused by the content policy |
404 | not_found | only /v1/chat/completions and /v1/models exist |
413 | request_too_large | body over 8 MB |
429 | rate_limited · concurrency | slow down: rate or parallel limit reached |
503 | upstream_busy | model busy — retry in a few seconds |
Questions and answers
What does uncensored mean for this API?
It means the model is tuned to answer without content refusals for lawful adult, fictional, or controversial topics. It does not block standard adult content, but it does block sexual content involving minors. This makes it suitable for creative writing, roleplay, and research where strict filtering is undesirable.
Do I need a credit card to start?
No. You can sign up with just an email and password to receive $0.50 of trial credit valid for 7 days. This allows you to test the API without providing payment details. You can add credit later using crypto (USDT or USDC) if you wish to continue.
Is this the official Qwen API?
No, this is an independent service running an open-weight uncensored model on our own GPU servers. We are not affiliated with Alibaba or any other vendor. Our API is OpenAI-compatible, so it works with standard clients, but it is a distinct service with its own pricing and limits.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.