Ceti AI LLM Inference
Choose from a selection of many state-of-the art models including Deepseek, Llama and Qwen. Many of these models are offered in various GPU tensor parallelisation configurations if inference speed is a high priority. All of the models are offered in the standard OpenAI request format.
Ceti AI LLM Inference endpoints
| Method | Endpoint | Description |
|---|---|---|
| itrl-meta-llama-3-3-70b-instruct-awq-int4-1tp-1pp | ||
| POST |
/itrl-meta-llama-3-3-70b-instruct-awq-int4-1tp-1pp/chatCompletions /itrl-meta-llama-3-3-70b-instruct-awq-int4-1tp-1pp/v1/chat/completions |
Sends a message to the ibnzterrell/Meta-Llama-3.3-70B-Instruct-AWQ-INT4 model and receives a response. |
| nrlmgc-deepseek-r1-distill-llama-70b-w4a16-1tp-1pp | ||
| POST |
/nrlmgc-deepseek-r1-distill-llama-70b-w4a16-1tp-1pp/chatCompletions /nrlmgc-deepseek-r1-distill-llama-70b-w4a16-1tp-1pp/v1/chat/completions |
Sends a message to the neuralmagic/DeepSeek-R1-Distill-Llama-70B-quantized.w4a16 model and receives a response. |
| nrlmgc-deepseek-r1-distill-qwen-32b-w4a16-1tp-1pp | ||
| POST |
/nrlmgc-deepseek-r1-distill-qwen-32b-w4a16-1tp-1pp/chatCompletions /nrlmgc-deepseek-r1-distill-qwen-32b-w4a16-1tp-1pp/v1/chat/completions |
Sends a message to the neuralmagic/DeepSeek-R1-Distill-Qwen-32B-quantized.w4a16 model and receives a response. |
| qwen-qwen2-5-coder-32b-instruct-gptq-int4-1tp-1pp | ||
| POST |
/qwen-qwen2-5-coder-32b-instruct-gptq-int4-1tp-1pp/chatCompletions /qwen-qwen2-5-coder-32b-instruct-gptq-int4-1tp-1pp/v1/chat/completions |
Sends a message to the Qwen/Qwen2.5-Coder-32B-Instruct-GPTQ-Int4 model and receives a response. |
Ceti AI LLM Inference pricing
| Plan | Price | Rate limit | Quotas |
|---|---|---|---|
| BASIC | Free | 100 / second |
|
| PRO | $10 / month | — |
|
| ULTRA | $100 / month | — |
|