Softelv AI Logo Softelv AI
Cloud Models Catalog

World-Class AI Models
Ultra-Low Latency Inference

All models are accessible through a unified OpenAI-compatible endpoint. Benefit from high-throughput infrastructure, deterministic pricing, and instant failover.

Production-Ready Foundation Models

Optimized for speed, reasoning depth, multilingual accuracy, and enterprise reliability.

Flagship LLM · Low Latency

Llama 3.3 70B Fast

Meta's state-of-the-art 70B parameter model tuned for ultra-fast edge response. Excels at Arabic & English conversational fluency, summarization, and instruction following.

ID: llama-3.3-70b Context: 128K tokens Speed: ~85 t/s
Complex Reasoning · CoT

DeepSeek R1 Distill

Advanced reasoning model featuring autonomous Chain-of-Thought deliberation before outputting answers. Perfect for competitive math, algorithmic design, and rigorous logic.

ID: deepseek-r1 Context: 64K tokens Type: Reasoning
Software Engineering

Qwen 2.5 Coder 32B

Specialized code intelligence supporting 90+ programming languages. Generates production-ready refactors, unit tests, regex patterns, and architectural scaffolds.

ID: qwen-2.5-coder Context: 32K tokens Focus: Code & Architecture
Generative Imagery

FLUX.1 Schnell

Next-generation text-to-image synthesis engine producing photorealistic high-resolution visuals. Supports automated prompt enhancement with Arabic and English inputs.

ID: flux-1-schnell Resolution: Up to 1024x1024 Latency: ~1.8s
Multimodal Vision

Llama 3.2 11B Vision

Visual reasoning for OCR, chart analysis, invoice extraction, and image inspection. Accepts high-resolution images inline with text prompts.

ID: llama-3.2-11b-vision Context: 128K tokens Input: Images & Text
Nashmi Intelligent Router

Nashmi Smart Core

Dynamic multi-agent routing layer that inspects query semantics and directs queries to the optimal specialized model with automated fallback and caching.

ID: nashmi-core Auto-switch: Yes Failover: 99.9% Uptime

One Line to Switch Any Model

Switch between models effortlessly by changing the model parameter in the official OpenAI SDK.

# Install official OpenAI client: pip install openai
from openai import OpenAI

client = OpenAI(
    base_url="https://softelv.ai/v1",
    api_key="sk-nashmi-your-api-key"
)

# Call Llama 3.3 70B Fast in milliseconds
response = client.chat.completions.create(
    model="llama-3.3-70b",
    messages=[{"role": "user", "content": "What are the best practices for building scalable AI APIs?"}]
)
print(response.choices[0].message.content)