PARROTSIGHT
Production catalogue · 141 models

Model Gallery

141 production models behind one unified API — six PARROTSIGHT native security models plus the full frontier catalogue from OpenAI, Anthropic, Google, DeepSeek, Qwen and 25+ other providers. One key, one bill, one SLA.

141 models
NativeStable

SightVision Pro

ps-vision-pro

Deep visual understanding for enterprise video assets

P50 380ms·$1.20 / 1k calls·Up to 4K image / 128 frames per request
NativeStable

SightGuard Max

ps-guard-max

Threat intelligence and anomaly detection at network scale

P50 620ms·$2.80 / 1k calls·Up to 512 KB telemetry per request
NativeStable

SightShield

ps-shield-base

Multilingual content safety for regulated platforms

P50 95ms·$0.35 / 1k calls·Up to 32 KB text per request
NativeStable

SightParse Pro

ps-parse-pro

Document intelligence for government and enterprise

P50 450ms·$0.90 / 1k calls·Up to 20 pages per request
NativeBeta

SightLingua 7B

ps-lingua-7b

ASEAN multilingual LLM with OpenAI-compatible API

P50 55ms·$0.14 / 1M tokens·8,192 token context window
NativeStable

SightEmbed

ps-embed-base

Multilingual semantic embeddings for search and RAG

P50 35ms·$0.02 / 1M tokens·512 tokens per input, batch up to 96
OpenAIStable

GPT-6 Astra

openai/gpt-6-astra

GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon...

P50 52ms·$22.50 / 1M tokens·1.1M token context
OpenAIStable

GPT-6 Astra Pro

openai/gpt-6-astra-pro

GPT-6 Astra Pro is the same underlying model as GPT-6 Astra, served with reasoning.mode set to pro for higher-quality responses on complex tasks. Learn more in OpenAI's docs:…

P50 60ms·$22.50 / 1M tokens·1.1M token context
OpenAIStable

GPT-5.6 Luna Pro

openai/gpt-5.6-luna-pro

GPT-5.6 Luna Pro is the same underlying model as GPT-5.6 Luna, served with reasoning.mode set to pro for higher-quality responses on complex tasks. Learn more in OpenAI's docs:…

P50 91ms·$0.50 / 1M tokens·1.1M token context
OpenAIStable

GPT-5.6 Luna

openai/gpt-5.6-luna

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for...

P50 31ms·$0.50 / 1M tokens·1.1M token context
OpenAIStable

GPT-5.6 Terra Pro

openai/gpt-5.6-terra-pro

GPT-5.6 Terra Pro is the same underlying model as GPT-5.6 Terra, served with reasoning.mode set to pro for higher-quality responses on complex tasks. Learn more in OpenAI's docs:…

P50 31ms·$5.00 / 1M tokens·1.1M token context
OpenAIStable

GPT-5.6 Terra

openai/gpt-5.6-terra

GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic...

P50 83ms·$5.00 / 1M tokens·1.1M token context
OpenAIStable

GPT-5.6 Sol Pro

openai/gpt-5.6-sol-pro

GPT-5.6 Sol Pro is the same underlying model as GPT-5.6 Sol, served with reasoning.mode set to pro for higher-quality responses on complex tasks. Learn more in OpenAI's docs:…

P50 30ms·$4.50 / 1M tokens·1.1M token context
OpenAIStable

GPT-5.6 Sol

openai/gpt-5.6-sol

GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks...

P50 87ms·$4.50 / 1M tokens·1.1M token context
OpenAIStable

GPT Chat Latest

openai/gpt-chat-latest

GPT Chat Latest points to OpenAI's stable API alias chat-latest that always resolves to the latest Instant chat model used in ChatGPT. As OpenAI rolls out new Instant model updates...

P50 63ms·$12.50 / 1M tokens·400K token context
OpenAIStable

GPT-5.5 Pro

openai/gpt-5.5-pro

GPT-5.5 Pro is OpenAI’s high-capability model optimized for deep reasoning and accuracy on complex, high-stakes workloads. It features a 1M+ token context window (922K input, 128K output) with support for...

P50 50ms·$75.00 / 1M tokens·1.1M token context
OpenAIStable

GPT-5.5

openai/gpt-5.5

GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks. It features a 1M+ token...

P50 49ms·$12.50 / 1M tokens·1.1M token context
OpenAIStable

GPT-5.4 Image 2

openai/gpt-5.4-image-2

GPT-5.4 Image 2 combines OpenAI's GPT-5.4 model with state-of-the-art image generation capabilities from GPT Image 2. It enables rich multimodal workflows, allowing users to seamlessly move between reasoning, coding, and...

P50 73ms·$11.75 / 1M tokens·272K token context
OpenAIStable

GPT-5.4 Nano

openai/gpt-5.4-nano

GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports text and image inputs and is designed for low-latency...

P50 65ms·$0.51 / 1M tokens·400K token context
OpenAIStable

GPT-5.4 Mini

openai/gpt-5.4-mini

GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding,...

P50 35ms·$1.88 / 1M tokens·400K token context
OpenAIStable

GPT-5.4 Pro

openai/gpt-5.4-pro

GPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning capabilities for complex, high-stakes tasks. It features a 1M+ token context window (922K input, 128K...

P50 43ms·$75.00 / 1M tokens·1.1M token context
OpenAIStable

GPT-5.4

openai/gpt-5.4

GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K input, 128K output) with support for...

P50 40ms·$6.25 / 1M tokens·1.1M token context
AnthropicStable

Claude Fable 5.1

anthropic/claude-fable-5.1

Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual...

P50 39ms·$22.50 / 1M tokens·1M token context
AnthropicStable

Claude Opus 5

anthropic/claude-opus-5

Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis...

P50 61ms·$11.25 / 1M tokens·1M token context
AnthropicStable

Claude Sonnet 5

anthropic/claude-sonnet-5

Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max,...

P50 89ms·$4.50 / 1M tokens·1M token context
AnthropicStable

Claude Fable 5

anthropic/claude-fable-5

Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and...

P50 59ms·$22.50 / 1M tokens·1M token context
AnthropicStable

Claude Opus 4.8

anthropic/claude-opus-4.8

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token...

P50 78ms·$11.25 / 1M tokens·1M token context
AnthropicStable

Claude Opus 4.7

anthropic/claude-opus-4.7

Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on...

P50 70ms·$11.25 / 1M tokens·1M token context
AnthropicStable

Claude Sonnet 4.6

anthropic/claude-sonnet-4.6

Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with...

P50 56ms·$6.75 / 1M tokens·1M token context
AnthropicStable

Claude Opus 4.6

anthropic/claude-opus-4.6

Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective...

P50 58ms·$11.25 / 1M tokens·1M token context
AnthropicStable

Claude Opus 4.5

anthropic/claude-opus-4.5

Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. It offers strong multimodal capabilities, competitive performance across real-world…

P50 63ms·$11.25 / 1M tokens·200K token context
AnthropicStable

Claude Haiku 4.5

anthropic/claude-haiku-4.5

Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...

P50 84ms·$2.25 / 1M tokens·200K token context
AnthropicStable

Claude Sonnet 4.5

anthropic/claude-sonnet-4.5

Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...

P50 81ms·$6.75 / 1M tokens·1M token context
GoogleStable

Gemini 3.8 Flash

google/gemini-3.8-flash

Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.

P50 88ms·$1.69 / 1M tokens·1M token context
GoogleStable

Gemini 3.7 Flash

google/gemini-3.7-flash

Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...

P50 78ms·$1.69 / 1M tokens·1M token context
GoogleStable

Gemini 3.6 Flash

google/gemini-3.6-flash

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...

P50 78ms·$1.69 / 1M tokens·1M token context
GoogleStable

Gemini 3.5 Flash Lite

google/gemini-3.5-flash-lite

Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.

P50 78ms·$0.93 / 1M tokens·1M token context
GoogleStable

Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)

google/gemini-3.1-flash-lite-image

Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is Google's fastest, most cost-efficient Gemini image model, built for high-velocity developer pipelines and rapid-fire visual exploration. It delivers text-to-image generation...

P50 60ms·$0.63 / 1M tokens·66K token context
GoogleStable

Nano Banana 2 (Gemini 3.1 Flash Image)

google/gemini-3.1-flash-image

Gemini 3.1 Flash Image, a.k.a. "Nano Banana 2," is Google’s latest state of the art image generation and editing model, delivering Pro-level visual quality at Flash speed. It combines advanced...

P50 34ms·$1.25 / 1M tokens·131K token context
GoogleStable

Nano Banana Pro (Gemini 3 Pro Image)

google/gemini-3-pro-image

Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly improved multimodal reasoning, real-world grounding, and...

P50 34ms·$5.00 / 1M tokens·131K token context
GoogleStable

Gemini 3.5 Flash

google/gemini-3.5-flash

Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...

P50 77ms·$3.75 / 1M tokens·1M token context
GoogleStable

Gemini 3.1 Flash Lite

google/gemini-3.1-flash-lite

Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic...

P50 78ms·$0.63 / 1M tokens·1M token context
GoogleStable

Gemma 4 26B A4B

google/gemma-4-26b-a4b-it

Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...

P50 56ms·$0.16 / 1M tokens·262K token context
GoogleStable

Gemma 4 31B

google/gemma-4-31b-it

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...

P50 56ms·$0.17 / 1M tokens·262K token context
GoogleStable

Lyria 3 Pro Preview

google/lyria-3-pro-preview

Full-length songs are priced at $0.08 per song. Lyria 3 is Google's family of music generation models, available through the Gemini API. With Lyria 3, you can generate high-quality, 48kHz...

P50 32ms·$0.01 / 1M tokens·1M token context
GoogleStable

Lyria 3 Clip Preview

google/lyria-3-clip-preview

30 second duration clips are priced at $0.04 per clip. Lyria 3 is Google's family of music generation models, available through the Gemini API. With Lyria 3, you can generate...

P50 54ms·$0.01 / 1M tokens·1M token context
QwenStable

Qwen3.8 Flash

qwen/qwen3.8-flash

Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis.

P50 56ms·$0.27 / 1M tokens·1M token context
QwenStable

Qwen3.8 27B

qwen/qwen3.8-27b

Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be...

P50 54ms·$0.85 / 1M tokens·1M token context

Need a custom model?

Fine-tune an existing model or commission a bespoke build for your domain — from specification to production under a single SLA.

Talk to sales