Integrations
The inference layer plugs into your existing stack
500+ models from every major provider on one side; the OpenAI SDK, Vercel AI SDK, LangChain, LlamaIndex, your observability and automation tools on the other.
Model providers
OpenAI
AvailableGPT model family — chat, reasoning, vision and embeddings.
Anthropic
AvailableClaude models with tool use and long-context reasoning.
Google Gemini
AvailableGemini models — multimodal, long context and embeddings.
Meta Llama
AvailableLlama open-weight models via multiple hosted providers.
Mistral
AvailableMistral and Mixtral models for chat, code and embeddings.
DeepSeek
AvailableDeepSeek chat and reasoning models at economy pricing.
xAI Grok
BetaGrok models for chat and reasoning workloads.
Qwen
BetaQwen open-weight models for chat, code and vision.
Cohere
BetaCommand models plus rerank and embedding endpoints.
Amazon Bedrock
BetaBedrock-hosted models through your unified API key.
Perplexity
Coming soonSearch-grounded Sonar models with citations.
Moonshot Kimi
Coming soonKimi long-context chat and reasoning models.
SDKs & frameworks
OpenAI SDK compatibility
AvailablePoint any official OpenAI SDK at our base URL — no code changes.
Vercel AI SDK
AvailableFirst-class provider for streaming UI and tool calling.
LangChain
AvailableChat model integration for chains, agents and tools.
LlamaIndex
AvailableLLM and embedding backends for RAG pipelines.
Haystack
BetaGenerator components for Haystack pipelines.
Semantic Kernel
Coming soonConnector for Microsoft Semantic Kernel apps.
Observability
Langfuse
AvailableTrace requests, prompts and costs into Langfuse.
Helicone
BetaRequest logging and analytics passthrough.
Datadog
Coming soonLatency, error and spend metrics in your dashboards.
Automation
Zapier
BetaTrigger inference from thousands of connected apps.
n8n
BetaModel nodes for self-hosted workflow automation.
Make
Coming soonInference modules for Make scenarios.
Missing a model or provider you depend on?
We add providers continuously, and Enterprise plans support private and fine-tuned model endpoints behind the same API.