Products
One platform. Four capabilities your inference stack is missing.
Every capability runs on the same Unified Inference Layer — the same OpenAI-compatible API, the same live provider telemetry, the same credit ledger.
Unified API
Unified API
Change the base URL, keep your SDK. RootInference exposes a single OpenAI-compatible API that speaks to 500+ models from 60+ providers — streaming, tool calling, structured outputs and vision included. No provider-specific clients, no per-vendor billing, no rewrites.
Integrate once and reach every frontier and open model — switching providers becomes a string change.
curl https://api.rootinference.com/v1/chat/completionsRequest received for anthropic/claude-sonnet-4-5, streaming on
Resolved via unified schema — no client changes required
Tool call emitted: get_weather, city Berlin
First token in 212 ms at 96 tok/s
Usage metered: 1,412 in / 388 out, priced at 0.9 credits
Same request shape works on 500+ models
Smart Routing
Smart Routing
Define routing policies once and RootInference picks the best model and provider per request. Optimise for cost, latency or quality, set fallback chains, pin rules per API key, and A/B new models against production traffic — all without touching application code.
Cut inference spend and latency without sacrificing quality — routing decisions become policy, not code.
rootinference route --policy latency-firstChat completion received under latency-first policy
6 candidate providers serve the requested model class
Selected provider-2: p50 189 ms, $0.9/M input
Fallback chain armed: provider-4, then provider-7
Completed in 1.4 s for 2.1 credits, quality check passed
Estimated saving vs pinned frontier model: 63%
Reliability
Reliability Engine
Every provider has bad days — rate limits, degraded latency, outages. The Reliability Engine health-checks a pool of providers continuously and reroutes traffic automatically the moment one degrades, so your app stays up when any single vendor goes down.
Single-provider outages stop being your outages — redundancy becomes the default, not a project.
rootinference status --watchProvider pool checked: 9 healthy, 1 degraded
provider-3 degraded: error rate 7.2%, p95 latency 4.8 s
provider-3 removed from rotation
Traffic rerouted to provider-1 / provider-5 in 340 ms
Client impact: 0 failed requests
provider-3 re-admitted after 12 consecutive healthy checks
Observability
Usage Observability
Unified request logs, latency and cost dashboards, and spend attribution by key, team and feature — across every provider you use. Set budget alerts, export the ledger, and answer 'what did we spend on inference and why' in seconds instead of spreadsheet archaeology.
Turn opaque multi-vendor AI bills into attributed, alertable, exportable spend data.
rootinference usage --by team --last 30d30-day totals: 4.2M requests, 9.8B tokens, 41,205 credits
search-team used 18,900 credits, 71% on economy-tier models
agents-team used 14,050 credits at p95 latency 2.1 s
Alert raised: copilot key at 84% of monthly budget
Ledger exported to usage-2026-08.csv
Credits expiring in next 90 days: 0 of 62,300