RootInference

Products

One platform. Four capabilities your inference stack is missing.

Every capability runs on the same Unified Inference Layer — the same OpenAI-compatible API, the same live provider telemetry, the same credit ledger.

Unified API

Unified API

Change the base URL, keep your SDK. RootInference exposes a single OpenAI-compatible API that speaks to 500+ models from 60+ providers — streaming, tool calling, structured outputs and vision included. No provider-specific clients, no per-vendor billing, no rewrites.

Integrate once and reach every frontier and open model — switching providers becomes a string change.

Tracecurl https://api.rootinference.com/v1/chat/completions
  1. Request received for anthropic/claude-sonnet-4-5, streaming on

  2. Resolved via unified schema — no client changes required

  3. Tool call emitted: get_weather, city Berlin

  4. First token in 212 ms at 96 tok/s

  5. Usage metered: 1,412 in / 388 out, priced at 0.9 credits

  6. Same request shape works on 500+ models

Smart Routing

Smart Routing

Define routing policies once and RootInference picks the best model and provider per request. Optimise for cost, latency or quality, set fallback chains, pin rules per API key, and A/B new models against production traffic — all without touching application code.

Cut inference spend and latency without sacrificing quality — routing decisions become policy, not code.

Tracerootinference route --policy latency-first
  1. Chat completion received under latency-first policy

  2. 6 candidate providers serve the requested model class

  3. Selected provider-2: p50 189 ms, $0.9/M input

  4. Fallback chain armed: provider-4, then provider-7

  5. Completed in 1.4 s for 2.1 credits, quality check passed

  6. Estimated saving vs pinned frontier model: 63%

Reliability

Reliability Engine

Every provider has bad days — rate limits, degraded latency, outages. The Reliability Engine health-checks a pool of providers continuously and reroutes traffic automatically the moment one degrades, so your app stays up when any single vendor goes down.

Single-provider outages stop being your outages — redundancy becomes the default, not a project.

Tracerootinference status --watch
  1. Provider pool checked: 9 healthy, 1 degraded

  2. provider-3 degraded: error rate 7.2%, p95 latency 4.8 s

  3. provider-3 removed from rotation

  4. Traffic rerouted to provider-1 / provider-5 in 340 ms

  5. Client impact: 0 failed requests

  6. provider-3 re-admitted after 12 consecutive healthy checks

Observability

Usage Observability

Unified request logs, latency and cost dashboards, and spend attribution by key, team and feature — across every provider you use. Set budget alerts, export the ledger, and answer 'what did we spend on inference and why' in seconds instead of spreadsheet archaeology.

Turn opaque multi-vendor AI bills into attributed, alertable, exportable spend data.

Tracerootinference usage --by team --last 30d
  1. 30-day totals: 4.2M requests, 9.8B tokens, 41,205 credits

  2. search-team used 18,900 credits, 71% on economy-tier models

  3. agents-team used 14,050 credits at p95 latency 2.1 s

  4. Alert raised: copilot key at 84% of monthly budget

  5. Ledger exported to usage-2026-08.csv

  6. Credits expiring in next 90 days: 0 of 62,300