ChivoxAI
Runtime & operations

Operate speech assessment with confidence.

The same assessment engine reaches production through SDK, MCP or function calling. Runtime is how that traffic stays authenticated, budgeted, observable, private and scalable after you go live.

  • Scoped API keys
  • Hard usage caps
  • Zero audio retention
  • 99.95% enterprise SLA
ArchitectureOperational
SDK, MCP and API converge on the assessment engine, with a five-node runtime and operations layer underneath
99.95%
enterprise SLA
240 ms
p50 latency
0 sec
audio retention
00 / Architecture

Three ways in. One runtime underneath.

SDK, MCP and function calling all hit the same scoring engine. Runtime & operations is the layer that keeps production traffic safe, stable and controllable.

01 / Control

Control access and usage before traffic scales.

Separate environments, enforce limits, and route alerts without rebuilding the integration or adding a second operations stack. Runtime is how the service stays safe, stable and controllable after you go live.

Dark-mode API keys dashboard for Dev, Staging and Production on a night desk
01 / Environment access

Separate environments without changing your integration

Use scoped keys for development, staging, and production. Rotate access safely while every environment keeps the same endpoint and response contract.

Scoped keysSafe rotation
Spend limits gauge and monthly cap controls on a dark developer desk
02 / Usage protection

Set enforceable limits before traffic scales

Assign a monthly cap to each key. When a limit is reached, the API returns a structured 429 so your product can degrade gracefully.

Per-key capStructured 429
Proactive alert thresholds and webhook toggles on a night operations desk
03 / Proactive alerts

Catch usage risk before users feel it

Notify engineering and operations by email or webhook as usage approaches a limit, leaving time to investigate, increase capacity, or adjust traffic.

Email alertsWebhook thresholds
02 / Observe

See issues before they reach users.

Track usage by key, inspect latency and error reasons, and understand which tools are driving traffic from the dashboard or export API.

  • Per-key usage
  • Latency percentiles
  • Tool breakdown
  • Structured error reasons
Explore MCP monitoring docs
Usage and latency charts for the speech assessment API on dual monitors
03 / Trust

Built for sensitive audio and production traffic.

Keep audio ephemeral while relying on a runtime already operating at billions of evaluations per year.

Stateless speech assessment streaming and zero audio retention
Zero-retention streamingTTL 0s

Audio in. Assessment JSON out.

Audio is scored in memory, never used for training, and not stockpiled by the scoring runtime.

Global speech assessment API status and latency
Production scale9.2B+ / year

A runtime built for peak traffic.

9.2B+ evaluations per year, p50 latency of 240 ms, and a 99.95% uptime SLA on the enterprise tier.

Runtime & operations FAQ

Questions teams ask before launch.

Can I set usage limits for each speech assessment API key?

Yes. Each key can have its own monthly limit and environment scope. When the hard cap is reached, calls return a structured 429 response instead of failing ambiguously.

Does Chivox retain audio used for speech assessment?

Streaming audio is scored in memory and is not retained for model training. Your application receives structured assessment JSON without creating an additional stored audio copy in the scoring runtime.

What production monitoring is available?

The dashboard shows per-key usage, latency percentiles, tool breakdowns, and error reasons. Teams can also export usage data through the API and route threshold alerts through email or webhooks.

Ready to put speech assessment into production?

Start MCP or function calling with free credits, or talk with the team about SDK/API licensing, volume, security, and enterprise SLA requirements.

Start MCP / function calling