A model call is only one component of the production system
Once AI reaches production, quality is not the only concern. Teams need to understand availability, latency, cost, provider limits and what happened inside a specific request. Without a shared telemetry layer, incidents are reconstructed from scattered logs and provider dashboards.
Clyvel was built around this operating problem: a common layer between applications and AI providers that can capture request, usage, cost, latency and incident context at the organization level.
- Provider, model and route.
- Token or usage volume.
- Average and tail latency.
- Errors by provider and type.
- Cost by organization, product or workflow.
Cost needs attribution, not just a monthly total
A monthly bill tells finance what was spent but does not explain why. Production operations benefit from attributing usage to features, organizations, agents, models or workflows. That makes it possible to find expensive behavior and evaluate whether an optimization changes the business cost profile.
Budgets also need alerts before the period ends. Visibility into burn rate is more useful than discovering an overrun after the fact.
- Cost by feature or workflow.
- Cost by tenant or customer.
- Provider and model mix.
- Budget and burn-rate alerts.
- Cost relative to volume and quality.

Latency and reliability need dimensions
Average latency can hide a long slow tail. Percentiles such as p95 become more useful when segmented by provider, model, endpoint and operation. The same applies to errors: a healthy global average can hide a critical workflow failing on one upstream provider.
Observability becomes operational when a team can move from an aggregate metric to a concrete request and inspect the evidence.
- Average and p95 latency.
- Error rate by provider and model.
- Timeouts and retries.
- Incident windows.
- Drill-down from dashboard to request evidence.
Evaluation answers a different question than telemetry
A system can have excellent uptime and still produce poor answers. Evaluation needs representative cases, stable acceptance criteria and regression checks when prompts, models or context change. Some workflows can use automated graders; others need human review because the cost of a subtle error is higher.
Mature AI operations connect technical telemetry with quality evaluation so model changes can be judged on cost, latency, reliability and usefulness together.
- Representative evaluation set.
- Task-specific quality criteria.
- Regression comparison before changes ship.
- Human review for high-impact cases.
- Quality, latency and cost viewed together.
Frequently asked questions
What should an LLM application monitor?
Requests, errors, average and p95 latency, usage, cost, provider, model and quality metrics tied to the actual task.
Why is the AI provider dashboard not enough?
It usually lacks application context such as tenant, feature, workflow, incident, agent and the relationship between spend and business behavior.
What is the difference between observability and evaluation?
Observability explains technical system behavior. Evaluation measures whether model output meets the quality criteria of the task.
Operate AI from evidence, not isolated averages
Clyvel provides an operational layer for requests, cost, latency, incidents, evaluations and AI governance.
Explore Clyvel