OpenTelemetry · AI-agent traces
Spancache stores the OTLP you already emit — compressed ~14×, queryable in under 150 ms, in your own cloud. Datadog and Langfuse bill per span and cap you at 15 days. This keeps all of it, for almost nothing.
No agent rewrite. Point your OTel exporter at one endpoint.
Log scale. Datadog/Langfuse scale per span; Spancache is a flat node + S3 on the compressed bytes. Measured list prices, Oct 2026.
The problem
Every model call, tool call, and retrieval is a span — dozens per user turn. Datadog bills ~$350 per million and holds 15 days; Langfuse ~$70 per million. The bill grows with the exact signal you need to keep. Spancache stores the same spans compressed in your own cloud, so cost barely moves.
At 50M spans/mo, flat-cost Spancache is ~700× under Datadog's 90-day tier — and keeps everything, not a 15-day window.
Spancache's cost is a node (~$50/mo) plus S3 on the compressed bytes — ~32 bytes a span. Keep 1M spans or 100M, 15 days or forever: the line stays flat because there's no hot-index tier to pay for. You query the compressed archive directly.
The speed
No decompress, no index rebuild. Queries run straight on the compressed segments, and latency tracks the matched set — not the size of the archive. These numbers hold as history grows.
How it works
If you emit OpenTelemetry GenAI spans — OpenLLMetry, OpenInference, the native SDKs — you're two environment variables away.
# point your exporter at Spancache export OTEL_EXPORTER_OTLP_TRACES_ENDPOINT=\ https://ingest.spancache.ai/v1/traces export OTEL_EXPORTER_OTLP_TRACES_HEADERS=\ "Authorization=Bearer $SPANCACHE_KEY" # that's it — traces land in seconds. # then query the whole history: tool=web_search status=error cost > 0.05 | stats sum(cost) by model
gen_ai.* spans in, derived cost / tokens / latency out.Data sovereignty
Spancache runs in your own account — the compute and the storage both live where your data lives. The control plane handles auth and billing and never sees a trace. Message content is opt-in and redactable at ingest. The console is code that runs in your browser against your cluster.
A reconstructed agent run — the span tree, per-span cost & latency, and (when captured) the conversation. Rendered from trace_id / parent_id.
Pricing
Hosted for speed, BYOC for scale and compliance. The compression works in your favor in both.