OpenTelemetry · AI-agent traces

Keep every trace your agents leave.
For a fraction of the bill.

Spancache stores the OTLP you already emit — compressed ~14×, queryable in under 150 ms, in your own cloud. Datadog and Langfuse bill per span and cap you at 15 days. This keeps all of it, for almost nothing.

No agent rewrite. Point your OTel exporter at one endpoint.

Monthly retention cost@ 10M spans/mo
Datadog Langfuse Spancache

Log scale. Datadog/Langfuse scale per span; Spancache is a flat node + S3 on the compressed bytes. Measured list prices, Oct 2026.

14.4×
compression on real
GenAI traces
<150ms
any query, 3M spans,
on the compressed store
56k/s
ingest through the
OTLP endpoint
$50→
flat, vs $1,050 on
Datadog at 3M spans

The problem

Per-span pricing punishes the thing agents do most.

Every model call, tool call, and retrieval is a span — dozens per user turn. Datadog bills ~$350 per million and holds 15 days; Langfuse ~$70 per million. The bill grows with the exact signal you need to keep. Spancache stores the same spans compressed in your own cloud, so cost barely moves.

Cost at 50M spans / month

At 50M spans/mo, flat-cost Spancache is ~700× under Datadog's 90-day tier — and keeps everything, not a 15-day window.

Flat, not metered.

Spancache's cost is a node (~$50/mo) plus S3 on the compressed bytes — ~32 bytes a span. Keep 1M spans or 100M, 15 days or forever: the line stays flat because there's no hot-index tier to pay for. You query the compressed archive directly.


The speed

The three questions on-call asks — under 150 ms.

No decompress, no index rebuild. Queries run straight on the compressed segments, and latency tracks the matched set — not the size of the archive. These numbers hold as history grows.

Query latency · 3,000,000 spans · single nodedirect on compressed segments

How it works

Same OTLP. One endpoint. Nothing to re-instrument.

If you emit OpenTelemetry GenAI spans — OpenLLMetry, OpenInference, the native SDKs — you're two environment variables away.

# point your exporter at Spancache
export OTEL_EXPORTER_OTLP_TRACES_ENDPOINT=\
  https://ingest.spancache.ai/v1/traces
export OTEL_EXPORTER_OTLP_TRACES_HEADERS=\
  "Authorization=Bearer $SPANCACHE_KEY"

# that's it — traces land in seconds.
# then query the whole history:
tool=web_search status=error
cost > 0.05 | stats sum(cost) by model
01 · Emit
Your agents, unchanged
Keep your existing OTel instrumentation. Spancache speaks native OTLP/HTTP — gen_ai.* spans in, derived cost / tokens / latency out.
02 · Store
Compressed, in your cloud
Spans land append-only, ~14× compressed, in your own S3. One flat-cost node — no per-span meter, no hot tier.
03 · Query
Without decompressing
Search, aggregate, and open any trace as a waterfall in <150 ms — across the full history, from a console or the API.

Data sovereignty

Your prompts never leave your cloud.

Spancache runs in your own account — the compute and the storage both live where your data lives. The control plane handles auth and billing and never sees a trace. Message content is opt-in and redactable at ingest. The console is code that runs in your browser against your cluster.

Trace · invoke_agent doc-rag26.0s · $0.015 · 8 spans

A reconstructed agent run — the span tree, per-span cost & latency, and (when captured) the conversation. Rendered from trace_id / parent_id.


Pricing

Start free. Pay flat. Or run it yourself.

Hosted for speed, BYOC for scale and compliance. The compression works in your favor in both.

Free
$0
For trying it on a real project.
  • 100k spans / month
  • 7-day retention
  • 1 project, hosted
  • Full console & API
Pro · hosted
$49 /mo + usage
Priced under Datadog & Langfuse by design.
  • Usage-based, compressed
  • 30–90 day retention
  • Alerting, seats, projects
  • Priority support
BYOC · enterprise
Flat
Your cloud. Your keys. Unlimited retention.
  • Runs in your account
  • Flat platform fee, not per-span
  • SSO, redaction, SLA
  • AWS · GCP · Azure · S3-compatible

Keep everything your agents do. Query any of it.