What it is
Langfuse captures traces of LLM applications as nested spans — one per model call, retrieval step and tool invocation — with inputs, outputs, latency and token counts attached. On top of the trace store it provides prompt management, datasets and evaluation runs, so a regression can be reproduced against recorded inputs. It is self-hostable and instrumented through SDKs or OpenTelemetry.
Best for
Teams that need production tracing and evaluation while keeping trace data inside their own infrastructure.
Where it falls short
Instrumenting deeply nested agent code takes deliberate effort, and self-hosting adds a database and ingestion service to operate.
Characteristics
- tracing
- evaluation
- prompt-management
- opentelemetry
This page carries no score, star rating or review count, and nothing about its placement in the directory was paid for. The outbound links above go to the product’s own domain with no referral parameters. Read the directory methodology for what that means in practice.
More in Observability
Other tools solving the same problem, so you can see what Langfuse is actually competing with.
Braintrust
FreemiumEvaluation and experiment tracking for AI
LangSmith
FreemiumTracing and evals with deep LangChain support