Langfuse evaluation
Evaluators
Live evaluator scores per run — deterministic code gates (output validity, tool-call excess) and sampled LLM-as-judge quality signals (hallucination, tool-call quality, context bloat, goal relevance).
Langfuse loading—
Live evaluator scores per run — deterministic code gates (output validity, tool-call excess) and sampled LLM-as-judge quality signals (hallucination, tool-call quality, context bloat, goal relevance).