Evaluation matrix

How good are your agent's answers?

Industry-standard metrics: faithfulness, hallucination (SelfCheckGPT sampling), precision and recall. Provide your agent's answer and the context it used — or let the model answer and self-check. Higher is better everywhere except hallucination.

Evaluate

Score distribution

No evaluations yet — run one on the left
0255075100

Recent evaluations

No evaluations yet.
AgentSwitch — prompt evaluation matrix. Scores blend deterministic heuristics with an LLM judge.