Industry-standard metrics: faithfulness, hallucination (SelfCheckGPT sampling), precision and recall. Provide your agent's answer and the context it used — or let the model answer and self-check. Higher is better everywhere except hallucination.