The Lifecycle of LLM-as-a-Judge for Large-Scale Recommendation Explanations
LLM-as-a-Judge, which leverages a large language model to evaluate natural language generated by another AI application or model, has become a standard, scalable approach for accelerating and extending costly human evaluation. However, most work treats a judge as a static artifact, evaluating it onc