Production Evaluation of Generated Content Quality
Was this section helpful?
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, Ion Stoica, 2023NeurIPS 2023 Datasets and Benchmarks TrackDOI: 10.48550/arXiv.2306.05685 - This paper introduces the LLM-as-a-Judge paradigm, evaluating its effectiveness for assessing open-ended text generation, relevant to the section's discussion of this automated metric.