The Fallacy of the Automated Arbiter Unpacking the Critical Biases of LLMs as Judges in Evaluation Frameworks

The rapid integration of Large Language Models (LLMs) into the infrastructure of modern evaluation—spanning from the grading of academic code to the ranking of peer-reviewed research—has been driven by the…