New Approaches to LLM Evaluation: AI-Assisted Benchmark Creation
As the selection of large language models (LLMs) continues to grow, traditional rankings often fall short in providing sufficient insight into how a specific model will perform on unique user tasks or what its operational costs will be. This challenge is driving researchers and companies to develop their own, more relevant benchmarks.
OenoBench: A Deep Dive into LLM Wine Knowledge
One such project, OenoBench, was developed to measure the depth of LLMs’ knowledge about wine. This extensive benchmark comprises 38,104 facts from open sources and 3,266 multiple-choice questions, against which sixteen different models were tested. The unique aspect of this approach lies in the active involvement of AI agents at every stage of the benchmark’s creation:
- Code Generation: Claude Code was utilized to write the necessary software.
- Question Generation: Five language models were responsible for creating the questions.
- Audit and Verification: Ten audit agents reviewed the generated questions for accuracy.
Human involvement in this process was minimal, focusing primarily on wine expertise. The total API cost amounted to approximately $800. The project uncovered interesting issues, such as a “guaranteed fallback” from the model’s memory for 17 out of 35 scrapers, as well as inaccuracies in evaluations where three LLM judges from different companies reached a unanimous yet incorrect conclusion. The study also indicated that the most complex questions generated by the models often turned out to be the most flawed.
Corporate Benchmarks: Tailoring LLMs for Internal Tasks
Beyond specialized benchmarks like OenoBench, companies are also developing internal evaluation systems to select the optimal LLM for their specific needs. In one instance, a laboratory needed to choose a model for an AI assistant that would interact with their electronic lab database. Since public rankings could not address these specific requirements, a custom benchmark of twenty-four questions based on real-world workflows was created. This test was run on DeepSeek, ChatGPT, and Claude models, evaluating nineteen combinations of models and reasoning effort levels across approximately 450 runs. This tailored approach enables precise determination of which model best meets specific corporate tasks and performance demands.
The emergence of AI-assisted benchmark creation, as exemplified by OenoBench, highlights a critical shift in LLM evaluation methodologies. The reported ‘guaranteed fallback’ and unanimous incorrect conclusions by LLM judges underscore the inherent biases and potential for hallucination, even in auditing processes. This necessitates further research into robust validation frameworks, especially as enterprises develop proprietary benchmarks tailored to specific operational contexts and data schemas, where accuracy and reliability are paramount for effective deployment.