As large language models (LLMs) gain momentum worldwide, there’s a growing need for reliable ways to measure their performance. Benchmarks that evaluate LLM outputs allow developers to track ...
Journal for Research in Mathematics Education. Monograph, Vol. 15, Psychometric Methods in Mathematics Education: Opportunities, Challenges, and Interdisciplinary Collaborations (2016), pp. 155-174 ...