LLM Evaluation
Resources for LLM evaluation.
Currently WIP!
General ML Evaluation
- https://www.argmin.net/p/machine-learning-evaluation-631 (Lecture notes by Ben Recht / argmin.net)
- https://github.com/huggingface/evaluation-guidebook (Practical insights and theoretical knowledge about LLM evaluation gathered while managing the Open LLM Leaderboard and designing lighteval)
Eval code and libraries
- https://www.awesomepython.org/?q=llm-evaluation (Awesome list of LLM evaluation code and libraries)
Videos
Links
https://www.evidentlyai.com/llm-guide/llm-evaluation
LLM Halucinations and thinking processes
- A comprehensive taxonomy of hallucinations in Large Language Models (Cossio, 2025) - https://arxiv.org/abs/2508.01781
- https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/
- https://openai.com/index/why-language-models-hallucinate/ & https://arxiv.org/abs/2509.04664
- https://machinelearning.apple.com/research/illusion-of-thinking