We benchmarked a multilingual LLM judge on Spanish and Catalan text against human labels using Jev from TypeSafe. The result: keep rubric questions in English to maintain accuracy.
Una guía práctica para integrar la evaluación de LLM en CI/CD, de modo que las regresiones de calidad se detecten antes de su lanzamiento y no después de que se acumulen los tickets de soporte.
How to measure whether an LLM application actually works, from traces and golden datasets through judge calibration, CI regression gating, and production monitoring.