Modeling what makes a test item hard, from the responses it collects and from the words it is made of.
Item difficulty modeling with transformers
Calibrating test items traditionally requires administering them to hundreds of respondents. Response-free item difficulty modeling asks instead how much of an item's difficulty can be predicted from its wording alone. I fine-tune transformer encoders end-to-end on item text, including the passage, question, and answer options of reading comprehension multiple-choice items, and study extensions that inject psychometric structure into the model: component-wise encoding of item parts, and multi-task learning with an auxiliary question-answering objective. My current work goes a step further and models the simulated response behavior of an "item reader" model, deriving difficulty from predicted response distributions rather than regressing on it directly.
The wording is split into tokens, each token becomes a vector of numbers, and the encoder's layers repeatedly mix information between tokens (attention) and transform each token (feed-forward) until the whole item is summarized in one vector, which a small regression head maps to a single difficulty number. Fine-tuning is what makes this work: for thousands of items with known difficulty, the gap between the predicted and the known value flows back through every layer and nudges the model's weights, so the encoder learns which features of wording matter for difficulty.
Too technical? The same idea in plain words.Seven short steps with small animations, no formulas.
This line of work is funded by my Charles University Grant Agency project Predicting knowledge test item difficulty from item wording with large language models (GA UK 332525, 2025 – 2027), where I am the principal investigator, and by the EduCoDe project (GAČR 25-16951S).
Interactive psychometrics: ShinyItemAnalysis and its modules
Together with Patrícia Martinková and Adéla Hladká I co-develop {ShinyItemAnalysis}, an open-source application for test and item analysis, and its modular architecture, which I designed. The {SIAtools} developer toolkit lets anyone create, test, and distribute their own modules as R packages. The framework and sample modules are described in our 2026 Psychometrika paper. In the EduTest project I also built standalone apps that let schools analyze their own Maturita and entrance-exam data from CZVV.
Item response theory methods
I work on parametrizations and estimation of the nominal response model and on marginal maximum likelihood estimation, with a focus on stable, interpretable parametrizations that also serve models predicting difficulty from item wording.
Evaluation of educational interventions
At Schola Empirica, a non-profit that brings evidence-based methods into Czech schools, I help design and analyze evaluations of its programs. Much of my work there is infrastructure: reproducible analysis and reporting workflows in R, built around the {reschola} package, to which I added the LimeSurvey API client that pulls survey data and variable labels straight into the analyses, and the templates that turn them into automated reports for schools and funders.
Projects and grants
2025 – 2027
Predicting knowledge test item difficulty from item wording with large language models
GA UK 332525 · Principal investigator
2025 – 2027
Comprehensive analysis of educational measurement data for understanding cognitive demands of test items (EduCoDe)
GAČR 25-16951S · Team member
2024 – 2028
Research of Excellence on Digital Technologies and Wellbeing (DigiWELL)
OP JAC CZ.02.01.01/00/22_008/0004583 · Team member
2021 – 2023
Advances in educational assessment: Analytical support for educational test development (EduTest)
TAČR ÉTA TL05000008 · Team member
2021 – 2023
Theoretical foundations of computational psychometrics
GAČR 21-03658S · Team member
2020
Testing, incentives, information: How to mobilize society's resources against the pandemic (CERGE-EI)
TAČR TP01010040 · Data analysis and programming
2020
Center for Educational Measurement and Psychometrics (CEMP)