Abstract

Large language models (LLMs) lower the practical barrier to generating error-sensitive feedback conditioned on a student’s actual incorrect answer, suggesting a plausible next step beyond question-level static feedback in textbook-embedded formative practice. In this study, we compare dynamic LLM-generated feedback to an existing static feedback approach for automatically generated fill-in-the-blank (FITB) cloze questions delivered alongside textbook content in an ereader platform to students learning in higher education contexts. During a randomized deployment, incorrect first attempts were assigned either to dynamic or static feedback, enabling intention-to-treat analysis. Secondary analyses examined complier average causal effects and realized feedback conditions. Results show that dynamic feedback did not outperform the static approach overall. Although dynamic feedback was associated with slightly lower answer reveal rates (a small persistence effect), it produced no detectable change in net target term recovery on the next action; the displaced sessions were absorbed primarily by incorrect retries rather than correct ones. Secondary decomposition analyses indicate that the strongest advantage over dynamic was observed for common answer feedback, which is associated with higher correct retry than dynamic at the feedback type level. We argue that this pattern is learning-science coherent: in textbook-embedded FITB practice designed to support reading comprehension and the doer effect, feedback that more directly constrains the answer space may be better aligned to the task than error-sensitive explanatory feedback.