Figure 5. Mapping Norman’s Gulfs of Execution and Evaluation to teacher-facing AI authoring. Adapted from Figure 2.1 in Norman ([Norman Donald, 2013](https://arxiv.org/html/#bib.bib56)). In our adaptation, the pedagogical Gulf of Execution captures the translation of a teacher’s pedagogical goal into chatbot configuration, while the pedagogical Gulf of Evaluation captures the interpretation of generated chatbot behavior relative to that original goal.Conceptual diagram mapping Norman's Gulf of Execution and Gulf of Evaluation to teacher-facing AI authoring. A teacher pedagogical goal is shown at the top and a teacher-facing AI authoring system at the bottom. The left side represents the pedagogical Gulf of Execution, showing how teachers translate pedagogical goals into configuration choices such as Purpose, Rules, and Persona. The right side represents the pedagogical Gulf of Evaluation, showing how teachers interpret generated chatbot behavior and assess whether it reflects their original pedagogical goals.In our study, the pedagogical Gulf of Execution captures the distance between a teacher’s intended pedagogical behavior and the configuration actions available for expressing it. Teachers had to translate instructional intentions into fields such as Purpose, Rules, and Persona, a process that was not always straightforward. The pedagogical Gulf of Evaluation captures the distance between generated behavior and the teacher’s ability to judge whether that behavior reflects the original pedagogical intention. This distinction is particularly important for generative AI, where a response may appear conversationally appropriate without fully realizing the intended instructional purpose. Together, these two gulfs show why configurable controls alone are insufficient: teacher-facing AI authoring must support both the expression of pedagogical intentions and the evaluation of whether those intentions are realized in system behavior. 出處: Will It Teach as Intended? How Teachers Configure Educational AI Chatbots(arXiv)