As generative artificial intelligence becomes an increasingly standard fixture in modern classrooms and teacher preparation programs, a central pedagogical question has emerged: how do future educators interpret, evaluate, and translate diverse feedback sources into coherent instructional designs? A recent empirical study published in Frontiers in Psychology explores this dynamic by examining how 22 South Korean pre-service mathematics teachers utilized a combination of AI-generated and peer-to-peer feedback while revising mathematical modeling tasks centered on school cafeteria food waste—a curriculum context closely aligned with United Nations Sustainable Development Goal 12 (SDG 12), which targets responsible consumption and production. The research sheds light on a phenomenon dubbed the "feedback translation gap," revealing that while future teachers actively engage with and selectively adopt feedback to make structural modifications to their teaching materials, these visible revisions do not automatically translate into measurable improvements in overall task-design quality. Background and Context of the Study The integration of generative AI tools, such as advanced large language models (LLMs), has rapidly transformed the feedback ecology of higher education and teacher training. Capable of delivering immediate, detailed, and individualized commentary on lesson plans and instructional tasks, AI offers a compelling solution to traditional constraints on instructor time. However, educational researchers have long emphasized that the value of feedback lies less in its speed or volume and more in how learners make sense of it and act upon it. This challenge is magnified in complex, open-ended pedagogical domains like sustainability-oriented mathematical modeling. Designing such tasks requires pre-service teachers (PSTs) to coordinate multiple rigorous competencies: grounding problems in authentic real-world contexts, formulating open inquiry questions, establishing viable data constraints, anticipating diverse student solution paths, and embedding robust validation procedures that move beyond simple answer-checking. When these tasks incorporate sustainability issues like food waste, they also demand the coordination of quantitative reasoning with environmental, social, and ethical considerations. To investigate how future teachers navigate these high demands, the study was embedded within a required undergraduate course on logic and essay writing in mathematics education at a South Korean university during the spring 2026 semester. Chronology and Research Design The study followed an exploratory, comparative mixed-methods sequence-comparison design across a structured seven-week instructional unit within a 15-week semester. The workflow unfolded through several distinct phases: Task Conceptualization: Participants drafted an initial mathematical modeling task focused on reducing school cafeteria food waste using a standardized design worksheet. The draft encompassed grade levels, core mathematical concepts, situation prompts, data sets, modeling assumptions, validation steps, and student action components. First Feedback Round: Participants were divided into two ordering conditions. An AI-to-peer group received feedback from a custom GPT built on advanced LLM architecture, while a peer-to-peer group exchanged written evaluations in class. Intermediate Revision: Following the first feedback round, participants recorded their responses, noted selective acceptances or rejections, and submitted an initial set of revisions. Second Feedback Round: Participants received the alternate feedback source (peers for the first group, AI for the second) and completed a second iteration of feedback analysis. Final Submission and Reflection: After incorporating insights from both sources, participants submitted their final revised task designs alongside structured written reflections detailing the perceived utility, limitations, and appropriate use of AI versus human peer feedback. Throughout this timeline, researchers evaluated the tasks using a rigorous consensus-coded framework divided into four analytical domains: Domain 1 (task-design quality evaluated across six subdimensions on a 0–12 scale), Domain 2 (feedback uptake and revision depth, categorized from surface edits to structural and reconceptual changes), Domain 3 (cognitive, affective, and behavioral functions of the feedback sources), and Domain 4 (validation-focused revision indicators). Key Findings: Selective Uptake and the Translation Gap The study evaluated 22 participants who met strict inclusion criteria by completing all phases of the instructional sequence, yielding a dataset comprising 8 participants in the AI-to-peer condition and 14 in the peer-to-peer-first condition. The empirical results challenged the assumption that learners passively accept external commentary. Across the board, participants demonstrated high levels of critical agency: Selective Acceptance: Out of 44 total feedback episodes, selective acceptance was recorded in 32 instances. A remarkable 20 out of 22 participants selectively accepted feedback in at least one round, modifying suggestions to fit their specific instructional goals rather than adopting them wholesale. Explicit rejections were rare, occurring only twice. Revision Depth: Revision activity was widespread. Seventeen of the 22 participants executed at least one structural or reconceptual task revision—such as reorganizing data constraints or reframing the core modeling question—with nine participants doing so across both feedback rounds. The Translation Gap: Despite active engagement and deep revisions, only five participants (22.7%) showed a positive composite gain in rated task-design quality (Domain 1). Scores remained unchanged for 10 participants and declined for 7. Approximately 70.6% of deep structural or reconceptual revisions failed to produce a net positive score improvement, illustrating that visible revision effort does not guarantee higher artifact quality. Divergent Roles of AI and Peer Feedback Participants’ qualitative reflections highlighted distinct cognitive, affective, and behavioral functions associated with each feedback source. AI-generated feedback was predominantly valued for its speed, technical specificity, and capacity to generate alternative perspectives. Furthermore, AI feedback more frequently prompted validation-focused revisions (observed in 50% of AI episodes compared to 36.4% of peer episodes). Conversely, peer feedback was prized for its contextual realism, alignment with student perspectives, and the value of mutual communication. While both sources required critical vigilance—with participants noting risks of AI hallucinations or contextual blindness alongside the variable specificity of peer comments—future teachers treated them as complementary assets rather than competing authorities. Implications for Teacher Education The study’s findings carry significant implications for the design of modern teacher-preparation curricula. Rather than treating artificial intelligence and peer collaboration as isolated evaluation tools, teacher educators are urged to scaffold evaluative judgment explicitly. The authors recommend implementing structured decision matrices where pre-service teachers explicitly log their uptake rationale, target design dimensions, and anticipated trade-offs. Additionally, incorporating mandatory "coherence audits" after every revision cycle can help future educators ensure that localized fixes—such as adding a new data constraint—do not inadvertently undermine the overall mathematical and contextual integrity of the task. By fostering critical source comparison, rigorous verification routines, and systematic coherence checking, teacher education programs can better prepare future mathematics educators to harness generative AI and peer networks responsibly, translating raw feedback into meaningful and effective instructional practice. Post navigation The relationship between colleagues’ social support and burnout: a scoping review of quantitative research