As artificial intelligence rapidly transforms language education and professional workflows, a new empirical study has shed light on how student interpreters navigate high-stakes bilingual tasks under intense temporal constraints. While previous research has largely explored human-AI interaction in self-paced environments such as L2 writing and translation post-editing, little has been understood about how learners manage AI-enabled systems in real-time, high-pressure contexts like computer-assisted consecutive interpreting (CACI).

Conducted by researchers investigating human-computer integration in language mediation, the study evaluated twenty-two native Chinese-speaking Master’s interpreting trainees as they performed bidirectional CACI tasks. By integrating automatic speech recognition (ASR) and machine translation (MT) features into a specialized dual-interface platform, the experiment captured granular behavioral metrics using eye-tracking, digital pen-recording, and voice-recording apparatuses. The findings reveal distinct behavioral profiles, significant stage-based shifts, and clear links between interaction strategies and interpreting quality.

Mapping Behavioral Profiles: How Students Engage with AI

To decode the complex cognitive dance between human interpreters and external AI support, the research team applied cluster analysis to process-oriented eye-tracking and pen-stroke data. This methodology yielded four distinct student-AI interaction archetypes:

  • Intensive Engagers (37.12% of overall observations): Characterized by heavy AI reliance, these participants allocated the vast majority of their visual fixation time to the AI display area while concentrating their handwriting notes within the same digital space. Their high regression rates and moderate processing depth pointed to a constant monitoring and verification strategy.
  • Fast Scanners (28.79% of observations): Defined by a visually dense yet shallow approach, this group exhibited high fixation frequencies alongside exceptionally short saccade amplitudes and low deep-fixation proportions. Their workflow reflected rapid, localized visual scanning rather than sustained, deep attention.
  • Traditionalists (20.45% of observations): These students maintained a workflow closely mirroring conventional consecutive interpreting. They registered the lowest AI reading proportions and virtually zero note-taking in the AI interface, relying heavily on traditional paper-style digital notepads and internal memory supported by high deep-fixation intervals.
  • Frequent Switchers (13.64% of observations): Distinguished by the highest frequency of visual switches between the AI support area and the notepad, this group demonstrated an active integration strategy. They balanced moderate AI reading proportions with broad visual exploration and high regression rates, coordinating multiple external resources simultaneously.

Experimental Chronology and Methodology

The empirical investigation was structured around rigorous baseline controls and standardized testing environments. Twenty-seven graduate students were initially recruited from top-tier university Master of Translation and Interpreting (MTI) programs in Beijing, with twenty-two completing the full data-collection pipeline. Participants were systematically divided into two comparative cohorts: Group A comprised eleven students who had completed a rigorous 16-week compulsory course in systematic CACI training, while Group B comprised eleven students with zero formal training in computer-assisted interpreting. Both cohorts possessed comparable foundational training, professional tenure, and English proficiency benchmarks, such as standardized Test for English Majors Grade Eight (TEM8) scores.

During the experimental phase, participants completed bidirectional interpretation tasks—translating from Chinese to English and English to Chinese—under three distinct AI support configurations: low-quality ASR (targeted at an 85% accuracy threshold), high-quality ASR (exceeding 98% accuracy), and combined ASR-plus-MT support. Each speech stimulus was calibrated to approximately one minute in duration, mirroring professional proficiency examinations and standard classroom segments. Tobii Glass 3 eye-trackers operating at 100 Hz monitored visual fixations and saccades, Huion digital tablets logged pen strokes, and iFlytek recording pens captured acoustic outputs. Following the task sessions, researchers conducted cued retrospection interviews to contextualize user intent, cognitive load, and problem-solving strategies.

Training Efficacy and Cross-Stage Dynamics

A critical discovery of the study concerns how behavioral patterns evolve between the input stage (listening and note-taking) and the output stage (note-reading and speech production) of consecutive interpreting. Overall, 58.3% of task samples maintained consistent interaction profiles across both stages, while 41.7% demonstrated dynamic shifts.

Fast Scanners proved exceptionally stable, demonstrating a 96.2% consistency rate across stages. Conversely, Traditionalists displayed high volatility, with only 18.6% maintaining a purely non-AI workflow across both input and output phases. When facing the acute time pressure and linguistic demands of speech production, many untrained Traditionalists abandoned their initial resistance and gravitated toward late-stage AI reliance, inadvertently triggering decision-making vacillation and cognitive overload.

Furthermore, formal pedagogical training served as a stabilizer for interaction strategies. Trained students in Group A demonstrated a notably higher pattern consistency rate (66.7%) compared to their untrained peers in Group B (50.0%). Structured instruction appeared to equip trained interpreters with a clearer metacognitive framework, enabling them to select predictable, efficient AI collaboration pathways rather than reacting haphazardly to fluctuating external text feeds.

Performance Outcomes and Quality Implications

To evaluate the real-world impact of these interaction patterns, researchers analyzed the acoustic properties and expert-graded quality metrics of the interpretation products. Interestingly, the data revealed a clear operational asymmetry: interaction patterns established during the input stage significantly predicted interpretation performance, whereas output-stage patterns showed no statistically significant relationship with final product quality.

Specifically, participants exhibiting the Intensive Engager profile during the input stage scored significantly lower in fluency of delivery (FluDel) and target language quality (TLQual) compared to all other behavioral groups. Expert raters observed that Intensive Engagers frequently became bogged down in real-time verification of ASR and MT outputs. This relentless monitoring consumed vital cognitive resources, leaving insufficient capacity for smooth target-speech organization and delivery.

In contrast, Frequent Switchers—who dynamically coordinated AI references with internal cognitive representations without sacrificing autonomy—achieved the most robust performance outcomes. Meanwhile, Fast Scanners and Traditionalists sustained intermediate performance levels, suggesting that multiple viable pathways to adequate task completion exist, provided the interpreter maintains efficient allocation of cognitive bandwidth. Information completeness scores showed no significant variance across clusters, hovering above 6 out of 8 points, which indicates that while content capture remained stable due to resource substitutability, the qualitative polish of delivery hinged heavily on the efficiency of input-stage processing.

Broader Industry and Pedagogical Implications

The findings carry profound implications for the future design of interpreter education curricula and computer-assisted interpreting (CAI) software development. As artificial intelligence embeds itself deeper into professional language mediation, educators face the challenge of moving beyond binary paradigms of absolute acceptance or outright rejection.

First, the documented heterogeneity of student-AI interaction underscores the necessity of individualized pedagogical guidance. Rather than enforcing a single standardized workflow, instructor-led training should foster metacognitive reflection, encouraging student interpreters to critically evaluate their own attentional distribution, recognize the cognitive costs of excessive AI monitoring, and develop strategic autonomy.

Second, software engineers designing enterprise interpreting systems must rethink interface architecture to accommodate stage-specific cognitive workflows. During the input phase, interfaces should optimize ASR readability—such as chunk-based spacing and low-confidence visual flags—to facilitate rapid information parsing. During the output phase, platforms should streamline data presentation to minimize visual and linguistic interference, ensuring that technological tools act as empowering collaborative partners rather than cognitive distractions in high-stakes professional environments.