Esperanto virtual reality (vr) and artificial intelligence (ai) integrated virtual training system

The ESL English speaking and interpretation virtual training system, which integrates VR and AI, uses eye-tracking and head posture data to construct a hierarchical auditory prediction profile and generates acoustic events in the high-confidence prediction dimension. This solves the problem that existing VR interpretation training systems cannot effectively train auditory demasking ability under unpredictable acoustic interference, and improves learners' English demasking ability under non-stationary conditions.

CN122367685APending Publication Date: 2026-07-10HUBEI POLYTECHNIC UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-17
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Existing VR interpretation training systems cannot effectively train learners' auditory demasking ability for weak-signal English texts under unpredictable acoustic interference, resulting in a gap between training and practical application.

Method used

The ESL English speaking and interpretation virtual training system, which integrates VR and AI, uses a data acquisition module to acquire learners' eye-tracking data and head posture data in real time, constructs a hierarchical profile of auditory prediction, and intervenes by creating acoustic events in the high-confidence prediction dimension, including spatial teleportation, rhythm disruption, and semantic conflict strategies, forming a closed-loop training process.

Benefits of technology

It improves learners' ability to unmask English source language under unpredictable acoustic interference. By dynamically matching the intensity and type of intervention, it prompts the auditory attention system to unmask again in each intervention, thus improving the training effect under non-stationary interference conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122367685A_ABST
    Figure CN122367685A_ABST
Patent Text Reader

Abstract

This invention discloses a VR and AI-integrated ESL (English Speaking and Interpretation) virtual training system, belonging to the field of virtual reality language training technology. The system includes a data acquisition module, a prediction and interpretation module, an intervention decision-making module, a sound field perturbation module, and a profile update module. In a VR virtual space sound field, the system collects learners' eye-tracking data and head posture data in real time, interpreting the real-time prediction intensity of three dimensions: spatial positioning, temporal rhythm, and semantic expectation, constructing a hierarchical auditory prediction profile. When the prediction intensity of a certain dimension is detected to be consistently higher than a preset intensity threshold and fluctuating below a preset fluctuation threshold within a preset time period, an intervention strategy matching that dimension is called from a preset intervention strategy library to create an acoustic event that conflicts with the learner's current prediction model at the natural semantic boundary of the source language. The profile is then updated and feedback is provided based on the adaptation and recovery indicators. This invention enables targeted training of the ability to de-mask English source language under non-stationary interference environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of virtual reality language training technology, and more specifically, to an ESL (English as a Second Language) speaking and interpretation virtual training system that integrates VR and AI. Background Technology

[0002] In the field of ESL (English as a Second Language) speaking and interpretation virtual training, the integration of VR and AI technologies has formed several technical routes.

[0003] CN117576982A discloses a spoken language training method based on ChatGPT. This method generates spoken language training results by creating a training model, inputting user training requirements into the training model, and then generating the results after output review and interactive training. Its focus is on using AI dialogue agents to automate the generation of spoken language training content and the training process.

[0004] CN120375864A discloses a multimodal interpreting training evaluation method and device. The method determines the semantic correctness and emotional tendency of the translation by performing standard text matching and sentiment score calculation on the text data, extracts multiple features from the audio data to determine the translation fluency, and evaluates the trainee's translation behavior based on semantic score, sentiment score and fluency score. Its focus is on the multidimensional automatic scoring of the interpreting output results.

[0005] The aforementioned technical solutions mainly revolve around two aspects: dialogue content generation and interpretation output evaluation. In the design of the training environment, the interference noise is typically preset background noise or interference sounds switched according to a predetermined program. The pattern of interference changes is predictable for learners. After a period of adaptation, learners actually develop a memorized adaptation to specific noise patterns, rather than a general auditory demasking ability for unpredictable interference. Therefore, when faced with sudden acoustic interference in real interpretation scenarios, such as a speaker suddenly moving away from the microphone, a neighboring representative interrupting, or unexpected changes in accent or speaking speed, existing training methods struggle to provide effective training support, resulting in a gap between training and practical application. Therefore, a virtual training system for ESL (English Speaking and Interpretation) that integrates VR and AI is proposed to address the above issues. Summary of the Invention

[0006] To overcome the aforementioned deficiencies of the prior art, embodiments of the present invention provide a VR and AI-integrated ESL English speaking and interpretation virtual training system. The problem to be solved is that existing VR interpretation training systems cannot effectively train learners' auditory demasking ability for weak signal English texts under unpredictable acoustic interference, resulting in a gap between training and actual practice.

[0007] To achieve the above objectives, the present invention provides a VR and AI-integrated ESL (English Speaking and Interpretation) virtual training system, including a data acquisition module, a prediction and interpretation module, an intervention decision module, a sound field perturbation module, and a profile update module.

[0008] The data acquisition module collects learners' eye-tracking and head posture data in real time when English translation source text is played in the VR virtual space sound field. This module uses the eye-tracking unit and inertial measurement unit built into the VR headset to acquire data, without the need for external physiological sensors, providing a behavioral data foundation for subsequent interpretation of auditory prediction states.

[0009] The prediction and interpretation module interprets the learner's real-time prediction intensity in spatial localization, temporal rhythm, and semantic expectation dimensions from the eye-tracking and head posture data, constructing a hierarchical auditory prediction profile for the current moment. This module transforms the auditory prediction state, which cannot be directly measured, into quantifiable behavioral indicators, enabling real-time representation of the learner's auditory cognitive processing.

[0010] The intervention decision-making module monitors the auditory prediction hierarchical profile. When it detects that the prediction intensity of at least one dimension is consistently higher than a preset intensity threshold and fluctuates below a preset fluctuation threshold within a preset time period, it identifies that dimension as the current high-confidence prediction dimension and calls an intervention strategy matching that dimension from a preset intervention strategy library. This module determines whether the learner has formed a relatively solidified auditory prediction model by monitoring the stability of the prediction intensity, thereby determining the timing and type of cognitive intervention.

[0011] According to the intervention strategy, the sound field perturbation module creates acoustic events that conflict with the learner's current high-confidence prediction model at the natural semantic boundary of the source language. This module performs prediction conflict operation at the natural node of cognitive processing—the semantic boundary—forcing the auditory attention system to correct prediction errors online without affecting the coherence of semantic understanding.

[0012] After the acoustic event of the conflict occurs, the profile update module calculates the learner's adaptation recovery index on the current high-confidence prediction dimension based on the eye-tracking data, and uses this index to correct the prediction strength of the corresponding dimension in the auditory prediction hierarchical profile, thereby updating the profile. This module quantifies the learner's cognitive recovery efficiency and feeds the intervention effect back into the profile.

[0013] The intervention decision module also receives updated auditory prediction hierarchical profiles, thus forming a closed-loop training process of data collection, behavior interpretation, intervention decision, sound field perturbation, and profile updating, enabling the intensity and type of subsequent interventions to be dynamically matched with the learner's current cognitive state.

[0014] Furthermore, the prediction interpretation module interprets the prediction strength in three dimensions in the following way: Based on the stability of the learner's gaze point locking onto the target speaker's sound source, and the degree of feedforward compensation of the learner's head rotation relative to the sound source movement, the predictive strength in the spatial localization dimension is interpreted. Gaze point anchoring reflects the degree of focus on the sound source's location from a visual attention perspective, while head feedforward compensation reflects the brain's forward predictive state of the sound source's spatial trajectory from a motion prediction perspective.

[0015] The predictive strength in the temporal rhythm dimension is interpreted based on the degree of synchronization between the learner's blink release and the prosodic boundary of the source language. The degree of synchronization between blink release and the prosodic boundary reflects the auditory system's unconscious predictive ability of the temporal structure of speech; a higher degree of synchronization indicates a more stable grasp of the pronunciation rhythm.

[0016] The prediction strength in the semantic expectation dimension is interpreted based on the fluctuation in learner saccade amplitude caused by semantic distractor words played in non-attentional locations. The fluctuation in saccade amplitude reflects the extent to which the semantic expectation model is impacted by irrelevant information; the smaller the fluctuation, the more stable the semantic expectation.

[0017] Furthermore, the stability of the gaze point locking onto the target sound source is determined by the proportion of time the angle between the gaze point direction and the target sound source direction remains within a preset spatial angle tolerance. The higher this proportion, the more focused the learner's spatial attention is on the target sound source direction, and the more stable the prediction of the spatial positioning dimension.

[0018] Furthermore, the degree of feedforward compensation for head rotation relative to sound source movement is determined based on the lead time of the cross-correlation peak between the head rotation angular velocity and the target sound source displacement angular velocity. When the lead time of the cross-correlation peak is positive, it indicates that the head movement leads the sound source displacement in time; the longer the lead time, the higher the degree of feedforward compensation.

[0019] Furthermore, the degree of fluctuation in saccade amplitude caused by the semantic interference word is specifically the degree of deviation of the saccade amplitude from the baseline level under undisturbed conditions, where the baseline level is collected and determined during the learner's undisturbed quiet listening phase. The smaller the deviation, the stronger the semantic expectation model's ability to resist disturbances from irrelevant semantic information.

[0020] Furthermore, the intervention strategies include spatial teleportation, rhythm disruption, and semantic conflict strategies. Among them: The spatial teleportation strategy is used to attack the prediction of the spatial positioning dimension by changing the spatial orientation of the sound source, thereby invalidating the learner's existing spatial prediction. The rhythm disruption strategy is used to attack the prediction of the temporal rhythm dimension by randomizing the pause gaps to invalidate the learner's existing temporal predictions. The semantic conflict strategy is used to attack the prediction of the semantic expectation dimension by introducing semantic opposition information to invalidate the learner's existing semantic predictions.

[0021] Furthermore, when the spatial teleportation strategy is executed, while maintaining the source speech volume and content unchanged, the virtual perceived position of the target speaker's sound source is instantaneously switched to a spatial orientation that is mirror-symmetrical to or deviates from the learner's current gaze direction by 90 to 180 degrees. During the switching process, the speech content continues to play continuously. When the learner perceives a sudden change in the sound source position, the auditory system needs to re-execute spatial unmasking to lock onto the target sound source.

[0022] Furthermore, when the rhythm disruption strategy is executed, while keeping the semantic content and spatial position of the source language unchanged, the duration of the silent gap between adjacent phrases in the source language is randomized, disrupting the inherent rhythmic beat of the source language. By disrupting the learner's prediction pattern of pause rhythm, the learner is forced to adapt to a non-stationary temporal structure.

[0023] Furthermore, when the semantic conflict strategy is executed, while maintaining the normal playback of the source language from the target speaker's voice source, an interfering speech stream is played from the learner's non-attentional location. The interfering speech stream contains phrases that conflict with the semantic expectations of the current source language context. By introducing semantic conflict information from the non-attentional location, the semantic expectations are directionally perturbed without interrupting the main speech stream.

[0024] Furthermore, the adaptation recovery indicators include a gaze recovery indicator reflecting the speed at which the gaze point re-locks onto the target sound source, a blink synchronization recovery indicator reflecting the speed at which blink patterns and rhythms resynchronize, and a saccade suppression recovery indicator reflecting the speed at which semantic interference is eliminated. These three indicators quantify the learner's cognitive adaptation efficiency after prediction conflict from spatial, temporal, and semantic dimensions, respectively.

[0025] The image update module denotes the prediction strength of the corrected dimension as... The adaptation recovery index is marked as Through calculation formula The predicted intensity is corrected to obtain the updated predicted intensity p', where and For the preset weighting coefficients, satisfy This update method weights and fuses the predicted strength before intervention and the adaptive performance after intervention, so that the profile update can maintain the continuity of historical state and reflect the latest adaptation results.

[0026] The technical effects and advantages of this invention are as follows: (1) The system interprets the real-time prediction intensity of three dimensions—spatial localization, temporal rhythm, and semantic expectation—from the learner's eye-tracking data and head posture data, and constructs a hierarchical auditory prediction profile. Since eye-tracking data and head posture data can be obtained through the native interface of the VR headset, the above process does not rely on external physiological sensors. When the prediction intensity of a certain dimension is detected to be continuously higher than the preset intensity threshold and the fluctuation is lower than the preset fluctuation threshold within a preset time period, it is determined that the learner has formed a relatively solidified auditory prediction model, thereby determining the timing of cognitive intervention and making the selection of intervention timing based on evidence.

[0027] (2) After identifying the high-confidence prediction dimension, the system selects an intervention strategy matching that dimension from a pre-set intervention strategy library, creating acoustic events that conflict with the learner's current prediction model at the natural semantic boundary of the source language. Specifically, for the spatial localization dimension, the system uses instantaneous sound source location shifting; for the temporal rhythm dimension, it uses pause interval randomization; and for the semantic expectation dimension, it uses semantic conflict speech interference. The non-repetitiveness and unpredictability of the above intervention strategies make it difficult for learners to form a memorized adaptation to specific interference patterns, which helps to prompt the auditory attention system to re-unmask in each intervention, thus providing a training method for unpredictable acoustic interference.

[0028] (3) After each intervention, the system extracts dimension-specific adaptation indicators from the eye-tracking data and uses fixation recovery speed, blink synchronization recovery rate, and irrelevant saccade suppression recovery rate to quantify the learner's cognitive recovery efficiency. Based on this, the auditory prediction hierarchical profile is updated with weights. The updated profile is fed back to the intervention decision-making stage, so that the intensity and type of subsequent interventions can be dynamically matched with the learner's current cognitive state, forming a complete closed-loop training process, which helps to improve the learner's ability to demystify English source language under non-stationary interference conditions. Attached Figure Description

[0029] Figure 1 This is a schematic diagram of the system module framework of the present invention.

[0030] Figure 2 This is a schematic diagram illustrating the system execution principle of the present invention. Detailed Implementation

[0031] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0032] Example 1 As attached Figures 1 to 2The system demonstrates a VR and AI-integrated ESL (English Speaking and Interpretation) virtual training system. By monitoring learners' eye and head movements during listening, it dynamically interprets their auditory prediction state and actively creates acoustic events that conflict with the current prediction model when the prediction reaches a confidence peak. This forces the auditory attention system to repeatedly undergo unmasking, reconstruction, and cognitive adaptation, continuously enhancing its ability to cope with non-stationary interference in a closed-loop feedback loop.

[0033] This system consists of a data acquisition module, a prediction and interpretation module, an intervention and decision-making module, a sound field perturbation module, and a profile update module. These five modules operate collaboratively in the sequence of "data acquisition - behavior interpretation - intervention and decision-making - sound field perturbation - adaptation assessment - profile update - re-decision-making," forming a complete closed loop. In a single virtual interpreting training course, this closed loop can be executed dozens of times, with each loop providing precise intervention to address the learner's weaknesses in their current predicted state.

[0034] During closed-loop operation, the data acquisition module outputs the acquired data to the prediction and interpretation module in units of frames; The prediction interpretation module interprets the prediction intensity in three dimensions from each frame of data, constructs an auditory prediction hierarchical profile, and outputs it to the intervention decision module. The intervention decision module determines whether intervention is needed based on the profile and sends a strategy instruction to the sound field disturbance module after making the decision. After the sound field perturbation module performs acoustic perturbation, the profile update module extracts adaptation indicators from the eye-tracking data and updates the profile. The updated profile is then fed back to the intervention decision-making module, thus forming a closed loop. The implementation details of each module in the system are explained below.

[0035] The specific details regarding data collection are as follows: During the process of playing English translation source text in the VR virtual space sound field, the system first obtains the learner's raw behavioral data in real time through the data acquisition module.

[0036] This module uses a VR headset as its hardware foundation. The VR headset has a built-in eye-tracking unit and an inertial measurement unit (IMU). The eye-tracking unit captures the gaze direction, pupil state, and blinking motion, while the IMU outputs three-dimensional head posture information. This module calls the device's native eye-tracking application programming interface (API) and IMU driver interface to synchronously collect eye-tracking data and head posture data with uniform timestamps at a preset sampling frequency. The sampling frequency can be set according to the VR device's hardware capabilities and actual needs; in one example, the sampling frequency is set to 60Hz.

[0037] In one example, eye-tracking data includes the fixation point 3D direction vector, fixation point spatial coordinates, blink occurrence marker, saccade amplitude, and peak saccade velocity at each sampling time. The fixation point 3D direction vector is read directly from the eye-tracking interface, represented as a unit vector with the learner's binocular center as the origin, indicating the current gaze direction in 3D space. The fixation point spatial coordinates are determined by intersecting the gaze direction with the virtual space geometry. The blink occurrence marker is a binary signal; it is set to 1 when the eye-tracking interface detects that the number of frames with eyelid closure has reached a preset threshold, indicating that blink suppression or eyelid closure is currently in effect; otherwise, it is set to 0. The saccade amplitude is calculated based on the angle between the fixation point direction vectors of two adjacent frames. The peak saccade velocity is obtained by dividing the saccade amplitude by the saccade duration.

[0038] Head pose data includes quaternions of the head in three-dimensional space. These quaternions are output by the inertial measurement unit (IMU) driver interface at the same frequency as the eye-tracking data. This module performs differentiation on the quaternion sequence to calculate the head's rotational angular velocity in three-dimensional space. After calculation, the head rotational angular velocity sequence is low-pass filtered to remove high-frequency jitter. The cutoff frequency of the filter can be set according to actual needs; in one example, it is set to 15Hz.

[0039] After completing the above acquisition and preprocessing, the module packages the eye-tracking data and head pose data at the same timestamp into a data frame, outputs them frame by frame in chronological order, forming a continuous structured data stream, and sends it to the prediction and interpretation module.

[0040] The specific interpretation of the prediction is as follows: The prediction and interpretation module receives the data stream from the data acquisition module frame by frame, extracting features related to auditory prediction behavior from each frame. This module needs to interpret three dimensions of auditory prediction intensity from the learner's eye movements and head movements: spatial localization, temporal rhythm, and semantic expectation. The prediction intensity of each dimension is reflected by a first behavioral indicator, a second behavioral indicator, and a third behavioral indicator, respectively. The calculation methods for these three behavioral indicators are explained below.

[0041] The first behavioral indicator corresponds to the spatial localization dimension and is used to quantify the learner's auditory prediction strength of the target speaker's sound source location in space. When a person is focused on listening to a sound source, not only will the gaze point tend to anchor in the direction of the sound source, but the head will also unconsciously feed forward to compensate for the movement of the sound source. Both of these behavioral manifestations can reflect the degree to which the brain predicts the spatial trajectory of the sound source.

[0042] When calculating the first-line metric, this module first calculates the spatial anchoring value. Using the learner's eye center as the vertex, it calculates the angle between the gaze point direction vector and the direction vector pointing to the target sound source frame by frame. The direction vector pointing to the target sound source is obtained in real-time based on the target sound source's position coordinates in VR space and the learner's eye center position coordinates. Whether this angle falls within a preset spatial angle tolerance range is used as the criterion for single-frame spatial anchoring. In one example, the spatial angle tolerance is 10°. When the angle is within 10°, the frame is marked as an anchored frame; when the angle exceeds 10°, the frame is marked as a non-anchored frame. The module maintains a sliding time window of 5 seconds and counts the ratio of anchored frames within the window to the total number of frames within the window; this ratio is denoted as R.

[0043] Simultaneously, the module also tracks the jump behavior of the gaze point relative to the target sound source direction. The rule for determining a jump event is as follows: when the angle changes from within 10° to more than 10°, and the number of frames continuously exceeding 10° reaches a preset frame number threshold, it is counted as a jump event; if the angle exceeds the tolerance range and then recovers to within 10° within a time less than the preset frame number threshold, it is considered a noise fluctuation and is not counted in the jump count. In one example, the continuous frame number threshold is set to 3 frames. The total number of jump events within the window is counted and divided by the window length to obtain the jump frequency F. The arithmetic mean of R and 1 / (1+F) is used as the spatial anchoring value, where R reflects the degree to which the gaze point is stably anchored in the direction of the sound source, and 1 / (1+F) reflects the gaze point's ability to resist jump interference.

[0044] Based on the calculated spatial anchoring value, this module also calculates the feedforward compensation amplitude. When the sound source moves regularly, the head often rotates slightly in the same direction before the sound source actually moves. The module performs cross-correlation calculations on the head rotation angular velocity sequence and the target sound source displacement angular velocity sequence within the current sliding time window. The target sound source displacement angular velocity sequence is obtained by frame-by-frame differencing of the sound source's position coordinate time sequence in VR space. The cross-correlation calculation is performed within the lag time interval, calculating the cross-correlation coefficient between the two sequences at each lag time.

[0045] The conventional calculation method for cross-correlation in discrete form can be expressed as follows: ,in This is a sequence of head rotation angular velocities. The target sound source displacement angular velocity sequence, The lag time is used. In one example, the lag time search interval is [-1s, 1s]. The global maximum cross-correlation coefficient is taken within the entire interval. If the maximum value is lower than the preset correlation threshold, it indicates that there is no significant predictive coupling relationship between head movement and sound source displacement. In this case, the feedforward compensation amplitude D is directly set to zero. If the maximum value is not lower than the correlation threshold, then the lag time corresponding to the maximum value is taken. In one example, the correlation threshold is set to 0.3. When When the value is positive, it indicates that the head movement precedes the sound source movement. Through mapping function Normalized to the 0-1 interval, the feedforward compensation amplitude D is obtained, where This is a preset positive constant. When When the value is non-positive, D is set to zero.

[0046] Finally, the spatial anchoring value and D are linearly weighted and summed, with each weight set to 0.5, to obtain the first row index. The value of the first row index ranges from 0 to 1, with values ​​closer to 1 indicating a higher predictive strength of the spatial positioning dimension.

[0047] As one implementation method, when the system is deployed on a VR all-in-one device with limited computing power, the first behavior indicator can also be characterized solely by the gaze point dwell time ratio R. The value of R can be directly taken as the indicator, omitting the cross-correlation calculation of the feedforward compensation amplitude.

[0048] The second behavioral indicator corresponds to the temporal rhythm dimension and is used to quantify the learner's predictive strength regarding the prosodic structure and pause rhythm of the source language. When listening to coherent speech, blinking is suppressed during the high-cognitive-load focused listening phase. However, at the speaker's natural pauses, the cognitive load decreases briefly, blinking suppression is released, and blinking is released. Therefore, the degree of synchronization between the blinking release point and the prosodic boundary can reflect the auditory system's predictive ability regarding the temporal rhythm of speech.

[0049] Before calculating the second line index, this module first performs offline prosodic annotation on the source language. The prosodic annotation uses an automatic boundary detection algorithm based on audio envelope and fundamental frequency trajectory. This algorithm scans the acoustic features of the source language audio, detects abrupt changes in energy and fundamental frequency between adjacent speech units, and stores the times of phrase boundaries and sentence boundaries in the form of a timestamp list, with each timestamp marking a prosodic boundary.

[0050] During training, this module identifies blink release points based on the blink occurrence marker sequence in the eye-tracking data. The specific determination rule is: if the blink occurrence marker remains at 1 for a continuous duration exceeding the preset minimum eye-closing duration, and then jumps from 1 to 0, this jump moment is recorded as a valid blink release point. In one example, the minimum eye-closing duration is set to 50ms.

[0051] Then, a time matching window is formed by extending forward and backward by a preset duration, centered on each rhythm boundary timestamp. In one example, the window is extended forward and backward by 100ms each, resulting in a matching window width of 200ms. If the timestamp of a blink release point falls within any time matching window, that blink release is counted as a synchronization event.

[0052] Based on this, the ratio of the total number of synchronization events to the total number of effective blink releases within the entire training segment is calculated, and this ratio is used as the raw value of the second behavior index. The raw value is smoothed to eliminate inter-frame jitter, and then the smoothed value is linearly mapped to the 0 to 1 interval using the min-max normalization method to obtain the second behavior index.

[0053] When using min-max normalization, the minimum and maximum values ​​can be preset to initial ranges during system initialization and gradually updated to the actual observed minimum and maximum values ​​as training data accumulates. In one example, smoothing can be achieved using an exponential moving average with a smoothing factor of 0.3.

[0054] As an alternative implementation, the width of the timing window can also be adaptively adjusted according to the real-time speech rate of the source speaker. The average interval duration of the most recent prosodic boundaries is taken as the speech rate measure, and the timing window width maintains a constant ratio to this average interval duration, while setting a lower and upper limit for the window width. In one example, the average interval duration of the most recent 5 prosodic boundaries is used, with a lower limit of 100ms and an upper limit of 300ms for the window width. Regardless of the specific value of the window width, the rule for determining whether the blink release point is a synchronous event remains unchanged.

[0055] The third behavioral indicator corresponds to the semantic expectation dimension, used to quantify learners' ability to actively predict the semantic content of the source language. Once the brain has established a forward prediction of the semantic content of the current text, sudden semantically irrelevant or conflicting information will trigger unconscious attentional capture, manifested as a momentary increase in saccade amplitude. The smaller the perturbation in saccade amplitude, the more robust the semantic expectation model.

[0056] The calculation of this metric depends on the occurrence of semantic interference events. When the sound field perturbation module executes the semantic conflict strategy and the first frame of the interfering speech stream begins playback, it sends an event notification with a timestamp to the prediction and interpretation module. The prediction and interpretation module uses this timestamp as a reference to start the calculation of the third-line metric. During the period without receiving new semantic conflict event notifications, the third-line metric maintains the value obtained from the previous event calculation.

[0057] Before formal training begins, this module first acquires an interference-free baseline during a quiet listening phase. The quiet listening phase works as follows: only the target speaker's original speech is played in the VR virtual space sound field, without any interference or predictive conflict operations. The learner maintains a natural listening state in the VR space, without performing any interpretation tasks, for a preset acquisition duration. In one example, the acquisition duration is 30 seconds. During this period, saccade amplitude values ​​are recorded frame by frame. To eliminate the contamination of baseline data by active saccades, frames with saccade amplitudes exceeding a preset saccade threshold are discarded, retaining only data on small saccades under fixational maintenance. In one example, the saccade threshold is 5°. The mean saccade amplitude is calculated using the retained frames. and standard deviation , as a baseline parameter.

[0058] In another implementation, and Data was collected separately before each training task, with a collection duration of 20 seconds. Historical baseline values ​​were not used to ensure that the baseline parameters were more closely matched with the learner's physiological state and attention level during the current training session.

[0059] Once the formal training phase begins, upon receiving an event notification of a semantic interference word, the maximum eye saccade amplitude within a preset observation window after the event's timestamp is recorded as . In one example, the observation window length is 500 ms. Calculation The degree of deviation relative to the baseline level, deviation value .right To perform a monotonically decreasing nonlinear mapping, in one example, the mapping function takes... :when A time mapping value of 1 indicates that the semantic expectation is completely unperturbed; when The time mapping value is 0.5; As the value continues to increase, it approaches 0. The mapping result is the third row indicator.

[0060] After the above three behavioral indicators are calculated, they are used as the real-time prediction strengths for the spatial positioning dimension, temporal rhythm dimension, and semantic expectation dimension, respectively, and are denoted as follows: , , The values ​​of all three are between 0 and 1. , , Together, they form a hierarchical auditory prediction profile for the current moment, which is continuously output to the intervention decision module in the form of data frames.

[0061] The specific details regarding intervention decisions are as follows: The intervention decision module continuously receives , , For time series analysis, three decisions need to be made at each judgment period: whether intervention is needed, which dimension to intervene in, and what intervention strategy to adopt. The specific implementation of these three decisions is explained step by step below.

[0062] First, determine whether intervention is needed. The criterion is that at least one dimension is identified as a dimension with high confidence in prediction. The module internally presets three decision parameters: intensity threshold Th, fluctuation threshold... and the length of the time period In one exemplary setup, Th is 0.75. Take 0.08, Take 8 seconds. For the dimension... ,in Take 1, 2, and 3 respectively, and at any current time, take the past... Within the duration Given a sequence, calculate its mean and standard deviation. When the mean is greater than Th and the standard deviation is less than Th... Time, dimension It was identified as a dimension where the learner currently has high confidence in prediction. This criterion indicates that the learner's predictions in this dimension not only remain at a high level, but the state has also become highly stable. The brain has formed a relatively solidified predictive model of the acoustic characteristics of this dimension, making this the optimal time to implement intervention.

[0063] Secondly, when multiple dimensions simultaneously meet the above criteria, one dimension is selected as the target of this intervention. In one implementation, the selection is based on a preset priority, with the priority order being: spatial positioning dimension takes precedence over temporal rhythm dimension, and temporal rhythm dimension takes precedence over semantic expectation dimension.

[0064] This prioritization is based on the hierarchical characteristics of auditory cognition: spatial localization is the foundational level of auditory perception, temporal rhythm is intermediate, and semantic anticipation belongs to a higher level of cognitive processing. Prioritizing intervention at the foundational level helps build a more robust auditory demasking ability. In another implementation, the module counts the cumulative number of times each dimension has been selected as a high-confidence prediction dimension and has received intervention in the current complete training task history, prioritizing the dimension with the fewest cumulative counts; if multiple dimensions have the same cumulative count, the aforementioned preset priority is used. The initial cumulative count for each dimension is 0, and after each intervention, the cumulative count of the intervened dimension increases by 1.

[0065] Finally, after determining the intervention dimensions, the module matches the corresponding intervention strategies from its internally maintained intervention strategy library and determines the execution timing. The intervention strategy library stores three sets of one-to-one correspondences in the form of a mapping table: spatial positioning dimension corresponds to spatial teleportation strategies, temporal rhythm dimension corresponds to rhythm disruption strategies, and semantic expectation dimension corresponds to semantic conflict strategies. Based on the identified high-confidence prediction dimensions, the module directly extracts the strategies that match those dimensions from the strategy library.

[0066] Regarding the timing of intervention, the timing is controlled by semantic boundary trigger markers. The intervention decision module simultaneously receives semantic boundary trigger markers from the AI ​​semantic analysis unit. The AI ​​semantic analysis unit takes the source language audio stream as input, performs real-time automatic speech recognition on the audio, and obtains the text transcription result of the source language; then it performs dependency parsing on the text transcription result to identify clause boundaries and sense group segmentation points; and outputs semantic boundary trigger markers at clause termination positions and sense group completion positions. The reason for choosing to perform intervention at the semantic boundary is that the semantic boundary itself is a natural node in the learner's cognitive processing. Applying a conflict event at this point can create maximum cognitive conflict without disrupting the coherence of semantic understanding.

[0067] When the intervention decision module detects that a semantic boundary trigger marker is valid and a high-confidence prediction dimension exists, it immediately sends the selected strategy along with the execution command to the sound field perturbation module. If no valid semantic boundary trigger marker appears after identifying a high-confidence prediction dimension, the module enters a waiting state and continues to receive marker signals until the next semantic boundary appears.

[0068] The intervention decision-making module runs continuously using a sliding time window, with each decision using the most recent time window. within seconds , , Data. The updated predicted intensity value is automatically included in the sliding window for the calculation of mean and standard deviation in the next judgment period without the need for additional triggering signals, thus forming a seamless connection with the subsequent portrait update process.

[0069] The specific details regarding the sound field disturbance are as follows: After receiving policy instructions from the intervention decision module, the sound field perturbation module performs a prediction conflict operation at the specified semantic boundary of the source language. This creates an acoustic event that conflicts with the learner's established prediction model on the high-confidence prediction dimension, allowing the learner's brain to detect the prediction error signal and thus actively initiate the reallocation of auditory attention and online correction of the prediction model. Depending on the selected strategy, the specific execution method is divided into three scenarios.

[0070] In the first scenario, the selected strategy is a spatial teleportation strategy, corresponding to the attack spatial localization dimension. This module keeps the source speech volume and content completely unchanged, only modifying the three-dimensional position coordinate parameters of the target sound source in the spatial audio engine to change the perceived orientation of the sound source. The angle between the new orientation and the learner's current gaze direction must be within the range of 90° to 180° to ensure that the new orientation exceeds the learner's central visual field. The current gaze direction is determined based on the gaze direction vector of the most recent frame of eye-tracking data.

[0071] In one implementation, the switching transition time is 20ms, and the target angle is 135°. Within the 20ms transition window, the spatial audio engine uses a spherical linear interpolation algorithm to update the spatial coordinates of the sound source frame by frame from the original orientation to the new orientation. After the transition, the sound source position stabilizes in the new orientation, and the audio content continues to play continuously during the switching process.

[0072] In another implementation, the new orientation is taken at a 180° angle to the current gaze direction, that is, the target sound source is switched to the area directly behind the learner, and the switching transition time is shortened to 10ms.

[0073] In the second scenario, the selected strategy is a rhythm disruption strategy, corresponding to the attack on the temporal rhythm dimension. This module maintains the semantic content and spatial location of the source language unchanged, only applying random perturbations to the duration of silent gaps between adjacent phrases in the audio. First, the entire audio signal is scanned using an energy threshold detection method, marking segments with audio energy continuously below a preset energy threshold as gap segments, and extracting the original duration of each gap segment. Then, a random multiplier is generated independently for each gap segment. , Random sampling is performed within a preset sampling range according to a uniform distribution.

[0074] In one example, the sampling range is 0.6 to 1.4. To ensure the unpredictability of rhythm changes, a constraint rule is introduced for the pseudo-random sequence generator: when three consecutive gaps are given a large multiplier, the multiplier direction of the third gap is forcibly adjusted. The large multiplier is defined as... or The adjustment method is to adjust the third gap. The value is replaced by its symmetrical value relative to 1.0, for example... Replace with , Replace with Then, the duration of each interval segment was modified to... .

[0075] The duration modification of the gap segment waveform is implemented using the WSOLA algorithm. WSOLA is a time-domain duration adjustment algorithm based on waveform similarity superposition. It takes the original gap segment waveform and the target duration as input. The algorithm searches for similar waveform segments matching the target duration in the original waveform at certain step sizes, and adjusts the total length of the output waveform by superimposing or extracting these similar segments, thereby achieving duration stretching or compression without changing the speech pitch. All modified gap segments are replaced in the audio stream in real time according to their original order and output to the spatial audio rendering pipeline for learners to listen to.

[0076] In another implementation, The sampling range was expanded to 0.5 to 1.5, and the stretching and compression operations were implemented using the PSOLA algorithm. The PSOLA algorithm is a time-domain duration adjustment algorithm based on pitch synchronization. It takes the original speech signal and the target duration as input, first marks the pitch period of the signal, and then performs superposition or decimation processing on the waveform at the pitch period granularity, outputting a speech signal whose duration changes while the pitch and timbre remain unchanged.

[0077] The third scenario involves a semantic conflict strategy, corresponding to the attack on the semantic expectation dimension. This module, while maintaining the normal playback of the target speaker's voice source, introduces a stream of interfering speech containing semantically conflicting content from the learner's current non-attentional location. In one implementation, the non-attentional location is defined as a spatial region with an angle exceeding 120° from the learner's current gaze direction, i.e., the learner's side-rear or rear. A spatial location point within this region is selected as the interfering voice source location, and a stream of interfering speech of a preset length is played at the same volume level as the target speaker's voice source. In one example, the length of the interfering speech stream is 1 second. The start time of the interfering speech stream playback is precisely aligned with the time of the semantic boundary trigger marker.

[0078] Meanwhile, when the first frame of the interfering speech stream begins to play, the sound field perturbation module sends a semantic conflict event notification to the prediction and interpretation module. This notification carries a timestamp of the start of the interference, which is used by the prediction and interpretation module to calculate the third behavior index.

[0079] The content that interferes with the audio stream is generated by combining offline construction with online matching.

[0080] In the offline phase, a speech fragment library containing semantically opposite word groups to the context of each topic is constructed by performing semantic analysis on the corpus in advance. Specifically, each text in the training corpus is semantically labeled and classified into topics, and the most frequently occurring keyword groups are extracted for each topic; then, for each keyword group, opposing word groups with opposite or contradictory meanings are constructed to form a conflict word group library; then, the text-to-speech engine is used to synthesize each group of words in the conflict word group library into speech fragments one by one, and these speech fragments are indexed and associated with the corresponding semantic boundary positions in the source text.

[0081] During the online phase, the AI ​​semantic analysis unit outputs semantic boundary trigger markers and extracts the core predicate and object from the most recent complete clause preceding the semantic boundary as context keywords. During retrieval, the semantic category of the core predicate and the entity type of the object are used as joint query conditions to filter out conflicting word groups with opposite semantic categories and consistent entity types from the index. If multiple matching results exist, the audio segment corresponding to the most frequent set of conflicting word groups in the corpus is selected for playback first.

[0082] For example, when the source language is describing a company's profit growth, the AI ​​semantic analysis unit extracts the core predicate "growth" and the object "profit," retrieves the conflict phrase library to obtain the corresponding audio segments of semantically contradictory words such as "decline" or "loss," and plays them synchronously from a non-attentional location, directly challenging the learner's established semantic expectation model.

[0083] In another implementation, the criteria for determining the non-attentional orientation are relaxed to include an angle of more than 90° with the gaze direction, the length of the interfering speech stream is increased to 2 seconds, and a conflicting word phrase is placed at the positions of the first 500 seconds and the last 500 seconds of the interfering speech stream, with a 1-second interval in between, to form a bimodal semantic perturbation effect.

[0084] The details regarding the portrait update are as follows: After the sound field perturbation module completes the prediction conflict operation, the profile update module extracts adaptation indicators corresponding to the intervened dimension from the eye-tracking data to quantify the learner's cognitive recovery efficiency after experiencing prediction conflict. There are three types of adaptation indicators, each corresponding to one of the three prediction dimensions. Only the adaptation indicator for the currently intervened dimension is calculated each time, while the prediction strength of the other dimensions remains unchanged. The calculation methods for the three adaptation indicators are explained below.

[0085] The adaptation metrics corresponding to the spatial localization dimension include gaze recovery latency and gaze convergence speed. The gaze recovery latency is timed from the moment the conflict prediction operation begins, and this moment is precisely recorded as a system timestamp. The angle between the learner's gaze direction vector and the new location direction vector of the target sound source is continuously tracked. When this angle first re-enters the spatial angular tolerance range, the time elapsed is recorded; this is the gaze recovery latency, denoted as . .

[0086] The gaze convergence velocity is measured after the gaze point enters the spatial angular tolerance range: the time series of Euclidean distances between the spatial coordinates of the gaze point and the new spatial coordinates of the target sound source are recorded during a preset measurement period. In one example, the preset measurement period is 2 seconds.

[0087] If the gaze point jumps out of the spatial angle tolerance range again during the measurement period, the data segment before the jump is truncated for subsequent fitting; if the length of the data segment that stably remains within the tolerance range is less than the preset minimum effective duration, the adaptation index is deemed invalid, and the previous effective value is used in subsequent update calculations. In one example, the minimum effective duration is 0.5s. The distance sequence within the effective data segment is fitted to a first-order exponential decay model, and the decay constant obtained from the fitting is taken as the gaze convergence rate, denoted as . .

[0088] Will and By performing a weighted summation with each weight set to 0.5, we obtain the adaptation index corresponding to the spatial positioning dimension. As an alternative implementation, the gaze convergence speed can also be directly taken as the slope of the linear regression of the Euclidean distance sequence between the gaze point and the target sound source within 1 second after the gaze point enters the spatial angle tolerance range, and the absolute value of this slope is taken as... Similarly with We obtain by weighted summation .

[0089] The adaptation metric corresponding to the temporal rhythm dimension is the blink synchronization pattern recovery rate. The module records the blink synchronization ratio at the most recent consecutive prosodic boundaries before the rhythm disruption event occurs, and takes the average of these prosodic boundary synchronization ratios as a reference value. In one example, the number of consecutive prosodic boundaries is taken as 5. After the rhythm disruption event occurs, starting from the first prosodic boundary after the event, the blink synchronization ratio on each boundary is counted. When the blink synchronization ratio of two consecutive prosodic boundaries is not less than 5, the blink synchronization ratio is calculated. When the first of these two consecutive boundaries is taken as the recovery boundary, the total number of prosodic boundaries traversed from the first boundary after the event to this recovery boundary is counted and denoted as . .Pick The reciprocal of this value is used as the blink synchronization mode recovery rate, i.e., the adaptation index. .

[0090] As another implementation scheme, the reference value The calculation range was expanded from 5 prosodic boundaries before the event to 10 prosodic boundaries before the event, and the average blink synchronization ratio at these 10 prosodic boundaries was taken as the mean. .

[0091] The adaptation metric for the semantic expectation dimension is the irrelevant saccade suppression recovery rate. The module records the occurrence time of semantic conflict events and continuously tracks the saccade amplitude deviation value for each subsequent frame from that moment. .when After a sustained decline from the post-conflict high, when it first falls below a preset decline threshold, the time elapsed between the moment the conflict occurred and this decline point is recorded. In one example, the fallback threshold is set to 1.0. The reciprocal of this value is used as the recovery rate of irrelevant saccade suppression, i.e., the adaptation index. As an alternative implementation, the fallback threshold can be relaxed from 1.0 to 2.0.

[0092] After obtaining the adaptation index for the intervened dimension, the profile update module uses this index to update the prediction strength of the corresponding dimension. For the intervened dimension... The predicted strength before the update and corresponding adaptation indicators The components are merged according to preset weights, and the fusion formula is as follows: ,in and The preset weighting coefficients sum to 1. For the uninterrupted dimension, its predictive strength remains unchanged, i.e. .

[0093] In one implementation, Take 0.7, A value of 0.3 indicates that the updated predicted strength retains 70% of historical inertia and incorporates 30% of recent adaptation. In another implementation, Take 0.8, Using 0.2 makes the update process smoother.

[0094] Adaptation indicators Before being integrated, the values ​​need to be normalized to match the predicted intensity's range of 0 to 1. The normalization process is as follows: for each dimension, maintain a global historical maximum observation value since the initial training. Each time a new adaptation metric is acquired, divide the metric by the corresponding dimension's historical maximum observation value to obtain the normalized value. If the normalized value is greater than 1, it is truncated to 1. During the system's initial training, the historical maximum observation value for each dimension is preset to 0.1 as the initial value. After each training iteration, if a newly observed adaptation metric value exceeds the currently stored historical maximum value, the new value replaces the old one. When the system is reset or switched to a new learner, the historical maximum observation value reverts to the initial value of 0.1.

[0095] Calculated , , This refers to the updated predicted intensity, and together these three factors constitute the updated auditory prediction hierarchical profile. This updated profile is then output to the intervention decision-making module. Since the intervention decision-making module continuously runs using a sliding time window, the updated... , , The value is automatically included in the sliding time window in the next judgment period, participating in the calculation of the mean and standard deviation. Thus, it can enter a new round of identification and intervention of high-confidence prediction dimensions without additional triggering signals. At this point, the system has completed a complete closed loop cycle from data collection, behavior interpretation, intervention decision-making, sound field perturbation to adaptation assessment and profile update.

[0096] In a typical virtual English interpreting training course, the aforementioned closed loop can be repeated dozens of times. With the accumulation of training sessions, the learner's auditory prediction profile shows an overall optimization trend, with each dimension gradually improving after targeted intervention: fixation recovery latency gradually shortens, fixation convergence speed gradually accelerates, blink synchronization pattern recovers faster after rhythmic perturbations, and irrelevant saccade suppression after semantic perturbations is also more rapid. Ultimately, in real interpreting scenarios facing unpredictable acoustic interference, the learner's auditory system can more quickly demask and re-lock onto the target sound source, significantly improving English source language comprehension and interpreting performance under weak signal conditions.

[0097] The above descriptions are merely some specific implementations of the technical solution of this invention. The specific implementation details of each module can be combined with each other without conflict. For the specific threshold values, window lengths, and weight coefficients involved in each module, those skilled in the art can make adaptive adjustments according to the actual application scenario and hardware conditions, and such adjustments do not depart from the protection scope of the claims of this invention.

Claims

1. A virtual training system for ESL (English as a Second Language) speaking and interpretation that integrates VR and AI, characterized in that, include: The eye-tracking and head-movement data acquisition module is used to collect learners' eye-tracking data and head posture data in real time when playing English interpretation source text in the sound field of VR virtual space. The auditory prediction behavior interpretation module is used to interpret the learner's real-time prediction intensity in the spatial positioning dimension, temporal rhythm dimension and semantic expectation dimension from the eye tracking data and head posture data, and construct the auditory prediction hierarchical profile at the current moment. The targeted intervention decision module is used to monitor the auditory prediction layer profile. When it is detected that the prediction intensity of at least one dimension is continuously higher than the preset intensity threshold and the fluctuation is lower than the preset fluctuation threshold within a preset time period, the dimension is identified as the current high-confidence prediction dimension, and the targeted acoustic intervention strategy matching the dimension is called from the preset intervention strategy library. A three-dimensional sound field perturbation execution module is used to generate acoustic events that conflict with the learner's current high-confidence prediction model at the natural semantic boundary of the source language, according to the targeted acoustic intervention strategy. The adaptation assessment and profile update module is used to calculate the learner's adaptation recovery index on the current high-confidence prediction dimension based on the eye-tracking data after the occurrence of the acoustic event of the conflict, and use the index to correct the prediction intensity of the corresponding dimension in the auditory prediction hierarchical profile to achieve profile update. The targeted intervention decision module is also used to receive the updated auditory prediction hierarchical profile for the next identification of high-confidence prediction dimensions.

2. The VR and AI-integrated ESL (English Speaking and Interpretation) virtual training system according to claim 1, characterized in that, The auditory prediction behavior interpretation module is specifically used for: Based on the stability of the learner's gaze point locking onto the target speaker's sound source, and the degree of feedforward compensation of the learner's head rotation relative to the sound source movement, the predicted intensity in the spatial localization dimension is interpreted. Based on the degree of synchronization between the learner's blink release time and the prosodic boundary of the source language, the prediction intensity in the temporal rhythm dimension is interpreted; Based on the degree of fluctuation caused by semantic interference words played in non-attentional locations to learners' eye saccade amplitude, the prediction strength on the semantic expectation dimension is interpreted.

3. The VR and AI-integrated ESL English speaking and interpretation virtual training system according to claim 2, characterized in that, The stability of the gaze point locking onto the target sound source is determined by the proportion of time the angle between the gaze point direction and the target sound source direction remains within a preset spatial angle tolerance.

4. The VR and AI-integrated ESL English speaking and interpretation virtual training system according to claim 2, characterized in that, The degree of feedforward compensation of the head rotation relative to the sound source movement is specifically determined based on the lead time of the cross-correlation peak between the head rotation angular velocity and the target sound source displacement angular velocity.

5. The VR and AI-integrated ESL English speaking and interpretation virtual training system according to claim 2, characterized in that, The degree of fluctuation in saccade amplitude caused by the semantic interference word is specifically the degree of deviation of the saccade amplitude from the baseline level in an undisturbed environment, wherein the baseline level is collected and determined during the learner's undisturbed quiet listening phase.

6. The VR and AI-integrated ESL (English Speaking and Interpretation) virtual training system according to claim 1, characterized in that, The targeted acoustic intervention strategies include: spatial teleportation strategy, rhythm disruption strategy, and semantic conflict strategy; Specifically, the spatial teleportation strategy is used to attack the prediction of the spatial positioning dimension, the rhythm disruption strategy is used to attack the prediction of the temporal rhythm dimension, and the semantic conflict strategy is used to attack the prediction of the semantic expectation dimension.

7. The VR and AI-integrated ESL English speaking and interpretation virtual training system according to claim 6, characterized in that, When the spatial teleportation strategy is executed, it is used to instantly switch the virtual perceived position of the target speaker's sound source to a spatial orientation that is mirror-symmetrical to or deviates from the learner's current gaze direction by 90 to 180 degrees, while keeping the source speech volume and content unchanged.

8. The VR and AI-integrated ESL English speaking and interpretation virtual training system according to claim 6, characterized in that, When the rhythm disruption strategy is executed, it is used to randomize the duration of silent gaps between adjacent phrases in the source language while keeping the semantic content and spatial position of the source language unchanged, so as to disrupt the inherent rhythmic beat of the source language.

9. The VR and AI-integrated ESL English speaking and interpretation virtual training system according to claim 6, characterized in that, When the semantic conflict strategy is executed, it is used to play an interfering speech stream from the learner's non-attentional location while maintaining the normal playback of the source language from the target speaker's voice source. The interfering speech stream contains phrases that conflict with the semantic expectations of the current source language context.

10. The VR and AI-integrated ESL (English Speaking and Interpretation) virtual training system according to claim 1, characterized in that, The dimension-specific adaptation indices include: a gaze recovery index reflecting the speed at which the gaze point re-locks onto the target sound source; a blink synchronization recovery index reflecting the speed at which blink patterns and rhythms resynchronize; and a saccade suppression recovery index reflecting the speed at which semantic interference is eliminated. The adaptation assessment and profile update module is specifically used to, denoted as p, the predicted strength of the dimension to be corrected, and a, the adaptation recovery index, and calculate using the formula... The predicted intensity is corrected to obtain the updated predicted intensity p', where and For the preset weighting coefficients, satisfy .

Citation Information

Patent Citations

  • Spoken language training method and device based on ChatGPT, electronic equipment and medium

    CN117576982A

  • Multi-modal interpretation training evaluation method and evaluation device

    CN120375864A