Interactive English teaching method and system based on AI vision

By analyzing students' voices and mouth movements in real time and generating dynamic feedback markers, the system solves the problems of delayed and inaccurate feedback in complex learning behaviors in existing systems, and optimizes personalized teaching paths and improves learning efficiency.

CN120672282AInactive Publication Date: 2025-09-19HUNAN INST OF INFORMATION TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510770074.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-09-19
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing AI vision-based interactive English teaching systems cannot effectively handle the interaction between voice and visual information when dealing with complex learning behaviors. The feedback is delayed or inaccurate, and it cannot provide flexible and adaptive responses, which limits its application effect in complex scenarios.

Method used

By obtaining the students' voice and mouth image frames, calculating the pronunciation of phonemes and the duration of mouth opening and closing, determining the offset and generating an offset marker, screening the sentence sets with high distribution frequency, analyzing the consistency between the active interval of behavioral signals and the timing of the triggering frames of the motive words, generating a language action response linkage sequence, evaluating the synchronization between voice output and visual actions, and outputting the individual English interactive expression matching results.

Benefits of technology

It achieves real-time and accurate feedback and interaction, improves learning efficiency and feedback accuracy, ensures efficient alignment of language expression with standard templates, and optimizes personalized teaching paths.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120672282A_ABST
    Figure CN120672282A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent education, in particular to an interactive English teaching method and system based on AI vision, and the method comprises the steps: obtaining the voice and mouth dynamic image frame segments of a student, extracting an offset frame segment to generate a semantic motion disjunction interval, analyzing the time sequence consistency of a behavior signal and a motion cause word triggering frame segment, and calculating the motion matching degree. Time synchronism of the speech output and the visual action is evaluated. According to the invention, through real-time analysis of the voice and mouth dynamic image of the student, the alignment offset of the pronunciation and the mouth action is identified, and in combination with behavior signals such as eye fixation and head rotation of the student, the interactivity between the language action and the behavior of the learner is comprehensively evaluated, so that the efficient alignment of the language expression and the standard template is ensured; by analyzing the distribution frequency and the matching degree of the voice behaviors, a personalized teaching path is optimized. In addition, delay frames and contour offset are deeply analyzed, the synchronism of voice and visual actions is enhanced, and the real-time feedback and interaction effect in the learning process is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent education technology, and in particular to an interactive English teaching method and system based on AI vision. Background Art

[0002] The field of intelligent education technology encompasses a technical system that utilizes artificial intelligence-related methods to assist, optimize, or reconstruct the educational process. The core content is to dynamically perceive, analyze, and generate strategies for teaching objects, teaching content, teaching behaviors, and teaching environments through artificial intelligence methods such as computer vision, speech recognition, and natural language processing, thereby achieving personalized, interactive, and sustainable learning support. Intelligent education not only covers the digitization and networking of educational content, but also involves key links such as real-time monitoring and feedback generation of learning behaviors, as well as dynamic regulation of learning paths and resource recommendations. Artificial intelligence perception and decision-making capabilities are widely embedded in the teaching process to achieve dynamic matching between teaching resources, teaching content, and teaching behaviors, thereby supporting multimodal learning interaction, adaptive learning rhythm, and precise teaching path planning.

[0003] Among them, the interactive English teaching method and system based on AI vision refers to an English language teaching scenario that uses computer vision technology to identify learners' visual behavioral features such as facial movements, lip shape changes and body postures, and performs matching analysis based on preset teaching tasks or standard language expression templates, thereby completing dynamic recognition and interactive feedback in the learning process. It covers multiple sub-tasks such as English speech recognition training, semantic understanding exercises and language expression behavior correction. Specifically, the learner's facial image data is collected by a visual acquisition device, and key visual features are extracted through an image processing model. Then, cross-validation and comparison are performed in combination with the text semantic information in the language teaching content to construct an association judgment path between individual behavior and target language expression. Relying on the human posture estimation model and lip shape recognition sub-model constructed in the image processing module, a refined analysis of the learner's movements during pronunciation is achieved. By comparing with standard samples, the learner's current language expression behavior is classified and recorded to achieve feedback processing of learning behavior and generation of training control logic.

[0004] Existing technologies lack real-time feedback and dynamic adjustment capabilities, particularly when faced with complex learning behaviors, and are unable to effectively handle the interaction between speech and visual information. Their over-reliance on static models and pre-set tasks prevents timely identification and adjustment of learners' pronunciation and movement deviations, resulting in delayed or inaccurate feedback. Furthermore, existing systems have limited ability to perceive and analyze learner behavior, failing to provide flexible and adaptive responses in multimodal interactive environments, thus limiting their effectiveness in complex scenarios. Summary of the Invention

[0005] In order to solve the technical problems existing in the prior art, the embodiments of the present invention provide an interactive English teaching method and system based on AI vision. The technical solution is as follows:

[0006] The interactive English teaching method based on AI vision includes the following steps:

[0007] S1: Obtain student speech and mouth image frames during oral interactive teaching, calculate phoneme pronunciation and mouth opening and closing duration, determine the alignment offset between phoneme pronunciation and mouth opening and closing, extract offset frame segments, and generate speech action corresponding offset identifiers;

[0008] S2: Calculate the phoneme distribution ratio based on the sentences covered by the offset frame segments in the offset identifier corresponding to the speech action, select a sentence set whose distribution frequency ratio exceeds the preset task participation ratio, and generate a teaching pronunciation structure mapping set;

[0009] S3: calling the sentence segments in the teaching pronunciation structure mapping set, analyzing the consistency between the active interval of the student's behavior signal and the trigger frame of the motivation word, screening the behavior segments linked with the signal change path and the semantic driving logic, and generating a language action response linkage sequence;

[0010] S4: extracting action record frames in the language action response linkage sequence, calculating the matching degree between the student and the teacher's actions, screening action frames with matching degrees exceeding a threshold, and generating imitation expression behavior results;

[0011] S5: Extract the student's language frames and action sequences in the imitation expression behavior results, analyze the integrity of the phoneme chain alignment, evaluate the time synchronization between the speech output starting point and the visual action response, and output the individual English interactive expression matching results.

[0012] As a further solution of the present invention, the speech action corresponding offset identifier is specifically the phoneme pronunciation duration, mouth opening and closing duration, offset frame segment starting point and end point, and phoneme and mouth action alignment status; the teaching pronunciation structure mapping set includes a set of covered sentences, a set of target phonemes, a phoneme distribution frequency ratio, and a task preset participation ratio; the language action response linkage sequence includes the student behavior signal active interval, the motivation word trigger frame segment timing consistency, the signal change path, and the semantically driven behavior segment; the imitation expression behavior result includes the student and teacher action matching degree, the action matching degree threshold, the action frame segment start and end time, and the matching action frame segment set; the individual English interactive expression matching result includes the phoneme chain integrity, the synchronization of the speech output starting point and the visual action response, the delay frame and the contour offset value, and the matching range.

[0013] As a further solution of the present invention, the steps for obtaining the offset identifier corresponding to the voice action are as follows:

[0014] S101: Obtain the pronunciation duration of the phoneme in the student's speech segment and the duration of the mouth opening and closing in the dynamic image, calculate the absolute value of the time difference between the phoneme start frame and the mouth movement start point, and generate a phoneme mouth timing benchmark set;

[0015] S102: Based on the phoneme mouth timing benchmark set, use the formula:

[0016]

[0017] Calculate the dynamic offset coefficient D c , the dynamic offset coefficient D c and the preset offset determination threshold θ d Compare and filter D c >θ d Frame segments, generate offset frame segment sets;

[0018] Where ΔS is the time difference between the phoneme and the mouth start, T ph is the pronunciation duration of the phoneme, T mo is the duration of mouth opening and closing, S m is the temporal stability of mouth movements, I ph is the logarithmic ratio of phoneme pronunciation intensity, θ d is the offset determination threshold;

[0019] S103: calling the offset frame segment set, merging consecutive offset frames according to the time window, extracting the interval exceeding the preset merging time threshold, and associating the phoneme label and the image number to establish the voice action corresponding offset identifier.

[0020] As a further solution of the present invention, the steps for obtaining the teaching pronunciation structure mapping set are:

[0021] S201: calling the offset frame segment in the offset identifier corresponding to the voice action, associating the teaching target label and the corresponding phoneme bound to the sentence text covered by the offset frame segment, counting the proportion of the number of occurrences of each phoneme in the offset frame segment, and generating a phoneme distribution frequency set;

[0022] S202: Based on the phoneme distribution frequency set, use the formula:

[0023]

[0024] Calculate the pronunciation deviation index D ij , screening pronunciation deviation index D ij Exceeding the preset threshold R j statements, generating a set of high deviation statements;

[0025] Among them F ij is the frequency ratio of phoneme i in sentence j, ΔT rel is the relative time offset, Sscore I is the tongue position stability score. ij is the energy logarithm ratio;

[0026] S203: calling the high deviation sentence set, clustering sentences according to the teaching target label, calculating the mean deviation index in each cluster group, screening cluster groups whose mean exceeds the teaching stage adaptation threshold, and generating a teaching pronunciation structure mapping set.

[0027] As a further solution of the present invention, the steps for obtaining the language action response linkage sequence are:

[0028] S301: calling the sentence segments marked in the teaching pronunciation structure mapping set, synchronously collecting the eye gaze direction change value, head rotation amplitude and mouth opening and closing frequency of the student during pronunciation, aligning the speech frame segments and the behavior signal sequence according to the time axis, and generating a standardized behavior signal parameter set;

[0029] S302: Based on the standardized behavior signal parameter set, use the formula:

[0030]

[0031] Calculate and obtain dynamic linkage score C mno , traverse the calculation results of all frame segments, filter the frame segments whose dynamic linkage scores are greater than the linkage threshold, and generate a dynamic linkage score set;

[0032] Among them, E m ′ represents the normalized value of eye gaze direction change, H n ′ represents the normalized value of the head rotation amplitude, M o ′ represents the normalized value of mouth opening and closing frequency, w E ,w H ,w M Respectively represent the contribution of eyes, head, and mouth to the linkage score, W dyn represents the dynamic regulatory factor, S sem ′ represents the semantic trigger density, σ stb ′ represents the path stability coefficient;

[0033] S303: Calling the dynamic linkage scoring set, classifying the frames by semantic labels, calculating the mean score and standard deviation within each frame category, screening the frame groups whose mean exceeds a preset mean threshold and whose standard deviation is lower than a standard deviation threshold, merging adjacent frames by timestamp, and generating a language action response linkage sequence.

[0034] As a further solution of the present invention, the steps for obtaining the result of the imitation expression behavior are:

[0035] S401: calling the language action response linkage sequence, extracting the timestamp, direction vector, outline coordinates and start and end frame delay of the action record frame segment, aligning the teacher action benchmark data according to the frame segment number, and generating an action record alignment set;

[0036] S402: Based on the action record alignment set, use the formula:

[0037]

[0038] Calculate the matching degree between the student and teacher's actions, traverse all frames, filter out frames with matching degrees exceeding the set threshold, and generate an action matching degree set;

[0039] Among them, D s is the student direction vector, D t is the teacher reference direction vector, IoU(P s ,P t ) is the intersection-over-union ratio of the student and teacher contour coordinates, ΔL is the start-end frame delay difference, and w D 、w C are direction matching weight and contour matching weight respectively;

[0040] S403: calling the action matching degree set, screening the frame segments whose matching degree exceeds the set matching degree threshold, merging the continuous frame segments according to the timestamp, and generating the imitation expression behavior result.

[0041] As a further solution of the present invention, the steps for obtaining the matching results of individual English interactive expressions are as follows:

[0042] S501: extracting the phoneme timestamp sequence and action response frame number of the language behavior frame segment in the imitation expression behavior result, calculating the difference between adjacent phoneme timestamps, marking the frame segments with the difference lower than the difference threshold, and generating the phoneme chain alignment;

[0043] S502: Based on the phoneme chain alignment, the formula is used:

[0044]

[0045] Calculate the voice-action coordination matching degree, traverse all frame segments, and generate dynamic offset matching coefficients;

[0046] Where ΔF d is the delay frame difference, F max is the maximum allowable delay frame difference, ΔA c is the contour offset area difference, A ref is the reference contour area, ΔS t is the time synchronization difference, T max is the maximum tolerable time difference, A p is the phoneme chain alignment;

[0047] S503: Calling the frame segments whose collaborative matching degree exceeds the threshold in the dynamic offset matching coefficient, merging the frame segment groups whose timestamps are continuous and whose synchronization difference is lower than the synchronization difference threshold, and generating the individual English interactive expression matching result.

[0048] An interactive English teaching system based on AI vision, comprising:

[0049] The speech action acquisition module acquires the speech signals and mouth image frames during students' oral interaction, calculates the pronunciation of phonemes and the duration of mouth opening and closing, analyzes the alignment offset between phoneme pronunciation and mouth opening and closing, extracts the offset frame segments, and generates offset identifiers corresponding to speech actions;

[0050] The speech action mapping module calculates the distribution ratio of phonemes according to the sentences covered by the offset identifier corresponding to the speech action, selects the sentence set whose distribution frequency ratio exceeds the preset participation ratio of the task, and generates a teaching pronunciation structure mapping set;

[0051] The behavior signal analysis module calls the sentence segments in the teaching pronunciation structure mapping set, analyzes the consistency between the active interval of the student's behavior signal and the timing of the trigger frame of the motivation word, selects the behavior segment in which the signal change path is linked with the semantic driving logic, and generates a language action response linkage sequence;

[0052] The action matching analysis module extracts the action record frames in the language action response linkage sequence, calculates the matching degree between the student and the teacher's actions, selects the action frames whose matching degree exceeds a threshold, and generates the imitation expression behavior results;

[0053] The interactive expression evaluation module extracts the student's language frames and action sequences from the imitation expression behavior results, analyzes the alignment integrity of the phoneme chain, evaluates the time synchronization between the starting point of the speech output and the visual action response, and outputs the individual English interactive expression matching results.

[0054] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:

[0055] In the present invention, by real-time analysis of the student's voice and mouth dynamic images, the alignment offset between pronunciation and mouth movements is accurately identified, and a dynamic feedback mark is generated. Combined with behavioral signals such as student eye gaze and head rotation, the interactivity of language movements and learner behavior can be comprehensively evaluated to ensure efficient alignment of language expression with standard templates. By analyzing the distribution frequency and matching degree of voice behavior, the sentence set with the most training value is screened out to optimize the personalized teaching path. In addition, in-depth analysis of delayed frames and contour offsets can enhance the synchronization of voice and visual movements, improve real-time feedback and interactive effects during the learning process, and thus greatly improve learning efficiency and feedback accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 is a flow chart of the method of the present invention;

[0057] Figure 2 This is a flow chart for obtaining the offset identifier corresponding to the voice action of the present invention;

[0058] Figure 3 A flowchart for obtaining a teaching pronunciation structure mapping set according to the present invention;

[0059] Figure 4 This is a flow chart for obtaining a language action response linkage sequence according to the present invention;

[0060] Figure 5 This is a flow chart for obtaining the results of the imitation expression behavior of the present invention;

[0061] Figure 6 This is a flow chart for obtaining matching results of individual English interactive expressions in the present invention. DETAILED DESCRIPTION

[0062] The technical solution of the present invention is described below in conjunction with the accompanying drawings.

[0063] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as an "exemplary" in the present invention should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of the word "exemplary" is intended to present concepts in a concrete manner. Furthermore, in the embodiments of the present invention, "and / or" can mean both or either of the two.

[0064] In the embodiments of the present invention, the terms "image" and "picture" may be used interchangeably. It should be noted that, when the distinction between them is not emphasized, their intended meanings are the same. The terms "of," "corresponding," and "corresponding" may be used interchangeably. It should be noted that, when the distinction between them is not emphasized, their intended meanings are the same.

[0065] In the embodiments of the present invention, sometimes a subscript such as W1 may be written as a non-subscript such as W1. When the difference is not emphasized, the meanings to be expressed are the same.

[0066] In order to make the technical problems, technical solutions and advantages to be solved by the present invention clearer, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.

[0067] See also Figure 1 The present invention provides a technical solution: an interactive English teaching method based on AI vision, comprising the following steps:

[0068] S1: Obtain student speech segments and mouth dynamic image frames in the oral interactive teaching session, calculate the phoneme pronunciation duration and mouth opening and closing duration, call the difference between the mouth movement starting point and the phoneme starting frame, determine whether the alignment between phoneme pronunciation and mouth opening and closing is offset, extract multiple offset frame segments to form a semantic action disjoint interval, and generate a speech action corresponding offset identifier;

[0069] S2: Extract the teaching targets and corresponding phoneme sets bound to the covered sentences in the task paragraph based on the offset frame segments in the offset identifiers corresponding to the speech actions, calculate the distribution frequency ratio of the target phonemes in the offset frame segments, select the sentence sets whose distribution frequency ratio exceeds the preset participation ratio of the task, and generate the teaching pronunciation structure mapping set;

[0070] S3: Call the sentence segments indicated in the teaching pronunciation structure mapping set, combine the direction of eye gaze changes, head rotation amplitude and mouth opening and closing frequency when the students pronounce, analyze the temporal consistency and response trend between the active interval of behavioral signals and the triggering frame segments of the motivation words, select the behavioral segments with the linkage between signal change paths and semantic driving logic, and generate a language action response linkage sequence;

[0071] S4: Extract the action record frames in the language action response linkage sequence, calculate the matching degree between the student action and the teacher action in terms of direction, contour overlap, and start and end frame delay, select the action frames whose action matching degree exceeds the set matching degree threshold, and generate the imitation expression behavior results;

[0072] S5: Extract the student's language behavior frame segments and action response sequences from the imitation expression behavior results, determine whether the phoneme chain is completely aligned based on language continuity, analyze whether the delay frames and contour offset values ​​in the action sequence exceed the matching range set by the current task, evaluate the time synchronization between the starting point of the speech output and the visual action response, and output the individual English interactive expression matching results.

[0073] The offset identifiers corresponding to speech actions are specifically the phoneme pronunciation duration, mouth opening and closing duration, offset frame start and end points, and the alignment status of phonemes and mouth actions; the teaching pronunciation structure mapping set includes the coverage sentence set, target phoneme set, phoneme distribution frequency ratio, and task preset participation ratio; the language action response linkage sequence includes the student behavior signal active interval, the motivation word trigger frame segment timing consistency, the signal change path, and the semantic-driven behavior segment; the imitation expression behavior results include the student and teacher action matching degree, the action matching degree threshold, the action frame segment start and end time, and the matching action frame segment set; the individual English interactive expression matching results include the phoneme chain integrity, the synchronization between the speech output starting point and the visual action response, the delay frame and contour offset value, and the matching range.

[0074] See also Figure 2, the steps to obtain the offset identifier corresponding to the voice action are:

[0075] S101: Obtain the pronunciation duration of the phoneme in the student's speech segment and the duration of the mouth opening and closing in the dynamic image, calculate the absolute value of the time difference between the phoneme start frame and the mouth movement start point, and generate a phoneme mouth timing benchmark set;

[0076] The speech segment is segmented into 25ms / frames by the audio framing tool, and the energy peak of each frame is extracted to locate the phoneme starting point (e.g. phoneme / / The starting frame is the 120th frame, timestamp t ph = 4.8s), and simultaneously use image processing tools to detect the opening and closing degree of the mouth area (the sudden increase in the mouth height from 10% to 30% is determined as the starting action, and the timestamp is t mo =4.85s), the absolute value of the calculated time difference is |ΔS| = |t ph -t mo |=0.05s, phoneme pronunciation duration T ph = 0.35s (frames 120-134), mouth opening and closing duration T mo = 0.40s (frames 121-137), the parameters are classified and stored according to the phoneme labels to form a structured data set.

[0077] S102: Based on the phoneme mouth timing benchmark set, use the formula:

[0078]

[0079] Calculate the dynamic offset coefficient D c , the dynamic offset coefficient D c and the preset offset determination threshold θ d Compare and filter D c >θ d Frame segments, generate offset frame segment sets;

[0080] Where ΔS is the time difference between the phoneme and the mouth start, T ph is the pronunciation duration of the phoneme, T mo is the duration of mouth opening and closing, S m is the temporal stability of mouth movements, I ph is the logarithmic ratio of phoneme pronunciation intensity, θ d is the offset determination threshold;

[0081] Phoneme / / as an example, ΔS=0.05s, T ph =0.35s, T mo =0.40s, mouth movement time stability S m =50s -1(Standard deviation of time interval between adjacent frames σ=0.02s, S m =1 / σ), the logarithmic ratio of phoneme intensity I ph =0.8 (reference energy E0 = 0.1Pa, measured peak value E = 0.3Pa, calculated I ph =10log 10 (E / E0) = 4.77dB, normalized to [0,1]), substitute into the formula:

[0082]

[0083] This result shows that D c =0.0091 is less than the preset threshold θ d =0.01, indicating that the phoneme pronunciation and mouth movement alignment of this frame segment have not significantly deviated, so it is not included in the offset frame segment set. If a frame segment parameter is ΔS = 0.12s, T ph =0.30s, T mo =0.28s, S m =40s -1 , I ph =0.6, then:

[0084]

[0085] This result shows that D c =0.032 exceeds the threshold θ d , indicating that there is an anomaly in the timing alignment of phonemes and mouth movements. This frame segment needs to be added to the offset frame segment set to provide a data basis for the subsequent merging of disjointed intervals.

[0086] S103: Calling the offset frame segment set, merging consecutive offset frames according to the time window, extracting the interval exceeding the preset merging time threshold, and associating the phoneme label and image number to establish the voice action corresponding offset identifier;

[0087] Assume that the merging time threshold τ w = 0.3s, for example, if the timestamps of consecutive offset frames are [5.0, 5.1], [5.1, 5.25], and [5.3, 5.5], the merged interval is [5.0, 5.25] (duration 0.25s < τ w , removed) and [5.3,5.5] (duration 0.2s<τ w , removed), if the consecutive offset frames are [6.0,6.15], [6.15,6.4], the merged interval is [6.0,6.4] (duration 0.4s>τ w ), associate the phoneme label / ∫ / with the image frame number (frame 240-256), generate a disjoint interval record table and offset identification code (such as ID_Offset_001), and only retain those with a duration exceeding τ through merging and filtering.w The interval of semantic action disjunction is used to avoid short-term error interference, and the final generated semantic action disjunction interval is directly used to construct the corresponding offset mark of speech action, providing precise positioning for teaching correction.

[0088] See also Figure 3 , the steps for obtaining the teaching pronunciation structure mapping set are:

[0089] S201: calling the offset frame segment in the offset identifier corresponding to the voice action, associating the offset frame segment with the teaching target label and the corresponding phoneme bound to the sentence text, counting the occurrence ratio of each phoneme in the offset frame segment, and generating a phoneme distribution frequency set;

[0090] Use the audio framing tool to split the speech segment at 25ms / frame, extract the offset frame segment (such as the timestamp [5.0s, 5.4s]), call the speech recognition interface to obtain the sentence text corresponding to this period "Please pass the red pen", associate the label "bilabial clarity training" in the teaching database with the bound phonemes / p / and / b / , count the number of times the phoneme / p / appears in the offset frame segment 4 times, and the total number of phonemes in the sentence is 20, calculate the frequency ratio F ij =4 / 20=0.2, phoneme / b / appears 2 times, F ij = 0.1, associate phoneme labels with frequencies and store them as a structured data table. For example, the offset frame segment [6.2s, 6.6s] corresponds to the sentence “Shesells seashells”, bound to the phoneme / ∫ / , which appears 6 times, with a total of 18 phonemes. ij =6 / 18≈0.33, the above data are classified into phoneme distribution frequency sets according to phoneme classification.

[0091] S202: Based on the phoneme distribution frequency set, use the formula:

[0092]

[0093] Calculate the pronunciation deviation index D ij , screening pronunciation deviation index D ij Exceeding the preset threshold R j statements, generating a set of high deviation statements;

[0094] Among them F ij is the frequency ratio of phoneme i in sentence j, ΔT rel is the relative time offset, S score I is the tongue position stability score. ij is the energy logarithm ratio;

[0095] Take the phoneme / ∫ / as an example to get F ij =0.33, calculate the time offset parameter ΔT ij=|actual pronunciation start time - theoretical start time|, collect the pronunciation start time of the phoneme three times in the sentence [6.21s, 6.35s, 6.40s], the theoretical start time [6.20s, 6.30s, 6.40s], and calculate ΔT ij =[0.01s, 0.05s, 0.00s], mean ΔT ij = 0.02s, total sentence duration T j =2.5s, relative time offset ΔT rel = 0.02 / 2.5 = 0.008, the standard deviation of tongue position coordinates when pronouncing / ∫ / is obtained by tongue tracking equipment ij =1.8mm(preset σ max =3.0mm), calculate the stability score S score =1-1.8 / 3.0=0.4, collect phoneme energy peak E ij =0.25Pa (base E0 = 0.1Pa), calculate the energy logarithmic ratio I ij =log(0.25 / 0.1)≈0.397, substitute into the formula:

[0096]

[0097] Preset threshold R j =0.005 (based on the D of 95% of the normal pronunciation sentences in the historical data ij ≤0.005), because 0.0066>0.005, the sentence is judged as a high deviation sentence, and the sentence ID, phoneme label and D ij Store the high deviation statement set;

[0098] This result shows that D ij =0.0066 exceeds the threshold R j , indicating that the phoneme / ∫ / has timing deviation and insufficient stability in this sentence, and needs to be included in the high deviation sentence set.

[0099] S203: calling a high-deviation sentence set, clustering sentences according to the teaching target label, calculating the mean deviation index within each cluster group, screening cluster groups whose mean exceeds the teaching stage adaptation threshold, and generating a teaching pronunciation structure mapping set;

[0100] Taking the teaching goal “fricative coherence training” as an example, the cluster contains sentence ID_001 (D ij =0.0066), ID_002(D ij =0.0058), ID_003(D ij =0.0049), calculate the mean , set the threshold θ according to the teaching stage j (Primary stage θ j=0.005, advanced stage θ j =0.003), if the current stage is the primary stage, because 0.0058>0.005, it is determined that the cluster group needs to be trained intensively, all sentences and phoneme labels in the group are extracted, and a mapping relationship table is generated. For example, the mean of the cluster group in the advanced stage ,θ j =0.003. Since 0.0027 < 0.003, it is determined that no additional training is required and only the mapping set is recorded. Finally, the teaching pronunciation structure mapping set is output, which includes the target phonemes, related sentences and training priority tags. The results show that: Exceeding the primary stage threshold θ j , indicating that the overall deviation of the teaching target group is relatively high, and it needs to be marked as high-priority training content, and a teaching pronunciation structure mapping set should be directly generated to provide precise positioning for pronunciation correction.

[0101] See also Figure 4 ,The steps for obtaining the language action response linkage sequence are:

[0102] S301: Calling the sentence segments marked in the teaching pronunciation structure mapping set, synchronously collecting the eye gaze direction change value, head rotation amplitude and mouth opening and closing frequency of the student during pronunciation, aligning the speech frame segments with the behavior signal sequence according to the time axis, and generating a standardized behavior signal parameter set;

[0103] For example, for the sentence "Please pass the red pen", the original data of the eye gaze direction change value is 0.5°, -0.3°, and 0.7°, the head rotation amplitude is 3.2mm and 4.1mm, and the mouth opening and closing frequency is 11 times / s and 13 times / s. Based on the historical data mean (eye mean 0.2°, head mean 3.5mm, mouth mean 12 times / s) and standard deviation (eye standard deviation 0.4°, head standard deviation 0.6mm, mouth standard deviation 1.2 times / s), the standardized value is calculated:

[0104] Eyes: E m '=(0.5-0.2) / 0.4=0.75, (-0.3-0.2) / 0.4=-1.25, (0.7-0.2) / 0.4=1.25;

[0105] Head: H n ′=3.2 / 3.5=0.914, 4.1 / 3.5=1.171;

[0106] Mouth: M o ′=11 / 12=0.917, 13 / 12=1.083;

[0107] The standardized value is aligned with the speech frame segment timestamp (eg, frame 120 corresponds to 5.0s) and stored as a structured data table with fields including frame number, timestamp, standardized eye value, standardized head value, and standardized mouth value.

[0108] S302: Based on the standardized behavior signal parameter set, the formula is used:

[0109]

[0110] Calculate and obtain dynamic linkage score C mno , traverse the calculation results of all frame segments, filter the frame segments whose dynamic linkage scores are greater than the linkage threshold, and generate a dynamic linkage score set;

[0111] Among them, E m ′ represents the normalized value of eye gaze direction change, H n ′ represents the normalized value of the head rotation amplitude, M o ′ represents the normalized value of mouth opening and closing frequency, w E ,w H ,w M Respectively represent the contribution of eyes, head, and mouth to the linkage score, W dyn represents the dynamic regulatory factor, S sem ′ represents the semantic trigger density, σ stb ′ represents the path stability coefficient;

[0112] Collect 100 sets of behavioral parameters (eyes, head, mouth) from historical data and calculate their variance contributions respectively. Assuming that the variance of eye parameters is Total variance Contribution Head parameter variance Contribution 0.2 / 0.4=0.5, mouth parameter variance Contribution 0.08 / 0.4=0.2, normalized contribution distribution weight w E =0.3, w H =0.5, w M =0.2, calculated based on the kurtosis of the behavioral signal W dyn =K B / 2, extract the eye normalization value sequence (e.g. E for frames 120-135) m ′=[0.75,-1.25,1.25]), calculate the mean μ=(0.75-1.25+1.25) / 3=0.25, calculate the fourth-order central moment μ4=[(0.75-0.25) 4 +(-1.25-0.25) 4 +(1.25-0.25) 4 ] / 3≈6.066, second-order central moment μ2=[(0.75-0.25)2 +(-1.25-0.25) 2 +(1.25-0.25) 2 ] / 3≈3.3125, kurtosis , then W dyn =0.553 / 2≈0.276, the number of occurrences of the motivation word N c =2, total sentence duration T s =2.5s, then S sem ′=N c / T s =2 / 2.5=0.8, signal path time interval T p =[0.2s, 0.3s, 0.25s], mean time interval μ Tp =(0.2+0.3+0.25) / 3≈0.25s, standard deviation Std(T p )=0.0707s,σ stb '=1 / 0.0707≈14.14, taking frame 120 as an example, the parameter value is: E m ′=0.75、H n ′=0.914、M o ′=0.917,w E =0.3, w H =0.5, w M =0.2,W dyn =0.276, S sem ′=0.8、σ stb ′=14.14, substitute into the formula to calculate:

[0113] Molecular part:

[0114] 0.3×0.75+0.5×0.914+0.2×|0.917-1|≈0.6986;

[0115] Denominator:

[0116]

[0117] Dynamic linkage rating:

[0118] C mno =0.6986 / 7.844≈0.089;

[0119] The result shows that the dynamic linkage score of 0.089 is less than the threshold of 0.1 (the threshold is set based on the top 20% quantiles of the linkage scores in historical data being 0.095-0.105, with 0.1 being the screening line), indicating that the linkage strength between the eye gaze, head rotation, and mouth opening and closing behaviors and the semantic logic in the current frame segment (frame 120) does not meet the standard, and therefore is not included in the dynamic linkage score set.

[0120] S303: Calling the dynamic linkage scoring set, classifying the frames by semantic labels, calculating the mean and standard deviation of the scores within each frame category, selecting the frame groups whose mean exceeds a preset mean threshold and whose standard deviation is lower than a standard deviation threshold, merging adjacent frames by timestamp, and generating a language-action response linkage sequence;

[0121] For example, the scores of the semantic label "request action trigger" associated frame segment group are 0.15, 0.22, and 0.18, with an average , standard deviation σ=0.029, linkage threshold 0.15 is based on the top 20% quantile of historical data (the average score of the top 20 in 100 groups of data is 0.14-0.16), and stability threshold 0.03 is set according to signal fluctuation tolerance. If it is greater than 0.15 and σ = 0.029 and less than 0.03, the timestamps 5.2s-5.5s and 5.6s-5.8s are merged into a continuous interval 5.2s-5.8s, which is marked as linkage sequence ID_001;

[0122] The results show that the linkage strength between the behavioral signals and semantic logic of this frame segment group is stable and significant, which meets the teaching response requirements. Therefore, it is merged into a continuous time interval as an effective component of the final language action response linkage sequence.

[0123] See also Figure 5 , the steps to obtain the results of imitating expression behavior are:

[0124] S401: Calling the language action response linkage sequence, extracting the timestamp, direction vector, contour coordinates and start and end frame delay of the action record frame segment, aligning the teacher action benchmark data according to the frame segment number, and generating an action record alignment set;

[0125] When extracting the timestamp, the start time 5.2 seconds and the end time 5.5 seconds of the frame segment ID_001 are obtained, and the student direction vector 0.32, -0.15, the contour coordinate point set (12, 45), (15, 48), and the start and end frame delay 3 frames are extracted. The teacher reference direction vector 0.30, -0.12, the contour coordinate point set (11, 44), (14, 47), and the start and end frame delay 1 frame are extracted. The direction vector difference is calculated to be 0.02, -0.03, the contour coordinate intersection-union ratio is the overlapping area 25 pixels / the union area 30 pixels = 0.83, and the start and end frame delay difference is 3-1 = 2. The fields stored as a structured data table include frame number, timestamp, direction vector difference, intersection-union ratio, and delay difference to generate an action record alignment set.

[0126] S402: Based on the action record alignment set, use the formula:

[0127]

[0128] Calculate the matching degree between the student and teacher's actions, traverse all frames, filter out frames with matching degrees exceeding the set threshold, and generate an action matching degree set;

[0129] Among them, D s is the student direction vector, D t is the teacher reference direction vector, IoU(P s ,P t ) is the intersection-over-union ratio of the student and teacher contour coordinates, ΔL is the start-end frame delay difference, and w D 、w C are direction matching weight and contour matching weight respectively;

[0130] 100 groups of valid action frames were extracted from the teaching data, and the influence of direction difference, contour overlap, and delay difference on teaching score were statistically analyzed. It is assumed that the variance contribution of direction difference V D =0.18, V contribution of profile coincidence variance C =0.12, total variance V 总 =0.3, then w D =V D / V 总 =0.6, w C =V C / V 总 =0.4;

[0131] Parameter calculation and substitution: If the student direction vector D s =[0.32,-0.15], teacher benchmark D t =[0.30,-0.12], difference vector ΔD = [0.02,-0.03], modulus Reciprocal ||ΔD||-1 =1 / 0.036≈27.78;

[0132] Assume that the area of ​​the student's outline coordinates is A s =30 pixels 2 , the area of ​​the teacher's contour coordinate region A t =28 pixels 2 , the overlapping area A 重叠 =25 pixels 2 , then the intersection over union ratio IoU = A 重叠 / (A s +A t -A 重叠 )=25 / (30+28-25)=25 / 33≈0.75; Student start and end frame delay L s = 3 frames, teacher benchmark L t = 1 frame, then ΔL = |3-1| = 2 frames;

[0133] Enter the formula to calculate:

[0134] Molecular part: w D ||ΔD|| -1 +w C ·IoU=0.6×27.78+0.4×0.758=16.67+0.303=16.973;

[0135] Denominator:

[0136] Matching calculation:

[0137] Normalization: If the maximum matching degree in historical data is 10, normalize M pqr =7.03 / 10=0.703. If the matching threshold is 0.75, the result shows that the matching degree 0.703 is lower than the threshold 0.75. It is determined that the current frame segment (ID_001) does not meet the imitation expression requirements and is not included in the action matching set.

[0138] S403: calling the action matching set, screening the frame segments whose matching degree exceeds the set matching degree threshold, merging the continuous frame segments according to the timestamp, and generating the imitation expression behavior result;

[0139] Assume that the matching degree of frame segment ID_002 is 0.82, corresponding to the timestamp of 5.6 seconds to 5.8 seconds, and the matching degree of frame segment ID_003 is 0.79, corresponding to the timestamp of 5.8 seconds to 6.0 seconds. Merge the adjacent timestamps of 5.6 seconds to 6.0 seconds into a continuous interval, store it as a structured record field including the start time, end time, and the matching degree average of 0.805, and generate the imitation expression behavior result.

[0140] See also Figure 6 , the steps to obtain the matching results of individual English interactive expressions are:

[0141] S501: extracting the phoneme timestamp sequence and action response frame number of the language behavior frame segment in the imitation expression behavior result, calculating the difference between adjacent phoneme timestamps, marking the frame segments with the difference lower than the difference threshold, and generating the phoneme chain alignment;

[0142] Extract the phoneme timestamp sequence [5.2s, 5.35s, 5.53s, 5.75s] of frame segment ID_001, calculate the adjacent timestamp differences as 0.15s, 0.18s, and 0.22s, set the phoneme continuity threshold to 0.2 seconds, compare 0.15s ≤ 0.2s, 0.18s ≤ 0.2s, and 0.22s > 0.2s, mark the continuous frames 5.2s-5.53s corresponding to the first two differences (0.15s and 0.18s), calculate the continuous frame ratio as 2 / 3 ≈ 0.67, store it in the field [frame segment ID_001, alignment 0.67], and generate the phoneme chain alignment.

[0143] S502: Based on the phoneme chain alignment, the formula is:

[0144]

[0145] Calculate the voice-action coordination matching degree, traverse all frame segments, and generate dynamic offset matching coefficients;

[0146] Where ΔF d is the delay frame difference, F max is the maximum allowable delay frame difference, ΔA c is the contour offset area difference, A ref is the reference contour area, ΔS t is the time synchronization difference, T max is the maximum tolerable time difference, A p is the phoneme chain alignment;

[0147] Assume that the student action frame number is 3 (corresponding to the timestamp 5.3 seconds), the speech start frame number is 1 (corresponding to the timestamp 5.2 seconds), F max =5 frames, ΔF d =|3-1|=2 frames, the student contour coordinate point set is [(12,45),(15,48)], A s = width × height = (15-12) × (48-45) = 3 × 3 = 9 pixels 2 , the teacher's benchmark contour coordinate point set is [(11,44),(14,47)], A t =(14-11)×(47-44)=3×3=9 pixels 2, calculate ΔA c =|A s-A t| =|9-9|=0 pixels 2 , A ref =9 pixels 2 ;

[0148] Assume that the speech start timestamp is 5.2 seconds and the action response timestamp is 5.3 seconds, ΔS t =|5.3-5.2|=0.1 seconds, T max = 0.2 seconds. If the phoneme timestamp sequence of frame segment ID_001 is [5.2s, 5.35s, 5.53s, 5.75s], and the adjacent interval difference is [0.15s, 0.18s, 0.22s], and the phoneme continuity threshold is 0.2 seconds, the number of frames with continuous difference ≤ 0.2 seconds is 2 (0.15s, 0.18s), and the total number of frames is 3, A p =2 / 3≈0.67.

[0149] Substitute the formula and calculate step by step:

[0150] Molecular part:

[0151] Denominator:

[0152] Voice-action synergy matching results:

[0153] Normalization processing:

[0154] Theoretical maximum value (numerator extreme value): Assuming ΔF d =5 frames, ΔA c =0, then the numerator = 5 / 5 + 0 = 1, the denominator = 0.1 / 0.2 + 0 = 0.5, the theoretical maximum = 1 / 0.5 = 2, then the normalized score: 0.482 / 2 ≈ 0.241 (0-1 interval);

[0155] If the matching threshold is set to 0.4, the current score 0.241<0.4, and the frame segment ID_001 is judged to be substandard and marked as a low match in the dynamic offset matching coefficient. The result shows that the speech-action coordination matching degree of the frame segment ID_001 is 0.241 (after normalization), which is significantly lower than the threshold of 0.4, indicating that the student has problems with action delay (2 frames) and insufficient phoneme continuity (alignment 0.67) in this frame segment, resulting in the synchronization of speech and action not meeting the standard.

[0156] S503: calling the frame segments whose collaborative matching degree exceeds the threshold in the dynamic offset matching coefficient, merging the frame segment groups whose timestamps are continuous and whose synchronization difference is lower than the synchronization difference threshold, and generating the individual English interactive expression matching result;

[0157] Traversing the matching coefficient set, frame segment ID_002 scores 0.82 (timestamp 5.6s-5.8s) and ID_003 scores 0.79 (timestamp 5.8s-6.0s). Calculate the timestamp continuity (5.8s-5.8s=0 seconds≤0.15 seconds) and merge them into the interval 5.6s-6.0s. Calculate the mean synchronization difference as |5.6-5.65|+|5.8-5.85|+|6.0-6.05|=0.05+0.05+0.05=0.15 seconds≤0.15 seconds. Store this in the field [start 5.6s, end 6.0s, match 0.805] to generate the individual English interactive expression matching results.

[0158] An interactive English teaching system based on AI vision, which includes:

[0159] The speech action acquisition module acquires the speech signals and mouth image frames during students' oral interaction, calculates the pronunciation of phonemes and the duration of mouth opening and closing, analyzes the alignment offset between phoneme pronunciation and mouth opening and closing, extracts the offset frame segments, and generates offset identifiers corresponding to speech actions;

[0160] The speech action mapping module calculates the distribution ratio of phonemes based on the sentences covered by the speech action corresponding offset identifier, selects the sentence set whose distribution frequency ratio exceeds the preset participation ratio of the task, and generates a teaching pronunciation structure mapping set;

[0161] The behavioral signal analysis module calls the sentence segments in the teaching pronunciation structure mapping set, analyzes the consistency between the active interval of the student's behavioral signal and the timing of the trigger frame of the motivation word, selects the behavioral segments with the linkage between the signal change path and the semantic driving logic, and generates a language action response linkage sequence;

[0162] The action matching analysis module extracts action record frames from the language action response linkage sequence, calculates the matching degree between the student and the teacher's actions, selects action frames with matching degrees exceeding a threshold, and generates imitation expression behavior results;

[0163] The interactive expression evaluation module extracts the student's language frames and action sequences from the results of the imitation expression behavior, analyzes the alignment integrity of the phoneme chain, evaluates the time synchronization between the starting point of the speech output and the visual action response, and outputs the individual English interactive expression matching results.

[0164] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. The interactive English teaching method based on AI vision is characterized by: The following steps are involved: S1: Obtain student speech and mouth image frames during oral interactive teaching, calculate phoneme pronunciation and mouth opening and closing duration, determine the alignment offset between phoneme pronunciation and mouth opening and closing, extract offset frame segments, and generate speech action corresponding offset identifiers; S2: Calculate the phoneme distribution ratio based on the sentences covered by the offset frame segments in the offset identifier corresponding to the speech action, select a sentence set whose distribution frequency ratio exceeds the preset task participation ratio, and generate a teaching pronunciation structure mapping set; S3: calling the sentence segments in the teaching pronunciation structure mapping set, analyzing the consistency between the active interval of the student's behavior signal and the trigger frame of the motivation word, screening the behavior segments linked with the signal change path and the semantic driving logic, and generating a language action response linkage sequence; S4: extracting action record frames in the language action response linkage sequence, calculating the matching degree between the student and the teacher's actions, screening action frames with matching degrees exceeding a threshold, and generating imitation expression behavior results; S5: Extract the student's language frames and action sequences in the imitation expression behavior results, analyze the integrity of the phoneme chain alignment, evaluate the time synchronization between the speech output starting point and the visual action response, and output the individual English interactive expression matching results.

2. The AI ​​vision-based interactive English teaching method according to claim 1, characterized in that: The speech action corresponding offset identifier is specifically the phoneme pronunciation duration, mouth opening and closing duration, offset frame segment start and end points, and phoneme and mouth action alignment status; the teaching pronunciation structure mapping set includes a set of covered sentences, a set of target phonemes, a phoneme distribution frequency ratio, and a task preset participation ratio; the language action response linkage sequence includes the student behavior signal active interval, the motivation word trigger frame segment timing consistency, the signal change path, and the semantically driven behavior segment; the imitation expression behavior result includes the student and teacher action matching degree, the action matching degree threshold, the action frame segment start and end time, and the matching action frame segment set; the individual English interactive expression matching result includes the phoneme chain integrity, the synchronization between the speech output starting point and the visual action response, the delay frame and the contour offset value, and the matching range.

3. The AI ​​vision-based interactive English teaching method according to claim 1, characterized in that: The steps to obtain the offset identifier corresponding to the voice action are as follows: S101: Obtain the pronunciation duration of the phoneme in the student's speech segment and the duration of the mouth opening and closing in the dynamic image, calculate the absolute value of the time difference between the phoneme start frame and the mouth movement start point, and generate a phoneme mouth timing benchmark set; S102: Based on the phoneme mouth timing benchmark set, use the formula: Calculate the dynamic offset coefficient D c , the dynamic offset coefficient D c and the preset offset determination threshold θ d Compare and filter D c >θ d Frame segments, generate offset frame segment sets; Where ΔS is the time difference between the phoneme and the mouth start, T ph is the pronunciation duration of the phoneme, T mo is the duration of mouth opening and closing, S m is the temporal stability of mouth movements, I ph is the logarithmic ratio of phoneme pronunciation intensity, θ d is the offset determination threshold; S103: calling the offset frame segment set, merging consecutive offset frames according to the time window, extracting the interval exceeding the preset merging time threshold, and associating the phoneme label and the image number to establish the voice action corresponding offset identifier.

4. The AI ​​vision-based interactive English teaching method according to claim 1, characterized in that: The steps for obtaining the teaching pronunciation structure mapping set are: S201: calling the offset frame segment in the offset identifier corresponding to the voice action, associating the teaching target label and the corresponding phoneme bound to the sentence text covered by the offset frame segment, counting the proportion of the number of occurrences of each phoneme in the offset frame segment, and generating a phoneme distribution frequency set; S202: Based on the phoneme distribution frequency set, use the formula: Calculate the pronunciation deviation index D ij , screening pronunciation deviation index D ij Exceeding the preset threshold R j statements, generating a set of high deviation statements; Among them F ij is the frequency ratio of phoneme i in sentence j, ΔT rel is the relative time offset, S score I is the tongue position stability score. ij is the energy logarithm ratio; S203: calling the high deviation sentence set, clustering sentences according to the teaching target label, calculating the mean deviation index in each cluster group, screening cluster groups whose mean exceeds the teaching stage adaptation threshold, and generating a teaching pronunciation structure mapping set.

5. The AI ​​vision-based interactive English teaching method according to claim 1, characterized in that: The steps to obtain the language action response linkage sequence are: S301: calling the sentence segments marked in the teaching pronunciation structure mapping set, synchronously collecting the eye gaze direction change value, head rotation amplitude and mouth opening and closing frequency of the student during pronunciation, aligning the speech frame segments and the behavior signal sequence according to the time axis, and generating a standardized behavior signal parameter set; S302: Based on the standardized behavior signal parameter set, use the formula: Calculate and obtain dynamic linkage score C mno , traverse the calculation results of all frame segments, filter the frame segments whose dynamic linkage scores are greater than the linkage threshold, and generate a dynamic linkage score set; Among them, E m ′ represents the normalized value of eye gaze direction change, H n ′ represents the normalized value of the head rotation amplitude, M o ′ represents the normalized value of mouth opening and closing frequency, w E ,w H ,w M Respectively represent the contribution of eyes, head, and mouth to the linkage score, W dyn represents the dynamic regulatory factor, S sem ′ represents the semantic trigger density, σ stb ′ represents the path stability coefficient; S303: Calling the dynamic linkage scoring set, classifying the frames by semantic labels, calculating the mean score and standard deviation within each frame category, screening the frame groups whose mean exceeds a preset mean threshold and whose standard deviation is lower than a standard deviation threshold, merging adjacent frames by timestamp, and generating a language action response linkage sequence.

6. The AI ​​vision-based interactive English teaching method according to claim 1, characterized in that: The steps to obtain the results of imitating expression behavior are: S401: calling the language action response linkage sequence, extracting the timestamp, direction vector, outline coordinates and start and end frame delay of the action record frame segment, aligning the teacher action benchmark data according to the frame segment number, and generating an action record alignment set; S402: Based on the action record alignment set, use the formula: Calculate the matching degree between the student and teacher's actions, traverse all frames, filter out frames with matching degrees exceeding the set threshold, and generate an action matching degree set; Among them, D s is the student direction vector, D t is the teacher reference direction vector, IoU(P s ,P t ) is the intersection-over-union ratio of the student and teacher contour coordinates, ΔL is the start-end frame delay difference, and w D 、w C are direction matching weight and contour matching weight respectively; S403: calling the action matching degree set, screening the frame segments whose matching degree exceeds the set matching degree threshold, merging the continuous frame segments according to the timestamp, and generating the imitation expression behavior result.

7. The AI ​​vision-based interactive English teaching method according to claim 1, characterized in that: The steps to obtain the matching results of individual English interactive expressions are as follows: S501: extracting the phoneme timestamp sequence and action response frame number of the language behavior frame segment in the imitation expression behavior result, calculating the difference between adjacent phoneme timestamps, marking the frame segments with the difference lower than the difference threshold, and generating the phoneme chain alignment; S502: Based on the phoneme chain alignment, the formula is used: Calculate the voice-action coordination matching degree, traverse all frame segments, and generate dynamic offset matching coefficients; Where ΔF d is the delay frame difference, F max is the maximum allowable delay frame difference, ΔA c is the contour offset area difference, A ref is the reference contour area, ΔS t is the time synchronization difference, T max is the maximum tolerable time difference, A p is the phoneme chain alignment; S503: Calling the frame segments whose collaborative matching degree exceeds the threshold in the dynamic offset matching coefficient, merging the frame segment groups whose timestamps are continuous and whose synchronization difference is lower than the synchronization difference threshold, and generating the individual English interactive expression matching result.

8. The interactive English teaching system based on AI vision is characterized by: The system is used for the AI ​​vision-based interactive English teaching method according to any one of claims 1 to 7, and the system comprises: The speech action acquisition module acquires the speech signals and mouth image frames during students' oral interaction, calculates the pronunciation of phonemes and the duration of mouth opening and closing, analyzes the alignment offset between phoneme pronunciation and mouth opening and closing, extracts the offset frame segments, and generates offset identifiers corresponding to speech actions; The speech action mapping module calculates the distribution ratio of phonemes according to the sentences covered by the offset identifier corresponding to the speech action, selects the sentence set whose distribution frequency ratio exceeds the preset participation ratio of the task, and generates a teaching pronunciation structure mapping set; The behavior signal analysis module calls the sentence segments in the teaching pronunciation structure mapping set, analyzes the consistency between the active interval of the student's behavior signal and the timing of the trigger frame of the motivation word, selects the behavior segment in which the signal change path is linked with the semantic driving logic, and generates a language action response linkage sequence; The action matching analysis module extracts the action record frames in the language action response linkage sequence, calculates the matching degree between the student and the teacher's actions, selects the action frames whose matching degree exceeds a threshold, and generates the imitation expression behavior results; The interactive expression evaluation module extracts the student's language frames and action sequences from the imitation expression behavior results, analyzes the alignment integrity of the phoneme chain, evaluates the time synchronization between the starting point of the speech output and the visual action response, and outputs the individual English interactive expression matching results.

Citation Information

Cited By

  • Vocal music remote interactive teaching system based on cloud

    CN121309554A

  • Cloud-based vocal music remote interactive teaching system

    CN121309554B

  • Teaching resource recommendation method and system based on natural language processing

    CN121724811A