Intelligent evaluation and error correction guidance method for pronunciation mode of English pronunciation

By acquiring individual learning state vectors to generate diagnostic focus vectors and performing adaptive kinematic diagnostic processing, combined with text and 3D animation feedback, the problem of insufficient personalization and immediacy in existing technologies is solved, realizing personalized and real-time speech evaluation and error correction guidance.

CN121583286APending Publication Date: 2026-02-27YANGZHOU POLYTECHNIC INST
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511774677.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-28
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

Existing speech evaluation and correction systems lack personalization and immediacy, and cannot dynamically adjust the focus and depth of diagnosis based on learners' long-term progress or short-term performance fluctuations. Feedback instructions are singular and lack specificity, making it difficult to achieve efficient pronunciation habit reshaping.

Method used

By acquiring the individual learning state vector, a diagnostic focus vector is generated, which includes error type weights and detection sensitivity thresholds. Adaptive kinematic diagnostic processing is then performed, and feedback instructions are generated by combining text prompts and 3D animations to dynamically adjust the diagnostic strategy and feedback modality.

Benefits of technology

It enables personalized voice evaluation and error correction guidance, responds to long-term trends and real-time state fluctuations, provides refined feedback, and improves the immediacy and efficiency of diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121583286A_ABST
    Figure CN121583286A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent evaluation and error correction guidance method for an English pronunciation mode, relates to the technical field of voice interaction, and seamlessly fuses a macroscopic teaching strategy and a microscopic diagnosis tactics by constructing a self-adaptive closed-loop system of'long-term state-real-time strategy-multi-dimensional diagnosis-layered feedback '. By dynamically adjusting the diagnosis focus, the detection rate and the analysis precision of intractable errors and high-priority errors are improved; secondly, through multi-factor fusion decision and hierarchical feedback, the pertinence and effectiveness of error correction guidance are enhanced, and especially for error types needing fine action simulation, the learning efficiency is improved compared with a traditional method; therefore, the voice evaluation system can remodel the pronunciation habit of the user in a prospective and personalized manner, the learning period of the user is shortened, the contusion feeling of the user is reduced, and the overall learning experience and effect are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of voice interaction, in particular to an intelligent evaluation and error correction guidance method for English voice pronunciation patterns. BACKGROUND

[0002] With the continuous development of human-computer interaction technology, voice interaction has become the core bridge connecting people and intelligent devices. In language learning, intelligent customer service and human-computer collaboration application scenarios, the system not only needs to "understand" the voice content of the user, but also needs to accurately evaluate the "quality" of the voice, especially the accuracy of pronunciation, and provide effective guidance. Therefore, how to build an intelligent evaluation and error correction system that can simulate the thinking of expert-level language teachers has become a research direction that is concerned and has important application value in the field of voice analysis and recognition technology.

[0003] The existing technology mainly faces the following challenges and limitations in realizing the intelligent evaluation and error correction guidance of pronunciation patterns: 1. Static and non-personalized diagnosis strategy: existing voice evaluation methods use fixed diagnosis models and intervention thresholds, treating all users' errors equally. For example, Chinese patent application CN120766718A discloses a training method that generates evaluation results by comparing actual voice with a preset model, but its evaluation criteria and diagnosis logic are relatively fixed in one training task. This way, it cannot dynamically adjust the focus and depth of diagnosis according to the learner's long-term progress state or short-term performance fluctuations, leading to possible over-intervention for "novices" and insufficient intervention for "veterans", lacking personalized teaching wisdom. 2. Single and lagging feedback instructions: in the error correction guidance link, existing technologies often only provide single-mode feedback based on evaluation results, such as text prompts or standard sound playback. This feedback method fails to fully consider the specific causes of errors, historical stubbornness, and the acceptance of different error types by guidance methods. For example, simple text prompts often have little effect on tongue position errors that require fine muscle control. This separation of feedback and diagnosis makes the guidance information lack of pertinence and immediate operability, making it difficult to achieve efficient pronunciation habit remodeling. SUMMARY

[0004] The purpose of the present application is to provide an intelligent evaluation and error correction guidance method for English voice pronunciation patterns to solve the problems raised in the background art.

[0005] To achieve the above purpose, the present application provides the following technical solutions: An intelligent evaluation and error correction guidance method for English voice pronunciation patterns, the specific steps comprising: S1: obtaining an individual learning state vector, the individual learning state vector being generated based on historical pronunciation data of a user and containing at least a knowledge point mastery rate parameter for representing a mastery rate of the user on a preset pronunciation knowledge point and an error mode solidification index for representing a stubborn degree of the user in repeating a specific pronunciation error mode; S2: obtaining a real-time speech signal of a current pronunciation of the user; S3: based on the individual learning state vector, performing mapping processing through a preset teaching strategy matrix to generate a diagnosis focus vector, the diagnosis focus vector containing at least a group of error type weights for adjusting analysis priorities of different pronunciation error types and at least one detection sensitivity threshold for triggering different diagnosis depths; S4: applying the diagnosis focus vector to perform adaptive kinematic diagnosis processing on the real-time speech signal to determine a deviation degree of a pronunciation organ kinematic parameter corresponding to the real-time speech signal relative to a preset standard pronunciation model, the kinematic diagnosis processing including: performing weighted analysis on acoustic characteristics of the real-time speech signal according to the error type weights, and triggering quantitative calculation of the deviation degree according to the detection sensitivity threshold; S5: generating a feedback instruction based on a result of the kinematic diagnosis processing.

[0006] Compared with the prior art, the beneficial effects of the present application are: by analyzing the historical pronunciation data of the user, an individual learning state vector containing a "knowledge point mastery rate" and an "error mode solidification index" is constructed. This directly responds to the lack of long-term strategy in the prior art.

[0007] A teaching strategy matrix generated through offline reinforcement learning is introduced, which can map the "individual learning state vector" representing long-term trends into a diagnosis focus vector in real time. The diagnosis focus vector contains a group of dynamic error type weights and detection sensitivity thresholds, which are used to influence the proportion of subsequent computing resources invested in the error types that the user needs to focus on most; further, the diagnosis focus is secondarily dynamically modulated by introducing an instant performance parameter, so that the diagnosis strategy not only responds to long-term trends, but also captures the instant state fluctuations within a single practice period. After obtaining a refined diagnosis result, the instant severity of deviation, historical persistence and visual teaching suitability of the error itself are comprehensively considered, and the most effective feedback mode is adaptively selected from a hierarchical instruction library containing text prompts and three-dimensional animation demonstrations. The fundamental problem of single, lagging and lack of targeted feedback instructions in the prior art is solved. BRIEF DESCRIPTION OF DRAWINGS

[0008] Figure 1 The technical roadmap of the method of the present application.

[0009] Figure 2 A technical logic diagram for steps S1 to S4 of the present application.

[0010] Figure 3 A technical logic diagram for step S5 of the present application.

[0011] Figure 4 An execution flow chart of the intelligent evaluation and error correction guidance method of the present application. DETAILED DESCRIPTION

[0012] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0013] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application, but the present application can also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the concept of the present application, therefore the present application is not limited to the specific embodiments disclosed below.

[0014] Embodiment one: Please refer to Figures 1 to 4 The present application provides a technical solution: An intelligent evaluation and error correction guidance method for English pronunciation patterns, comprising the following steps: S1: obtaining an individual learning state vector, the individual learning state vector being generated based on historical pronunciation data of a user, and at least containing a knowledge point mastery rate parameter for representing the user's mastery rate of a preset pronunciation knowledge point, and an error pattern solidification index for representing the user's stubbornness of repeating a specific pronunciation error pattern; S2: obtaining a real-time speech signal of the current pronunciation of the user; S3: based on the individual learning state vector, mapping processing is performed through a preset teaching strategy matrix to generate a diagnosis focus vector, the diagnosis focus vector at least containing a group of error type weights for adjusting the analysis priority of different pronunciation error types, and at least one detection sensitivity threshold for triggering different diagnosis depths; S4: applying the diagnosis focus vector, performing adaptive kinematic diagnosis processing on the real-time speech signal to determine the deviation degree of the kinematic parameters of the pronunciation organs corresponding to the real-time speech signal relative to a preset standard pronunciation model, the kinematic diagnosis processing including: performing weighted analysis on the acoustic characteristics of the real-time speech signal according to the error type weights, and triggering quantitative calculation of the deviation degree according to the detection sensitivity threshold; S5: generating a feedback instruction based on the result of the kinematic diagnosis processing.

[0015] ForFigure 1 It needs to be explained that: Figure 1 The core content is presented in the form of an infographic, showcasing the complete information processing loop of the invention's method. The process begins with an "execution object" representing the user, inputting real-time voice signals. This process is visualized as a stylized sound wave emitted from the user's mouth and directed towards the microphone, corresponding to step S2. The information flow represented by this signal enters the processing stage labeled "1. Multidimensional State Assessment and Strategy Generation." This stage integrates steps S1 and S3: first, step S1 is executed to obtain an "individual learning state vector" representing the user's long-term learning trends; then, step S3 is executed to map this vector through a preset teaching strategy matrix, generating a dynamic diagnostic strategy. The brain-embedded chart icon below this stage symbolizes the formation of intelligent strategies based on long-term data.

[0016] This strategy was subsequently used in the "2. Adaptive Acoustic Diagnostic Analysis" stage, which executes step S4. The icon of a magnifying glass examining sound waves visually illustrates how the system, based on the strategy generated in S3, performs focused, variable-depth kinematic diagnostic analysis on the speech signal acquired in S2 to determine the degree of deviation of the vocal organs. The diagnostic results proceed to the next two stages, which together constitute the complete execution process of step S5. In the "3. Multi-Factor Fusion Feedback Decision" stage, the icon of multiple arrows converging into a gear vividly represents the system fusing information from multiple dimensions, such as deviation severity and historical persistence, to drive a precise decision model and determine the type and intensity of the required feedback. Subsequently, in the "4. Hierarchical Instruction Generation and Output" stage, based on the decision results of the previous stage, the final feedback instructions are generated and output. The icons below, showing the text document and 3D head model side-by-side, clearly demonstrate the adaptive selection from the hierarchical instruction library, outputting text or 3D animation instructions, thus completing a technically clear and highly visualized evaluation and correction loop.

[0017] for Figure 4 It should be noted that the execution flow of the intelligent evaluation and error correction guidance method is as follows: Step S1 is executed to obtain the "individual learning state vector" based on the user's historical data. Simultaneously, step S2 is executed to obtain the real-time speech signal of the user's current pronunciation. Figure 4 The sound wave symbol 1 represents this. Next, the method executes steps S3 and S4, the core processing steps of which are... Figure 4This is embodied in the logical functional area represented by label 2. Specifically, in step S3, the system generates a "diagnostic focus vector" based on the "individual learning state vector" obtained in step S1. Subsequently, in step S4, the system applies this "diagnostic focus vector" to perform adaptive kinematic diagnostic processing on the real-time speech signal 1. The core of this process is to compare it with the data in the pre-existing standard pronunciation model database 3, thereby accurately calculating the "degree of deviation" of the kinematic parameters of the speech organs. After determining the "degree of deviation," the method proceeds to step S5, namely, "generating feedback instructions." Figure 4 The structured "animation instruction" data package 4 shown is a specific and preferred form of feedback instruction generated in step S5. This "animation instruction" 4 has a sophisticated internal data structure, containing a sequence of keyframes 5 arranged chronologically. Each keyframe in the sequence contains a set of precise parametric data fields 6, such as three-dimensional coordinates (XYZ) for defining position, rotation for defining pose, and deformation parameters (M) for defining shape. The feedback instruction generated in step S5 needs to be presented to the user. Figure 4 As shown, parameterized data field 6 is used to drive a three-dimensional simulation entity 7 of a vocal organ. There is a clear control relationship between parameterized data field 6 and specific movable components in the simulation entity 7, such as the lips 8, tongue 9, and jaw 10. Through this control relationship, the system can drive the simulation entity 7 to perform a visual demonstration of correct vocalization movements. This visual guidance result is ultimately presented to the user to assist them in the next round of practice, thus linking steps S1 to S5 to form a complete and personalized "evaluation-analysis-guidance" technical loop.

[0018] The core of this embodiment lies in a closed-loop adaptive speech evaluation and error correction guidance method. By acquiring an individual learning state vector representing the user's long-term learning history, and mapping this vector through a preset teaching strategy matrix, a dynamic diagnostic focus vector is generated. This diagnostic focus vector is used to adjust the kinematic diagnostic processing of the user's current pronunciation in real time and adaptively. The teaching strategy matrix generated through offline training is a static mapping. If the user's long-term learning state is determined, the diagnostic focus vector obtained within a practice cycle remains fixed. This method cannot respond to the user's immediate performance fluctuations within a single practice cycle; the user may suddenly grasp a difficult point in the early stages of practice, or their performance may decline due to fatigue. Static diagnostic focus cannot capture such short-term dynamics, resulting in a delay in diagnostic strategy adjustment and affecting the immediacy and efficiency of intervention. Based on the above core technical features, this embodiment proposes the following dynamic modulation mechanism: Further explanation: Step S1 specifically includes: calculating the knowledge point mastery rate parameter by performing time series analysis on the pronunciation score, error type, practice frequency and progress rate in the user's historical pronunciation data; and calculating the error pattern solidification index by statistically analyzing the frequency and duration of specific pronunciation errors in a continuous practice cycle.

[0019] In this embodiment, a learning trajectory analysis engine is deployed on the cloud server. The learning trajectory analysis engine continuously receives and stores de-identified practice data uploaded by user terminals. An exponential moving average (EMA) model is applied to each user's pronunciation score record to smooth the data and calculate the rate of improvement; simultaneously, a hidden Markov model (HMM) is used to identify and track the occurrence of specific error patterns (including mispronouncing / θ / as / s / ), and the error pattern persistence index is quantified based on its persistence in the state transition probability matrix.

[0020] Further explanation: In step S4, the kinematic diagnostic processing specifically includes: The real-time speech signal is subjected to time-varying formant trajectory extraction to obtain the vocal tract transient formant spectrum peak shift characterizing vowel pronunciation; and the real-time speech signal is subjected to plosive energy analysis to obtain the consonant plosive energy release rate characterizing consonant pronunciation. The degree of deviation of the kinematic parameters of the vocal organs is determined by comparing the shift of the transient resonance peak of the vocal tract and the energy release rate of the consonant plosive with the standard vocal model.

[0021] In this embodiment, the user terminal application embeds a kinematic diagnostic engine. This engine extracts the time-varying trajectory of formants from real-time speech signals using a combination of Linear Predictive Coding (LPC) and Short-Time Fourier Transform (STFT). Then, it calculates the first derivative of this trajectory in the vowel core segment to obtain the "transient resonance spectrum peak shift (Hz / ms)". Simultaneously, the engine quantifies the "consonant plosive energy release rate" by analyzing the rising slope of the energy envelope curve of plosive segments. It should be noted that the construction method of the "standard pronunciation model" follows the following standardized process: The goal of the standard pronunciation model is to provide a reference parameter vector for each target phoneme, including the "standard tract transient resonance peak shift" and the "standard consonant plosive energy release rate". Twenty certified professional broadcasters (ten men and ten women) whose native language is the target language were recruited as standard pronunciation providers. All recordings were conducted in an acoustically anechoic chamber conforming to ISO-3745 standards, with background noise below 20 dB. Newman U87 condenser microphones were used for digital recording at 48 kHz and 24-bit depth. Each broadcaster was required to repeat five times a Harvard sentence list containing all target phonemes, designed for speech balance, at a normal, clear speaking speed. For each piece of expert speech data collected, the signal processing described in this embodiment of the invention (including the same pre-emphasis coefficient, frame length, frame shift, LPC order, etc.) was used to calculate instance values ​​for the "tract transient resonance peak shift" and "consonant plosive energy release rate" corresponding to each target phoneme. To eliminate individual differences and the randomness of single pronunciations, a Gaussian Mixture Model (GMM) is used to model all parameter instance values ​​for the same phoneme. GMM better captures the complex shape of parameter distributions, thus providing a more statistically representative benchmark. For each target phoneme, a single-component Gaussian model is trained using the set of all extracted instances of "transient resonance peak shift" as input, and the mean of this Gaussian distribution is determined as the "standard transient resonance peak shift" for that phoneme. The same operation is performed on the "consonant plosive energy release rate". The final generated "standard pronunciation model" is stored as a queryable data table. Each row of this table corresponds to a target phoneme, and the columns contain the phoneme's identifier, as well as the corresponding benchmark values ​​for "standard transient resonance peak shift" and "standard consonant plosive energy release rate".

[0022] Further explanation: In step S3, the preset teaching strategy matrix is ​​generated through offline training using a reinforcement learning model; The training process of the reinforcement learning model uses the individual learning state vectors of a group of anonymous users as the state space, the parameter combination of the diagnostic focus vector as the action space, and the average decline rate of the error pattern solidification index of the anonymous user group within a preset time period as the reward function.

[0023] Further explanation: Before executing step S4, obtain real-time performance parameters that characterize the user's real-time pronunciation performance in the current practice cycle; Based on the real-time performance parameters, the generated diagnostic focus vector is adjusted using a preset modulation function to generate the final diagnostic focus vector; Specifically, step S4 involves applying the final diagnostic focus vector to perform adaptive kinematic diagnostic processing on the real-time speech signal.

[0024] Further explanation: The modulation function is configured to: reduce the weight of the corresponding error type in the final diagnostic focus vector when the real-time performance parameter indicates that the user's performance is better than the historical average level; conversely, increase the weight of the corresponding error type in the final diagnostic focus vector when the real-time performance parameter indicates that the user's performance is worse than the historical average level.

[0025] Further explanation: In the absence of the individual learning state vector, the S4 step is performed using a preset default diagnostic focus vector, where the error type weights of the default diagnostic focus vector are equal.

[0026] The following is a detailed description of the implementation of the above content. In specific embodiments of the present invention, the following key technical parameters are involved, and their definitions and determination methods are as follows: The individual learning state vector is a multi-dimensional vector used to comprehensively represent the user's long-term historical state and core characteristics in pronunciation learning. It consists of at least the following two core parameters: Key knowledge point: the rate parameter, denoted as V. mastery This parameter quantifies a user's learning efficiency for a specific pronunciation knowledge point (including specific phonemes or linking rules). Its physical meaning is the rate of change of the user's pronunciation score with the increase of effective practice sessions.

[0027] Error mode solidification index, denoted as I fossil This parameter quantifies the persistence of a specific pronunciation error pattern (including pronouncing / θ / as / s / ) in a user's pronunciation habits. Physically, it reflects the combined frequency and persistence of this error pattern in recent practice.

[0028] The diagnostic focus vector is a control vector, generated by the cloud server based on the user's individual learning state vector and sent to the user's terminal to guide the local kinematic diagnostic engine. Its components include at least: Error type weight, denoted as W error This refers to the weight values ​​corresponding to several preset pronunciation error types (including vowel tongue position errors, consonant aspiration errors, etc.). The error type weight values ​​are used to adjust the diagnostic engine's analysis priority and resource allocation for the corresponding error type.

[0029] The detection sensitivity threshold, denoted as Tmd, represents one or more values ​​used in kinematic diagnosis to determine whether the motion parameters of the vocal organs constitute a significant deviation.

[0030] The immediate performance parameter, denoted as Pjbx, quantifies a user's immediate performance level within the current single practice cycle. Physically, it represents the deviation of a user's recent pronunciation score from their long-term average score.

[0031] The final diagnostic focus vector is a control vector generated locally on the user terminal by dynamically adjusting the "diagnostic focus vector" through "real-time performance parameters" and ultimately used to guide the diagnosis of the current pronunciation.

[0032] The transient resonance peak shift of the vocal tract, denoted as ΔFsd: This parameter is used to quantify the rate of change in vocal tract morphology during vowel articulation. Its physical meaning is the first derivative of the formant frequency during the core phase of articulation with time, measured in Hertz per millisecond. In this embodiment, the formant frequency includes both the first and second formants. The plosive energy release rate, denoted as Rnsf, is used to quantify the force and crispness of plosive articulation. Physically, it represents the rate of acoustic energy accumulation and release within a short time window from closure to release of the plosive.

[0033] 1.1) Traditional scoring systems for generating individual learning state vectors only provide single-score results, failing to reveal user learning trends and distinguish between accidental errors and deeply ingrained habits. This module aims to address this issue by deeply mining user historical data to construct a quantitative model that comprehensively and dynamically reflects their learning characteristics. It introduces time series analysis to model the user's score history to extract their learning rate; simultaneously, it draws inspiration from Hidden Markov Models (HMMs) in pattern recognition to model the user's error sequences, identifying and quantifying persistent error patterns.

[0034] This module is deployed on a cloud server and runs as a "learning trajectory analysis engine". All configuration items required for calculation, including those defining the "recent" time window length and the smoothing coefficient of the EMA algorithm, are read from spreadsheet files and loaded through a data processing library. The calculation process is as follows: 1.11) Regarding the knowledge point mastery rate parameter V masteryThe determination of the model: The Exponential Moving Average (EMA) model is selected to calculate the user's average score. All historical pronunciation score sequences of the user, sorted by time, are obtained for a specific knowledge point. A smoothing coefficient is set to control the degree of smoothing. This smoothing coefficient determines the rate at which the weight of historical data decreases. In this embodiment, the smoothing coefficient is set to 0.2. The technical consideration for setting this value is to achieve a technical balance between "rapidly responding to recent performance changes" and "avoiding excessive disturbances caused by single, accidental score fluctuations." The EMA value of the first time point is initialized to the actual score of that point. Starting from the second time point, the EMA value of each time point is calculated sequentially. The calculation logic is as follows: multiply the actual score of the current time point by the smoothing coefficient to obtain the first part; multiply the EMA value of the previous time point by (1 minus the smoothing coefficient) to obtain the second part; add the first part and the second part to obtain the EMA value of the current time point. The final calculated EMA value sequence is used as the user's smoothed score curve, and the time slope of the smoothed score curve in the most recent evaluation period (in this embodiment, the evaluation period is the most recent ten effective practice sessions) is calculated. This slope is determined as the knowledge point mastery rate parameter V. mastery .

[0035] 1.12) For the error mode solidification index I fossil Determination: A Hidden Markov Model (HMM) was chosen to identify persistent errors. For a specific error pattern (including / θ / being mistakenly issued as / s / ), the user's historical practice records were converted into an observation sequence consisting of two states: "correct" and "incorrect". An HMM model containing two hidden states ("skilled state" and "fixed error state") was constructed. The model was trained using the Baum-Welchalgorithm to estimate the state transition probability matrix and the emission probability matrix. From the trained state transition probability matrix, the probability value of the "fixed error state" transitioning to itself was extracted, i.e., P(fixed error|fixed error). A time weighting factor was set to adjust for the time decay effect. The role of the time weighting factor is to give higher weight to recently occurring persistent errors than those occurring in the distant future. In this embodiment, the value of the time weighting factor was set to 0.95. The technical consideration for setting this value is to ensure that the model focuses on the persistent problems that the user most needs to solve at present, while not completely ignoring historical habits. The self-transition probability value of the "fixed error state" is multiplied by a decay function based on the most recent occurrence time of the error, and the final product is determined as the error mode fixation index I. fossil In this embodiment, the decay function is calculated using a time weighting factor.

[0036] 1.13) It should be noted that before performing iterative training using the forward-backward algorithm, the three core probability matrices of the HMM are initialized as follows: Determining the initial state probability distribution: The initial probability of the "proficient state" is set to 0.4. The initial probability of the "fixed error state" is set to 0.6. The technical consideration behind this setting is that, without any prior knowledge, it is assumed that the user's initial state is more likely to be unproficient or prone to errors when encountering a new or difficult knowledge point.

[0037] Regarding the determination of the state transition probability matrix: the probability of transitioning from a "proficient state" to a "proficient state" is set to 0.7; the probability of transitioning from a "proficient state" to a "perpetuated error state" is set to 0.3. The probability of transitioning from a "perpetuated error state" to a "perpetuated error state" is set to 0.8; the probability of transitioning from a "perpetuated error state" to a "proficient state" is set to 0.2. The technical consideration behind this setting is that it reflects the learning habit: when in a proficient state, there is a greater probability of maintaining proficiency, while when in a perpetuated error state, there is a greater probability of continuing to make mistakes, and the difficulty of transitioning from an error state to a correct state is greater than that of transitioning from a correct state to an error state.

[0038] To determine the emission probability matrix: The total number of "correct" and "incorrect" observations is counted across all observation sequences. The probability of emitting a "correct" observation in the "skilled state" is set to 0.9, and the probability of emitting an "incorrect" observation is set to 0.1. The probability of emitting a "correct" observation in the "fixed error state" is set to 0.3, and the probability of emitting an "incorrect" observation is set to 0.7. The technical consideration behind this setting is to define a strong correlation between the hidden state and the observed value: the "skilled state" is prone to producing correct pronunciation, while the "fixed error state" is prone to producing incorrect pronunciation.

[0039] 1.2) Explanation of the Generation and Dynamic Modulation of the Teaching Strategy Matrix: After accurately acquiring the user's long-term learning status, the next core issue is how to transform this "diagnostic result" into an effective teaching intervention strategy. Simple "IF-THEN" rules cannot cope with complex and ever-changing combinations of user states and are difficult to achieve optimal intervention. The core of this module models the generation process of teaching strategies as a Markov decision process, aiming to maximize the reward represented by the user's long-term learning gains by generating a diagnostic focus vector through a series of decisions. Furthermore, the concept of feedback control from cybernetics is introduced to construct a two-layer control structure: the teaching strategy matrix generated by the reinforcement learning model provides macroscopic, strategic feedforward control signals, while the dynamic modulation mechanism provides microscopic feedback adjustment signals based on immediate performance.

[0040] It is important to emphasize that the core creative contribution of this invention is that it abstracts and models the problem of personalized teaching intervention into a general framework of Markov Decision Process (MDP). Within this framework, any algorithm that can effectively solve this MDP to derive an optimal or suboptimal mapping strategy from "state" (individual learning state vector) to "action" (diagnostic focus vector) falls within the scope of this invention. The implementation based on Deep Q-Networks (DQN) is a preferred path for solving this MDP based on value functions.

[0041] In another alternative embodiment, a policy gradient-based reinforcement learning model can also be used. For example, a policy network can be constructed whose input layer also receives a normalized "individual learning state vector," and whose output layer uses a Softmax activation function to directly output the probability of selecting each discretized "diagnostic focus vector." The core idea of ​​its training process is: after a complete practice sequence, based on the total discounted reward of the sequence, the network parameters are updated through gradient ascent to directly increase the probability of the occurrence of "state-action" paths that lead to higher total rewards.

[0042] 1.21) The offline training of the teaching strategy matrix is ​​as follows: A Deep Q-Network (DQN) is selected, which can effectively utilize historical data through the experience replay mechanism, improving learning efficiency and stability. The specific training steps are as follows: A training dataset is constructed by collecting a large number of anonymous users' "individual learning state vector" sequences and their corresponding subsequent learning effect data. The input (state) of the DQN model is defined as the normalized "individual learning state vector." The output (action) of the DQN model is defined as a discretized set of "diagnostic focus vectors." This embodiment includes dividing the weights and thresholds into five levels each, forming a finite action space. A reward function is defined, whose calculation logic is as follows: obtain the user's "error pattern consolidation index" at time point T1, and then obtain the same "error pattern consolidation index" at the subsequent time point T2; calculate the difference between the two index values; divide this difference by the time interval from T1 to T2. The negative of this result is used as the reward value. Iterative training is performed using the DQN algorithm with experience replay and a target network until the model's Q-value converges. The trained DQN model is then solidified. For any "individual learning state vector" in the state space, the action with the largest Q-value output by the model is determined as the optimal "diagnostic focus vector" corresponding to that state. All these optimal "state-action" pairs together constitute the teaching strategy matrix.

[0043] 1.31) Regarding the dynamic modulation of the diagnostic focus vector, the following explanation is provided: The rule for determining the instantaneous performance parameter Pjbx is as follows: Set the sliding window size N used to define the "instantaneous" range. windowSliding window size N window Its function is to determine the most recent practice count to be included in the calculation. In this embodiment, the sliding window size N window The value is set to five. The technical consideration behind this setting is to reflect the user's short-term performance promptly while avoiding overreaction to a single mistake. The value is used to obtain the user's most recent N... window The pronunciation score is calculated for each pronunciation. The sliding window size N is then calculated. window The arithmetic mean of the scores is used to obtain the "instantaneous average score". The long-term EMA score for the corresponding knowledge point is then obtained. The "long-term EMA score" is subtracted from the "instantaneous average score", and the difference is normalized. The result is determined as the instantaneous performance parameter Pjbx. A positive parameter indicates exceptional recent performance, while a negative parameter indicates poor recent performance.

[0044] 1.32) For the generation of the final diagnostic focus vector: Obtain the baseline diagnostic focus vector generated by the teaching strategy matrix mapping from the cloud. Obtain the real-time performance parameter Pjbx calculated according to the above steps. Set a modulation gain coefficient K to control the adjustment amplitude. mod Modulation gain coefficient K mod Its function is to amplify or reduce the impact of real-time performance on the baseline strategy. In this embodiment, the modulation gain coefficient K mod The value is set to 0.3. For each "error type weight" in the baseline diagnostic focus vector, the following adjustment logic is performed: multiply the "real-time performance parameter" by the "modulation gain coefficient" to obtain an adjustment amount; then, subtract this adjustment amount from 1 to obtain a modulation factor; finally, multiply the original error type weight by this modulation factor to obtain the adjusted new weight. This logic ensures that the weight decreases when performance is good and increases when performance is poor. The diagnostic focus vector adjusted by the above steps is determined as the final diagnostic focus vector.

[0045] 1.4) Explanation of Adaptive Kinematic Diagnostic Processing: This module aims to inversely deduce the kinematic characteristics of invisible articulatory organs through deep physical-level analysis of speech signals, thereby achieving transparency and causal diagnosis of articulation movements. The core idea of ​​this module originates from the "source-filter" theory in acoustic phonetics. This theory states that speech signals are quasi-periodic signals generated by vocal cord vibration (sound source), modulated by the vocal tract composed of the oral cavity, nasal cavity, etc. The geometry of the vocal tract determines the position of the formants; therefore, by analyzing the dynamic changes of the formants, changes in the shape of the vocal tract, i.e., the movement of the articulatory organs, can be deduced. This module runs locally on the user terminal. It receives the final diagnostic focus vector as control input and processes the real-time speech signal.

[0046] 1.41) Determining the transient resonance peak shift ΔFsd of the vocal tract: Linear predictive coding (LPC) is selected to estimate the formants. LPC can fit the vocal tract response well with a set of prediction coefficients, thereby calculating the formant frequencies. Specifically, pre-emphasis processing is performed on the input 16-bit quantized real-time speech signal with a sampling rate of 16 kHz. This step is implemented through a first-order high-pass digital filter, and its calculation logic is: subtract the product of the previous sampling point's value and the pre-emphasis coefficient from the value of the current sampling point. In this step, the pre-emphasis coefficient needs to be set. The function of this parameter is to enhance the energy of the high-frequency part of the speech signal to compensate for the natural spectral tilt during the phonation process, thereby making the subsequent formant estimation more accurate in the high-frequency region. In this embodiment, the value of the pre-emphasis coefficient is set to 0.97.

[0047] The pre-emphasized signal is framed. A frame length Lcd and a frame shift Swd are set. The frame length Lcd is used to extract a signal segment that is long enough to contain several fundamental frequency cycles, yet short enough to satisfy the short-term stationarity assumption of the speech signal. In this embodiment, this parameter is set to 400 sampling points, corresponding to 25 milliseconds. The frame shift Swd determines the degree of overlap between adjacent frames to ensure the continuity of the analysis. In this embodiment, this parameter is set to 160 sampling points, corresponding to 10 milliseconds.

[0048] A window function is applied to each frame of the signal; a Hamming window is specifically chosen to obtain clearer and more accurate spectral estimation results during LPC analysis. LPC analysis is then performed on each windowed frame of the signal. This step uses the autocorrelation method to calculate the linear prediction coefficients. In this step, the order P of the LPC model is set. lpc This parameter determines the order of the predictor used to fit the vocal tract response. Its value should be approximately equal to the sampling rate (kHz) plus 2 to 4. In this embodiment, for a sampling rate of 16 kHz, the order P... lpcThe value of is set to 18. This provides sufficient model complexity to accurately fit the first five formants of the vocal tract, while avoiding spurious peaks or model instability caused by excessively high order. A root-finding operation is performed on the obtained LPC coefficient polynomial to obtain a set of complex roots. All roots located within the unit circle of the complex plane with a positive imaginary part are selected. For each selected root, its argument is calculated, multiplied by the sampling rate, and then divided by twice pi to obtain the corresponding formant frequency. A dynamic programming-based tracking algorithm is used to connect the calculated formant frequency points of each frame into a smooth trajectory. The objective function of this algorithm is to find a global path that minimizes the sum of squares of the differences between the formant frequencies of adjacent frames on the path, while penalizing paths with excessive frequency jumps. Through the formant tracking algorithm, the formants of each frame are connected into a continuous trajectory. The core segment of vowel pronunciation is identified, and the time series of the first and second formant trajectories within this segment are extracted. For the extracted resonance peak trajectory time series, its first difference is calculated, that is, the value of the next moment minus the value of the previous moment, and the average value in the core segment is obtained. This average value is determined as the transient resonance spectrum peak shift ΔFsd of the vocal tract.

[0049] 1.42) It should be further explained that the specific implementation steps of the dynamic programming pursuit algorithm are as follows: A two-dimensional data structure is obtained as input, where the first dimension represents the time frame number, and the second dimension stores all candidate formant frequency values ​​obtained by the LPC root-finding method within that time frame. A comprehensive cost function is defined for the algorithm's recursive process. This function consists of two parts: a transition cost and a discontinuity penalty. The calculation logic for the transition cost is as follows: for a path transitioning from the i-th formant (frequency Fi) of the previous frame to the j-th formant (frequency Fj) of the current frame, its transition cost is the square of the difference between Fi and Fj. This cost is used to encourage the algorithm to choose a smoother, more continuous trajectory.

[0050] In this embodiment, the discontinuity penalty is a preset, relatively large constant value. This penalty is introduced into the cost calculation in the following two cases: when the formant trajectory begins in the current frame (i.e., there is no corresponding formant in the previous frame), or when it is interrupted in the current frame (i.e., there is no connected formant in the next frame). The purpose of this penalty is to make the algorithm tend to find the longest possible continuous trajectory, avoiding the generation of meaningless fragmented trajectory segments.

[0051] Further explanation is needed: the discontinuity penalty is determined based on the following: the discontinuity penalty should be greater than the square of the largest possible formant frequency difference between two frames in normal, continuous pronunciation. The preferred calculation logic is as follows: On a verification speech dataset containing multiple phoneme transitions, the square of the formant frequency difference between all adjacent frames is calculated, forming a "transition cost" distribution. The 99th quantile of this distribution is calculated, i.e., a value is found such that 99% of the transition costs are less than this value. This 99th quantile value is multiplied by a safety gain factor, and the product is determined as the value of the "discontinuity penalty". In this embodiment, the safety gain factor is set to 2. The technical consideration for setting this value is to ensure that the penalty value can both suppress almost all normal transition costs and avoid overflow in numerical calculations due to excessive value. If the 99th quantile of the transition cost obtained through the above analysis is X, then the final value of the "discontinuity penalty" is the product of X and 2.

[0052] Initialize a cumulative cost matrix and a backtracking pointer matrix of the same size as the input data structure. Starting from the second frame, for each candidate formant point j in the current frame, traverse all candidate formant points i in the previous frame and calculate the comprehensive cost of transitioning from point i to point j, which is the cumulative cost of point i in the previous frame plus the transition cost between the two. Find the node i in the previous frame that minimizes this comprehensive cost. min Store the minimum integrated cost in the cumulative cost matrix of the current frame point j, and set i min Store the backtracking pointer matrix for the current frame point j. Simultaneously, considering the trajectory at the beginning of this frame, calculate the discontinuity penalty value and compare it with the minimum comprehensive cost calculated above; the smaller value is taken as the final cumulative cost. After completing the calculation for all frames, find the node with the smallest value in the cumulative cost matrix of the last frame. Using this node as the endpoint, backtrack frame by frame from back to front according to the backtracking pointer matrix to construct the formant trajectory with the minimum global cost. Perform the above process for the first and second formants respectively to obtain two independent, smooth formant trajectories.

[0053] 1.5) Determination of the consonant plosive energy release rate Rnsf: The input real-time speech signal is segmented into frames, and the short-time energy of each frame is calculated. Through energy change detection, the critical point at the end of the plosive occlusion segment and the beginning of the plosive segment is located. Starting from this critical point, a short-time energy sequence within a subsequent short window of thirty milliseconds is extracted. The linear regression slope of this energy sequence is calculated, and this slope is determined as the consonant plosive energy release rate Rnsf.

[0054] 1.6) Calculation of Weighted Deviation: From the standard pronunciation model library, load the standard "offset" and "energy release rate" baseline values ​​corresponding to the currently pronounced phonology. Subtract the user-calculated "offset" from the baseline value to obtain the vowel deviation; subtract the "energy release rate" from the baseline value to obtain the consonant deviation. From the final diagnostic focus vector, retrieve the weights corresponding to "vowel tongue position error" and "consonant aspiration error". Multiply the vowel deviation by its corresponding weight to obtain the weighted vowel deviation; multiply the consonant deviation by its corresponding weight to obtain the weighted consonant deviation. Combine the weighted vowel deviation and weighted consonant deviation to obtain the final overall deviation. This value will be compared with the detection sensitivity threshold in the final diagnostic focus vector to trigger the generation of subsequent feedback instructions.

[0055] It should be further explained that before calculating the rising slope of the energy envelope curve, the following automatic critical point detection steps are performed: For the short-time energy sequence of the speech signal, its first-order difference is calculated to obtain the energy change rate sequence. A silence energy threshold is set. This parameter distinguishes between speech segments and non-speech segments. In this embodiment, the value of the silence energy threshold is dynamically determined by multiplying the average energy of the input signal over the first 100 milliseconds by a gain factor of 2. The short-time energy sequence is traversed from back to front to find the first time point where the energy value is lower than the silence energy threshold, which is recorded as the occlusion end point. Within a short search window of 50 milliseconds after the occlusion end point, the maximum peak point in the energy change rate sequence is found. The time point corresponding to this maximum peak point is ultimately determined as the critical point.

[0056] 1.7) The overall computing process in this embodiment is a closed-loop system that integrates cloud-based strategic planning and terminal tactical execution: From a cloud perspective: The learning trajectory analysis engine continuously analyzes users' historical data, generating and updating each user's "individual learning state vector." Based on this vector, the personalized teaching state modulator generates a "benchmark diagnostic focus vector" through a "teaching strategy matrix" trained by reinforcement learning, and stores it, awaiting synchronization with the terminal.

[0057] From the user's perspective: When a user begins practice, the application synchronizes the latest "benchmark diagnostic focus vector" from the cloud. The user pronounces words, and the application captures the "real-time speech signal." Simultaneously, dynamic modulation calculates "instantaneous performance parameters" based on recent pronunciation scores and uses them to fine-tune the "benchmark diagnostic focus vector" in real time, generating the "final diagnostic focus vector." The kinematic diagnostic engine applies this "final diagnostic focus vector" to perform weighted and adaptive diagnostic processing on the "real-time speech signal," calculating the specific degree of deviation in kinematic parameters. Finally, based on the diagnostic results, highly personalized and targeted feedback instructions are generated and presented to the user. Furthermore, the results of this practice session are uploaded to the cloud to update the user's historical database, thus initiating the next cycle and enabling continuous adaptation and evolution of the system.

[0058] It is necessary to further clarify the following regarding the "knowledge point mastery rate parameter V". mastery "The slope of the user's historical score curve is obtained through time series analysis, and its value range covers positive, zero, and negative intervals. When the 'knowledge point mastery rate parameter V'..." mastery When the value of "" is significantly positive, it indicates that the user's learning curve is in a steep upward phase, meaning the user is in a stage of rapid absorption and skill improvement of the knowledge point. When the value of this parameter approaches zero, it indicates that the user's learning performance has entered a plateau phase, and progress has stagnated. When the value of this parameter is negative, it indicates that the user's performance is showing a downward trend, which may be caused by fatigue, confusion, or continued use of incorrect methods.

[0059] Regarding "Error Pattern Consolidation Index I" fossil ": This parameter is obtained by statistically modeling the persistence and frequency of error occurrence sequences, and its output value falls in the [0,1] interval after normalization.

[0060] When "Error Pattern Solidification Index I" fossil The closer the value of "" is to 1, the more stubborn and automatic the specific pronunciation error has become in the user's pronunciation habits, requiring strong targeted intervention. When the value of this parameter is closer to 0, it indicates that the error is more likely to be an occasional and random mistake, and has not yet formed a stable error pattern.

[0061] Regarding "Error Type Weight W" error This parameter is a core component of the "diagnostic focus vector," and its value is normalized to the [0,1] interval. When a specific error type corresponds to the "error type weight W..." errorThe closer the value of "" is to 1, the greater the proportion of computing resources the characterization diagnostic system will allocate, and the higher the priority it will be to perform in-depth and refined kinematic analysis of the acoustic features related to this error type. Conversely, the closer the value is to 0, the more basic and resource-efficient the characterization will be in monitoring this error type.

[0062] Regarding the "real-time performance parameter Pjbx": This parameter is obtained by comparing the difference between the user's short-term performance and long-term average performance. Its output value falls in the range [-1,1] after normalization.

[0063] When the value of the "Immediate Performance Parameter Pjbx" is significantly positive, approaching 1, it indicates that the user's performance in the current practice cycle is significantly better than their historical average, indicating a technical state such as "sudden enlightenment." When the value of this parameter approaches 0, it indicates that the user's current performance is in line with their historical average. When the value of this parameter is significantly negative, approaching -1, it indicates that the user's current performance is significantly worse than their historical average, indicating that they are facing immediate and sudden difficulties.

[0064] 1.8) The core advantage of this invention lies in its two-layer adaptive logic: "strategic adjustment" based on long-term learning status and "fine-tuning" based on immediate performance. This two-layer logic is achieved through the intrinsic relationship between the following parameters.

[0065] The strategic impact of the "individual learning state vector" on the "diagnostic focus vector" is analyzed and represented by the teaching strategy matrix: the input parameter is the error pattern solidification index I. fossil Knowledge point mastery rate parameter V mastery The output parameter is the error type weight W. error Detection sensitivity threshold Tmd; Error Mode Consolidation Index I fossil Error type weight W error There is a positive correlation. When other parameters remain unchanged, the "error mode solidification index I" shows a positive correlation. fossil The increase of "" will cause the "teaching strategy matrix" generated by the reinforcement learning model to output a higher corresponding "error type weight W". error The reinforcement learning model of this invention is trained using the "average decline rate of the error pattern solidification index" as the reward function. Its final convergence strategy must prioritize addressing the factors that most hinder the decline of the index. A high "error pattern solidification index I" fossil The core obstacle is precisely this kind of error. The optimal strategy for maximizing long-term learning benefits is for the model to autonomously learn to focus diagnostic resources on the most persistent errors.

[0066] Knowledge point mastery rate parameter V masteryIt is positively correlated with the detection sensitivity threshold Tmd. When other parameters remain constant, the "knowledge point mastery rate parameter V" is... mastery An increase in the value of "" (i.e., a change from negative to positive) will lead to a higher "detection sensitivity threshold Tmd" output by the "teaching strategy matrix," meaning the diagnostic system becomes relatively "insensitive." When the user is in a rapid learning phase, the knowledge point mastery rate parameter V... mastery At high levels, excessive intervention in minor deviations can disrupt the learning rhythm and undermine the process of building fluency in pronunciation. In this case, a better teaching strategy is to tolerate minor, unstructured deviations and only point out significant errors that affect comprehension. Conversely, when users hit a plateau or regress, the rate of knowledge acquisition (parameter V) decreases. mastery When the value is low or negative, it indicates that the diagnostic threshold cannot distinguish subtle errors. In this case, the diagnostic threshold must be lowered to become more "sensitive" in order to provide more refined feedback to help users overcome bottlenecks.

[0067] The impact of "real-time performance parameters" on the "final diagnostic focus vector" is analyzed using a modulation function: The input parameter is the immediate performance parameter Pjbx; the output parameter is the error type weight in the final diagnostic focus vector; the adjustment direction of the immediate performance parameter Pjbx and the error type weight is negatively correlated. When the immediate performance parameter Pjbx is positive, the modulation function will temporarily reduce the weight of the corresponding error type; conversely, when the immediate performance parameter Pjbx is negative, the modulation function will temporarily increase the weight of the corresponding error type. This solves the delay problem caused by only long-term strategic adjustments. For a user's error, it is not serious in the long run, and the error pattern solidification index I... fossil The value is low, but it appears repeatedly in the current exercise, causing the immediate performance parameter Pjbx to be negative. The modulation mechanism of this invention can immediately capture this short-term trend and instantaneously increase the diagnostic weight of the error, achieving "instant error correction". Conversely, for persistent errors that have existed for a long time, the error mode solidification index I... fossil If the value is high, and the user suddenly performs well in this exercise, the immediate performance parameter Pjbx will be positive, which can immediately reduce the intervention weight and provide positive incentives, avoiding the frustration caused by excessive repetitive guidance.

[0068] This invention further provides an adaptive feedback instruction generation method based on multidimensional state assessment. This method follows the kinematic diagnostic results of the vocal organs output in step S4. Its core technical feature is that the feedback instructions are generated adaptively from a pre-set instruction library containing different intervention depths through a two-stage, multi-factor fusion dynamic decision-making model. This model comprehensively considers the immediate severity of the deviation, its historical persistence, the visual teaching suitability of the error itself, and the overall teaching strategy transmitted by the upper-level strategy module, thereby achieving the accuracy, personalization, and maximization of teaching effectiveness in the feedback.

[0069] Further explanation: In step S5, the feedback instruction is adaptively selected and generated from a preset instruction library based on the degree of deviation of the kinematic parameters of the vocal organ; The instruction library includes at least: a first type of instruction for low deviation, wherein the first type of instruction is a text prompt; and a second type of instruction for high deviation, wherein the second type of instruction is an animation instruction that drives the three-dimensional simulated solid structure of the vocal organs to perform motion demonstration.

[0070] Further explanation: The process of selecting and generating feedback instructions includes: An intervention necessity gating step is used to determine whether feedback instructions need to be generated based on the current pronunciation deviation; and, An intervention modality selection step is used to select between a first type of instruction and a second type of instruction when the determination result of the intervention necessity gating step is that a feedback instruction needs to be generated.

[0071] Further explanation: The decision-making basis of the intervention modality selection module is based on a feedback intervention score; the calculation of the feedback intervention score incorporates at least the following three technical elements: The magnitude of the kinematic deviation vector, which represents the severity of the current deviation; the deviation persistence index, which represents the historical persistence of the error pattern; and the error diagnosticability score, which represents whether the error type is suitable for animated demonstration.

[0072] The following is a detailed implementation description of the above content: This embodiment introduces a "deviation persistence index," extending feedback decision-making from "point" evaluation to "line" evaluation, enabling the identification and targeted intervention of persistent errors that evolve over time and exhibit memory effects. Specifically, the technical effects are as follows: By constructing a two-stage cascaded structure of "intervention necessity gating" and "intervention modality selection," the decision-making process is decomposed into two levels, achieving logical decoupling and refined processing. Operations are performed directly on the multi-dimensional "kinematic deviation vector," and an "error diagnosticability score" is introduced, improving the information completeness of the decision. To implement the feedback instruction generation mechanism of this invention, the following key parameters are defined and obtained.

[0073] 3.1) The "kinematic deviation vector" is a multi-dimensional vector used to accurately and quantitatively represent the deviation of the user's current pronunciation from the standard pronunciation model in multiple kinematic dimensions. In this embodiment, the vector contains two dimensions.

[0074] The output is directly from the diagnostic processing module in step S4 above. Its calculation model is as follows: obtain the "transient resonance peak shift of the vocal tract" calculated in S4, and use it as the first component of the kinematic deviation vector.

[0075] Obtain the deviation value of the "consonant blast energy release rate" calculated by S4, and use it as the second component of the kinematic deviation vector. Combine the two components to form the final "kinematic deviation vector".

[0076] 3.2) The Deviation Persistence Index is a scalar value. Its core idea originates from the exponential moving average theory in signal processing, used to characterize the temporal continuity and persistence of deviations of a specific error type in recent practice history. Its value, after normalization, falls within the [0,1] interval; the higher the value, the more persistent the deviation. The Deviation Persistence Index is calculated by the state update module after each practice session. Its calculation model is as follows: Obtain the magnitude of the "kinematic deviation vector" calculated in the current practice cycle, denoted as the current deviation value. Read the historical value of the "Deviation Persistence Index" stored for the same error type after the end of the previous practice cycle from the storage unit. Obtain the preset "forgetting factor," used to control the decay rate of historical data. This parameter adjusts the influence weight of historical deviations on the current index. In this embodiment, the forgetting factor is set to 0.8. The technical consideration for setting this value is that it can better reflect the short-term trend of deviation within five consecutive practice cycles, capturing persistence without being overly influenced by outdated historical data. The historical index value is multiplied by the "forgetting factor" to obtain the historical weighted value. A preset "current weight factor" is obtained; this parameter defines the weight of the current deviation, and its value is the difference between 1 and the "forgetting factor," i.e., 0.2. The current deviation value is multiplied by the "current weight factor" to obtain the current weighted value. The historical weighted value and the current weighted value are added together, and the result is the updated "deviation persistence index," which is then written back to the storage unit.

[0077] 3.21) It should be further explained that the value of the "forgetting factor" directly determines the duration of the "deviation persistence index's" "memory" of historical errors. If the value is too high, the index will be slow to react to errors that have already been corrected; if the value is too low, it will be unable to effectively identify truly persistent error patterns. Therefore, finding the optimal value of the "forgetting factor" to most accurately predict the user's future error trends is a key technical problem. The core idea of ​​determining the optimal value of the "forgetting factor" in this embodiment is that the effective "deviation persistence index," at its current value, has the strongest positive correlation with the user's actual error rate within a short future time window. Through systematic parameter scanning and correlation analysis experiments, the value of the "forgetting factor" that maximizes this predictive ability is found and determined. This embodiment determines the optimal value of the "forgetting factor" in the range [0.5, 0.99]. A longitudinal pronunciation learning dataset containing at least fifty anonymous users, each with at least ten practice days, is collected. The dataset must contain the original signal, timestamp, and "instantaneous deviation amplitude" for a specific error type, diagnosed by the S4 method of this invention. Specifically, a candidate value set for the "forgetting factor" is set, starting from 0.5 and increasing in increments of 0.05 to 0.95, forming a test sequence containing ten candidate values. For each candidate "forgetting factor" value, a complete "deviation persistence index" time series is calculated for each user and each error type in the dataset. The "future error rate" is defined as the evaluation metric. For any point t in the time series, its corresponding "future error rate" is calculated as the percentage of times the "instantaneous deviation amplitude" for that error type exceeds a preset low threshold of 0.2 within the next five practice days represented by the time window [t+1, t+5]. For each candidate "forgetting factor" value, the Pearson correlation coefficient between its generated "deviation persistence index" time series and the corresponding "future error rate" time series is calculated. All candidate "forgetting factor" values ​​and their corresponding correlation coefficients are compared. The candidate value that produces the largest positive correlation coefficient is determined as the optimal "forgetting factor" of this invention.

[0078] 3.3) Error Diagnosability Score: This parameter is a preset scalar value used to characterize the extent to which a specific type of pronunciation error is suitable for visualization and correction through 3D animation. Its value falls within the range [0,1], with higher values ​​indicating greater suitability for animation demonstration. This score is pre-configured and stored in a structured spreadsheet file on an external data carrier, achieving decoupling between the core algorithm logic and specific teaching knowledge. In this embodiment, for errors mainly involving macroscopic physical morphology such as tongue position and lip shape (including vowel formant shift), the "Error Diagnosability Score" is set to a higher 0.9; while for errors mainly involving internal, less visually perceptible aspects such as airflow control and vocal cord vibration (including plosive energy release), the "Error Diagnosability Score" is set to a lower 0.3.

[0079] It should be further explained that the implementation example of objectively calculating the "error diagnosticability score" for a specific type of pronunciation error is as follows: Forty learners were recruited who were pre-diagnosed with the target pronunciation error via step S4. The participants were completely randomly assigned to group A1 (control group) and group B1 (experimental group), with twenty participants in each group. All participants underwent a baseline pronunciation test containing the target error phoneme, and their initial error rate was recorded. Both groups performed the same fifty targeted exercises. When the target error was detected: group A1 received only a pre-set text prompt for the first type of instruction; group B1 received only a pre-set 3D animation demonstration corresponding to the second type of instruction. After the intervention exercise, all participants again underwent a final pronunciation test with the same content as the baseline test, and their final error rate was recorded.

[0080] For each subject in group A1, subtract the endpoint error rate from the baseline error rate, and then divide by the baseline error rate to obtain the subject's "error improvement rate A1". The arithmetic mean of the improvement rates for all subjects in group A1 is then taken to obtain the "average improvement rate of group A1".

[0081] Perform the same procedure as for "A1 group average improvement rate" on group B1 to obtain "B1 group average improvement rate".

[0082] To quantify the relative performance improvement brought by animation guidance, the following calculations are performed: Determine if the "average improvement rate of group B1" is greater than the "average improvement rate of group A1". If so, subtract the "average improvement rate of group A1" from the "average improvement rate of group B1" to obtain an "improvement rate difference"; then, divide this "improvement rate difference" by the "average improvement rate of group A1" to obtain an initial "relative improvement gain". If the "average improvement rate of group B1" is not greater than the "average improvement rate of group A1", the initial "relative improvement gain" is directly set to zero, indicating that in this experiment, animation guidance did not show an effect superior to text guidance. Map the "relative improvement gain" to the scoring range [0,1] and perform reasonable saturation processing, executing the following steps: Determine if the "relative improvement gain" is greater than 1. If so, set its value to 1. This step aims to handle the extreme case where the improvement effect of group B1 far exceeds that of group A1, uniformly mapping it to the highest score. If not, its value remains unchanged. The final output of this step is the "error diagnosticability score" determined as the target error type. This experimental paradigm and computational model ensures that the scores directly and objectively reflect the actual teaching effectiveness of the animation instruction. For errors with a diagnosticability score of 0.9, the physical meaning is that animation instruction provides significantly greater additional learning benefits compared to text instruction; while an error score of 0.3 indicates that the advantages of animation instruction are not obvious, and the system should prioritize the lighter text feedback.

[0083] 3.4) Teaching strategy vector, denoted as V strategy This parameter, output from the teaching strategy matrix in step S3, is used to convey macro-level, personalized teaching guidance strategies to downstream feedback instructions. In this embodiment, the teaching strategy vector includes the level of intervention aggressiveness.

[0084] The intervention aggressiveness level, denoted as Ldj, is a scalar value in the interval [0,1]. Its function is to regulate the sensitivity and intervention intensity of the overall control feedback system. When the intervention aggressiveness level approaches 1, it indicates that a more aggressive error correction strategy should be adopted, intervening even for minor deviations; when the intervention aggressiveness level approaches 0, it indicates that the system should adopt a more lenient strategy, intervening only for severe deviations. This value is dynamically generated by S3 based on the "knowledge point mastery rate parameter" representing the user's overall learning status, specifically including: The input is the "knowledge point mastery rate parameter V" determined in the previous steps. mastery The following two key internal control parameters are introduced and defined: A plateau threshold is set to define whether a user's learning state has entered a "plateau" or "slight regression" phase. This is a small positive value greater than zero, used to classify even very slow progress as a plateau requiring active intervention. In this embodiment, the specific value is 0.1. This value is set to avoid misjudging weak positive slopes caused by computational noise as effective progress, even if they are meaningless. Only when the mastery rate exceeds this threshold is the user considered to have entered a true, protected state of "progress."

[0085] A maximum progress rate reference value is set. This parameter defines the upper limit of the maximum progress rate that is considered "ideal" or "saturated." Its function is to normalize the "knowledge point mastery rate parameter" within the open interval to a controllable range for calculation. In this embodiment, the specific value is 5.

[0086] The "knowledge point mastery rate parameter V" mastery Based on whether it exceeds the "plateau period judgment threshold", different calculation paths are selected: Path A2: If the "Knowledge Point Mastery Rate Parameter" is less than or equal to the "Plateau Period Judgment Threshold", the system determines that the user is in a state requiring active intervention. Under this path, the final value of "Intervention Aggression Level Ldj" is directly set to 1.

[0087] Path B2: If the "Knowledge Point Mastery Rate Parameter" is greater than the "Plateau Judgment Threshold," the system determines that the user is in a progress state requiring protection. Under this path, the following consecutive sub-steps are executed: Sub-step B21: Compare the "Knowledge Point Mastery Rate Parameter" with the "Maximum Progress Rate Reference Value". If the former is greater than the latter, set a temporary "Rate Value Used for Calculation" as the "Maximum Progress Rate Reference Value"; otherwise, set the "Rate Value Used for Calculation" to the original value of the "Knowledge Point Mastery Rate Parameter". This step ensures that subsequent calculations will not produce abnormal results due to extreme progress rates.

[0088] Sub-step B22: Using the "rate value for calculation" obtained in "sub-step B21" as the dividend and the "maximum progress rate reference value" as the divisor, perform a division operation to obtain a "normalized progress ratio". This ratio value is in the range [0,1], reflecting the degree of the current progress rate relative to the ideal maximum rate.

[0089] Sub-step B23: Subtract the "normalized progress ratio" obtained in "sub-step B22" from the constant "1", and the difference is determined as the final value of the "intervention radicalism level".

[0090] 4.0) Furthermore, to avoid problems such as inappropriate feedback timing or mismatched feedback formats, based on the "intervention aggressiveness level," the sensitivity of feedback is dynamically adjusted according to the user's learning stage, in order to comprehensively judge whether a deviation is "instantaneously severe" or "long-term persistent"; specifically including: 4.01) The first stage, the principle of intervention necessity gating steps: This technique aims to filter out minor, random deviations that learners can correct themselves without systemic intervention, thus avoiding over-feedback. Its core idea is to establish a dynamic "trigger gate." Subsequent feedback selection is only initiated when the severity of the deviation "crosses" this threshold determined by the current teaching strategy. Specifically, this includes: Obtain the input "kinematic deviation vector" and calculate its Euclidean modulus to obtain the "instantaneous deviation amplitude," which represents the total deviation. Obtain the "intervention aggressiveness level" parameter from the input "teaching strategy vector." Obtain a preset "basic intervention threshold," which is defined as the trigger threshold under the most tolerant strategy. In this embodiment, the basic intervention threshold is set to 0.5. Subtract the value of "intervention aggressiveness level" from 1 to obtain the adjustment coefficient. Multiply the "basic intervention threshold" and the adjustment coefficient to obtain the "dynamic intervention threshold." This calculation logic ensures that the higher the "intervention aggressiveness level," the lower the dynamic intervention threshold, and the more sensitive the system.

[0091] Compare the "instantaneous deviation amplitude" with the "dynamic intervention threshold". If the former is greater than the latter, the "intervention necessity gating step" outputs an "intervention required" signal and activates the "intervention modality selection step" in the second stage; otherwise, an "intervention not required" signal is output and the current feedback process terminates.

[0092] 4.02) The second stage, the technical principle of intervention modality selection: After determining the need for intervention, it aims to solve the problem of "how to intervene." Its core idea is to calculate a comprehensive "feedback intervention score" by non-linearly fusing multiple dimensions of deviation features. This score can more comprehensively evaluate the most suitable intervention modality for the current error than any single indicator, including text or animation. Specifically, it includes: Obtain three input parameters: "Instantaneous Deviation Amplitude," "Deviation Persistence Index," and "Error Diagnosability Score." To ensure that parameters with different dimensions and value ranges can be compared and integrated on a uniform scale, perform the following normalization operation: Normalize the "instantaneous deviation amplitude": Obtain a preset "amplitude reference upper limit," which defines the level of deviation amplitude that can be considered saturated. In this embodiment, the upper limit is set to 2. Divide the "instantaneous deviation amplitude" by the "amplitude reference upper limit" to obtain a quotient; determine if the quotient is greater than 1. If yes, take 1 as the result; otherwise, take the quotient as the result. The output of this step is the "normalized deviation amplitude."

[0093] Normalize the "deviation from persistence index": Using the same logic as the "normalized deviation magnitude", obtain the preset "persistence index reference upper limit" of 1.5, and perform division and saturation processing on the "deviation from persistence index" to obtain the "normalized persistence index".

[0094] A set of preset "fusion weights" for adjusting the importance of each parameter is obtained, and this weighting is stored in the aforementioned external spreadsheet file. In this embodiment, the weight of "instantaneous deviation amplitude" is set to 0.4, the weight of "deviation persistence index" is set to 0.5, and the weight of "error diagnosticability score" is set to 0.1. This weighting configuration indicates that persistent errors that have been present for a long time require stronger intervention than occasional serious errors.

[0095] The obtained "normalized deviation amplitude" is multiplied by its corresponding weight of 0.4. The "normalized persistence index" is multiplied by its corresponding weight of 0.5. The "error diagnosticability score" is multiplied by its corresponding weight of 0.1. The above three calculation results are added together to obtain the final "feedback intervention score". A preset "modal selection threshold" is obtained. This parameter is used to distinguish between the first type of instruction and the second type of instruction. In this embodiment, this value is set to 0.7. The "feedback intervention score" is compared with the "modal selection threshold". If the former is greater than the latter, the second type of instruction (animation instruction) is selected; otherwise, the first type of instruction (text prompt) is selected.

[0096] In another embodiment, a relatively simplified implementation of feedback instruction generation is provided. In this implementation, the "degree of deviation" is represented as a single comprehensive scalar value. The specific implementation steps are as follows: First, obtain the "kinematic deviation vector" output from step S4; second, obtain a set of preset simplified weights for calculating the comprehensive degree of deviation (assigning equal weights to each component of the vector); multiply the absolute value of each component of the "kinematic deviation vector" with its corresponding simplified weight, and then add all products to obtain a comprehensive "degree of deviation" scalar value. Subsequently, compare this "degree of deviation" scalar value with a preset "first threshold" of 0.4. If it is less than the first threshold, no feedback is generated. If it is greater than the first threshold, continue comparing it with a preset "second threshold" of 0.8. If it is less than the "second threshold," select and generate a first type of instruction (text prompt) from the instruction library; if it is greater than or equal to the "second threshold," select and generate a second type of instruction (animation instruction). Although this simplified embodiment does not include the complex factors such as timing analysis in the aforementioned embodiments, it still achieves the core technical concept of "adaptively selecting and generating feedback from the two-level instruction library according to the degree of deviation," and therefore also falls within the protection scope of this invention.

[0097] 4.1) This embodiment dynamically adjusts the feedback sensitivity through an "intervention aggressiveness level," providing more intensive guidance for beginners and a more relaxed environment for advanced learners. By integrating instantaneous, historical, and error type features, it can accurately identify error patterns that most require strong intervention methods like animation, including examples of persistent tongue position errors suitable for animation demonstration, thus avoiding wasting animation resources on unsuitable or unnecessary scenarios.

[0098] 4.2) In this embodiment of the invention, the complete calculation process of step S5 is as follows: Receive the "kinematic deviation vector" from step S4. Invoke the state update module, combine historical values ​​with the current deviation, and update and store the "deviation persistence index" for this error type. Load the "error diagnosticability score" corresponding to this error type, as well as all preset parameters required for "intervention necessity gating" and "intervention modality selection" from the external configuration file, including the basic intervention threshold, fusion weights, etc. Receive the "teaching strategy vector" from step S3 and extract the "intervention aggression level". All parameters are sent to the "intervention necessity gating step". Calculate the "dynamic intervention threshold" and compare it with the "instantaneous deviation amplitude" to determine whether to continue. If it is decided to continue, the "instantaneous deviation amplitude", "deviation persistence index", and "error diagnosticability score" are sent to the "intervention modality selection step".

[0099] The "Intervention Modality Selection Module Step" calculates the "Feedback Intervention Score" using a weighted fusion algorithm and compares it with the "Modality Selection Threshold," ultimately deciding whether to "select the first type of instruction" or "select the second type of instruction." Based on this decision, the system retrieves and generates corresponding text prompts or animation instruction data streams from a pre-set instruction library, completing the feedback process. If the decision is for the second type of instruction, the system will retrieve the animation instruction corresponding to the current error (including the / r / sound with an excessively low tongue position). This instruction contains a set of keyframe data used to drive the tongue model in the 3D simulation structure of the articulatory organ from the current incorrect position to the target correct position, and outputs this data stream to the rendering engine.

[0100] In this invention, a pre-constructed three-dimensional simulation entity structure representing standard human vocal organs is utilized. The entity structure includes at least a segmented tongue model that can be driven by external data, as well as associated movable parts such as the jaw and lips. The tongue model possesses a skeletal binding or equivalent deformer system, enabling it to respond to external commands and accurately simulate key vocalization movements such as curling, raising, and lowering. This entity structure itself can be constructed using any three-dimensional modeling software (including Blender, Maya, etc.) and techniques (including polygon mesh modeling) well-known to those skilled in the art; the construction methods are common knowledge and will not be elaborated upon here.

[0101] If the decision is a type 2 instruction, the animation instruction corresponding to the current error will be retrieved. The animation instruction contains a set of keyframe data streams conforming to a predetermined data structure. In this embodiment, the keyframe data stream is a time-ordered data sequence, and each data unit contains at least: [timestamp, target bone / deformer identifier, target position coordinates, target rotation quaternion]. This data stream is output to a rendering module. The rendering module is configured to parse the aforementioned keyframe data stream and, using animation techniques known in the art such as linear or spline interpolation, calculate the intermediate state of the tongue model in each rendered frame, thereby generating a continuous animation on the display interface that smoothly transitions from the current error position to the correct target position. The rendering module is implemented based on common graphics APIs (including WebGL, OpenGL, DirectX, etc.).

[0102] The following detailed implementation instructions apply to the above content: "Feedback intervention score", after normalization, its output value range is strictly limited to the [0,1] interval.

[0103] When the "feedback intervention score" approaches 1, it indicates that the decision model of this invention determines that the current pronunciation error scenario requires a high-intensity, visual intervention. This directly corresponds to an increased tendency to select the second type of instruction (animation instruction) from the instruction library. Conversely, when the score approaches 0, it indicates that the model determines that the current error only requires a low-intensity, suggestive intervention, corresponding to a tendency to select the first type of instruction (text prompt).

[0104] The magnitude of the kinematic deviation vector, representing the normalized deviation amplitude, indicates the instantaneous severity of the user's current pronunciation action deviating from the standard model in physical space. When other parameters remain constant, an increase in this magnitude leads to a monotonically increasing "feedback intervention score." A more severe pronunciation distortion inherently requires clearer and more intuitive guidance for correction, and animation instructions can provide this high-information-dimensional guidance. This positive correlation design accurately maps the real-world logic that "the more severe the problem, the more powerful the solution required."

[0105] The Deviation Persistence Index, representing the Normalized Persistence Index, quantifies the persistence of a specific error pattern in a user's recent practice history. When other parameters remain constant, an increase in the Deviation Persistence Index leads to a significant increase in the Feedback Intervention Score. The underlying logic is that occasional, random errors attract the user's attention through simple text prompts; however, a recurring, persistent error indicates that low-intensity interventions have failed and must be upgraded to a more impactful and memorable intervention modality such as animated demonstrations.

[0106] Error diagnosticability scoring is a priori knowledge that characterizes the extent to which a particular error type is suitable for instruction through animation. When other parameters remain constant, a higher error diagnosticability score will tend to generate a higher feedback intervention score, even with moderate deviations. This is reasonable because not all errors are suitable for visual instruction. Prioritizing animation intervention resources for error types that benefit most from visualization (including tongue and lip shape errors), while "de-weighting" unsuitable error types (including airflow control), reflects optimized and intelligent allocation of instructional resources, avoiding ineffective interventions.

[0107] Furthermore, to quantitatively verify the technical advantages of the method described in this invention, a computational environment was built to simulate the user's pronunciation learning process. In this environment, the method of this invention (experimental group) was compared with a prior art method (control group) that judges solely based on instantaneous deviation amplitude. The logic of the control group is as follows: when the normalized deviation amplitude is greater than a fixed high threshold of 0.6, an animation command is output; otherwise, a text command is output.

[0108] The experiment designed six representative pronunciation error scenarios to examine the decision-making rationality and teaching efficiency of the two methods in different situations. The experimental results are recorded in Table 1 below: Table 1: Comparison of decision-making behavior and teaching efficiency between the method of this invention and the control group method

[0109] The teaching efficiency improvement index, denoted as IEII, is calculated as follows: IEII = [(Error Improvement Rate (Invention) - Error Improvement Rate (Control Group)) / Error Improvement Rate (Control Group)] × 100%. Here, "Error Improvement Rate" is defined as the percentage decrease in the average deviation of the error in the subsequent three practice sessions after receiving one feedback instruction. This index quantifies the relative gain of the invention in improving teaching efficiency compared to the control group. This parameter quantitatively characterizes the efficiency advantage of the invention's method in promoting user error correction in a specific scenario, compared to the baseline control group method. A larger positive value indicates a more effective decision-making process.

[0110] Comparing Scenario 1 and Scenario 2: In Scenario 1 (severe but accidental), the control group output animation due to deviation amplitude (0.8) exceeding the threshold, resulting in over-intervention. This invention comprehensively considers a low persistence index (0.1), calculating an intervention score (0.46) lower than the modality selection threshold (0.7 in this embodiment), thus correctly selecting a lighter text instruction. The final result is that the error improvement rate (30.1%) of this invention is slightly lower than the control group (35.0%), with an IEII of -14.0%. This indicates that by avoiding unnecessary strong intervention, this invention saves computational resources and user attention while slightly sacrificing the highest single-time improvement rate, achieving higher overall efficiency.

[0111] In Scenario 2 (mild but persistent), the control group repeatedly output invalid text commands because the deviation amplitude (0.3) did not reach the threshold, resulting in an error improvement rate of only 20.0%. This invention, by weighting with a high persistence index (0.9), increased the intervention score to 0.66, showing a clear trend towards upgrading to animation intervention, and thus achieved an error improvement rate of 29.0%. Its teaching efficiency improvement index reached +45.0%. This comparison powerfully demonstrates that this invention, by introducing a "deviation persistence index," can effectively distinguish between "novice's occasional mistakes" and "experienced's persistent errors," thereby enabling more targeted intervention decisions.

[0112] Comparing Scenario 3 and Scenario 4: In Scenario 3 (severe and persistent lip-reading errors), all input parameters pointed to strong intervention. Both the present invention and the control group correctly selected the animation instructions, and the teaching efficiency improvement index was +5.0%, which was basically the same, demonstrating the consistency of decision-making in high-risk scenarios.

[0113] However, in scenario 4 (severe but accidental airflow error), the control group again erroneously output animation due to a high deviation amplitude (0.8), with an error improvement rate of only 25.0%. This invention, on the other hand, "downweighted" the results with a low "error diagnosticability score" (0.2), resulting in a final intervention score of only 0.39, thus selecting text instructions and achieving an improvement rate of 33.8%. This resulted in a teaching efficiency improvement index of +35.2%. This comparison demonstrates that this invention, by introducing an "error diagnosticability score," achieves intelligent and maximized allocation of intervention resources (animation), avoiding ineffective visual teaching on "unsuitable" errors.

[0114] Scenario 6 (moderate and persistent tongue positioning error) is a typical critical decision-making scenario. The control group selected text because the deviation amplitude (0.5) did not reach the threshold, resulting in an improvement rate of 22.0%. This invention, by fusing a higher persistence index (0.7) and a diagnostic score (0.9), calculated an intervention score as high as 0.64, approaching the threshold for triggering animation (0.7), achieving an improvement rate of 31.2%. This resulted in a teaching efficiency improvement index of +41.8%. This demonstrates that the multi-factor fusion model of this invention, compared to the "step-like" decision-making of a single threshold, provides a continuous and smooth output, more precisely reflecting the cumulative process of risk.

[0115] In another embodiment, the interval division of the "feedback intervention score" and the corresponding strategy selection are determined scientifically through the "objective threshold derivation" methodology. Specifically, this includes: Selection and Justification of Effectiveness Indicator (VEI): In this embodiment, the "Probability of Error Correction After a Single Intervention (PICP)" is selected as the VEI. This indicator directly quantifies the probability that a user will successfully correct an error in the next attempt after receiving specific feedback. It is directly related to the immediate effectiveness of the teaching intervention. Through regression analysis of the experimental data, the "Feedback Intervention Score" and the VEI "Probability of Error Correction After a Single Intervention" show a significant S-shaped curve relationship. Specifically, in the low-value region of the "Feedback Intervention Score," PICP increases slowly with the increase of the score; in the medium-value region, PICP rises rapidly; and in the high-value region, the growth of PICP tends to level off again, approaching saturation. The point of maximum curvature of the S-shaped curve, that is, the point where the absolute value of the second derivative of the curve reaches its maximum, is denoted as the modality selection threshold Tmode, marking the critical point where "the marginal benefit of increasing the intensity of intervention begins to diminish." Through calculation of the experimental data, the modality selection threshold Tmode≈0.70 is obtained. The technical meaning of this threshold is that when the "feedback intervention score" is below 0.70, the improvement in PICP brought about by upgrading from text instructions to animation instructions is significant; however, when the score exceeds 0.70, further increasing the intensity of intervention will no longer result in an economical increase in PICP.

[0116] Level 1 intervention (text prompt): Low generalized cost, low computational overhead, and minimal disruption to users. The intervention's effectiveness is within the low to medium score range.

[0117] Secondary intervention (animated demonstration): This approach is generally costly, involves significant rendering overhead, and requires the user to maintain focused attention. However, its effectiveness far exceeds that of text prompts in the high-scoring range.

[0118] The activation condition for secondary intervention (animation) is set after the "feedback intervention score" exceeds the modality selection threshold Tmode by 0.70. The optimization of this strategy mapping is that it ensures that high-cost secondary interventions are activated only within the range where "the benefits are maximized," while low-cost primary interventions are used in the low-score range where the benefits are not significant. This leads to the identification of the following practical application ranges with a solid data foundation: In the low intervention range, the feedback intervention score ∈ [0, 0.70]: its upper boundary of 0.7 is determined by the modality selection threshold Tmode, where the marginal benefit of the intervention begins to diminish. Initiate Level 1 intervention (text prompt). This strategy, through cost-benefit analysis, has been proven to be the optimal strategy for balancing risk control and system costs at this stage.

[0119] The high intervention range has a feedback intervention score ∈ (0.70, 1.0]: its lower boundary of 0.70 is determined by the modality selection threshold Tmode, which indicates the most cost-effectiveness of strong interventions. Initiate the secondary intervention (animation instruction). This strategy is reserved for addressing error scenarios that have been confirmed by data as the most needed and most likely to benefit from visual teaching.

[0120] It should be noted that all calculation formulas in this application employ regression analysis, including but not limited to machine learning algorithms, to deeply analyze the collected parameters and identify their natural trends and interrelationships. Specialized software, such as Python's Scikit-learn library or the R language, is used to automatically generate mathematical models that match the data. Then, cross-validation and other methods are used to objectively evaluate the model performance, and continuous feedback and optimization are combined to ensure that the created formulas truly reflect the inherent laws of the data, thereby guaranteeing their effectiveness and accuracy. In all calculation formulas in this application, the parameters in each formula undergo dimensionless processing within a consistent range to ensure that different physical quantities are compared on the same scale; dimensionless processing techniques include, but are not limited to, min-max-normalization and Z-score standardization. The algorithm of this invention is implemented as a Python script. Before executing the core logic, the program first executes a data loading module (e.g., using the widely used pandas library in Python) configured to read the aforementioned spreadsheet file and load its contents into the program's working memory (e.g., a DataFrame data structure). Subsequent algorithm steps will directly query and retrieve the required configuration parameters from this in-memory data structure.

[0121] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for intelligent evaluation and error correction guidance of English pronunciation patterns, characterized in that, The specific steps include: S1: Obtain an individual learning state vector, which is generated based on the user's historical pronunciation data and includes at least a knowledge point mastery rate parameter to characterize the user's mastery rate of preset pronunciation knowledge points, and an error pattern solidification index to characterize the persistence of the user's repetition of specific pronunciation error patterns. S2: Acquire the real-time voice signal of the user's current pronunciation; S3: Based on the individual learning state vector, a pre-set teaching strategy matrix is ​​used for mapping to generate a diagnostic focus vector. The diagnostic focus vector includes at least a set of error type weights for adjusting the analysis priority of different pronunciation error types, and at least one detection sensitivity threshold for triggering different diagnostic depths. S4: Apply the diagnostic focus vector to perform adaptive kinematic diagnostic processing on the real-time speech signal to determine the degree of deviation of the kinematic parameters of the speech organs corresponding to the real-time speech signal from the preset standard speech model. The kinematic diagnostic processing includes: performing weighted analysis on the acoustic features of the real-time speech signal according to the error type weight, and triggering the quantitative calculation of the degree of deviation according to the detection sensitivity threshold. S5: Based on the results of the kinematic diagnostic processing, generate feedback instructions.

2. The intelligent evaluation and error correction guidance method for English speech pronunciation patterns according to claim 1, characterized in that: The steps of S1 specifically include: calculating the knowledge point mastery rate parameter by performing time series analysis on the pronunciation score, error type, practice frequency and progress rate in the user's historical pronunciation data; and calculating the error pattern solidification index by statistically analyzing the frequency and duration of specific pronunciation errors in continuous practice cycles.

3. The intelligent evaluation and error correction guidance method for English speech pronunciation patterns according to claim 2, characterized in that: In step S4, the kinematic diagnostic process specifically includes: The real-time speech signal is subjected to time-varying formant trajectory extraction to obtain the vocal tract transient formant spectrum peak shift characterizing vowel pronunciation; and the real-time speech signal is subjected to plosive energy analysis to obtain the consonant plosive energy release rate characterizing consonant pronunciation. The degree of deviation of the kinematic parameters of the vocal organs is determined by comparing the shift of the transient resonance peak of the vocal tract and the energy release rate of the consonant plosive with the standard vocal model.

4. The intelligent evaluation and error correction guidance method for English speech pronunciation patterns according to claim 3, characterized in that: In step S3, the preset teaching strategy matrix is ​​generated through offline training using a reinforcement learning model; The training process of the reinforcement learning model uses the individual learning state vectors of a group of anonymous users as the state space, the parameter combination of the diagnostic focus vector as the action space, and the average decline rate of the error pattern solidification index of the anonymous user group within a preset time period as the reward function.

5. The intelligent evaluation and error correction guidance method for English speech pronunciation patterns according to claim 4, characterized in that: Before performing step S4, obtain real-time performance parameters that characterize the user's real-time pronunciation performance in the current practice cycle; Based on the real-time performance parameters, the generated diagnostic focus vector is adjusted using a preset modulation function to generate the final diagnostic focus vector; Specifically, step S4 involves applying the final diagnostic focus vector to perform adaptive kinematic diagnostic processing on the real-time speech signal.

6. The intelligent evaluation and error correction guidance method for English speech pronunciation patterns according to claim 5, characterized in that: The modulation function is configured to: reduce the weight of the corresponding error type in the final diagnostic focus vector when the real-time performance parameter indicates that the user's performance is better than the historical average; conversely, increase the weight of the corresponding error type in the final diagnostic focus vector when the real-time performance parameter indicates that the user's performance is worse than the historical average.

7. The intelligent evaluation and error correction guidance method for English speech pronunciation patterns according to claim 6, characterized in that: If the individual learning state vector cannot be obtained, step S4 is performed using a preset default diagnostic focus vector, wherein the error type weights of the default diagnostic focus vector are equal.

8. The intelligent evaluation and error correction guidance method for English speech pronunciation patterns according to claim 7, characterized in that: In step S5, the feedback instruction is adaptively selected and generated from a preset instruction library based on the degree of deviation of the kinematic parameters of the vocal organ. The instruction library includes at least: a first type of instruction for low deviation levels, wherein the first type of instruction is a text prompt; And a second type of instruction for high deviation, which is an animation instruction that drives the three-dimensional simulated solid structure of the vocal organs to perform motion demonstration.

9. The intelligent evaluation and error correction guidance method for English speech pronunciation patterns according to claim 8, characterized in that: The process of selecting and generating feedback instructions includes: An intervention necessity gating step is used to determine whether feedback instructions need to be generated based on the current pronunciation deviation; and, An intervention modality selection step is used to select between a first type of instruction and a second type of instruction when the determination result of the intervention necessity gating step is that a feedback instruction needs to be generated.

10. The intelligent evaluation and error correction guidance method for English speech pronunciation patterns according to claim 9, characterized in that: The decision-making basis of the intervention modality selection module is based on the feedback intervention score; the calculation of the feedback intervention score integrates at least the following three technical elements: The magnitude of the kinematic deviation vector, which represents the severity of the current deviation; the deviation persistence index, which represents the historical persistence of the error pattern; and the error diagnosticability score, which represents whether the error type is suitable for animated demonstration.

Citation Information

Patent Citations

  • Voice auditory language training method and system based on artificial intelligence

    CN120766718A