A Cantonese pronunciation correction system based on AI speech recognition

CN122575410APending Publication Date: 2026-08-14GUANGDONG DANCE & DRAMA VOCATIONAL COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-20
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

现有声学比对系统只能告知用户“哪里错了”,却完全无法解释“错在哪里、如何纠正”

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122575410A_ABST
    Figure CN122575410A_ABST
Patent Text Reader

Abstract

This invention discloses a Cantonese pronunciation error correction system based on AI speech recognition, belonging to the field of Cantonese pronunciation error correction technology. It includes a multi-channel speech signal acquisition and front-end enhancement module, an acoustic feature decoupling extraction module based on self-supervised learning, a differentiable pronunciation physiological motion inverse estimation module, a spatiotemporal registration module for the motion trajectory of speech organs, a standard Cantonese pronunciation dynamic template library management module, a multi-dimensional physiological deviation quantification analysis module, a three-dimensional virtual duct dynamic simulation and rendering module, an augmented reality interaction and correction force field visualization module, a multimodal error correction feedback strategy generation module, a personalized pronunciation profile and learning path planning module, and a closed-loop pronunciation adaptability and calibration module. This invention can transform abstract pronunciation errors into visualized organ-level guidance, solving the technical problem of "knowing the error but not knowing how to correct it."
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of Cantonese pronunciation correction technology, and in particular to a Cantonese pronunciation correction system based on AI speech recognition. Background Technology

[0002] With increasingly frequent regional economic and cultural exchanges, the demand for non-native speakers to learn Cantonese is gradually increasing. However, the Cantonese phonological system is complex, possessing more than six tones and a rich variety of vowels and entering tone codas. Its articulation mechanism involves the fine coordinated movements of the tongue position, lip shape, soft palate, and glottis. For learners whose native language is Mandarin or another language, mastering the precise posture of these articulatory organs is a major challenge.

[0003] Current Cantonese pronunciation learning and correction techniques can be mainly divided into two categories: The first type is the traditional "listen-imitation" teaching model. This model relies on teachers or pre-recorded audio for auditory demonstration, and learners practice through repeated imitation and subjective feeling. The fundamental flaw of this method is that learners cannot directly "see" the actual movement of their oral articulatory organs, resulting in opaque and imprecise error correction information, low learning efficiency, and difficulty in mastering the subtle differences in higher-level consonants and vowels.

[0004] The second category consists of existing AI speech evaluation and correction systems based on acoustic features. These systems extract the Mel-frequency cepstral coefficients and linear predictive coding acoustic features from the user's speech, compare them with a standard acoustic template, calculate a confidence score, and highlight low-scoring parts in red to indicate "inaccurate pronunciation." However, this type of technology has a fundamental limitation: the "one-to-many" mapping problem between acoustic features and physiological movements. Different combinations of articulatory organs can produce extremely similar acoustic effects. Existing acoustic comparison systems can only tell users "where the mistake is," but cannot explain "what the mistake is and how to correct it." This "knowing what but not why" feedback leaves learners helpless even when faced with red markings, hindering effective and rapid pronunciation correction.

[0005] To address the aforementioned shortcomings, this invention proposes a Cantonese pronunciation correction system based on AI speech recognition. This system transcends the traditional "acoustic feature comparison" level, delving into the physiological movement of pronunciation. It inversely infers the user's vocal organ movements through a differentiable physiological model of pronunciation, then performs spatiotemporal registration with a standard dynamic template. Finally, augmented reality technology visualizes the abstract intraoral movements, quantifies deviations, and intuitively guides the user's organs to the correct position in the form of dynamic "force lines." This mechanism completely solves the problem of "black box operation" in pronunciation correction, providing Cantonese learners with a learning tool possessing physiological-level comprehension and intuitive guidance. Summary of the Invention

[0006] The purpose of this invention is to address the shortcomings of existing technologies by proposing a Cantonese pronunciation correction system based on AI speech recognition.

[0007] To achieve the above objectives, the present invention adopts the following technical solution: A Cantonese pronunciation error correction system based on AI speech recognition includes a multi-channel speech signal acquisition and front-end enhancement module, an acoustic feature decoupling extraction module based on self-supervised learning, a differentiable pronunciation physiological motion inverse estimation module, a speech organ motion trajectory spatiotemporal registration module, a standard Cantonese pronunciation dynamic template library management module, a multi-dimensional physiological deviation quantitative analysis module, a three-dimensional virtual duct dynamic simulation and rendering module, an augmented reality interaction and correction force field visualization module, a multimodal error correction feedback strategy generation module, a personalized pronunciation profile and learning path planning module, and a closed-loop pronunciation adaptability and calibration module. The multi-channel voice signal acquisition and front-end enhancement module is used to acquire Cantonese voice signals emitted by users and perform noise reduction, echo cancellation and sound source enhancement processing. The self-supervised learning-based acoustic feature decoupling extraction module is used to extract decoupled acoustic features related to the pronunciation content from the enhanced speech signal. A differentiable phonation physiological motion reverse deduction module, with a built-in differentiable phonation physiological model, is used to reverse deduce the continuous motion trajectory parameters of the phonation organs in real time based on the acoustic characteristics. The spatiotemporal registration module for the motion trajectory of the vocal organs is used to perform dynamic temporal warping and nonlinear spatial registration between the continuous motion trajectory parameters and the standard templates in the standard Cantonese pronunciation dynamic template library. The multi-dimensional physiological deviation quantification analysis module is used to calculate the deviation vector between the user's actual vocal trajectory and the standard template at each time slice and each vocal organ after nonlinear spatial registration; The 3D virtual audio channel dynamic simulation and rendering module is used to generate and render semi-transparent 3D virtual audio channel animations of the user and the standard template in real time based on the continuous motion trajectory parameters and standard templates. The augmented reality interaction and correction force field visualization module is used to overlay the three-dimensional virtual vocal tract animation on the screen, highlight physiological deviations with a heat map, and dynamically overlay the "correction force line" animation that guides the vocal organs to move to the correct position. A multimodal error correction feedback strategy generation module is used to generate error correction feedback instructions that include visual, auditory and tactile cues based on the deviation vector. The personalized pronunciation profile and learning path planning module is used to record the deviation vector sequence of the user's pronunciation in each attempt, assess the shortcomings in pronunciation ability, and dynamically plan targeted learning content and error correction paths. The closed-loop pronunciation adaptation and calibration module is used to record the user's response to error correction feedback commands and the correction effect, and dynamically adjust the inverse parameters of the differentiable pronunciation physiological model and the local weights of the standard template.

[0008] As a further aspect of the present invention, the differentiable articulation physiological model in the differentiable articulation physiological motion reverse deduction module is a geometric and biomechanical model of the articulation organs, including the tongue, lips, mandible, soft palate, and glottis, constructed based on MRI and EMA data. The reverse deduction process adopts the gradient descent method, which is realized by the difference between the acoustic features generated by the acoustic feature decoupling extraction module and the model's forward synthesized speech. Finally, the coordinates, angles, and opening / closing parameters of each organ are continuously output.

[0009] As a further aspect of the present invention, the dynamic time warping performed by the spatiotemporal registration module for the vocal organ movement trajectory adopts a dynamic programming algorithm based on multidimensional acceleration features to eliminate the temporal misalignment between the user's pronunciation and the standard template in terms of speech rate and rhythm; the nonlinear spatial registration adopts a thin-plate spline function to deform the user's organ movement trajectory to the spatial coordinate system of the standard template.

[0010] As a further aspect of the present invention, the deviation vector output by the multi-dimensional physiological deviation quantification analysis module specifically includes the displacement vector of each vocal organ in three-dimensional space, the motion velocity deviation vector, and the coupling coordination deviation value between vocal organs.

[0011] As a further aspect of the present invention, the standard Cantonese pronunciation dynamic template library management module stores standard templates marked with the International Phonetic Alphabet and the Cantonese Pinyin scheme, covering all initials, finals, tones and tone sandhi scenarios. Each standard template consists of time-series three-dimensional organ motion trajectory data and multimodal acoustic features.

[0012] As a further aspect of the present invention, in the augmented reality interaction and correction force field visualization module, the "correction force line" is a Bezier curve with dynamic flowing particles and directional arrows pointing from the user's current incorrect organ position to the correct position of the standard template. Its color and thickness change dynamically with the magnitude of the deviation vector. The greater the deviation, the redder the force line, the thicker the line, and the faster the dynamic particle flow speed.

[0013] As a further aspect of the present invention, the auditory prompts generated by the multimodal error correction feedback strategy generation module are pre-recorded targeted pronunciation demonstrations, the demonstration text of which is automatically generated based on the deviation vector; the tactile prompts generate vibrations of specific frequencies and patterns through wearable devices to prompt the user that the tongue position is too high or the glottal tension is abnormal.

[0014] As a further aspect of the present invention, in the closed-loop pronunciation adaptability and calibration module, dynamically adjusting the local weight of the standard template specifically involves: for specific phonemes or specific physiological deviations that the user repeatedly encounters and finds difficult to correct, the system reduces the confidence weight of the corresponding standard template segment and integrates the successful pronunciation trajectories after multiple corrections by the user to generate a personalized and achievable "progressive target template".

[0015] As a further aspect of the present invention, a virtual organ dynamic simulation configuration module is also included. This module allows creators of instructional content to manually adjust the initial posture, range of motion, and biomechanical parameters of any vocal organ in the virtual vocal tract to simulate various vocal defects or pathological states, and to generate training cases for specific needs.

[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention reverse-engineers organ movement trajectories using a differentiable pronunciation physiological model, combines spatiotemporal registration and multidimensional deviation quantification, and utilizes augmented reality to display semi-transparent 3D vocal tract animation and dynamic "correction force lines" to transform abstract pronunciation errors into visualized organ-level guidance, solving the technical problem of "knowing the mistake but not knowing how to correct it." Simultaneously, by generating progressive target templates through closed-loop adaptive calibration, it effectively reduces learning difficulty and frustration, significantly improving the accuracy and learning efficiency of Cantonese pronunciation correction.

[0017] 2. This invention includes a virtual organ dynamic simulation configuration module, allowing creators of teaching content to manually adjust the initial posture, range of motion, and biomechanical parameters of any vocal organ within the virtual vocal tract. This simulates vocal tract movement under various vocal defects or pathological conditions, thereby generating training cases tailored to specific needs. This design not only serves regular Cantonese learning but can also be applied to scenarios such as speech pathology rehabilitation, dialect comparison teaching, and speech art creation, greatly expanding the system's application flexibility and scope of applicability. Attached Figure Description

[0018] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof.

[0019] Figure 1 This is a system block diagram of a Cantonese pronunciation correction system based on AI speech recognition proposed in this invention. Detailed Implementation

[0020] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0021] Example 1 Reference Figure 1 A Cantonese pronunciation error correction system based on AI speech recognition includes a multi-channel speech signal acquisition and front-end enhancement module, an acoustic feature decoupling extraction module based on self-supervised learning, a differentiable pronunciation physiological motion inverse estimation module, a spatiotemporal registration module for the motion trajectory of the speech organs, a standard Cantonese pronunciation dynamic template library management module, a multi-dimensional physiological deviation quantitative analysis module, a three-dimensional virtual duct dynamic simulation and rendering module, an augmented reality interaction and correction force field visualization module, a multimodal error correction feedback strategy generation module, a personalized pronunciation profile and learning path planning module, and a closed-loop pronunciation adaptability and calibration module.

[0022] The system operation is first performed by the multi-channel speech signal acquisition and front-end enhancement module. In practical applications, this system can be deployed on an augmented reality headset with an integrated microphone. When a user wears the augmented reality headset and starts a Cantonese learning application, the module begins acquiring the speech signal of the user reading a specified Cantonese word at a sampling rate of 16kHz. The module integrates an adaptive filtering algorithm that can eliminate steady-state and non-steady-state noise in the environment in real time, and uses beamforming technology to enhance sound sources from the user's direction while suppressing interference from other directions. After processing by this module, a clean, high signal-to-noise ratio single-channel or dual-channel speech stream is obtained.

[0023] The self-supervised learning-based acoustic feature decoupling extraction module receives the aforementioned clean, high signal-to-noise ratio single-channel or dual-channel speech stream. At the core of this module is a deep neural network pre-trained on a large-scale multi-speaker Cantonese corpus, based on contrastive learning and a Transformer architecture. This deep neural network is designed to decompose the speech signal into two independent latent representation vectors: a "content representation" vector, primarily encoding pronunciation content information related to phonemes and tones; and a "speaker representation" vector, encoding redundant information unrelated to individual timbre, emotion, and the physiological movements of pronunciation.

[0024] The differentiable phonation physiological motion reverse engineering module starts working after obtaining the decoupled content acoustic features. This module internally contains a "differentiable phonation physiological model" built based on the finite element method and a large amount of real nuclear magnetic resonance and electromagnetic phonation instrument data. This model comprises three main components: an organ parameterized model defining the parameters of the tongue, lips, mandible, soft palate, and glottis; a forward acoustic simulator that calculates the vocal tract shape and solves for sound waves; and an optimization solver.

[0025] The specific reverse engineering process is as follows: The optimization solver first initializes a set of organ control parameters. These parameters are then fed into a forward acoustic simulator to calculate a synthesized acoustic feature. This synthesized feature is compared with the "content representation" extracted from the user's speech, i.e., a loss function is calculated. The optimization solver uses automatic differentiation to efficiently calculate the gradient of the loss function with respect to each organ control parameter, i.e., how the organ parameters need to be fine-tuned to make the synthesized acoustic feature closer to the user's real features. Subsequently, the optimization solver updates the organ parameters once using an optimization algorithm based on the gradient information. This process is repeated dozens or even hundreds of times at millisecond speeds until the difference between the synthesized acoustic feature and the user's target feature is less than a preset threshold. Finally, this iterative convergence process outputs a high-dimensional time-series vector, which precisely describes the continuous changes over time of the parameters of the user's seven key points on the tongue, the convexity coefficients of the upper and lower lips, the elevation angle of the soft palate, and the tightness of the glottis during this pronunciation process.

[0026] The vocal organ movement trajectory spatiotemporal registration module receives the continuously changing trajectory data of the aforementioned users and retrieves the corresponding standard templates for phonemes from the standard Cantonese pronunciation dynamic template library management module. The standard Cantonese pronunciation dynamic template library stores standard templates annotated with the International Phonetic Alphabet and the Cantonese Pinyin scheme, covering all initials, finals, tones, and tone sandhi scenarios. Each standard template is accompanied by detailed timeline annotations.

[0027] Registration is performed in two steps: The first step is temporal registration. Since the user's pronunciation speed cannot be exactly the same as the standard template, the module uses an algorithm based on dynamic time warping. It compares not only position but also higher-order features such as motion acceleration to find the optimal alignment path between the two time series, thus accurately mapping the user's "slow tongue forward movement" to the "rapid tongue forward movement" in the standard template on the time axis. The second step is spatial registration. Because everyone's oral cavity size and tongue shape differ, the module uses a thin-plate spline function algorithm to gently "stretch" and "distort" the user's trajectory nonlinearly, aligning its key anatomical landmarks with the spatial coordinate system of the standard template. This "physiological difference-free" registration technique ensures that subsequent deviation calculations focus only on the "texture" and "morphology" differences in the pronunciation movements, rather than individual size differences.

[0028] After registration is complete, the multi-dimensional physiological deviation quantification analysis module begins operation. At each time sampling point, this module calculates the difference vector between the registered user organ position and the standard template position. Its output is not a single score or heatmap, but a multi-dimensional physical deviation field. Specifically, this includes: positional deviation, motion velocity deviation, and coupling coordination deviation. These quantified deviation vectors are sent in real time to the 3D virtual channel dynamic simulation and rendering module and the multimodal error correction feedback strategy generation module. Upon receiving the data, the former calls the underlying graphics rendering engine to quickly construct two semi-transparent 3D virtual channel models of different colors: one blue representing the user's current actual pronunciation posture, and the other gold representing the standard pronunciation posture for that phoneme. These two models synchronously "perform" the fine deformations of the tongue, palate, and pharynx structures within the vocal tract in real time.

[0029] The interactive component is handled by the Augmented Reality Interaction and Corrective Force Field Visualization module. When the user views the interface through the augmented reality headset, this module precisely anchors the aforementioned 3D virtual sound channel model around the user's face. At this point, the user can see a "ghostly" blue tongue moving inside their mouth, while a golden "mentor tongue" demonstrates the correct movements.

[0030] Furthermore, the module generates a dynamic "correction force line" based on the deviation vector: a Bezier curve with dynamic flowing particles and directional arrows emanates from the tip of the blue tongue, meandering towards the correct target position of the golden tongue. The color of this force line changes dynamically; the greater the deviation, the redder the force line and the faster the dynamic flowing particles move.

[0031] The multimodal error correction feedback strategy generation module generates supplementary auditory and tactile feedback based on the specific properties of the deviation vector;

[0032] Specifically, when the module detects that the user's tongue position is always too high, it will provide a synthesized voice prompt. At the same time, the wearable device worn by the user will generate a subtle, downward-facing tactile vibration at the corresponding position. If improper opening and closing of the soft palate is detected, resulting in excessive nasalization, the neck strap may generate a pulse-like vibration in the larynx.

[0033] All the data generated by the above process, including the complete deviation vector sequence for each pronunciation, user response time, and correction trajectory, will be recorded in the personalized pronunciation profile and learning path planning module; Specifically, this module constructs a "pronunciation ability feature matrix" for each user. The horizontal axis represents all the initials, finals, and tones in Cantonese, while the vertical axis represents the tongue, lips, and palate articulation organs. By analyzing this matrix, the system can diagnose the user's "weaknesses."

[0034] The closed-loop pronunciation adaptation and calibration module monitors the entire system's "prediction-correction-revision" loop. If the system detects that after multiple force-guided exercises on a certain phoneme, the root mean square error of a certain deviation no longer decreases in five consecutive pronunciations, the system will not further frustrate the user. Instead, it triggers a slowdown mechanism: dynamically reducing the rigor of the standard template and weighting it with the user's most successful pronunciation trajectories to generate a "progressive target template." This new template is positioned between the user's current level and the absolute standard. The force-guided target point will then switch to the new, closer "progressive target." Once the user easily reaches this target, the system gradually increases the weight of the standard template, guiding the user smoothly towards the final standard, thus effectively solving the problems of frustration and bottlenecks in learning.

[0035] Example 2 Reference Figure 1 This invention relates to a Cantonese pronunciation correction system based on AI speech recognition, which integrates a virtual organ dynamic simulation configuration module. This module allows creators of instructional content to manually adjust the initial posture, range of motion, and biomechanical parameters of any vocal organ in the virtual vocal tract through a graphical parameter panel.

[0036] Specifically, a virtual model can be deliberately created where the "tongue tip elevation is limited to 30% of the normal range" to simulate the pathological pronunciation of "lisping" for students and demonstrate its acoustic consequences. Simultaneously, teachers can create novel pronunciation movements not found in Cantonese, allowing students to imitate and learn, for training in creativity and sound effects design. All new models set in the configuration module can be exported as "custom pronunciation characters" and shared with other users via the cloud. This greatly enriches the system's application scenarios, providing a powerful tool for linguistic research, special education, and speech art creation.

Claims

1. A Cantonese pronunciation correction system based on AI speech recognition, characterized in that, It includes a multi-channel speech signal acquisition and front-end enhancement module, an acoustic feature decoupling and extraction module based on self-supervised learning, a differentiable speech physiological motion inverse estimation module, a speech organ motion trajectory spatiotemporal registration module, a standard Cantonese pronunciation dynamic template library management module, a multi-dimensional physiological deviation quantitative analysis module, a three-dimensional virtual tract dynamic simulation and rendering module, an augmented reality interaction and correction force field visualization module, a multimodal error correction feedback strategy generation module, a personalized pronunciation profile and learning path planning module, and a closed-loop pronunciation adaptability and calibration module; The multi-channel voice signal acquisition and front-end enhancement module is used to acquire Cantonese voice signals emitted by users and perform noise reduction, echo cancellation and sound source enhancement processing. The self-supervised learning-based acoustic feature decoupling extraction module is used to extract decoupled acoustic features related to the pronunciation content from the enhanced speech signal. A differentiable phonation physiological motion reverse deduction module, with a built-in differentiable phonation physiological model, is used to reverse deduce the continuous motion trajectory parameters of the phonation organs in real time based on the acoustic characteristics. The spatiotemporal registration module for the motion trajectory of the vocal organs is used to perform dynamic temporal warping and nonlinear spatial registration between the continuous motion trajectory parameters and the standard templates in the standard Cantonese pronunciation dynamic template library. The multi-dimensional physiological deviation quantification analysis module is used to calculate the deviation vector between the user's actual vocal trajectory and the standard template at each time slice and each vocal organ after nonlinear spatial registration; The three-dimensional virtual audio channel dynamic simulation and rendering module is used to generate and render semi-transparent three-dimensional virtual audio channel animations of the user and the standard template in real time based on the continuous motion trajectory parameters and standard templates. The augmented reality interaction and correction force field visualization module is used to overlay the three-dimensional virtual vocal tract animation on the screen, highlight physiological deviations with a heat map, and dynamically overlay the "correction force line" animation that guides the vocal organs to move to the correct position. A multimodal error correction feedback strategy generation module is used to generate error correction feedback instructions that include visual, auditory and tactile cues based on the deviation vector. The personalized pronunciation profile and learning path planning module is used to record the deviation vector sequence of the user's pronunciation in each attempt, assess the shortcomings in pronunciation ability, and dynamically plan targeted learning content and error correction paths. The closed-loop pronunciation adaptation and calibration module is used to record the user's response to error correction feedback commands and the correction effect, and dynamically adjust the inverse parameters of the differentiable pronunciation physiological model and the local weights of the standard template.

2. The Cantonese pronunciation correction system based on AI speech recognition according to claim 1, characterized in that, The differentiable articulation physiological model in the differentiable articulation physiological motion reverse inference module is a geometric and biomechanical model of the articulation organs, including the tongue, lips, mandible, soft palate, and glottis, constructed based on MRI and EMA data. The reverse inference process adopts the gradient descent method, which is realized by the difference between the acoustic features generated by the acoustic feature decoupling extraction module and the model's forward synthesized speech. Finally, the coordinates, angles, and opening parameters of each organ are continuously output.

3. The Cantonese pronunciation correction system based on AI speech recognition according to claim 1, characterized in that, The dynamic time warping performed by the spatiotemporal registration module for the vocal organ movement trajectory adopts a dynamic programming algorithm based on multidimensional acceleration features to eliminate the temporal misalignment between the user's pronunciation and the standard template in terms of speech rate and rhythm; the nonlinear spatial registration adopts a thin plate spline function to deform the user's organ movement trajectory to the spatial coordinate system of the standard template.

4. The Cantonese pronunciation correction system based on AI speech recognition according to claim 1, characterized in that, The deviation vector output by the multi-dimensional physiological deviation quantification analysis module specifically includes the displacement vector, motion velocity deviation vector, and coupling coordination deviation value of each vocal organ in three-dimensional space.

5. A Cantonese pronunciation correction system based on AI speech recognition according to claim 1, characterized in that, The standard Cantonese pronunciation dynamic template library management module stores standard templates marked with the International Phonetic Alphabet and the Cantonese Pinyin scheme, covering all initials, finals, tones, and tone sandhi scenarios. Each standard template consists of time-series three-dimensional organ motion trajectory data and multimodal acoustic features.

6. A Cantonese pronunciation correction system based on AI speech recognition according to claim 1, characterized in that, In the augmented reality interaction and correction force field visualization module, the "correction force line" is a Bezier curve with dynamic flowing particles and directional arrows that points from the user's current incorrect organ position to the correct position of the standard template. Its color and thickness change dynamically with the magnitude of the deviation vector. The greater the deviation, the redder the force line, the thicker the line, and the faster the dynamic particle flow speed.

7. A Cantonese pronunciation correction system based on AI speech recognition according to claim 1, characterized in that, The auditory cues generated by the multimodal error correction feedback strategy generation module are pre-recorded targeted pronunciation demonstrations, the demonstration text of which is automatically generated based on the deviation vector; the tactile cues generate vibrations of specific frequencies and patterns through wearable devices to alert the user to excessive tongue position or abnormal glottal tension.

8. A Cantonese pronunciation correction system based on AI speech recognition according to claim 1, characterized in that, In the closed-loop pronunciation adaptation and calibration module, the local weight of the standard template is dynamically adjusted as follows: for specific phonemes or specific physiological deviations that the user repeatedly encounters and finds difficult to correct, the system reduces the confidence weight of the corresponding standard template segment and integrates the successful pronunciation trajectories after multiple corrections by the user to generate a personalized and achievable "progressive target template".

9. A Cantonese pronunciation correction system based on AI speech recognition according to claim 1, characterized in that, It also includes a virtual organ dynamic simulation configuration module, which allows creators of instructional content to manually adjust the initial posture, range of motion, and biomechanical parameters of any vocal organ in the virtual vocal tract to simulate various vocal defects or pathological states, and to generate training cases for specific needs.