Speech rehabilitation training and dynamic feedback system and method based on tone-gesture mapping

The speech rehabilitation training system based on tone-gesture mapping solves the problem of hearing-impaired patients being unable to accurately perceive tone differences, realizes personalized and dynamic feedback guidance, and improves the tone pronunciation ability and speech expression ability of hearing-impaired patients.

CN120656637APending Publication Date: 2025-09-16HANGZHOU NORMAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510766283.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-03-06
Filing Date
2025-06-10
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Hearing-impaired patients are unable to accurately perceive tone differences through auditory feedback. Existing gesture-assisted systems lack dynamic visualization, personalized, and timely guidance of tone features. There is also a lack of specialized and real-time quantitative evaluation and feedback methods for Chinese tone training.

Method used

A speech rehabilitation training system based on tone-gesture mapping is adopted. Through the audio recognition module, tone standard deviation analysis module and gesture encoding module, combined with dynamic time warping algorithm and root mean square error calculation, the alignment and deviation analysis of the user's tone fundamental frequency curve and the standard tone fundamental frequency curve are achieved, providing personalized and dynamic feedback guidance.

Benefits of technology

It achieves efficient and accurate training of the tone pronunciation ability of hearing-impaired patients, provides real-time and visual feedback, and improves their verbal expression and social communication abilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120656637A_ABST
    Figure CN120656637A_ABST
Patent Text Reader

Abstract

The invention relates to a speech rehabilitation training and dynamic feedback system and method based on tone-gesture mapping, and the system comprises a processor, a display screen, an audio input device, an audio output device, an audio recognition module, a tone standard deviation analysis module, a detection result output module, and a gesture coding module. The audio recognition module comprises a target word access module, an audio acquisition module and an audio analysis module, and the target word access module is used for storing and calling standard tones of target words; the audio acquisition module acquires user audio information; the audio analysis module is used for extracting acoustic features of a user and aligning a user tone with a standard tone by using a dynamic time warping algorithm; the tone standard degree deviation analysis module performs normalization processing, and calculates the deviation distance between the user tone and the standard tone, the standard deviation degree and the user pronunciation accuracy; and the gesture coding module is used for gesture guidance action display regulation and control. The method is high in efficiency, high in adaptability and easy to identify and expand.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to a speech rehabilitation training and dynamic feedback system and method based on tone-gesture mapping, which is mainly suitable for speech and meaning communication between hearing-impaired patients or between hearing-impaired patients and ordinary people, as well as rehabilitation training for hearing-impaired patients. Background Art

[0002] Chinese is a typical tonal language. The differences in the four tones (level, yang, rising, and departing) have a significant impact on semantic differences. For example, the differences between "bā" (eight), "bá" (pull), "bǎ" (put), and "bà" (dad) are entirely dependent on the tones. Unlike "oral expression" in language, speech needs to be "clear, fluent, and rhythmically pronounced," and the correct mastery of tones is an important part of this.

[0003] Relevant research shows that, based on the principle of mental anchoring, in language teaching, gestures are actually the inducement between the psychological state and pronunciation behavior during pronunciation. Fixed, systematic, and regular gesture reinforcement can have a positive effect on language learners' pronunciation ability. Specifically, it has the same effect on intonation.

[0004] With the development of rehabilitation technology and the popularization of hearing aids (such as cochlear implants), the demands of hearing-impaired patients for speech training have shifted from "basic speech recognition" to "precise tone mastery" to improve speech naturalness and social communication skills.

[0005] However, existing speech rehabilitation technologies mainly rely on auditory feedback, which has limited effects on hearing-impaired patients (such as hearing-impaired children and cochlear implant users); existing Chinese tone rehabilitation training mainly relies on the therapist's auditory judgment, which is highly subjective and inefficient; some systems attempt to use gesture-assisted training, but they mostly focus on imitating the mouth shape or pronunciation part, and lack quantitative correlation with Chinese tone characteristics. For existing tone recognition systems, the pronunciation feedback and auxiliary correction for hearing-impaired patients are very limited, and cannot adapt to the speech rehabilitation needs of hearing-impaired patients. According to statistics, the majority of patients with speech disorders caused by hearing impairment have abnormal tone problems, and mastering the correct tone pronunciation is an important part of better mastering Mandarin Chinese.

[0006] Traditional speech rehabilitation technology is mainly designed for people with normal hearing, with auditory feedback as the core, and does not fully consider the perceptual needs of hearing-impaired patients. In addition, other multimodal feedback devices require complex sensors, such as tactile feedback - wearable devices, which have certain requirements on hardware costs and are difficult to popularize and apply.

[0007] The dynamic changes in the fundamental frequency (F0) of Chinese tones (such as the rising slope of Yangping 35) require high-precision algorithms. Early technologies, lacking open-source tools, struggled to implement real-time analysis, preventing real-time feedback during rehabilitation training and resulting in low rehabilitation efficiency. Furthermore, existing gesture-assisted systems are mostly empirically designed (e.g., mimicking lip shape) and lack a quantitative relationship between tone characteristics and gesture movement parameters (such as speed and trajectory angle). Furthermore, interdisciplinary research (phonetics and kinematics) is lacking, making it difficult for hearing-impaired patients to accurately express their training needs and for others to understand their intentions. Summary of the Invention

[0008] The technical problem solved by this application is to overcome the above-mentioned deficiencies in the prior art and provide a speech rehabilitation training and dynamic feedback system based on tone-gesture mapping and a method for using the same, which solves the following technical problems: 1. Hearing-impaired patients cannot accurately perceive tone differences through auditory feedback; 2. Existing gesture assistance systems lack dynamic visualization of tone features, as well as personalized and timely guidance. 3. There is a lack of specialized and real-time quantitative evaluation and feedback methods for Chinese tone training.

[0009] The technical solution adopted by the present application to solve the above technical problems includes: a speech rehabilitation training and dynamic feedback system based on tone-gesture mapping, including a processor, a display screen, an audio input device, and an audio output device. The processor is connected to the display screen, the audio input device, and the audio output device. It is characterized in that it is also provided with an audio recognition module, a tone standard deviation analysis module, a detection result output module, and a gesture encoding module. The audio recognition module includes a target word access module, an audio acquisition module, and an audio analysis module. The target word access module is used for storing and calling the standard tone fundamental frequency data of the target word; The audio acquisition module collects valid user audio information through the audio input device and transmits it to the audio analysis module together with the standard audio information of the current target word based on the five-degree marking method in the target word access module; the audio analysis module obtains the user's tone fundamental frequency curve through user acoustic feature extraction, aligns the voice signal duration of the user's tone fundamental frequency with the standard tone fundamental frequency using the dynamic time warping algorithm, and then transmits it to the tone standard deviation analysis module and the detection result output module; the tone standard deviation analysis module normalizes the user's tone fundamental frequency data using the T value method, and the normalization formula is as follows: T=(Log x-Log b) / (Log a-Log b)×5, where x is the current fundamental frequency, a is the maximum value of all fundamental frequencies of the user, and b is the minimum value of all fundamental frequencies of the user. The deviation distance between the user's tone fundamental frequency curve and the standard tone fundamental frequency curve is calculated by the root mean square error calculation method, and then the standard deviation degree between the user's tone fundamental frequency and the standard tone fundamental frequency expressed as a percentage is calculated and sent to the gesture encoding module. Then, the user's pronunciation accuracy is calculated and sent to the detection result output module; the gesture encoding module is used to display and regulate the gesture guidance action according to the user's standard deviation degree. The gesture guidance action is set according to the gesture action coding rules and displayed in the form of a digital human; the processor is connected to the audio recognition module, the tone standard deviation analysis module, the detection result output module, and the gesture encoding module and controls them to operate according to the process. The processor is provided with a target word standard tone fundamental frequency database and a training result database.

[0010] The present application is also provided with a user information entry or selection button, a target word selection or addition button, a pronunciation demonstration button, a gesture action demonstration button, a continue training button, an exit current training button, another target word training button, and an end training button. The user information entry or selection button is used for confirmation after selecting or entering user information; the target word selection or addition button is used to call the target word access module to select or add target words; the pronunciation demonstration button is used to play the pronunciation demonstration audio once; the gesture action demonstration button is used to play the gesture action demonstration video once; the continue training button is used for repeated training when the user wants to continue the current target word training; the exit current training button is used to end the current target word training; the another target word training button is used when the current user intends to change a target word for training, and the end training button is used for the user to end the current training and exit.

[0011] The present application may also provide a visual rendering module, which is connected to the processor and is used to store and set the appearance of the digital human.

[0012] The user acoustic feature extraction described in this application adopts the Mel frequency cepstral coefficient method to identify and extract acoustic features.

[0013] The technical solution adopted by the present application to solve the above technical problems also includes: an operating method of the above speech rehabilitation training and dynamic feedback system, which is characterized by comprising the following steps: S1 system initialization and user configuration S11 system initialization; S12 User information entry or selection; S13 target word selection or addition; S2 Pre-training preparation S21 displays the system interface for the user; During the pre-training preparation process, pressing the user training button will enter step S3; S3 User Training S31 user audio collection; S32 user audio analysis; S33 Tone deviation analysis and feedback; S34 Repeat training or make a stage summary: If you want to continue training, press the Continue Training button in step S341, go to step S31, and repeat the above steps S31, S32, and S33; if you want to end the current target word training, press the Exit Current Training button; S342 stage summary; If the user wants to train other target words, press the other target word training button and go to step S13 to continue training; if the user wants to end the training, press the end training button to end the user training and go to step S12.

[0014] The S2 step described in this application also includes an S22 gesture action demonstration step, in which a gesture action demonstration video is played each time the gesture action demonstration button is pressed.

[0015] The S2 step described in this application also includes an S23 pronunciation demonstration step, in which the pronunciation demonstration audio is played each time the pronunciation demonstration button is pressed.

[0016] This application specifically leverages the tonal characteristics of Mandarin Chinese, the tone-gesture anchoring principle, and general tone recognition techniques to efficiently and accurately analyze the tone pronunciation of hearing-impaired patients and provide personalized, dynamic, and timely guidance. This overcomes the limitations of traditional gesture-assisted systems and traditional speech rehabilitation techniques for the hearing-impaired. Furthermore, based on phonetic principles and incorporating kinematic characteristics, it expands the application scope of existing tone training, making it more universal, easier to identify, and more easily scalable, offering both guiding and educational benefits for the rehabilitation training of hearing-impaired patients. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 This is a structural framework diagram of the speech rehabilitation training and dynamic feedback system according to an embodiment of the present application.

[0018] Figure 2 This is the overall idea diagram of the speech rehabilitation training and dynamic feedback system for this application.

[0019] Figure 3 This is a flow chart of the audio acquisition module of the speech rehabilitation training and dynamic feedback system of this application.

[0020] Figure 4 This is a flow chart of the audio analysis module of the speech rehabilitation training and dynamic feedback system of this application.

[0021] Figure 5 This is a flow chart of the tone standard deviation analysis module of the speech rehabilitation training and dynamic feedback system of this application.

[0022] Figure 6 This is a flow chart of the method of using the speech rehabilitation training and dynamic feedback system in this application. DETAILED DESCRIPTION

[0023] The present application will be further described in detail below with reference to the accompanying drawings and examples. The following examples are intended to explain the present application but the present application is not limited to the following examples.

[0024] See also Figures 1 to 6The embodiment of the present application is a speech rehabilitation training and dynamic feedback system for hearing-impaired patients, which mainly includes an audio recognition module for the user's voice (including an audio acquisition module, an audio analysis module, and a target word access module), a tone standard deviation analysis module, a detection result output module, a gesture encoding module, a visual rendering module, and a processor, a display screen, an audio input device (such as a microphone), and an audio output device (such as a speaker) in the prior art. The processor is connected to the audio recognition module, the tone standard deviation analysis module, the detection result output module, the gesture encoding module, and the visual rendering module and controls the operation of these modules. The processor is provided with a target word standard tone fundamental frequency database and a user's training result database, as well as a conventional input keyboard (or a simplified keyboard, at least containing numeric keys and a confirmation key).

[0025] The audio acquisition module within the audio recognition module is responsible for collecting valid audio information, specifically Mandarin Chinese audio information, including the specific tones of individual Mandarin Chinese characters or words. Relevant research indicates that the most important information for tone recognition is changes in sound frequency. The essence of tone recognition is the discrimination of sound frequency, acoustically manifested as changes in the fundamental frequency (F0) and its harmonic components. After effective acoustic analysis and reliability and validity testing, the system can convey speech semantic information and matching user tone information. This audio information is simultaneously sent to the audio analysis module for comparative analysis.

[0026] The audio analysis module analyzes the audio information obtained by the audio acquisition module. This audio information is identified and acoustically characterized using MFCC (Mel-Frequency Cepstral Coefficient) technology. The user's voice pitch (Pitch Fractal Frequency) curve of the real-time collected user audio information is then aligned with the standard pitch using the DTW (Dynamic Time Warping) algorithm. The waveform of the user's pitch pitch is then displayed on the display screen, where it is displayed as a red dynamic curve. The waveform is then sent to the Tone Standardization Deviation Analysis Module for deviation analysis. The display screen simultaneously displays the existing standard pitch pitch curve for the target word (from the target word standard pitch pitch database, shown as a green dashed line) and the real-time collected user's pitch pitch curve (i.e., the user's pitch pitch, shown as a red dynamic curve).

[0027] The target word access module is used to store and call individual characters, words, word pinyin, and standard tone waveforms under standard tones. The standard tone waveform is based on the specific fundamental frequency change pattern of different characters and words. The fundamental frequency curve of the specific character or word is extracted and analyzed using existing technology, and combined with the tone rules of Mandarin Chinese - this application uses the five-degree marking method to form the target word standard tone fundamental frequency data (waveform curve) and save it in the target word standard tone fundamental frequency database of the processor for easy access at any time. The target word standard tone fundamental frequency data is sent to the detection result output module, which presents it to the user in the form of a standard tone waveform diagram. The standard tone waveform diagram is marked with a green dotted line.

[0028] The tone standard deviation analysis module is used to match and analyze the user's tone pitch curve (the user's tone signal and the standard tone signal have been time-aligned using the DTW technique) with the standard tone pitch curve. This matching analysis first normalizes the pitch data using the T-value method, mapping it to the range of 0 and 5 to eliminate differences between speakers and facilitate subsequent analysis. The normalization formula is as follows: T = (Log x - Log b) / (Log a - Log b) × 5, where x is the current pitch of the user, a is the maximum pitch of all the user's pitches, and b is the minimum pitch of all the user's pitches. The deviation distance between the user's tone pitch data and the standard pitch pitch data is then calculated using existing deviation distance calculation methods (such as the Root Mean Square Error (RMSE) algorithm), resulting in quantified data. The quantified result is then converted into a percentage (deviation distance / average pitch of the standard pitch) × 100%. The quantified result is the standard deviation (deviation percentage) between the user's tone fundamental frequency data and the standard tone fundamental frequency data, which can be used to calculate the user's pronunciation accuracy percentage. User pronunciation accuracy = 100 - standard deviation.

[0029] The gesture encoding module is used to display and control the intensity of the gesture guidance action based on the standard deviation. The gesture guidance action intensity is matched to the standard deviation using an existing matching algorithm. The gesture guidance action intensity is expressed as a percentage, with 100% being the standard gesture guidance action intensity.

[0030] The gesture encoding rules are shown in the following table: Gesture action coding rule table Tone type (numbers in brackets are based on fifth notation) Gesture trajectory Movement characteristics (the data in brackets are the initial standard gesture guidance movement intensity, which can be adjusted by the user according to their own situation) Yinping (55) Horizontal straight line Uniform motion (speed 0.4m / s) Hinata (35) 30° diagonal line to the upper right Accelerated motion (0.2 → 0.6 m / s²) rising tone (214) V-shaped track (base angle 60°) First decelerate and then accelerate (0.5 → 0.1 → 0.5 m / s²) Falling tone (51) 45° diagonal line at the lower right Deceleration (0.6 → 0.2 m / s²) The visual rendering module is used to store and set the appearance of the digital human (a virtual human image in existing technology, directly displayed on a display screen), including but not limited to rendering the head, arms, hands, and their motion states, to ensure a natural appearance and natural, clear, and smooth movements. For example, this can be achieved using existing technologies such as anime character design. The user can select the appearance of the digital human according to their needs. The resulting digital human's motion state is controlled by the gesture encoding module and is simultaneously sent to the detection result output module.

[0031] The detection result output module is used to display a comparison chart of the standard tone curve and the user's tone curve collected in real time, the user's pronunciation accuracy and corresponding evaluation, and the digital human's gesture guidance movements. At the same time, the training results are stored in the processor database for subsequent analysis. The result comparison chart refers to a screenshot of the sound wave graph area after the standard tone curve and the user's tone curve are fully aligned. The corresponding evaluation is given according to the following standards: User pronunciation accuracy evaluate 95%-100% excellent 80%-94% good 0%-79% Still need practice The embodiment of the present application constructs a speech rehabilitation training system based on tone-gesture mapping, which helps to visually guide the tone training of hearing-impaired patients and provide real-time standard deviation analysis results, thereby promoting the improvement of their tone pronunciation ability. It is suitable for school-age hearing-impaired children and hearing-impaired adults with speech rehabilitation needs.

[0032] This system is also equipped with some operation (virtual) buttons for the operation control of speech rehabilitation training and dynamic feedback system, including user information entry or selection button, target word selection or addition button, pronunciation demonstration button, gesture demonstration button, continue training button, exit current training button, other target word training button, end training button. The user information entry or selection button is used for confirmation after selecting or entering user information; the target word selection or addition button is used to call the target word access module to select or add target words; the pronunciation demonstration button is used to play the pronunciation demonstration audio once; the gesture demonstration button is used to play the gesture demonstration video once; the continue training button is used for repeated training when the user wants to continue the current target word training; the exit current training button is used to end the current target word training; the other target word training button is used when the current user intends to change a target word for training, and the end training button is used for the user to end the current training and exit.

[0033] The specific implementation points of the embodiments of this application are as follows: 1. System initialization and target word selection and loading Target word access module: Pre-store Chinese characters in standard tones (such as "eight", "pull", "hold", "dad") and their corresponding standard tone waveforms. The standard waveforms are generated based on the five-degree marking method (prior art). For example, the high level tone (55) corresponds to a flat waveform for the fundamental frequency curve, the rising tone (35) corresponds to a rising diagonal waveform, the falling-rising tone (214) corresponds to a V-shaped waveform, and the falling tone (51) corresponds to a falling diagonal waveform. The waveform data is marked with a green dashed line on the display screen. User configuration: According to the rehabilitation needs of the user (such as a hearing-impaired child), select training vocabulary from the target word library and set the initial gesture action intensity (default 0).

[0034] 2. Audio acquisition and analysis Audio acquisition module: Real-time collect the audio signal of the user's pronunciation (such as "eight") through a microphone. After filtering out non-speech noise, transmit the effective audio to the audio analysis module. Acoustic feature extraction: Use MFCC (Mel Frequency Cepstral Coefficients) technology to extract the static feature parameters of the user's audio, and calculate the dynamic feature parameters through differential calculation to form the user's tone fundamental frequency curve (red dynamic curve). Time alignment: Use the DTW (Dynamic Time Warping) algorithm to align the fundamental frequency curve of the user's pronunciation with the standard tone waveform in time to eliminate individual differences in pronunciation duration. 3. Tone deviation analysis and quantification Normalization processing: Map the user's fundamental frequency data to the range of 0 - 5 through the T-value method. The formula is: T = (Log x - Log b) / (Log a - Log b) × 5, where x is the user's current fundamental frequency, a is the maximum value of all the user's fundamental frequencies, and b is the minimum value of all the user's fundamental frequencies.

[0035] Deviation calculation: Use the RMSE (Root Mean Square Error) algorithm to calculate the deviation distance between the user's fundamental frequency curve and the standard curve, and convert it into the standard deviation degree expressed in percentage form (such as a deviation of 15%). 4. Gesture encoding and dynamic feedback Gesture control: Adjust the gesture action intensity according to the percentage of the standard deviation degree. For example, if the standard deviation degree is 15%, then the gesture guidance action intensity is adjusted to 15% (the standard gesture guidance action intensity is consistent with the standard deviation degree, and the variation range is 0 - 100%. When the user's audio is exactly the same as the standard audio, the digital human does not make gesture actions). <​​​​​​​Rising tone (214): V-shaped trajectory (bottom angle 60°), the speed first decelerates to 0.1 m / s² and then accelerates. Falling tone (51): Diagonal deceleration movement at 45° to the lower right (acceleration 0.6 → 0.2 m / s²). 5. Visual Rendering and Result Display Digital human image: The display screen presents a pre-rendered virtual digital human image (such as a cartoon character), and its arm and hand movements are adjusted in real time according to the output of the gesture coding module. Comparison display: The screen displays the standard tone waveform (green dotted line), the user's real-time fundamental frequency curve result (red dynamic curve), and the user's pronunciation accuracy rate (such as "User pronunciation accuracy rate: 85%") in different regions. Gesture guidance: The digital human synchronously demonstrates standard gesture actions (such as the arm accelerating upward when pronouncing the rising tone), and the user can follow the practice. 6. Storage of Training Results Data storage: Each training result (including audio, deviation data, and gesture parameters) is saved to the training result database of the processor for subsequent analysis. Real-time feedback scenario application of the embodiments of this application: During the rehabilitation training process, when the user pronounces "ba" (rising tone 35), the system detects in real time that the rising slope of the fundamental frequency is insufficient (standard deviation degree 20%), then the intensity of the gesture action is 20%, and it prompts "Please increase the tone rising speed". After the user adjusts the pronunciation according to the feedback, the deviation drops to 8%, and the intensity of the gesture action is 8%, and the system gives a "good" evaluation.

[0037] See Figure 6 , the usage method of the speech rehabilitation training and dynamic feedback system of this application is as follows: S1. System Initialization and User Configuration S11 System initialization: Ensure that the hardware devices required by the system (such as processors, microphones, display screens, etc.) are connected normally, and the software system has been installed and completed the initialization settings.

[0038] S12 User information entry or selection: If the user information is not in the system, enter the basic information of the hearing-impaired user (such as age, degree of hearing impairment, rehabilitation training goals, etc.) into the system, and adjust the initial gesture action intensity according to the specific situation of the user (the default is 100%); if the user information has been entered, select the user from the system database, and the user's relevant information is automatically extracted.

[0039] S13 Target word selection or addition: Select or add vocabulary (such as single words, two-word words, etc.) suitable for the user's current training stage from the target word standard tone frequency database through the target word access module. All trained users have saved the target words used in the user's training. You only need to select the target word and do not need to add new target words.

[0040] S2. Preparation before training S21 displays the system interface for the user: displays the system interface containing the user's target words for this training. The rehabilitation therapist or digital human can introduce the system's operating interface to the user, including the standard tone curve display area, real-time fundamental frequency curve display area, gesture guidance action display area, and deviation percentage prompt area, etc., to help the user understand the function and role of each part.

[0041] S22 Gesture Demonstration: A rehabilitation therapist or the system's built-in digital human will demonstrate standard gestures corresponding to various tones (e.g., horizontal straight-line uniform motion for yinping, accelerated motion of a 30° diagonal line in the upper right corner for yangping), and explain the motion characteristics of each gesture and its relationship with the tone, helping the user establish a preliminary understanding of tone-gesture mapping. This step is optional and can be repeated. This application provides a gesture demonstration button, which plays a gesture demonstration video each time it is pressed. S23 Pronunciation Demonstration: The therapist or the system's audio output device plays standard pronunciation audio and displays the corresponding tone curve and hand gestures, allowing the user to listen and observe multiple times to feel the correspondence between tone changes and hand gestures. This step is optional and can be repeated. This application provides a pronunciation demonstration button, and each time the pronunciation demonstration button is pressed, the pronunciation demonstration audio is played once; Pressing the user training button in any of steps S21, S22, and S23 will enter step S3. There is no distinction between steps S22 and S23, and they can be performed simultaneously or one before the other. S3. User Training S31 User Audio Collection: The user attempts to pronounce the target vocabulary according to the system prompts or the therapist's instructions. The audio collection module collects the user's vocal audio signals in real time, filters them, and transmits them to the audio analysis module. S32 User Audio Analysis: The audio analysis module first extracts user acoustic features (MFCC method) to form the user's tone fundamental frequency curve, and then uses the DTW algorithm to time-align it with the standard tone waveform; S33 Tone Deviation Analysis and Feedback: The Tone Standard Deviation Analysis module calculates the deviation between the user's pronunciation and the standard tone and displays it on the screen as a percentage. The Gesture Coding module adjusts the intensity of the gesture guidance action based on the degree of deviation. The digital human avatar simultaneously demonstrates the adjusted gesture action, providing intuitive visual feedback to the user. S34 Repeat training or make a stage summary: If you want to continue training at S341, press the Continue Training button and go to step S31. Repeat the above steps S31, S32, and S33 to gradually approach the standard tone. During the practice, the user can repeat the pronunciation multiple times. The system will update the deviation percentage and gesture intensity in real time to help the user continuously optimize the pronunciation effect. If you want to end the current target word training, press the Exit Current Training button; S342 Periodic Summary: After each vocabulary training session, the system automatically calculates the user's training results, including average deviation percentage, pronunciation accuracy, and other indicators, and generates a training report. The therapist can provide targeted guidance and corrections to the user's pronunciation problems based on the training report, and adjust subsequent training plans and target vocabulary. If the user wants to train other target words, he presses the other target word training button and goes to step S13 to continue training; if he presses the end training button, the user training is ended and goes to step S12.

[0042] 4. Notes Training frequency and duration: It is recommended that users train for a certain amount of time each day, but each session should not be too long to avoid fatigue. The training time and frequency can be reasonably arranged based on the user's age, recovery progress, and concentration level.

[0043] Individual differences: Since each hearing-impaired user has different levels of hearing loss, language foundation, and learning ability, the training process should be personalized according to the user's specific situation to avoid a one-size-fits-all training model.

[0044] Psychological support: During the training process, it is important to provide users with adequate encouragement and support to help them build confidence and overcome psychological barriers. Patiently guide users through difficulties and setbacks that arise during training to avoid anxiety and resistance.

[0045] By using the above system, hearing-impaired users can, with the assistance of the system, gradually master the tone pronunciation rules of Mandarin Chinese, improve their verbal expression ability, and ultimately reach a level where they can express their true meaning more accurately and be understood by others, thereby better integrating into social life and improving their quality of life.

[0046] The main improvements of this application are as follows: First, a tone-gesture dynamic mapping model, which converts tone deviation distance into gesture motion parameters; 2. One-stop recognition and real-time dynamic feedback of tone information for hearing-impaired users; 3. Differentiated training methods based on the degree of tone achievement.

Claims

1. A speech rehabilitation training and dynamic feedback system based on tone-gesture mapping, comprising a processor, a display screen, an audio input device, and an audio output device, wherein the processor is connected to the display screen, the audio input device, and the audio output device, and is characterized in that It is also provided with an audio recognition module, a tone standard deviation analysis module, a detection result output module, and a gesture encoding module. The audio recognition module includes a target word access module, an audio acquisition module, and an audio analysis module. The target word access module is used to store and call the standard tone fundamental frequency data of the target word; the audio acquisition module collects valid user audio information through an audio input device and transmits it to the audio analysis module together with the standard audio information of the current target word based on the five-degree marking method in the target word access module; the audio analysis module obtains the user tone fundamental frequency curve through user acoustic feature extraction, and uses the dynamic time warping algorithm to align the voice signal duration of the user tone fundamental frequency with the standard tone fundamental frequency, and then transmits it to the tone standard deviation analysis module and the detection result output module; the tone standard deviation analysis module normalizes the user tone fundamental frequency data by the T value method, and the normalization formula is as follows: T=(Log x-Log b) / (Log a-Log b)×5, where x is the current fundamental frequency, a is the maximum value of all fundamental frequencies of the user, and b is the minimum value of all fundamental frequencies of the user. The deviation distance between the user's tone fundamental frequency curve and the standard tone fundamental frequency curve is calculated by the root mean square error calculation method, and then the standard deviation degree between the user's tone fundamental frequency and the standard tone fundamental frequency expressed as a percentage is calculated and sent to the gesture encoding module. Then, the user's pronunciation accuracy is calculated and sent to the detection result output module; the gesture encoding module is used to display and control gesture guidance actions according to the user's standard deviation degree. The gesture guidance actions are set according to the gesture action coding rules and displayed in the form of a digital human; the processor is connected to the audio recognition module, the tone standard deviation analysis module, the detection result output module, and the gesture encoding module, and the processor is provided with a target word standard tone fundamental frequency database and a training result database.

2. The speech rehabilitation training and dynamic feedback system based on tone-gesture mapping according to claim 1 is characterized by: It also has a user information input or selection button, a target word selection or addition button, a pronunciation demonstration button, a gesture action demonstration button, a continue training button, an exit current training button, another target word training button, and an end training button. The user information input or selection button is used to confirm after selecting or entering user information; the target word selection or addition button is used to call the target word access module to select or add target words; The pronunciation demonstration button is used to play the pronunciation demonstration audio once; the gesture demonstration button is used to play the gesture demonstration video once; the continue training button is used for repeated training when the user wants to continue the current target word training; the exit current training button is used to end the current target word training; the other target word training button is used when the current user intends to change the target word for training, and the end training button is used for the user to end the current training and exit.

3. The speech rehabilitation training and dynamic feedback system based on tone-gesture mapping according to claim 1 or 2, characterized in that: A visual rendering module is also provided, which is connected to the processor and is used for storing and setting the appearance of the digital human.

4. The method for using the speech rehabilitation training and dynamic feedback system based on tone-gesture mapping according to claim 2, characterized in that The following steps are involved: S1 system initialization and user configuration S11 system initialization; S12 User information input or selection; S13 target word selection or addition; S2 Pre-training preparation S21 displays the system interface for the user; During the pre-training preparation process, pressing the user training button will enter step S3; S3 User Training S31 user audio collection; S32 user audio analysis; S33 Tone deviation analysis and feedback; S34 Repeat training or make a stage summary: If you want to continue training, press the Continue Training button in step S341, go to step S31, and repeat the above steps S31, S32, and S33; if you want to end the current target word training, press the Exit Current Training button; S342 stage summary; If the user wants to train other target words, press the other target word training button and go to step S13 to continue training. If the user wants to end the training, press the end training button to end the user training and go to step S12.

5. The method for using the speech rehabilitation training and dynamic feedback system based on tone-gesture mapping according to claim 4 is characterized by: The step S2 also includes a gesture action demonstration step S22, in which a gesture action demonstration video is played each time the gesture action demonstration button is pressed.

6. The speech rehabilitation training and dynamic feedback system based on tone-gesture mapping according to claim 4 is characterized by: The step S2 also includes a pronunciation demonstration step S23, in which the pronunciation demonstration audio is played each time the pronunciation demonstration button is pressed.