Voice auditory language training method and system based on artificial intelligence
Through the artificial intelligence system, personalized evaluation and analysis of the trainee's voice data is carried out, which solves the problems of the inability to make personalized adjustments and the influence of noise in traditional language training methods, realizes personalized training plans and real-time feedback, and improves the training effect and convenience.
Patent Information
- Application Number
- CN202510991891.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-07-18
AI Technical Summary
Traditional language training methods are unable to adjust training plans for individual needs, lack automated diagnosis and feedback, make it difficult to detect pronunciation problems, environmental noise affects training results, and lack personalized correction strategies.
The artificial intelligence system obtains the trainee's voice data, conducts personalized evaluation and analysis, locates key pronunciation nodes, generates personalized training plans, combines noise suppression technology to ensure data purity, and provides real-time feedback and adjustments.
It realizes the generation of personalized training plans, accurately diagnoses pronunciation disorders, improves training effects, ensures data purity, dynamically adjusts training difficulty, and improves training convenience and participation.
Smart Images

Figure CN120766718A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of language training technology, and in particular to an artificial intelligence-based sound auditory language training method and system. Background Art
[0002] Traditional methods usually adopt a unified training program, which cannot be adjusted individually according to the pronunciation characteristics of each trainee. The training content and difficulty are often fixed and cannot be fine-tuned according to the specific pronunciation quality differences of the trainees. Therefore, it may not meet the needs of different trainees. In addition, traditional methods rely more on manual evaluation or simplified algorithms to judge the pronunciation quality, lacking sufficient data analysis and automated diagnosis capabilities. The trainee's pronunciation problems may not be discovered and located in a timely and accurate manner, resulting in the problem being ignored or misdiagnosed. For example, it is impossible to deeply analyze the specific aspects of pronunciation and ignore potential physiological or organic pronunciation disorders. In addition, traditional training methods usually lack continuous evaluation and real-time feedback, and the training progress and difficulty adjustment are relatively fixed. When the trainee When progress is made at a certain stage, the training difficulty will not be automatically increased, which may cause the progress of training to stagnate. In addition, the adjustment of the training plan lacks specificity, making it difficult to achieve truly personalized training. Moreover, traditional methods often cannot effectively suppress environmental noise, resulting in the inability to ensure the purity of the data. External noise and other interference factors may affect the quality of voice data, thereby reducing the training effect. Traditional methods are usually relatively simple in processing noise and reverberation, and fail to achieve the level of optimizing training data. In addition, traditional methods cannot analyze the key nodes in pronunciation in an automated way, lack analysis of pronunciation abnormalities and detection of organic disorders, and trainees may not be able to obtain accurate pronunciation disorder diagnosis reports in a timely manner, nor do they have targeted training plans to correct these problems. Summary of the Invention
[0003] To achieve the above objectives, the present invention provides the following technical solution: a method for training speech and hearing based on artificial intelligence, comprising:
[0004] Obtaining a preset language training program for a trainee, controlling a voice acquisition device to perform pre-training based on the preset training program, and obtaining actual voice data generated by the trainee during the pre-training process;
[0005] Determining a preset voice model that the trainee should achieve during the training process, evaluating the trainee's pronunciation quality based on the actual voice data and the preset voice model, and generating a first evaluation result or a second evaluation result;
[0006] If the result is the first evaluation, it means that the trainee's pronunciation meets the preset standard and can enter the next stage of intensive training;
[0007] If the result is the second evaluation, a correlation analysis is performed on the trainee's pronunciation links to locate the key pronunciation nodes that cause the pronunciation quality to be substandard;
[0008] Performing pronunciation abnormality analysis on the key pronunciation nodes, and generating a pronunciation disorder diagnosis report if the posterior probability that the key pronunciation nodes have an organic pronunciation disorder is greater than a preset probability threshold;
[0009] If the posterior probability that the key pronunciation node has an organic pronunciation disorder is not greater than a preset probability threshold, a personalized training adjustment plan is generated and sent to the trainee's terminal device.
[0010] Preferably, evaluating the trainee's pronunciation quality based on the actual voice data and the preset voice model to generate a first evaluation result or a second evaluation result includes:
[0011] Mapping the actual speech data and the preset speech model into the vector space to obtain a speech feature vector space;
[0012] Extracting voiceprint feature parameters of the actual speech data and reference feature parameters of the preset speech model according to the speech feature vector space, wherein the voiceprint feature parameters include fundamental frequency, formant frequency, speech rate, and sound intensity;
[0013] Calculating the difference between the characteristic parameters of the actual voice data and the preset voice model on the same pronunciation unit based on the voiceprint characteristic parameters and the reference characteristic parameters;
[0014] The difference of the characteristic parameters is calculated by using a dynamic time warping algorithm to calculate the pronunciation timing offset and a cosine similarity algorithm to calculate the voiceprint feature similarity;
[0015] The feature difference is compared with a preset difference threshold. If the feature difference is less than the preset difference threshold, a first evaluation result is generated; if the feature difference is not less than the preset difference threshold, a second evaluation result is generated.
[0016] Preferably, if the second evaluation result is obtained, a correlation analysis is performed on the trainee's pronunciation links to locate the key pronunciation nodes that cause the pronunciation quality to be substandard, including:
[0017] If the result is the second evaluation, extracting pronunciation units whose characteristic difference exceeds a preset difference threshold in the actual speech data, and marking the pronunciation units as abnormal pronunciation segments;
[0018] Acquire physiological parameters of the trainee during the pronunciation process, wherein the physiological parameters include tongue position coordinates, lip opening and closing degree, and airflow pressure.
[0019] Preferably, if the result is the second evaluation, then correlation analysis is performed on the trainee's pronunciation links to locate key pronunciation nodes that cause the pronunciation quality to be substandard, further comprising:
[0020] Performing time-series alignment on the abnormal pronunciation segment and the physiological parameter, and analyzing the abnormal value of the physiological parameter corresponding to the abnormal pronunciation segment;
[0021] Constructing a vocal organ motion model, inputting the abnormal value of the physiological parameter into the vocal organ motion model, and determining the deviation of each vocal organ motion trajectory from the standard trajectory;
[0022] The pronunciation organ movement nodes whose deviation is greater than a preset deviation threshold are marked as key pronunciation nodes that cause the pronunciation quality to be substandard.
[0023] Preferably, the key pronunciation nodes are analyzed for pronunciation anomalies. If the posterior probability of the key pronunciation nodes having an organic pronunciation disorder is greater than a preset probability threshold, a pronunciation disorder diagnosis report is generated, including:
[0024] Acquiring historical pronunciation data of the trainee; extracting abnormal pronunciation pattern features from the historical pronunciation data; wherein the abnormal pronunciation pattern features include deviation and feature difference;
[0025] Constructing a neural network-based pronunciation disorder diagnosis model, wherein the input parameters of the pronunciation disorder diagnosis model include abnormal pronunciation pattern characteristics;
[0026] Inputting the abnormal pronunciation pattern features corresponding to the key pronunciation nodes into the pronunciation disorder diagnosis model, and outputting the posterior probability that the key pronunciation nodes have organic pronunciation disorders;
[0027] If the posterior probability is greater than a preset probability threshold, a pronunciation disorder diagnosis report is generated; wherein, the pronunciation disorder diagnosis report includes the pronunciation organs corresponding to the key pronunciation nodes.
[0028] Preferably, if the posterior probability that the key pronunciation node has an organic dysphonia is not greater than a preset probability threshold, generating a personalized training adjustment plan and sending the personalized training adjustment plan to the trainee's terminal device, including:
[0029] Constructing a knowledge base including a pronunciation correction case library and training parameter optimization rules, wherein the case library includes correction strategies corresponding to different types of pronunciation defects;
[0030] Extracting abnormal pattern features of voiceprint characteristics of key pronunciation nodes, and generating feature coding labels based on the abnormal pattern features of voiceprint characteristics;
[0031] Matching correction strategies and training parameter adjustment rules are retrieved from the knowledge base based on the feature encoding labels.
[0032] Preferably, if the posterior probability that the key pronunciation node has an organic dysphonia is not greater than a preset probability threshold, generating a personalized training adjustment plan and sending the personalized training adjustment plan to the trainee's terminal device, further comprising:
[0033] Generating a personalized training adjustment plan including pronunciation action guidance, number of repetitions and training duration suggestions based on the correction strategy and the training parameter adjustment rules;
[0034] The personalized training adjustment plan is pushed to the trainee through a mobile terminal application and is synchronously updated to the control module of the voice training device.
[0035] Preferably, controlling the voice acquisition device to perform pre-training based on the preset training scheme and obtaining actual voice data generated by the trainee during the pre-training process includes:
[0036] The original speech signal of the trainee during the pre-training process is collected by a speech collection device, and the original speech signal is pre-emphasized, framed and windowed;
[0037] A deep learning noise reduction model is used to suppress background noise and eliminate reverberation on the processed original speech signal to obtain a pure speech signal;
[0038] Performing endpoint detection and voice activity detection on the clean speech signal to segment out valid pronunciation segments;
[0039] The effective pronunciation segment is stored as a structured voice data file, and a spectrogram and voiceprint waveform of the actual voice data are generated by a voice visualization tool.
[0040] Preferably, the method further comprises:
[0041] Controlling the voice training device to execute a dynamic training mode based on the personalized training adjustment scheme;
[0042] When the first evaluation result appears three times in a row, the training stage promotion mechanism is triggered and a higher difficulty training module is entered.
[0043] The artificial intelligence-based sound, hearing and language training system is applicable to the artificial intelligence-based sound, hearing and language training method described above, including:
[0044] A voice collection unit, the voice collection unit is used to obtain a preset language training program of the trainee, control the voice collection device to perform pre-training based on the preset training program, and obtain actual voice data generated by the trainee during the pre-training process;
[0045] a pronunciation evaluation unit, the pronunciation evaluation unit being configured to determine a preset speech model that the trainee should achieve during the training process, evaluate the trainee's pronunciation quality based on the actual speech data and the preset speech model, and generate a first evaluation result or a second evaluation result;
[0046] An intensive training unit, wherein if the first evaluation result is obtained, it indicates that the trainee's pronunciation meets the preset standard and the trainee can enter the next stage of intensive training;
[0047] a pronunciation locating unit configured to, if the result is the second evaluation, perform a correlation analysis on the trainee's pronunciation links to locate key pronunciation nodes that cause the pronunciation quality to be substandard;
[0048] an obstacle diagnosis unit configured to perform an anomaly analysis on the key pronunciation node and generate an anomaly diagnosis report if the posterior probability that the key pronunciation node has an organic anomaly is greater than a preset probability threshold;
[0049] A training adjustment unit is used to generate a personalized training adjustment plan if the posterior probability of the existence of an organic pronunciation disorder in the key pronunciation node is not greater than a preset probability threshold, and send the personalized training adjustment plan to the trainer's terminal device.
[0050] Compared with the prior art, the present invention has the following beneficial effects:
[0051] (1) The present invention can generate a personalized training adjustment plan for each trainee by obtaining the trainee's actual voice data and comparing it with the preset voice model. This plan not only takes into account the pronunciation characteristics of each trainee, but also can be accurately optimized according to the differences in their pronunciation quality, thereby enhancing the pertinence of the training. If an abnormality is found in the pronunciation quality assessment, further pronunciation abnormality analysis can be performed to locate specific key pronunciation nodes, and possible pronunciation disorders can be analyzed based on physiological parameters. With the assistance of the neural network model, it can automatically diagnose whether there is an organic pronunciation disorder and provide the trainee with a professional pronunciation disorder diagnosis report.
[0052] (2) The present invention combines speech acquisition, noise suppression, dynamic time warping and cosine similarity technologies to make speech data processing more accurate, better avoid environmental noise and other interference factors, ensure the purity of data during training, and thus improve the quality of training; and if the trainee reaches the preset standard at a certain stage, he or she can enter a more difficult training module to gradually improve his or her pronunciation ability. For unsatisfactory pronunciation, the problem can be accurately found through the correlation analysis of the pronunciation link and the pronunciation organ movement model, and targeted training can be implemented;
[0053] (3) Through continuous evaluation and feedback, the present invention can dynamically adjust the training plan according to the trainee's progress. For example, if the trainee performs well at a certain stage, the training difficulty will be automatically increased to ensure that the training does not stagnate. Moreover, through mobile terminal applications, personalized training plans can be pushed to the trainee's device in real time, greatly improving the convenience and real-time nature of training, while enhancing the trainee's sense of participation and accomplishment. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 A schematic diagram of the steps of the overall method in one embodiment of the present invention;
[0055] Figure 2 FIG. 1 is a schematic diagram of the system architecture structure of the overall system in one embodiment of the present invention.
[0056] In the figure: 1. Voice collection unit; 2. Pronunciation evaluation unit; 3. Reinforcement training unit; 4. Pronunciation positioning unit; 5. Obstacle diagnosis unit; 6. Training adjustment unit. DETAILED DESCRIPTION
[0057] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0058] For example 1, please refer to Figure 1 The present invention provides a technical solution: a sound auditory language training method based on artificial intelligence, comprising:
[0059] S1. Obtaining a preset language training plan for the trainee, controlling a voice acquisition device to perform pre-training based on the preset training plan, and obtaining actual voice data generated by the trainee during the pre-training process;
[0060] S2. Determine a preset speech model that the trainee should achieve during the training process, evaluate the trainee's pronunciation quality based on the actual speech data and the preset speech model, and generate a first evaluation result or a second evaluation result;
[0061] S3. If the result is the first evaluation, it means that the trainee's pronunciation meets the preset standard and can enter the next stage of intensive training;
[0062] S4. If the result is the second evaluation, then a correlation analysis is performed on the trainee's pronunciation links to locate the key pronunciation nodes that cause the pronunciation quality to be substandard;
[0063] S5. Analyze the pronunciation abnormalities of the key pronunciation nodes. If the posterior probability of an organic pronunciation disorder in the key pronunciation nodes is greater than a preset probability threshold, generate a pronunciation disorder diagnosis report.
[0064] S6. If the posterior probability that the key pronunciation node has an organic pronunciation disorder is not greater than a preset probability threshold, a personalized training adjustment plan is generated and sent to the trainee's terminal device.
[0065] It should be noted that the trainee is undergoing speech training with the aim of improving his pronunciation, especially the problems that may exist on certain specific syllables; assuming that the trainee's goal is to correctly pronounce the two sounds sh and ch; S1: Obtain the trainee's preset language training program and conduct pre-training. In this stage, the trainee first sets his training goal with the system, such as learning how to correctly pronounce sh and ch; Based on this goal, a voice acquisition device (such as a microphone) is controlled to record his voice; In the pre-training stage, the trainee starts to pronounce, and the system collects the actual voice data generated when he pronounces; S2: Determine the preset voice model to be achieved and perform pronunciation quality evaluation. The system has A preset standard model, which defines the correct way to pronounce sh and ch (for example, the coordination of lips, tongue and breathing); at this stage, the system will compare the trainee's actual pronunciation with the preset standard to evaluate the quality of his pronunciation; the system will give two evaluation results: the first evaluation result: the trainee's pronunciation is consistent with the standard; the second evaluation result: there is a gap between the trainee's pronunciation and the standard; S3: If it is the first evaluation result; if the system's evaluation result is the first evaluation result, it means that the trainee's pronunciation has met the standard, and he can enter the next stage of intensive training to continue practicing and consolidating the correct pronunciation skills; S4: If it is the second evaluation result; if the system's evaluation result is the first evaluation result, it means that the trainee's pronunciation has met the standard, and he can enter the next stage of intensive training to continue practicing and consolidating the correct pronunciation skills; If it is the second evaluation result, it means that the trainee's pronunciation still has problems; the system will analyze which pronunciation links have problems (such as unclear pronunciation of sh); it will find specific key pronunciation nodes through correlation analysis, for example, the trainee's tongue position may be incorrect when pronouncing sh; S5: Pronunciation abnormality analysis and pronunciation disorder diagnosis; Next, the system will conduct a more in-depth analysis of these key pronunciation nodes; for example, the system will detect whether there is an organic pronunciation disorder (such as tongue movement problems or oral structure problems); If the analysis results show that the trainee's pronunciation problem is due to some physiological reasons, and the probability of this problem occurring exceeds the preset threshold (for example, a probability of 70%), the system A pronunciation disorder diagnosis report will be generated to inform the trainee that he may have a pronunciation disorder and recommend further medical examination or treatment; S6: Personalized training adjustment plan; if the analysis results show that the trainee's pronunciation problem does not meet the standards of an organic disorder, the system will generate a personalized training adjustment plan based on the specific problem; for example, if the system finds that the trainee has a problem with the tongue position when pronouncing sh, the system may recommend that he pay attention to the movement of the tongue when pronouncing, or provide specific pronunciation training exercises (such as strengthening the control of the tongue position by imitating the correct pronunciation); this personalized plan will be sent to the trainee through a terminal device (such as a mobile phone, computer, etc.) to help him improve his pronunciation during training.
[0066] In an optional embodiment, the trainee's pronunciation quality is evaluated based on the actual speech data and the preset speech model to generate a first evaluation result or a second evaluation result, including:
[0067] Mapping the actual speech data and the preset speech model into a vector space to obtain a speech feature vector space;
[0068] Extracting voiceprint feature parameters of actual speech data and baseline feature parameters of a preset speech model based on the speech feature vector space, wherein the voiceprint feature parameters include fundamental frequency, formant frequency, speech rate, and sound intensity;
[0069] Calculate the difference between the characteristic parameters of the actual voice data and the preset voice model on the same pronunciation unit based on the voiceprint characteristic parameters and the reference characteristic parameters;
[0070] Among them, the difference of feature parameters is calculated by dynamic time warping algorithm to calculate the pronunciation timing offset and the voiceprint feature similarity is calculated by cosine similarity algorithm;
[0071] The feature difference is compared with a preset difference threshold. If the feature difference is less than the preset difference threshold, a first evaluation result is generated; if the feature difference is not less than the preset difference threshold, a second evaluation result is generated.
[0072] It should be noted that the actual voice data and the information of the preset voice model are mapped into a vector space; in this vector space, each voice signal will be converted into a vector to represent the different features of the voice; this vector space will include various parameters related to the sound features, which will help the system compare the similarity between voices; from the voice data mapped to the vector space, the system will extract a set of key voiceprint feature parameters and the baseline feature parameters of the preset voice model; the following are the specific contents of these features: Pitch: refers to the basic frequency of the sound, which determines the pitch of the voice; Formant Frequencies: This is an important audio component in the voice, which is related to the sound quality and accent of the voice; Speech Speed: Rate): refers to the speed of pronunciation; Loudness: refers to the loudness of speech; compare the difference in feature parameters between the actual speech data and the preset speech model on the same pronunciation unit; the pronunciation unit here can be understood as a syllable or word in the language; the purpose of calculating these differences is to evaluate the deviation between the trainer's pronunciation and the standard pronunciation; the dynamic time warping algorithm is used to compare the similarity between two time series, and is usually used in fields such as speech and handwriting recognition; here, the DTW algorithm will be used to calculate the offset of the pronunciation timing, that is, the time difference between the trainer's pronunciation and the standard pronunciation; for example, the trainer's pronunciation may be slightly faster or slower than the standard pronunciation, and the DTW algorithm can adjust the time alignment of the two sequences to calculate the time deviation; the cosine similarity algorithm is used to calculate the difference between two vectors The cosine similarity will be used to calculate the similarity of the voiceprint features, that is, to compare the similarity between the trainee's voiceprint features and the standard voiceprint features; by calculating the angle between the feature vectors, the similarity between the two can be obtained, and the smaller the angle, the higher the similarity between the two; the calculated feature difference is compared with the preset difference threshold to determine whether the trainee's pronunciation meets the standard: if the feature difference is less than the preset difference threshold, it means that the trainee's pronunciation is very close to the standard pronunciation, and the system will generate a first evaluation result, indicating that the trainee's pronunciation meets the standard and can enter the next stage of training; if the feature difference is not less than the preset difference threshold, it means that there is a large gap in the trainee's pronunciation, and the system will generate a second evaluation result, which means that the trainee's pronunciation needs further improvement, and the system may provide a personalized adjustment plan.
[0073] In an optional embodiment, if the second evaluation result is obtained, a correlation analysis is performed on the trainee's pronunciation links to locate the key pronunciation nodes that cause the pronunciation quality to be substandard, including:
[0074] If the result is the second evaluation, the pronunciation units whose characteristic difference exceeds the preset difference threshold in the actual speech data are extracted and marked as abnormal pronunciation segments;
[0075] The physiological parameters of the trainees during the pronunciation process are obtained, including tongue position coordinates, lip opening and closing degree, and airflow pressure.
[0076] It should be noted that.
[0077] In an optional embodiment, if the second evaluation result is obtained, correlation analysis is performed on the trainee's pronunciation links to locate key pronunciation nodes that cause the pronunciation quality to be substandard, further comprising:
[0078] Align the abnormal pronunciation segments with the physiological parameters in time sequence and analyze the abnormal physiological parameter values corresponding to the abnormal pronunciation segments;
[0079] Constructing a vocal organ motion model, inputting abnormal physiological parameter values into the vocal organ motion model, and determining the deviation of each vocal organ motion trajectory from the standard trajectory;
[0080] The articulatory movement nodes whose deviation is greater than a preset deviation threshold are marked as key articulatory nodes that cause the pronunciation quality to be substandard.
[0081] It should be noted that when the feature difference is calculated and the second evaluation result is obtained, it means that the trainee's pronunciation is quite different from the standard pronunciation; at this time, these pronunciation units will be further analyzed, and the pronunciation units with feature differences exceeding the preset threshold will be extracted, and these will be marked as abnormal pronunciation segments; these segments are parts where there are obvious problems in the trainee's pronunciation, which may be due to factors such as inaccurate pronunciation, inconsistent intonation or speaking speed; the trainee's physiological parameters during the pronunciation process are recorded and obtained, and these parameters are used to further analyze the potential causes of pronunciation quality; specific physiological parameters include: tongue position coordinates: the position of the tongue during the pronunciation process; the position of the tongue is crucial to the clarity and accuracy of pronunciation, and tongue position errors may lead to some inaccurate pronunciations; lip opening: the degree of opening and closing of the lips, which affects the clarity of pronunciation; for example, if the lips are not completely closed during pronunciation, it may cause unclear or unclear pronunciation; airflow pressure: the size and pressure of the airflow, which usually affects the vibration and pronunciation of the vocal cords strength; unstable or uneven airflow may cause discontinuous sound or non-standard pronunciation; time-align the abnormal pronunciation segments with the corresponding physiological parameters, that is, synchronize these physiological parameters with the specific pronunciation moment in order to analyze the relationship between them; time-alignment ensures that the physiological changes of each pronunciation unit during pronunciation can match the specific speech characteristics, so as to analyze the source of the pronunciation quality problem; for example, the characteristic difference of a certain pronunciation unit is large, which may correspond to physiological abnormalities such as tongue position coordinate offset, insufficient lip opening or unstable airflow pressure; through time-alignment, it is possible to accurately identify which physiological parameters have abnormalities; after time-alignment, the abnormal values of the physiological parameters corresponding to the abnormal pronunciation segments will be specifically analyzed; these abnormal values indicate the deviation of the trainee's physiological parameters from the standard state during normal pronunciation; for example, tongue position deviation from the normal range, insufficient lip closure or abnormal airflow pressure will lead to distorted or unclear pronunciation;
[0082] In order to further analyze the specific impact of abnormal physiological parameters on the quality of pronunciation, a movement model of the pronunciation organs will be constructed; this model usually includes a description of the movement trajectory of the pronunciation organs (such as tongue, lips, vocal cords, etc.); by simulating and analyzing these movement trajectories, it can be determined whether the trainee's pronunciation organs move according to the standard trajectory; the standard trajectory refers to the ideal movement path of the pronunciation organs (such as tongue, lips, etc.) during normal and accurate pronunciation; after the movement model of the pronunciation organs is constructed, the abnormal values of the physiological parameters are input into this model to analyze the deviation of the movement trajectory of each pronunciation organ from the standard trajectory; the deviation refers to the deviation of the movement path of each pronunciation organ from the standard path during the pronunciation process of the trainee. The degree of difference; the movement of the vocal organs with a large deviation may be the root cause of inaccurate pronunciation; the greater the deviation, the more non-standard the trainee's vocal organ movement, which may lead to a decline in pronunciation quality; finally, the vocal organ movement nodes with a deviation greater than the preset deviation threshold are marked as key pronunciation nodes; these key pronunciation nodes are the most important moments in the pronunciation process, and the deviation of the vocal organs is most serious at these nodes, which directly affects the quality of the entire pronunciation; if the deviation of a certain vocal organ at a certain pronunciation node is greater than the threshold, this node is considered to be the key node that causes substandard pronunciation; these nodes require special attention, and the trainer needs to adjust the pronunciation method of these nodes.
[0083] In an optional embodiment, a pronunciation anomaly analysis is performed on key pronunciation nodes. If the posterior probability of an organic pronunciation disorder in the key pronunciation node is greater than a preset probability threshold, a pronunciation disorder diagnosis report is generated, including:
[0084] Acquiring historical pronunciation data of a trainee; extracting abnormal pronunciation pattern features from the historical pronunciation data; wherein the abnormal pronunciation pattern features include deviation and feature difference;
[0085] Constructing a neural network-based pronunciation disorder diagnosis model, wherein the input parameters of the pronunciation disorder diagnosis model include abnormal pronunciation pattern characteristics;
[0086] Input the abnormal pronunciation pattern features corresponding to the key pronunciation nodes into the pronunciation disorder diagnosis model, and output the posterior probability that the key pronunciation nodes have organic pronunciation disorders;
[0087] If the posterior probability is greater than the preset probability threshold, a pronunciation disorder diagnosis report is generated; wherein the pronunciation disorder diagnosis report includes the pronunciation organs corresponding to the key pronunciation nodes.
[0088] It should be noted that the trainee's past pronunciation data are collected; these data may come from recordings of the trainee's multiple pronunciations, pronunciation tests or other pronunciation evaluation processes; historical pronunciation data provides the trainee's long-term pronunciation characteristics, and serves as basic data for subsequent analysis; in the historical pronunciation data, abnormal pronunciation pattern features are extracted, which reflect possible problems in the trainee's pronunciation process, usually identified by analyzing pronunciation deviations or changes; abnormal pronunciation pattern features include the following two types: Deviation: refers to the degree of deviation between the trainee's pronunciation organs (such as tongue position, lip opening, airflow, etc.) and the standard pronunciation organ movement trajectory; the greater the deviation, the greater the degree of deviation, the greater the deviation of the pronunciation organ movement during the pronunciation process deviates from the standard path. The higher the degree, the more likely it is to indicate inaccurate or impaired pronunciation; Feature difference: This refers to the degree of difference between the trainee's pronunciation and the standard pronunciation, including differences in voice quality, pitch, timbre, etc.; If these differences exceed the range of normal fluctuations, it may indicate abnormal pronunciation; After extracting the abnormal pronunciation pattern features, the system will build a neural network model based on these features for pronunciation disorder diagnosis; The input of this model is the abnormal pronunciation pattern features. By learning these features, the neural network model can be trained to identify pronunciation disorders; The model may consider the following aspects: Input features: abnormal pronunciation features such as deviation and feature difference; Network structure: The neural network may include multiple layers, such as convolutional neural networks (CNN), long short-term memory network (LSTM), etc., are used to identify complex time series patterns; target output: the goal of the model is to diagnose pronunciation disorders, to determine whether the trainee's pronunciation has any obstacles, and whether it is an organic pronunciation disorder; then the abnormal pronunciation pattern features corresponding to the key pronunciation nodes (that is, the nodes that are abnormally significant during the pronunciation process) are input into the trained pronunciation disorder diagnosis model; key pronunciation nodes refer to the moments when the movement or characteristics of the pronunciation organs are seriously abnormal during the pronunciation process, and these nodes are crucial for diagnosing pronunciation disorders; by inputting the features of these nodes into the model, the system can further analyze whether the node shows signs of organic pronunciation disorders; through the neural network After processing the network model, the system will output a posterior probability for each key pronunciation node, that is, the probability that the key pronunciation node has an organic pronunciation disorder; the posterior probability is based on the model's training on historical data, combined with the analysis results of the current pronunciation pattern characteristics, to estimate the possibility of whether the node has an organic pronunciation disorder; if this probability is higher than the set threshold (usually a standard value derived from expert experience or data analysis), it can be considered that the node may have a pronunciation disorder; if the posterior probability is lower than the preset threshold, the pronunciation performance of the node is considered to be normal; if the posterior probability of the key pronunciation node exceeds the set probability threshold, that is, the model diagnoses the possibility of an organic pronunciation disorder, and the system will generate a pronunciation disorder diagnosis report;The report includes the following: Key articulatory nodes: The report will indicate which articulatory nodes are performing abnormally and how these nodes deviate from the normal pronunciation trajectory; Articulatory organs: The report will also list the articulatory organs involved, that is, which articulatory organs (such as tongue, lips, vocal cords, etc.) have problems at these nodes; this helps to understand which part of the articulatory organs has deviated from the normal trajectory, causing pronunciation quality problems.
[0089] In an optional embodiment, if the posterior probability of an organic dysphonia existing in a key pronunciation node is not greater than a preset probability threshold, a personalized training adjustment plan is generated and sent to a terminal device of the trainee, including:
[0090] Build a knowledge base that includes a pronunciation correction case library and training parameter optimization rules. The case library includes correction strategies corresponding to different types of pronunciation defects.
[0091] Extract the abnormal pattern features of the voiceprint characteristics of key pronunciation nodes, and generate feature coding labels based on the abnormal pattern features of the voiceprint characteristics;
[0092] Matching correction strategies and training parameter adjustment rules are retrieved from the knowledge base based on the feature encoding labels.
[0093] It should be noted that the pronunciation correction case library stores different types of pronunciation defects and their corresponding correction strategies; these defects may be tongue malposition, incorrect pronunciation position, insufficient airflow control, oral structure problems, etc.; for each pronunciation defect, the knowledge base has pre-stored effective correction strategies, such as: tongue position error: can be corrected through tongue training, imitation training, etc.; high or low pitch: can be adjusted by adjusting the pitch and sound quality of the voice; oral structure problems: if the pronunciation problem is caused by a structural disorder, physical correction or surgery of the pronunciation organs can be performed; these correction strategies are usually based on professional linguistic research, clinical experience and the latest technology of pronunciation correction, aiming to provide targeted treatment for each pronunciation defect; training The parameter optimization rules store the parameter adjustment rules related to the training and treatment process; these rules are used to optimize the pronunciation practice parameters during the training process; for example: training intensity (such as the frequency, duration, and volume of pronunciation training); attention to sound features (such as pitch, sound quality, pronunciation position, etc.); the way of simulating feedback (such as the feedback frequency and difficulty setting of the speech recognition system); during the pronunciation training or evaluation process, the system will analyze the voiceprint characteristics of the trainee; voiceprint characteristics refer to the pronunciation sound wave characteristics of each person, including multi-dimensional information such as timbre, tone, pitch, speaking speed, and speech duration; key pronunciation nodes refer to the time nodes that show significant abnormalities in the pronunciation process; for example, when pronouncing a specific syllable, the pronunciation organs There are deviations from standard movements, or changes in sound exceed the normal range; these nodes can usually reflect the main problems in the trainee's pronunciation process; the abnormal pattern characteristics of voiceprint features refer to the abnormal voiceprint features corresponding to these nodes; for example: the pronunciation pitch is too high or too low; the pronunciation timbre is distorted; the pronunciation speed is uneven; by analyzing the voiceprint features of pronunciation, the trainee's abnormal pronunciation patterns at these key pronunciation nodes can be identified and their corresponding features can be extracted; once the abnormal patterns of voiceprint features are extracted, the next step is to convert these features into a feature coding label; the label is a symbol or digital representation that can encode these abnormal patterns for subsequent processing; the generation of feature coding labels is a feature compression process that converts the original The complex voiceprint features are converted into a more concise form; these labels usually reflect information such as the type, severity, and frequency of pronunciation defects; for example, a trainer may generate the following labels: Label 1: The pitch is too low and the duration is too long; Label 2: The tone is distorted and the position of the pronunciation organs is deviated; Label 3: Insufficient airflow and unclear pronunciation; These labels can quantify the abnormal characteristics of the pronunciation process, which is convenient for subsequent diagnosis and treatment; Use the generated feature encoding labels to match them in the pre-built pronunciation correction case library; Based on the labels, the system will retrieve the relevant correction strategies and training parameter optimization rules; According to the labels, the relevant pronunciation defect types are matched, and then the corresponding correction strategies are retrieved from the case library;For example, if the label indicates that the pitch is too low, the system may recommend adjusting the frequency range of pronunciation or provide specific pronunciation exercises to help the trainee adjust the pitch; adjust the training parameters based on the pronunciation problem described in the label; if the label indicates insufficient airflow, the system may automatically increase the intensity of pronunciation training or adjust the volume and syllable duration.
[0094] In an optional embodiment, if the posterior probability of the presence of an organic dysphonia in a key pronunciation node is not greater than a preset probability threshold, generating a personalized training adjustment plan and sending the personalized training adjustment plan to the trainee's terminal device further includes:
[0095] Generate personalized training adjustment plans based on correction strategies and training parameter adjustment rules, including pronunciation action guidance, repetition times, and training duration recommendations;
[0096] The personalized training adjustment plan is pushed to the trainee through the mobile terminal application and synchronously updated to the control module of the voice training device.
[0097] It should be noted that.
[0098] In an optional embodiment, controlling the voice acquisition device to perform pre-training based on a preset training scheme and obtaining actual voice data generated by the trainee during the pre-training process includes:
[0099] The original speech signals of the trainees during the pre-training process are collected by the speech collection equipment, and the original speech signals are pre-emphasized, framed and windowed;
[0100] A deep learning noise reduction model is used to suppress background noise and eliminate reverberation on the processed original speech signal to obtain a pure speech signal;
[0101] Perform endpoint detection and voice activity detection on clean speech signals to segment valid pronunciation segments;
[0102] The valid pronunciation segments are stored as structured speech data files, and the spectrogram and voiceprint waveform of the actual speech data are generated through speech visualization tools.
[0103] It should be noted that pronunciation action guidance refers to detailed pronunciation action training guidance; for example, the system will provide specific suggestions on how to adjust the tongue position, airflow control, oral shape, and the pronunciation of syllables. For example, if the trainee has a tongue position deviation when pronouncing, the system may provide the following guidance: slightly raise the tongue to touch the gum position of the upper palate, or keep the tip of the tongue not touching the hard palate when pronouncing. These are all detailed actions for pronunciation correction; the number of repetitions refers to how many times each training action or pronunciation exercise needs to be performed; based on the severity of the pronunciation defect and individual differences, the system can automatically set the number of repetitions for each action or exercise; for example, for a difficult syllable to pronounce, the system may recommend repeating it 100 times a day to deepen the muscles. The system can also be used to adjust the training time, so as to improve the coordination of muscle memory and oral movements; the training time recommendation refers to the length of time for each training session or each time; the personalized training plan will set a reasonable training time according to the trainee's ability and needs to ensure that the training is effective without excessive fatigue; for example, if the trainee's speech problem is mild, the system may recommend 15 minutes of training; if the problem is more serious, the system may recommend 30 minutes to 1 hour of training, divided into multiple sessions; after the personalized training adjustment plan is generated, the next step is to push it to the trainee's device through the mobile terminal application; the key points of this process are: push the personalized training adjustment plan to the trainee through mobile devices such as smartphones, tablets or smart watches; when the personalized training adjustment plan is generated, the system will immediately This information will be pushed to the trainee to ensure that the trainee can train on time; the application will display the guidance of training movements through a graphical interface, and can provide voice, text or video feedback to ensure that the trainee can clearly understand each movement and operation step; the application can set training reminders to remind the trainee to train at regular intervals to ensure that the training process is not interrupted and the training plan is fully implemented; for example, the application may prompt: Today's pronunciation training begins, repeat the pronunciation of "shi" 10 times for 10 minutes; in addition to mobile terminals, the training system also includes voice training equipment, which may include voice training tools with microphones and speakers, or smart headphones, voice feedback devices, etc.; voice training equipment The control module refers to the core hardware and software part of the device, which controls the operation, feedback and adjustment of the device. The system will synchronize and update the personalized training adjustment plan to the control module of the device, including: Training plan synchronization: the personalized pronunciation training plan (such as pronunciation movements, number of repetitions and training duration, etc.) is transmitted to the voice training device; the control module sets the working mode of the voice training device according to these plans; for example, the device may set the frequency, pitch and other parameters of each pronunciation according to the synchronized plan; the control module of the training device can provide real-time feedback based on the trainee's pronunciation performance; if the trainee's pronunciation does not meet expectations, the device will provide corrective suggestions through acoustic feedback or image feedback (such as display screen or voice prompts);For example, the device can display tongue deviations during pronunciation or provide audio prompts to slightly raise the pitch. The control module records training progress and provides feedback to the system and the trainer. For example, the system tracks the duration of each training session and the number of movement repetitions completed, and continuously optimizes the training plan based on this data. The device adjusts training intensity based on real-time feedback to ensure the personalized nature of the training plan is implemented. For example, if the trainee performs well on a certain movement, the device may reduce the number of repetitions of that movement and add other more challenging exercises.
[0104] In an optional embodiment, the method further includes:
[0105] Controlling the voice training device to execute a dynamic training mode based on a personalized training adjustment plan;
[0106] When the first evaluation result appears three times in a row, the training stage promotion mechanism is triggered and a more difficult training module is entered.
[0107] For example 2, please refer to Figure 2 The present invention provides a technical solution: an artificial intelligence-based sound, auditory and language training system, which is applicable to the above-mentioned artificial intelligence-based sound, auditory and language training method, including:
[0108] The voice collection unit 1 is used to obtain the trainee's preset language training plan, control the voice collection device to perform pre-training based on the preset training plan, and obtain the actual voice data generated by the trainee during the pre-training process;
[0109] The pronunciation evaluation unit 2 is used to determine the preset speech model that the trainee should achieve during the training process, evaluate the trainee's pronunciation quality based on the actual speech data and the preset speech model, and generate a first evaluation result or a second evaluation result;
[0110] Intensive training unit 3, if the first evaluation result is obtained, it means that the trainee's pronunciation meets the preset standard and can enter the next stage of intensive training;
[0111] The pronunciation positioning unit 4 is used to perform a correlation analysis on the trainee's pronunciation links if the second evaluation result is obtained, and locate the key pronunciation nodes that cause the pronunciation quality to be substandard;
[0112] The disorder diagnosis unit 5 is used to analyze the pronunciation abnormalities of the key pronunciation nodes and generate a pronunciation disorder diagnosis report if the posterior probability of the key pronunciation node having an organic pronunciation disorder is greater than a preset probability threshold;
[0113] The training adjustment unit 6 is used to generate a personalized training adjustment plan if the posterior probability of the existence of an organic pronunciation disorder in a key pronunciation node is not greater than a preset probability threshold, and send the personalized training adjustment plan to the trainee's terminal device.
[0114] The embodiments of the present invention are described in detail above with reference to the accompanying drawings, but the present invention is not limited thereto. Various changes can be made within the scope of knowledge possessed by those skilled in the art without departing from the spirit of the present invention.
Claims
1. The method of sound and auditory language training based on artificial intelligence is characterized by ,include: Obtaining a preset language training program for a trainee, controlling a voice acquisition device to perform pre-training based on the preset training program, and obtaining actual voice data generated by the trainee during the pre-training process; Determining a preset voice model that the trainee should achieve during the training process, evaluating the trainee's pronunciation quality based on the actual voice data and the preset voice model, and generating a first evaluation result or a second evaluation result; If the result is the first evaluation, it means that the trainee's pronunciation meets the preset standard and can enter the next stage of intensive training; If the result is the second evaluation, a correlation analysis is performed on the trainee's pronunciation links to locate the key pronunciation nodes that cause the pronunciation quality to be substandard; Performing pronunciation abnormality analysis on the key pronunciation nodes, and generating a pronunciation disorder diagnosis report if the posterior probability that the key pronunciation nodes have an organic pronunciation disorder is greater than a preset probability threshold; If the posterior probability that the key pronunciation node has an organic pronunciation disorder is not greater than a preset probability threshold, a personalized training adjustment plan is generated and sent to the trainee's terminal device.
2. The method for sound and auditory language training based on artificial intelligence according to claim 1, characterized in that: Evaluating the pronunciation quality of the trainee according to the actual voice data and the preset voice model to generate a first evaluation result or a second evaluation result, including: Mapping the actual speech data and the preset speech model into the vector space to obtain a speech feature vector space; Extracting voiceprint feature parameters of the actual speech data and reference feature parameters of the preset speech model according to the speech feature vector space, wherein the voiceprint feature parameters include fundamental frequency, formant frequency, speech rate, and sound intensity; Calculating the difference between the characteristic parameters of the actual voice data and the preset voice model on the same pronunciation unit based on the voiceprint characteristic parameters and the reference characteristic parameters; The difference of the characteristic parameters is calculated by using a dynamic time warping algorithm to calculate the pronunciation timing offset and a cosine similarity algorithm to calculate the voiceprint feature similarity; The feature difference is compared with a preset difference threshold. If the feature difference is less than the preset difference threshold, a first evaluation result is generated; if the feature difference is not less than the preset difference threshold, a second evaluation result is generated.
3. The method for sound and auditory language training based on artificial intelligence according to claim 2, characterized in that: If the second evaluation result is obtained, a correlation analysis is performed on the trainee's pronunciation links to locate the key pronunciation nodes that cause the pronunciation quality to be substandard, including: If the result is the second evaluation, extracting pronunciation units whose characteristic difference exceeds a preset difference threshold in the actual speech data, and marking the pronunciation units as abnormal pronunciation segments; Acquire physiological parameters of the trainee during the pronunciation process, wherein the physiological parameters include tongue position coordinates, lip opening and closing degree, and airflow pressure.
4. The method for sound and auditory language training based on artificial intelligence according to claim 3, characterized in that: If the second evaluation result is obtained, a correlation analysis is performed on the trainee's pronunciation links to locate the key pronunciation nodes that cause the pronunciation quality to be substandard, further comprising: Performing time-series alignment on the abnormal pronunciation segment and the physiological parameter, and analyzing the abnormal value of the physiological parameter corresponding to the abnormal pronunciation segment; Constructing a vocal organ motion model, inputting the abnormal value of the physiological parameter into the vocal organ motion model, and determining the deviation of each vocal organ motion trajectory from the standard trajectory; The pronunciation organ movement nodes whose deviation is greater than a preset deviation threshold are marked as key pronunciation nodes that cause the pronunciation quality to be substandard.
5. The method for sound and auditory language training based on artificial intelligence according to claim 4, characterized in that: Performing pronunciation abnormality analysis on the key pronunciation nodes, and generating a pronunciation disorder diagnosis report if the posterior probability of the key pronunciation nodes having an organic pronunciation disorder is greater than a preset probability threshold, including: Acquiring historical pronunciation data of the trainee; extracting abnormal pronunciation pattern features from the historical pronunciation data; wherein the abnormal pronunciation pattern features include deviation and feature difference; Constructing a neural network-based pronunciation disorder diagnosis model, wherein the input parameters of the pronunciation disorder diagnosis model include abnormal pronunciation pattern characteristics; Inputting the abnormal pronunciation pattern features corresponding to the key pronunciation nodes into the pronunciation disorder diagnosis model, and outputting the posterior probability that the key pronunciation nodes have organic pronunciation disorders; If the posterior probability is greater than a preset probability threshold, a pronunciation disorder diagnosis report is generated; wherein, the pronunciation disorder diagnosis report includes the pronunciation organs corresponding to the key pronunciation nodes.
6. The method for sound and auditory language training based on artificial intelligence according to claim 5, characterized in that: If the posterior probability that the key pronunciation node has an organic dysphonia is not greater than a preset probability threshold, generating a personalized training adjustment plan and sending the personalized training adjustment plan to the trainee's terminal device, including: Constructing a knowledge base including a pronunciation correction case library and training parameter optimization rules, wherein the case library includes correction strategies corresponding to different types of pronunciation defects; Extracting abnormal pattern features of voiceprint characteristics of key pronunciation nodes, and generating feature coding labels based on the abnormal pattern features of voiceprint characteristics; Matching correction strategies and training parameter adjustment rules are retrieved from the knowledge base based on the feature encoding labels.
7. The method for sound and auditory language training based on artificial intelligence according to claim 6, characterized in that: If the posterior probability that the key pronunciation node has an organic dysphonia is not greater than a preset probability threshold, generating a personalized training adjustment plan and sending the personalized training adjustment plan to the trainee's terminal device, further comprising: Generating a personalized training adjustment plan including pronunciation action guidance, number of repetitions and training duration suggestions based on the correction strategy and the training parameter adjustment rules; The personalized training adjustment plan is pushed to the trainee through a mobile terminal application and is synchronously updated to the control module of the voice training device.
8. The method for sound and auditory language training based on artificial intelligence according to claim 7, characterized in that: Controlling the voice acquisition device to perform pre-training based on the preset training scheme and obtaining actual voice data generated by the trainee during the pre-training process includes: The original speech signal of the trainee during the pre-training process is collected by a speech collection device, and the original speech signal is pre-emphasized, framed and windowed; A deep learning noise reduction model is used to suppress background noise and eliminate reverberation on the processed original speech signal to obtain a pure speech signal; Performing endpoint detection and voice activity detection on the clean speech signal to segment out valid pronunciation segments; The effective pronunciation segment is stored as a structured voice data file, and a spectrogram and voiceprint waveform of the actual voice data are generated by a voice visualization tool.
9. The method for sound and auditory language training based on artificial intelligence according to claim 8, characterized in that: The method further comprises: Controlling the voice training device to execute a dynamic training mode based on the personalized training adjustment scheme; When the first evaluation result appears three times in a row, the training stage promotion mechanism is triggered and a higher difficulty training module is entered.
10. An artificial intelligence-based sound, auditory, and language training system, which is applicable to the artificial intelligence-based sound, auditory, and language training method according to any one of claims 1 to 9, characterized in that: include: A voice collection unit (1), the voice collection unit (1) is used to obtain a preset language training program of a trainee, control a voice collection device to perform pre-training based on the preset training program, and obtain actual voice data generated by the trainee during the pre-training process; A pronunciation evaluation unit (2), the pronunciation evaluation unit (2) is used to determine a preset speech model that the trainee should achieve during the training process, evaluate the trainee's pronunciation quality based on the actual speech data and the preset speech model, and generate a first evaluation result or a second evaluation result; An intensive training unit (3), wherein if the first evaluation result is obtained, it indicates that the trainee's pronunciation meets the preset standard and the trainee can enter the next stage of intensive training; A pronunciation positioning unit (4), wherein if the result is the second evaluation result, the pronunciation positioning unit (4) is used to perform a correlation analysis on the pronunciation links of the trainee and locate the key pronunciation nodes that cause the pronunciation quality to be substandard; An obstacle diagnosis unit (5) is used to perform pronunciation abnormality analysis on the key pronunciation node, and generate a pronunciation disorder diagnosis report if the posterior probability that the key pronunciation node has an organic pronunciation disorder is greater than a preset probability threshold; A training adjustment unit (6) is used to generate a personalized training adjustment plan if the posterior probability of the existence of an organic pronunciation disorder in the key pronunciation node is not greater than a preset probability threshold, and send the personalized training adjustment plan to the trainer's terminal device.
Citation Information
Patent Citations
Phoneme-level low-power consumption spoken language assessment and defect diagnosis method
CN103985392A
Personalized spoken foreign language learning system and method
CN105654785A
Language training method for simulating pronunciation and air flow change based on 3D tongue position model
CN113327483A
Intelligent pronunciation correction method and system thereof
CN113658584A
Training device and equipment for shaping language fluency of children and storage medium
CN115662242A
Cited By
Intelligent evaluation and error correction guidance method for pronunciation mode of English pronunciation
CN121583286A