Correction system, method and equipment for oral vocal training
By collecting and analyzing environmental noise parameters in real time, combining speech and visual data, adjusting weights for pronunciation evaluation, the recognition accuracy and evaluation deviation problems of the oral training system in noisy environments are solved, and more efficient training results are achieved.
Patent Information
- Application Number
- CN202510623291.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-08-12
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing oral vocal training system has insufficient accuracy in extremely noisy environments, and pronunciation evaluation relies on a single mode to lead to evaluation bias.
The environmental data analysis module collects noise parameters in real time, combines multimodal data fusion technology, adjusts the voice and visual judgment weights, conducts pronunciation accuracy evaluation, and provides targeted correction suggestions.
The clarity and accuracy of speech signal acquisition are improved in noisy environments, the comprehensiveness and environmental adaptability of pronunciation evaluation are enhanced, and the training effect and efficiency are improved.
Smart Images

Figure CN120472935A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech recognition technology, and in particular to a correction system, method and device for oral pronunciation training. Background Art
[0002] With the widespread adoption of speech recognition technology, particularly in education, healthcare, and intelligent assistants, users are placing higher demands on speech recognition accuracy and comprehensive pronunciation assessment. However, existing speech recognition systems still perform poorly in extremely noisy environments, with recognition accuracy significantly declining. Furthermore, traditional pronunciation assessment methods primarily rely on a single dimension of the speech signal and lack comprehensive evaluation, preventing users from receiving comprehensive pronunciation correction recommendations.
[0003] For example, the invention patent with announcement number CN108922563B announces a method for oral learning correction based on the visualization of deviant organ morphology and behavior. By comparing the phonemes, stress, pauses between words and intonation when the learner pronounces with the standard pronunciation, the learner's pronunciation accuracy and the deviation between the pronunciation organ behavior and the standard behavior are calculated, and then visually displayed to the learner. The main steps are S1. Collecting the pronunciation information of the learner and the standard pronunciation, preprocessing the collected signal, and extracting features; S2. Constructing a standard pronunciation organ morphology and behavior library for sentences, and mapping the pronunciation features of the standard pronunciation to the organ morphology and behavior library; S3. Calculating the similarity between the phonemes, stress, pauses and intonation of the learner's pronunciation and the standard pronunciation, calculating the deviation value of the organ behavior, and visually displaying it to the learner; S4. Score the learner's pronunciation based on the four indicators and provide feedback to the learner to improve learning efficiency.
[0004] For example, the invention patent with the announcement number CN113454717B announces a speech recognition device and method. The method for recognizing user speech includes the following steps: obtaining an audio signal divided into multiple frame units; determining the energy component for each filter group by applying a filter group distributed according to a preset scale to the spectrum of the audio signal divided into frame units; smoothing the determined energy component for each filter group; extracting the feature vector of the audio signal based on the smoothed energy component for each filter group; and recognizing the user speech in the audio signal by inputting the extracted feature vector into a speech recognition model.
[0005] However, during the implementation of the technical solutions described in the embodiments of this application, the present applicant discovered that the aforementioned technology suffers from at least the following technical issues: The correction system for spoken voice training suffers from insufficient speech recognition accuracy in extremely noisy environments. In noisy environments, ambient noise can severely interfere with the extraction and analysis of speech signals, resulting in the speech recognition system being unable to accurately recognize the user's speech content. Summary of the Invention
[0006] In view of the deficiencies in the prior art, the present invention provides a correction system, method and device for oral pronunciation training, which can effectively solve the problems involved in the above-mentioned background technology.
[0007] To achieve the above objectives, the present invention is implemented through the following technical solutions: The first aspect of the present invention provides a correction system for oral voice training, including: an environmental data analysis module, which is used to collect environmental noise parameters in real time, preliminarily judge whether the training environment meets the requirements based on the environmental noise parameters, and optimize the training environment.
[0008] The multimodal environmental adaptability module is used to collect optimized environmental noise parameters and adjust the speech judgment weight and visual judgment weight according to the optimized environmental noise parameters.
[0009] The spoken language accuracy judgment module is used to collect the pronunciation data and facial image data of the speaker, combine the voice judgment weight and visual judgment weight to evaluate the pronunciation accuracy, and provide targeted correction suggestions based on the evaluation results.
[0010] As a further method, the environmental noise parameters are used to preliminarily determine whether the training environment meets the requirements. The specific analysis process is as follows: the environmental noise parameters include sound noise data and image noise data, wherein the sound noise data includes decibel mean, noise spectrum peak and burst noise frequency, and the image noise data includes light intensity and suspended particulate matter concentration; the preset critical decibel mean, critical noise spectrum peak, critical burst noise frequency, reference light intensity, allowable light intensity deviation value and critical suspended particulate matter concentration are extracted from the oral training database; the sound noise data and image noise data are processed respectively to obtain a first interference quantification index and a second interference quantification index; and a preliminarily judgment is made on whether the training environment meets the requirements based on the first interference quantification index and the second interference quantification index.
[0011] As a further method, the sound noise data and the image noise data are processed separately to obtain a first interference quantification index and a second interference quantification index. The specific analysis process is: the decibel mean and the critical decibel mean, the noise spectrum peak and the critical noise spectrum peak, the burst noise frequency and the critical burst noise frequency are compared and weighted superimposed to obtain the first interference quantification index; the difference between the light intensity and the reference light intensity and the allowable light intensity deviation value, the suspended particulate matter concentration and the critical suspended particulate matter concentration are compared and weighted superimposed to obtain the second interference quantification index.
[0012] As a further method, a preliminary judgment is made on whether the training environment meets the requirements based on the first interference quantification index and the second interference quantification index. The specific analysis process is: obtaining the preset first interference threshold and the second interference threshold from the oral training database, and comparing the first interference quantification index and the second interference quantification index with the first interference threshold and the second interference threshold respectively; if the first interference quantification index is greater than or equal to the first interference threshold and the second interference quantification index is greater than or equal to the second interference threshold, then it is judged that the training environment does not meet the requirements; otherwise, a second judgment is made on the training environment.
[0013] As a further method, the optimized environmental noise parameters are collected, and the speech judgment weight and the visual judgment weight are adjusted according to the optimized environmental noise parameters. The specific analysis process is: the optimized first interference quantization index and the optimized second interference quantization index are obtained according to the optimized environmental noise parameters; the optimized first interference quantization index and the optimized second interference quantization index are summed to obtain the total interference quantization index, and the second interference quantization index is compared with the total interference quantization index to obtain the speech judgment weight; the first interference quantization index is compared with the total interference quantization index to obtain the visual judgment weight.
[0014] As a further method, the training environment is judged twice, and the specific analysis process is: the optimized first interference quantization index and the optimized second interference quantization index are compared with the first interference threshold and the second interference threshold respectively; if the optimized first interference quantization index is greater than or equal to the first interference threshold and the optimized second interference quantization index is less than the second interference threshold, the optimized first interference quantization index is subtracted from the first interference threshold, and the optimized second interference quantization index is corrected according to the difference to obtain the corrected second interference quantization index, and the corrected second interference quantization index is compared with the second interference threshold. If the corrected second interference quantization index is less than the second interference threshold, it is judged that the training environment meets the requirements, otherwise, it is judged that the training environment does not meet the requirements; if the optimized first interference quantization index is greater than or equal to the first interference threshold and the optimized second interference quantization index is less than the second interference threshold, it is judged that the training environment meets the requirements. If the quantitative index is less than the first interference threshold and the optimized second interference quantitative index is greater than or equal to the second interference threshold, the optimized second interference quantitative index is subtracted from the second interference threshold, and the optimized first interference quantitative index is corrected according to the difference to obtain the corrected first interference quantitative index, and the corrected first interference quantitative index is compared with the first interference threshold. If the corrected first interference quantitative index is less than the first interference threshold, it is judged that the training environment meets the requirements, otherwise, it is judged that the training environment does not meet the requirements; if the optimized first interference quantitative index is less than the first interference threshold and the optimized second interference quantitative index is less than the second interference threshold, it is judged that the training environment meets the requirements; when the training environment meets the requirements, pronunciation training is allowed; when the training environment does not meet the requirements, feedback reminders are given.
[0015] As a further method, the pronunciation accuracy evaluation is carried out, and the specific analysis process is as follows: the pronunciation data includes pitch, intensity and length, and the facial image data includes mouth opening, facial muscle movement amplitude and tongue position offset; the preset reference pitch, reference intensity, reference length, reference mouth opening, reference facial muscle movement amplitude, reference tongue position offset, allowed pitch deviation value, allowed intensity deviation value, allowed length deviation value, allowed mouth opening deviation value, allowed facial muscle movement amplitude deviation value and allowed tongue position offset deviation value are extracted from the oral training database; the difference between the pitch and the reference pitch is compared with the allowed pitch deviation value, the difference between the intensity and the reference intensity is compared with the allowed intensity deviation value, and the difference between the length and the reference length is compared with the allowed tone deviation value. The length deviation value, the difference between the mouth opening and closing degree and the reference mouth opening and closing degree and the allowed mouth opening and closing deviation value, the difference between the facial muscle movement amplitude and the reference facial muscle movement amplitude and the allowed facial muscle movement amplitude deviation value, the difference between the tongue position offset and the reference tongue position offset and the allowed tongue position offset deviation value are compared respectively to obtain a pronunciation accuracy evaluation value, which is used to quantitatively evaluate the degree of deviation between the current pronunciation of the speaker and the standard pronunciation; the pronunciation accuracy evaluation value is compared with a preset pronunciation accuracy evaluation threshold extracted from the oral training database; if the pronunciation accuracy evaluation value is greater than the pronunciation accuracy evaluation threshold, corrective training is performed; if the pronunciation accuracy evaluation value is greater than the pronunciation accuracy evaluation threshold, no additional operation is performed.
[0016] As a further method, the correction training is carried out, and the specific analysis process is as follows: according to the deviation of each indicator in the pronunciation accuracy assessment, the error type of the speaker in the pronunciation process is analyzed, and targeted training materials and training methods are selected from the training content library according to the error type; during the training process, the pronunciation data and facial image data of the speaker are collected in real time, the accuracy is assessed again, and feedback is provided in a timely manner based on the accuracy assessment value.
[0017] The second aspect of the present invention provides a method for a correction system for oral voice training, which is characterized by including: real-time collection of environmental noise parameters, preliminary judgment of whether the training environment meets the requirements based on the environmental noise parameters, and optimization of the training environment; collection of optimized environmental noise parameters, and adjustment of speech judgment weights and visual judgment weights based on the optimized environmental noise parameters; collection of pronunciation data and facial image data of the speaker, combining the speech judgment weights and visual judgment weights to evaluate pronunciation accuracy, and providing targeted correction suggestions based on the evaluation results.
[0018] A third aspect of the present invention provides a device for a correction system for oral pronunciation training, characterized in that it includes: a memory and a processor, the memory is used to store a computer program, and the processor is used to call the computer program.
[0019] Compared with the prior art, the embodiments of the present invention have at least the following beneficial effects: (1) The present invention provides a correction system, method and equipment for oral voice training, thereby achieving real-time monitoring and optimization of the training environment. By collecting environmental noise parameters and combining them with preset critical values and interference quantification indicators, it is possible to determine whether the training environment meets the requirements and dynamically adjust the training environment to ensure the clarity and accuracy of voice signal acquisition, thereby improving the training effect.
[0020] (2) The present invention collects key parameters of the speaker such as pitch, intensity, duration, mouth shape, facial muscle movement amplitude and tongue position deviation in real time, and combines them with preset reference values and allowable deviation values to generate a pronunciation accuracy evaluation value, quantify the degree of deviation between the speaker's current pronunciation and the standard pronunciation, and provide targeted correction suggestions based on the deviation, so as to help users quickly identify pronunciation problems and improve the efficiency of oral training.
[0021] (3) The present invention uses multimodal data fusion technology to combine voice data and facial image data, dynamically adjusts the voice judgment weight and visual judgment weight, and realizes a comprehensive evaluation of pronunciation accuracy. When the environmental noise is high, the system automatically increases the visual judgment weight and reduces the dependence on voice data, thereby avoiding the evaluation bias caused by the failure of a single mode, improving the environmental adaptability of the system, and enhancing the accuracy of pronunciation evaluation.
[0022] Of course, any product implementing the present invention does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 This is a schematic diagram of system module connections of the present invention.
[0024] Figure 2 Schematic diagram of the method of the present invention. DETAILED DESCRIPTION
[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0026] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.
[0027] Reference Figure 1 As shown, the first aspect of the present invention provides a correction system for oral voice training, including: an environmental data analysis module, which is used to collect environmental noise parameters in real time, preliminarily judge whether the training environment meets the requirements based on the environmental noise parameters, and optimize the training environment.
[0028] Specifically, a preliminarily judgment is made on whether the training environment meets the requirements based on the environmental noise parameters. The specific analysis process is as follows: the environmental noise parameters include sound noise data and image noise data, wherein the sound noise data includes decibel mean, noise spectrum peak and burst noise frequency, and the image noise data includes light intensity and suspended particulate matter concentration; the preset critical decibel mean, critical noise spectrum peak, critical burst noise frequency, reference light intensity, allowable light intensity deviation value and critical suspended particulate matter concentration are extracted from the oral training database; the first interference quantification index and the second interference quantification index are obtained by processing the sound noise data and the image noise data respectively; a preliminarily judgment is made on whether the training environment meets the requirements based on the first interference quantification index and the second interference quantification index.
[0029] In this embodiment, the decibel mean refers to the average noise level in the environment, reflecting the overall noise intensity; the noise spectrum peak refers to the peak frequency of the noise in the frequency domain, reflecting the main frequency component of the noise; the burst noise frequency refers to the frequency of occurrence of burst noise in the environment, reflecting the suddenness of the noise; the light intensity refers to the light intensity in the environment, reflecting the impact of light on image acquisition; the suspended particulate matter concentration refers to the concentration of suspended particulate matter in the air, reflecting the impact of air cleanliness on image clarity.
[0030] It should be understood that in this embodiment, the decibel mean value can be measured using a decibel meter; the noise spectrum peak value can be measured using a spectrum analyzer; the burst noise frequency can be measured using a noise event counter; the light intensity can be measured using an illuminance meter; and the suspended particulate matter concentration can be measured using a particle sensor. The critical decibel mean value, critical noise spectrum peak value, critical burst noise frequency, reference light intensity, allowable light intensity deviation, and critical suspended particulate matter concentration can be directly obtained from the spoken language training database and used to determine whether the environmental noise parameters exceed the standard.
[0031] It should be understood that in this embodiment, the first interference quantification index refers to the degree of interference of acoustic noise on speech signal acquisition, obtained by comparing acoustic noise data with a preset critical value. The second interference quantification index refers to the degree of interference of image noise on visual data acquisition, obtained by comparing image noise data with a preset critical value.
[0032] Specifically, the first interference quantification index and the second interference quantification index are obtained by processing the sound noise data and the image noise data respectively. The specific analysis process is: the decibel mean and the critical decibel mean, the noise spectrum peak and the critical noise spectrum peak, the burst noise frequency and the critical burst noise frequency are compared and weightedly superimposed to obtain the first interference quantification index; the difference between the light intensity and the reference light intensity and the allowable light intensity deviation value, the suspended particulate matter concentration and the critical suspended particulate matter concentration are compared and weightedly superimposed to obtain the second interference quantification index.
[0033] In a specific embodiment, the first interference quantification indicator and the second interference quantification indicator are obtained as follows: ; ; ; ; Where, represents the first interference quantification index, represents the second interference quantification index, represents the mean decibel value, Indicates the preset critical decibel mean, represents the peak of the noise spectrum, Indicates the preset critical noise spectrum peak, represents the burst noise frequency, Indicates the preset critical burst noise frequency, Indicates the light intensity, Indicates the preset reference light intensity, Indicates the preset allowable light intensity deviation value. Indicates the concentration of suspended particulate matter. Indicates the preset critical suspended particulate matter concentration, represents the preset decibel mean weighting factor, Represents the preset noise spectrum peak weight factor, Indicates the preset burst noise frequency weight factor, Represents the preset light intensity weight factor, Indicates the preset suspended particulate matter concentration weight factor.
[0034] 、 and When in use, the influence weights of the first interference quantification index process corresponding to the decibel mean, noise spectrum peak and burst noise frequency can be directly obtained from the spoken language training database, which respectively represent the numerical values of the influence degree of the decibel mean, noise spectrum peak and burst noise frequency on the first interference quantification index, and the corresponding relationship can be a pre-set mapping relationship. In actual applications, the signal of the data and the influence weights of the first interference quantification index process corresponding to the decibel mean, noise spectrum peak and burst noise frequency preset in the spoken language training database form a mapping set, and the signal of the data is input into the mapping set to obtain the influence weights of the first interference quantification index process corresponding to the decibel mean, noise spectrum peak and burst noise frequency, wherein the mapping relationship can be a one-to-one correspondence or a many-to-one relationship. In this embodiment, the value ranges of the decibel mean weight factor, the noise spectrum peak weight factor and the burst noise frequency weight factor are all limited to between 0 and 1, and the sum of the decibel mean weight factor, the noise spectrum peak weight factor and the burst noise frequency weight factor is 1.
[0035] and When in use, the influence weights of the second interference quantification index process corresponding to the light intensity and the suspended particle concentration can be directly obtained from the spoken language training database, which respectively represent the numerical values of the degree of influence of the light intensity and the suspended particle concentration on the second interference quantification index, and the corresponding relationship can be a pre-set mapping relationship. In actual applications, the signal of the data and the influence weights of the second interference quantification index process corresponding to the light intensity and the suspended particle concentration preset in the spoken language training database form a mapping set, and the signal of the data is input into the mapping set to obtain the influence weights of the second interference quantification index process corresponding to the light intensity and the suspended particle concentration, wherein the mapping relationship can be a one-to-one correspondence or a many-to-one relationship. In this embodiment, the value ranges of the light intensity weight factor and the suspended particle concentration weight factor are both limited to between 0 and 1, and the sum of the light intensity weight factor and the suspended particle concentration weight factor is 1.
[0036] In this embodiment, the first interference quantification index is used to quantitatively assess the degree of interference of the sound noise data on the training environment. A higher decibel mean, a higher noise spectrum peak, or a higher burst noise frequency indicates a higher degree of interference of the sound noise data on the training environment.
[0037] In this embodiment, the second interference quantification index is used to quantitatively assess the degree of interference of the image noise data on the training environment. The greater the deviation between the light intensity and the reference light intensity, or the higher the concentration of suspended particulate matter, the greater the degree of interference of the image noise data on the training environment.
[0038] The algorithm in this embodiment combines the decibel mean, noise spectrum peak, and burst noise frequency to comprehensively analyze and derive the first interference quantification indicator. In this formula, the decibel mean, noise spectrum peak, and burst noise frequency interact with each other. As the decibel mean increases, the intensity of background sound in the environment increases, making the noise spectrum peak more likely to reach a higher value. This is because a higher decibel mean indicates a greater concentration of sound energy across different frequencies, potentially causing noise energy in certain frequency bands to be prominent, forming high spectral peaks. As the decibel mean increases, burst noise is more likely to be triggered in environments with high background noise, resulting in an increase in burst noise frequency. This is because a noisy environment is more likely to trigger various sudden noise source activities. As the noise spectrum peak increases, the noise energy increases, further affecting the overall decibel mean and causing it to rise. Furthermore, higher noise spectrum peaks change the acoustic characteristics of the environment, making burst noise more likely to occur, thereby increasing the frequency of burst noise. As the frequency of burst noise increases, the decibel mean of the environment continues to rise, as burst noise continuously adds sound energy to the environment. Furthermore, frequent burst noise may stimulate resonance or other acoustic effects, leading to higher noise spectrum peaks. By comprehensively analyzing the decibel mean, noise spectrum peak and burst noise frequency, the first interference quantification index can be accurately obtained, which quantitatively reflects the degree of interference of the sound noise data in the training environment on oral voice training and the complexity and instability of the sound environment.
[0039] The algorithm in this embodiment combines light intensity and suspended particulate matter concentration to comprehensively analyze and derive a second interference quantification indicator. In this formula, light intensity and suspended particulate matter concentration interact with each other. As light intensity decreases, the environment becomes darker. Suspended particulate matter is less likely to be penetrated and scattered by light in darker environments, resulting in higher suspended particulate matter concentrations. This lack of light makes it difficult to clearly discern the actual distribution of particles, and the light-blocking effect of particles in darker environments is more pronounced. As light intensity decreases, visibility decreases. To maintain a certain level of visual clarity, the perception of particles in the environment may be increased, which is reflected in the data as an increase in suspended particulate matter concentration. As suspended particulate matter concentration increases, particles scatter and absorb more light, resulting in a decrease in light intensity. Because particles block and consume light, the amount of light energy reaching the observation point is reduced. Furthermore, high concentrations of suspended particulate matter can make the environment turbid, reducing the efficiency of light transmission and further reducing light intensity. By comprehensively analyzing the light intensity and suspended particulate matter concentration, the second interference quantification index can be accurately obtained, which quantitatively reflects the degree of interference of image noise data in the training environment on oral pronunciation training, as well as the quality of visual conditions in the training environment and the possible impact on image acquisition quality.
[0040] Specifically, a preliminary judgment is made on whether the training environment meets the requirements based on the first interference quantification index and the second interference quantification index. The specific analysis process is: obtain the preset first interference threshold and the second interference threshold from the oral training database, and compare the first interference quantification index and the second interference quantification index with the first interference threshold and the second interference threshold respectively; if the first interference quantification index is greater than or equal to the first interference threshold and the second interference quantification index is greater than or equal to the second interference threshold, then it is judged that the training environment does not meet the requirements; otherwise, a second judgment is made on the training environment.
[0041] It should be understood that in this embodiment, the preset first interference threshold and second interference threshold are directly obtained from the oral training database, and the first interference quantification index and the second interference quantification index are compared with the first interference threshold and the second interference threshold respectively. If the first interference quantification index is greater than or equal to the first interference threshold, and the second interference quantification index is greater than or equal to the second interference threshold, it indicates that the interference levels of the sound noise and image noise in the current training environment are beyond the acceptable range, then the training environment is judged to be not in compliance with the requirements, and a prompt is given to change the environment. If the condition that "the first interference quantification index is greater than or equal to the first interference threshold and the second interference quantification index is greater than or equal to the second interference threshold" is not met, that is, at least one interference quantification index is less than the corresponding interference quantification threshold, it indicates that the interference level of the training environment in a certain aspect is within an acceptable range, but it cannot be directly determined that the entire training environment fully meets the requirements. Therefore, a secondary judgment of the training environment is required to further evaluate whether the environment is suitable for oral voice training to ensure the reliability of subsequent pronunciation accuracy evaluation.
[0042] The multimodal environmental adaptability module is used to collect optimized environmental noise parameters and adjust the speech judgment weight and visual judgment weight according to the optimized environmental noise parameters.
[0043] Specifically, the optimized environmental noise parameters are collected, and the speech judgment weight and the visual judgment weight are adjusted according to the optimized environmental noise parameters. The specific analysis process is: the optimized first interference quantization index and the optimized second interference quantization index are obtained according to the optimized environmental noise parameters; the optimized first interference quantization index and the optimized second interference quantization index are summed to obtain the total interference quantization index, and the second interference quantization index is compared with the total interference quantization index to obtain the speech judgment weight; the first interference quantization index is compared with the total interference quantization index to obtain the visual judgment weight.
[0044] It should be understood that in this embodiment, after preliminary judgment and optimization of the training environment, the noise parameters of the current environment are collected. Based on the collected optimized environmental noise parameters, an optimized first interference quantization index and an optimized second interference quantization index are respectively obtained. The optimized first interference quantization index and the optimized second interference quantization index are summed to obtain a total interference quantization index, which comprehensively reflects the overall interference level of sound noise and image noise in the current training environment. The optimized second interference quantization index is divided by the total interference quantization index to obtain a speech judgment weighting factor. When the image noise interference level is higher than the sound noise interference level, it indicates that the image information in the environment may be less reliable. When performing pronunciation accuracy assessment, the weight of visual judgment should be relatively reduced, and the weight of speech judgment should be correspondingly increased. The optimized first interference quantization index is divided by the total interference quantization index to obtain a visual judgment weighting factor. When the sound noise interference level is higher than the image noise interference level, it indicates that the speech information in the environment may be significantly disturbed. When evaluating pronunciation accuracy, the weight of speech judgment should be relatively reduced, and the weight of visual judgment should be increased.
[0045] Specifically, a secondary judgment is made on the training environment, and the specific analysis process is: the optimized first interference quantization index and the optimized second interference quantization index are compared with the first interference threshold and the second interference threshold respectively; if the optimized first interference quantization index is greater than or equal to the first interference threshold and the optimized second interference quantization index is less than the second interference threshold, then the optimized first interference quantization index is subtracted from the first interference threshold, and the optimized second interference quantization index is corrected according to the difference to obtain the corrected second interference quantization index, and the corrected second interference quantization index is compared with the second interference threshold. If the corrected second interference quantization index is less than the second interference threshold, it is judged that the training environment meets the requirements; otherwise, it is judged that the training environment does not meet the requirements.
[0046] It should be understood that in this embodiment, the optimized first interference quantification index and the optimized second interference quantification index obtained by optimizing the environmental noise parameters are respectively compared with the first interference threshold and the second interference threshold obtained from the spoken language training database. When the optimized first interference quantification index is greater than or equal to the first interference threshold, and at the same time, the optimized second interference quantification index is less than the second interference threshold, it means that the degree of sound noise interference in the environment at this time exceeds the acceptable range, while the degree of image noise interference is within the acceptable range. First, the optimized first interference quantification index is subtracted from the first interference threshold to obtain the difference between the two. This difference reflects the degree to which the sound noise interference exceeds the acceptable range. Based on the obtained difference, the optimized second interference quantification index is weighted and corrected to obtain the corrected second interference quantification index. Compare the corrected second interference quantification index with the second interference threshold. If the corrected second interference quantification index is less than the second interference threshold, it means that after comprehensively considering the impact of the excess sound noise interference on the overall environment, the image noise interference level is still within an acceptable range. At this time, it is judged that the training environment meets the requirements; if the corrected second interference quantification index is greater than or equal to the second interference threshold, it means that after comprehensive consideration, the interference level of the overall environment still exceeds the acceptable range, and the training environment is judged not to meet the requirements.
[0047] Specifically, if the optimized first interference quantization index is less than the first interference threshold and the optimized second interference quantization index is greater than or equal to the second interference threshold, the optimized second interference quantization index is subtracted from the second interference threshold, and the optimized first interference quantization index is corrected according to the difference to obtain the corrected first interference quantization index, and the corrected first interference quantization index is compared with the first interference threshold. If the corrected first interference quantization index is less than the first interference threshold, it is judged that the training environment meets the requirements; otherwise, it is judged that the training environment does not meet the requirements.
[0048] It should be understood that in this embodiment, when the optimized first interference quantification index is less than the first interference threshold and the optimized second interference quantification index is greater than or equal to the second interference threshold, this indicates that the level of acoustic noise interference in the training environment is within an acceptable range, while the level of image noise interference exceeds the acceptable range. The optimized second interference quantification index is subtracted from the second interference threshold to obtain a difference between the two. This difference reflects the extent to which the image noise interference exceeds the acceptable range. Based on the obtained difference, the optimized first interference quantification index is weighted and corrected to obtain a corrected first interference quantification index. The corrected first interference quantification index is compared with the first interference threshold. If the corrected first interference quantification index is less than the first interference threshold, it indicates that after comprehensively considering the impact of the excess image noise interference on the overall environment, the acoustic noise interference level remains within an acceptable range. In other words, the overall interference level of the training environment meets the requirements after comprehensive adjustment. Therefore, the training environment is judged to meet the requirements. If the corrected first interference quantification index is greater than or equal to the first interference threshold, it indicates that after comprehensive consideration and adjustment, the overall interference level of the environment still exceeds the acceptable range. Therefore, the training environment is judged to meet the requirements.
[0049] Specifically, if the optimized first interference quantization index is less than the first interference threshold and the optimized second interference quantization index is less than the second interference threshold, the training environment is judged to meet the requirements; when the training environment meets the requirements, pronunciation training is allowed; when the training environment does not meet the requirements, feedback reminders are given.
[0050] It should be understood that in this embodiment, when the first interference quantification index after optimization is less than the first interference threshold, and at the same time, the second interference quantification index after optimization is less than the second interference threshold, it means that the degree of sound noise interference in the training environment is within an acceptable range, and the degree of image noise interference is also within an acceptable range. It is then judged that the training environment meets the requirements, that is, the environment is suitable for oral voice training.
[0051] In this embodiment, when the training environment meets the requirements, pronunciation training is allowed; when the training environment does not meet the requirements, feedback reminders are given to inform relevant personnel (such as the speaker) that there are problems with the current training environment, such as displaying prompt information on the system interface, issuing a sound alarm, etc.
[0052] The spoken language accuracy judgment module is used to collect the pronunciation data and facial image data of the speaker, combine the voice judgment weight and visual judgment weight to evaluate the pronunciation accuracy, and provide targeted correction suggestions based on the evaluation results.
[0053] Specifically, the pronunciation accuracy is evaluated, and the specific analysis process is as follows: the pronunciation data includes pitch, intensity and length, and the facial image data includes mouth opening, facial muscle movement amplitude and tongue position offset; the preset reference pitch, reference intensity, reference length, reference mouth opening, reference facial muscle movement amplitude, reference tongue position offset, allowed pitch deviation value, allowed intensity deviation value, allowed length deviation value, allowed mouth opening deviation value, allowed facial muscle movement amplitude deviation value and allowed tongue position offset deviation value are extracted from the oral training database; the difference between the pitch and the reference pitch is compared with the allowed pitch deviation value, the difference between the intensity and the reference intensity is compared with the allowed intensity deviation value, and the difference between the length and the reference length is compared with the allowed length deviation value. , the difference between the mouth opening and closing degree and the reference mouth opening and closing degree and the allowed mouth opening and closing deviation value, the difference between the facial muscle movement amplitude and the reference facial muscle movement amplitude and the allowed facial muscle movement amplitude deviation value, the difference between the tongue position offset and the reference tongue position offset and the allowed tongue position offset deviation value are compared respectively to obtain a pronunciation accuracy evaluation value, which is used to quantitatively evaluate the degree of deviation between the current pronunciation of the speaker and the standard pronunciation; the pronunciation accuracy evaluation value is compared with a preset pronunciation accuracy evaluation threshold extracted from the oral training database, if the pronunciation accuracy evaluation value is greater than the pronunciation accuracy evaluation threshold, correction training is performed; if the pronunciation accuracy evaluation value is greater than the pronunciation accuracy evaluation threshold, no additional operation is performed.
[0054] It should be understood that, in this embodiment, pitch refers to the pitch of the pronunciation; intensity refers to the sound pressure level of the pronunciation; duration refers to the pronunciation duration of each phoneme; mouth shape refers to the vertical distance between the lips; facial muscle movement amplitude refers to the movement amplitude of key facial points; tongue position offset refers to the offset distance of the tongue from the standard pronunciation position.
[0055] It should be understood that in this embodiment, the pitch can be measured by speech analysis software (such as Praat); the sound intensity can be measured by a decibel meter or speech analysis software; the sound length can be measured by speech analysis software (such as Audacity); the mouth opening and closing degree can be measured by a camera combined with a facial key point detection algorithm; the facial muscle movement amplitude can be measured by a facial motion capture device; the tongue position offset can be measured by an ultrasonic tongue position detector; the reference pitch, reference sound intensity, reference sound length, reference mouth opening and closing degree, reference facial muscle movement amplitude, reference tongue position offset, allowable pitch deviation value, allowable sound intensity deviation value, allowable sound length deviation value, allowable mouth opening and closing degree deviation value, allowable facial muscle movement amplitude deviation value and allowable tongue position offset deviation value can be directly obtained from the oral training database.
[0056] In a specific embodiment, the pronunciation accuracy evaluation value is obtained as follows: ; ; ; ; ; ; ; ; Where, represents the pronunciation accuracy evaluation value, Indicates the speech judgment value, Represents the visual judgment value, Indicates tone, Indicates the preset reference tone, Indicates the preset allowable pitch deviation value, Indicates sound intensity. Indicates the preset reference sound intensity. Indicates the preset allowable sound intensity deviation value. Indicates the length of the sound. Indicates the preset reference tone length. Indicates the preset allowable tone length deviation value. Indicates the degree of mouth opening and closing. Indicates the preset reference lip opening and closing degree. Indicates the preset allowable mouth opening and closing deviation value. Indicates the range of motion of facial muscles. Indicates the preset reference facial muscle movement range, Indicates the preset deviation value of the allowed facial muscle movement amplitude. Indicates tongue offset. Indicates the preset reference tongue offset, Indicates the preset allowable tongue offset deviation value, represents the preset pitch weight factor, Indicates the preset sound intensity weight factor, Indicates the preset length weight factor, Indicates the preset weight factor of the mouth opening and closing degree, Represents the preset facial muscle movement amplitude weight factor, Indicates the preset tongue offset weight factor, represents the speech judgment weight factor, represents the visual judgment weight factor, represents the first interference quantification index after optimization, Represents the optimized second interference quantification index.
[0057] 、 and When in use, the influence weights in the speech judgment value process corresponding to the pitch, intensity and length can be directly obtained from the spoken language training database, respectively representing the numerical values of the degree of influence of the pitch, intensity and length on the speech judgment value, and their corresponding relationships can be pre-set mapping relationships. In actual applications, the signal of the data and the influence weights in the speech judgment value process corresponding to the pitch, intensity and length preset in the spoken language training database form a mapping set, and the signal of the data is input into the mapping set to obtain the influence weights in the speech judgment value process corresponding to the pitch, intensity and length, wherein the mapping relationship mapping relationship can be a one-to-one correspondence or a many-to-one relationship. In this embodiment, the value ranges of the pitch weight factor, the intensity weight factor and the length weight factor are all limited to between 0 and 1, and the sum of the pitch weight factor, the intensity weight factor and the length weight factor is 1.
[0058] 、 and When in use, the influence weights in the visual judgment value process corresponding to the mouth shape opening and closing degree, the facial muscle movement amplitude and the tongue position offset can be directly obtained from the oral training database, which respectively represent the numerical values of the influence degree of the mouth shape opening and closing degree, the facial muscle movement amplitude and the tongue position offset on the visual judgment value, and their corresponding relationship can be a pre-set mapping relationship. In actual application, the signal of the data and the influence weights in the visual judgment value process corresponding to the mouth shape opening and closing degree, the facial muscle movement amplitude and the tongue position offset preset in the oral training database form a mapping set, and the signal of the data is input into the mapping set to obtain the influence weights in the visual judgment value process corresponding to the mouth shape opening and closing degree, the facial muscle movement amplitude and the tongue position offset, wherein the mapping relationship can be a one-to-one correspondence or a many-to-one relationship. In this embodiment, the value ranges of the mouth shape opening and closing weight factor, the facial muscle movement amplitude weight factor and the tongue position offset weight factor are all limited to between 0 and 1, and the sum of the mouth shape opening and closing weight factor, the facial muscle movement amplitude weight factor and the tongue position offset weight factor is 1.
[0059] and The weights of the speech and visual judgment values in the pronunciation accuracy evaluation process are the influence weights of the speech and visual judgment values, respectively indicating the degree of influence of the speech and visual judgment values on the pronunciation accuracy evaluation value. In this embodiment, the speech and visual judgment weight factors are both limited to a value range between 0 and 1, and the sum of the speech and visual judgment weight factors is 1.
[0060] In this embodiment, the speech judgment value is used to quantitatively assess the degree of deviation between the speaker's pronunciation data and the standard pronunciation. The greater the deviation between the pitch and the reference pitch, the greater the deviation between the intensity and the reference intensity, or the greater the deviation between the duration and the reference duration, the greater the deviation between the speaker's pronunciation data and the standard pronunciation.
[0061] In this embodiment, the visual judgment value is used to quantitatively assess the degree of deviation between the speaker's facial image data and the facial features corresponding to the standard pronunciation. The greater the deviation between the mouth shape opening and closing degree and the reference mouth shape opening and closing degree, the greater the deviation between the facial muscle movement amplitude and the reference facial muscle movement amplitude, or the greater the deviation between the tongue position offset and the reference tongue position offset, the greater the degree of deviation between the speaker's facial image data and the facial features corresponding to the standard pronunciation.
[0062] In this embodiment, the pronunciation accuracy evaluation value is used to quantitatively evaluate the degree of deviation between the speaker's current pronunciation and the standard pronunciation. A higher voice judgment value or a higher visual judgment value indicates a greater degree of deviation between the speaker's current pronunciation and the standard pronunciation.
[0063] The algorithm of this embodiment combines pitch, intensity, and duration to comprehensively analyze and obtain a speech judgment value. In this formula, pitch, intensity, and duration interact with each other. As pitch deviation increases, pronunciation may need to adjust the intensity to maintain clarity. For example, high-pitched sounds are often paired with strong sounds. At the same time, the duration may also be affected. High-pitched pronunciation sometimes has difficulty maintaining a long duration. Changes in intensity can alter the vibration state of the vocal cords and affect pitch, and also affect the duration. Strong sounds may be difficult to sustain. When the duration changes, the pitch is prone to fluctuations, and the speaker also needs to adjust the intensity to maintain the duration. By comprehensively analyzing pitch, intensity, and duration, a speech judgment value can be accurately obtained, which quantitatively reflects the degree of deviation in phonetic characteristics between the speaker's actual performance in speech production and the standard pronunciation, as well as the differences in the speaker's ability to control speech and coordinate the vocal organs.
[0064] The algorithm of this embodiment combines the degree of mouth opening and closing, the amplitude of facial muscle movement, and the offset of tongue position, and comprehensively analyzes to obtain a visual judgment value. In this formula, the degree of mouth opening and closing, the amplitude of facial muscle movement, and the offset of tongue position influence each other. As the degree of mouth opening and closing increases, the amplitude of facial muscle movement will also increase accordingly. For example, the opening and closing of the lips requires the participation of the lip muscles of the face, and the muscle stretching amplitude increases when the mouth shape is wide open. At the same time, the tongue position will also change, because different mouth shapes will affect the activity space and position of the tongue in the mouth. For example, when pronouncing some sounds with a larger opening, the tongue position moves back or lowers. As the amplitude of facial muscle movement increases, it will have a direct impact on the mouth shape. For example, the tension or relaxation of facial muscles will change the shape and size of the mouth shape. In addition, the movement of facial muscles will also affect the stability and flexibility of the tongue position. Because there is a certain connection between facial muscles and the muscle tissue in the mouth, the movement of facial muscles may drive changes in the muscles in the mouth, thereby affecting the tongue position. As tongue displacement increases, the mouth shape also adjusts. For example, when the tongue is positioned forward, the mouth shape changes to ensure accurate pronunciation. At the same time, changes in tongue position may also cause subtle adjustments in facial muscles, as changes in tongue position require the coordinated action of oral muscles, which in turn affects facial muscles. By comprehensively analyzing the degree of mouth opening and closing, the amplitude of facial muscle movement, and tongue displacement, a precise visual judgment value can be obtained, quantitatively reflecting the degree of deviation between the speaker's facial vocalization characteristics during pronunciation and those during standard pronunciation, as well as the differences in the speaker's ability to control facial and oral muscles for accurate pronunciation.
[0065] The algorithm in this embodiment combines the speech judgment value and the visual judgment value to comprehensively analyze and determine the pronunciation accuracy assessment value. In this formula, the speech judgment value and the visual judgment value influence each other. As the speech judgment value increases, it indicates that the speaker's pronunciation deviates significantly from the standard pronunciation. This may cause the speaker to make inappropriate adjustments to their facial movements to compensate for the speech deficiencies, thereby increasing the visual judgment value. For example, when the speaker's pitch is inaccurate, they may unconsciously change their mouth shape or the range of motion of their facial muscles in an attempt to make their pronunciation closer to the standard. However, this may cause their facial movements to deviate from the facial features of the standard pronunciation. As the visual judgment value increases, the speaker's facial features deviate more from the standard pronunciation, affecting the accuracy of the speech and causing the speech judgment value to increase. For example, inaccurate mouth opening and closing can affect airflow and vocal resonance, thereby changing speech characteristics such as intensity, pitch, and duration, deviating from the standard pronunciation. By comprehensively analyzing the speech judgment value and the visual judgment value, the pronunciation accuracy evaluation value can be accurately obtained, which quantitatively reflects the overall degree of deviation between the speaker's current pronunciation and the standard pronunciation, as well as the comprehensive impact of the speaker's coordination and matching degree in voice and facial pronunciation movements on pronunciation accuracy.
[0066] Specifically, correction training is carried out, and the specific analysis process is as follows: according to the deviation of various indicators in the pronunciation accuracy assessment, the types of errors made by the speaker in the pronunciation process are analyzed, and targeted training materials and training methods are selected from the training content library according to the error types; during the training process, the pronunciation data and facial image data of the speaker are collected in real time, the accuracy is assessed again, and timely feedback is provided based on the accuracy assessment value.
[0067] It should be understood that in this embodiment, after the speaker's pronunciation accuracy is assessed by the spoken language accuracy judgment module, the deviation between various indicators (such as pitch, intensity, duration, mouth shape opening, facial muscle movement amplitude, tongue position offset, etc.) and standard pronunciation parameters is determined. Based on these deviations, the system can conduct in-depth analysis to identify the specific types of errors the speaker has made during pronunciation. For example, if the deviation in the pitch indicator is large and significantly higher than the reference pitch, it can be determined that the speaker has made a high pitch error. If the deviation in mouth shape opening from the reference mouth shape opening indicates that the mouth shape opening is too small, it can be determined that the mouth shape error is insufficient. This detailed analysis can accurately identify the speaker's errors, providing clear direction for subsequent corrective training. Based on the analysis results of the pronunciation error type, the system selects targeted training materials and methods from the training content library. The training content library stores a rich and diverse training resource, with corresponding training content for different error types. For example, for errors involving high pitch, specific scale practice materials may be selected to help the speaker practice smooth transitions from low to high and back to low, helping them master correct pitch control. For issues with insufficient mouth opening and closing, training materials with demonstration lip movements may be provided to guide the speaker in exaggerating the mouth shape and increasing its opening and closing. Training methods are also selected based on the type of error and training materials, and may include imitation, comparison, and feedback exercises to improve training effectiveness and efficiency. During correction training, the system collects real-time pronunciation data (pitch, intensity, duration, etc.) and facial image data (mouth opening, facial muscle movement, tongue position deviation, etc.). This real-time data collection allows for accuracy assessment, helping to understand progress and changes in training. For example, if, after a period of training, the speaker's intensity indicators still deviate significantly from the standard, this indicates that the current training method may be ineffective and requires adjustment. Based on the re-evaluation of accuracy, the system will provide timely feedback to the speaker. The feedback includes an evaluation of the speaker's current pronunciation, such as pointing out which aspects have improved and which aspects still need further work; and providing specific suggestions for improvement, such as adjusting the pronunciation method, strengthening muscle training in a certain part, etc.
[0068] Reference Figure 2As shown, the second aspect of the present invention provides a method for a correction system for oral voice training, which is characterized by including: real-time acquisition of environmental noise parameters, preliminarily judging whether the training environment meets the requirements based on the environmental noise parameters, and optimizing the training environment; acquiring the optimized environmental noise parameters, adjusting the speech judgment weight and the visual judgment weight based on the optimized environmental noise parameters; acquiring the pronunciation data and facial image data of the speaker, combining the speech judgment weight and the visual judgment weight to evaluate the pronunciation accuracy, and providing targeted correction suggestions based on the evaluation results.
[0069] A third aspect of the present invention provides a device for a correction system for oral pronunciation training, characterized in that it includes: a memory and a processor, the memory is used to store a computer program, and the processor is used to call the computer program.
[0070] The oral training database is used to store various key data and information in the oral training field, including reference pitch, reference intensity, reference duration, reference mouth shape opening, reference facial muscle movement amplitude, reference tongue position deviation for standard pronunciation, as well as allowable deviations for pitch, intensity, duration, mouth shape opening, facial muscle movement amplitude, and tongue position deviation. The data in the oral training database is obtained from professional language research institutions, teaching practice data from senior language education experts, teaching research project data from language majors in universities, and teaching test data from large language training companies. This data provides an important data foundation and reference basis for oral voice training, pronunciation accuracy assessment, and the development of corrective training programs.
[0071] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not describe all details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made based on the contents of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can better understand and utilize the present invention. As long as they do not deviate from the structure of the present invention or exceed the scope defined by the present invention, they should fall within the scope of protection of the present invention.
Claims
1. A correction system for oral pronunciation training, characterized by: include: Environmental data analysis module, used to collect environmental noise parameters in real time, preliminarily determine whether the training environment meets the requirements based on the environmental noise parameters, and optimize the training environment; A multimodal environmental adaptability module is used to collect optimized environmental noise parameters and adjust the speech judgment weight and visual judgment weight according to the optimized environmental noise parameters; The spoken language accuracy judgment module is used to collect the pronunciation data and facial image data of the speaker, combine the voice judgment weight and visual judgment weight to evaluate the pronunciation accuracy, and provide targeted correction suggestions based on the evaluation results.
2. The correction system for oral pronunciation training according to claim 1, characterized in that: The preliminary judgment of whether the training environment meets the requirements based on the environmental noise parameters is as follows: The environmental noise parameters include sound noise data and image noise data, wherein the sound noise data includes decibel mean, noise spectrum peak and burst noise frequency, and the image noise data includes light intensity and suspended particulate matter concentration; Extracting preset critical decibel mean, critical noise spectrum peak, critical burst noise frequency, reference light intensity, allowable light intensity deviation value and critical suspended particulate matter concentration from the spoken language training database; Obtaining a first interference quantification index and a second interference quantification index according to the sound noise data and the image noise data respectively; A preliminary judgment is made based on the first interference quantification index and the second interference quantification index whether the training environment meets the requirements.
3. The correction system for oral pronunciation training according to claim 2, characterized in that: The first interference quantization index and the second interference quantization index are obtained by respectively processing the sound noise data and the image noise data. The specific analysis process is as follows: The decibel mean and the critical decibel mean, the noise spectrum peak and the critical noise spectrum peak, the burst noise frequency and the critical burst noise frequency are compared and weightedly superimposed to obtain a first interference quantification index; The difference between the light intensity and the reference light intensity, the allowable light intensity deviation value, the suspended particulate matter concentration and the critical suspended particulate matter concentration are compared and weightedly superimposed to obtain a second interference quantitative index.
4. The correction system for oral pronunciation training according to claim 2, characterized in that: The specific analysis process of performing a preliminary judgment on whether the training environment meets the requirements based on the first interference quantification index and the second interference quantification index is as follows: Obtaining a preset first interference threshold and a second interference threshold from a spoken language training database, and comparing the first interference quantification index and the second interference quantification index with the first interference threshold and the second interference threshold respectively; If the first interference quantification index is greater than or equal to the first interference threshold and the second interference quantification index is greater than or equal to the second interference threshold, it is determined that the training environment does not meet the requirements; otherwise, a second judgment is performed on the training environment.
5. The correction system for oral pronunciation training according to claim 1, characterized in that: The optimized environmental noise parameters are collected, and the voice judgment weight and the visual judgment weight are adjusted according to the optimized environmental noise parameters. The specific analysis process is as follows: Obtaining an optimized first interference quantization index and an optimized second interference quantization index according to the optimized environmental noise parameter; The optimized first interference quantization index and the optimized second interference quantization index are summed to obtain a total interference quantization index, and the second interference quantization index is compared with the total interference quantization index to obtain a speech judgment weight; The first interference quantification index is compared with the total interference quantification index to obtain a visual judgment weight.
6. The correction system for oral pronunciation training according to claim 3, characterized in that: The specific analysis process of the secondary judgment of the training environment is as follows: Comparing the optimized first interference quantization index and the optimized second interference quantization index with the first interference threshold and the second interference threshold respectively; If the optimized first interference quantization index is greater than or equal to the first interference threshold and the optimized second interference quantization index is less than the second interference threshold, the optimized first interference quantization index is subtracted from the first interference threshold, and the optimized second interference quantization index is corrected according to the difference to obtain a corrected second interference quantization index, and the corrected second interference quantization index is compared with the second interference threshold. If the corrected second interference quantization index is less than the second interference threshold, it is judged that the training environment meets the requirements; otherwise, it is judged that the training environment does not meet the requirements; If the optimized first interference quantization index is less than the first interference threshold and the optimized second interference quantization index is greater than or equal to the second interference threshold, the optimized second interference quantization index is subtracted from the second interference threshold, and the optimized first interference quantization index is corrected according to the difference to obtain a corrected first interference quantization index, and the corrected first interference quantization index is compared with the first interference threshold. If the corrected first interference quantization index is less than the first interference threshold, it is judged that the training environment meets the requirements; otherwise, it is judged that the training environment does not meet the requirements; If the first interference quantization index after optimization is less than the first interference threshold and the second interference quantization index after optimization is less than the second interference threshold, it is determined that the training environment meets the requirements; When the training environment meets the requirements, pronunciation training is allowed; When the training environment does not meet the requirements, feedback reminders will be given.
7. The correction system for oral pronunciation training according to claim 1, characterized in that: The pronunciation accuracy assessment is performed, and the specific analysis process is as follows: The pronunciation data includes pitch, intensity and duration, and the facial image data includes mouth opening and closing, facial muscle movement amplitude and tongue position offset; Extracting preset reference pitch, reference intensity, reference length, reference mouth opening, reference facial muscle movement amplitude, reference tongue position deviation, allowable pitch deviation, allowable intensity deviation, allowable length deviation, allowable mouth opening deviation, allowable facial muscle movement amplitude deviation, and allowable tongue position deviation from a spoken language training database; Compare the difference between the pitch and the reference pitch with the allowed pitch deviation value, the difference between the intensity and the reference intensity with the allowed intensity deviation value, the difference between the tone length and the reference tone length with the allowed tone length deviation value, the difference between the mouth shape opening and the reference mouth shape opening and the allowed mouth shape opening deviation value, the difference between the facial muscle movement amplitude and the reference facial muscle movement amplitude with the allowed facial muscle movement amplitude deviation value, and the difference between the tongue position offset and the reference tongue position offset with the allowed tongue position offset deviation value to obtain a pronunciation accuracy evaluation value, which is used to quantitatively evaluate the degree of deviation between the speaker's current pronunciation and the standard pronunciation; Comparing the pronunciation accuracy evaluation value with a preset pronunciation accuracy evaluation threshold extracted from the spoken language training database, and performing correction training if the pronunciation accuracy evaluation value is greater than the pronunciation accuracy evaluation threshold; If the pronunciation accuracy evaluation value is greater than the pronunciation accuracy evaluation threshold, no additional operation is performed.
8. The correction system for oral pronunciation training according to claim 1, characterized in that: The specific analysis process of the correction training is as follows: According to the deviation of various indicators in the pronunciation accuracy assessment, the types of errors made by the speaker in the pronunciation process are analyzed, and targeted training materials and training methods are selected from the training content library according to the error types; During the training process, the pronunciation data and facial image data of the speaker are collected in real time, the accuracy is evaluated again, and feedback is provided in a timely manner based on the accuracy evaluation value.
9. A method for applying the correction system for oral pronunciation training according to any one of claims 1 to 8, characterized in that: include: Collect environmental noise parameters in real time, make a preliminary judgment on whether the training environment meets the requirements based on the environmental noise parameters, and optimize the training environment; Collect optimized environmental noise parameters, and adjust the speech judgment weight and visual judgment weight according to the optimized environmental noise parameters; The pronunciation data and facial image data of the speaker are collected, and the pronunciation accuracy is evaluated by combining the voice judgment weight and visual judgment weight. Based on the evaluation results, targeted correction suggestions are provided.
10. A device for a correction system for oral pronunciation training, characterized in that include: A memory and a processor, wherein the memory is used to store a computer program, and the processor is used to call the computer program, including a correction system for oral pronunciation training according to any one of claims 1 to 8.
Citation Information
Patent Citations
Oral language learning correction method based on visualization of deviated organ morphology and behavior
CN108922563B
Speech recognition device and method
CN113454717B
Cited By
Mandarin pronunciation real-time correction method and system based on multi-modal streaming learning
CN121393456A
Mandarin pronunciation real-time correction method and system based on multi-modal streaming learning
CN121393456B