Artificial intelligence-based voice auditory language training method and system

By analyzing trainees' speech data using artificial intelligence, key pronunciation nodes are located and organic disorders are diagnosed, generating personalized training plans. This solves the problems of traditional language training methods, such as the inability to personalize adjustments and the influence of noise, and achieves efficient and accurate language training results.

CN120766718BActive Publication Date: 2026-02-06WUHAN CONSERVATORY OF MUSIC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510991891.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2026-02-06
Estimated Expiration
2045-07-18

AI Technical Summary

Technical Problem

Traditional language training methods cannot be personalized to the pronunciation characteristics of each trainee, lack data analysis and automated diagnostic capabilities, make it difficult to identify and correct pronunciation problems, and environmental noise affects the training effect.

Method used

By acquiring trainees' voice data, artificial intelligence is used to evaluate pronunciation quality and analyze anomalies, generating personalized training plans. Combined with voice acquisition, noise suppression, and dynamic time warping techniques, key pronunciation nodes are located and organic speech disorders are diagnosed.

Benefits of technology

It enables personalized training programs, improves the relevance and accuracy of training, ensures data purity, dynamically adjusts training difficulty, provides professional pronunciation disorder diagnosis and real-time feedback, and enhances training effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120766718B_ABST
    Figure CN120766718B_ABST
Patent Text Reader

Abstract

The application relates to the field of language training and discloses a voice auditory language training method and system based on artificial intelligence, which comprises the following steps: obtaining a preset language training scheme of a trainee, controlling a voice collection device to perform pre-training based on the preset training scheme, and obtaining actual voice data generated by the trainee in the pre-training process; determining a preset voice model that the trainee should reach in the training process, evaluating the pronunciation quality of the trainee according to the actual voice data and the preset voice model, and generating a first evaluation result or a second evaluation result. Through comparison and analysis of the actual voice data of the trainee and the preset voice model, the application can generate a personalized training adjustment scheme for each trainee, which not only considers the pronunciation characteristics of each trainee, but also can accurately optimize according to the differences in pronunciation quality, thereby enhancing the pertinence of the training.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of language training, in particular to a voice auditory language training method and system based on artificial intelligence. BACKGROUND

[0002] The traditional method usually adopts a unified training scheme, which cannot be adjusted individually according to the pronunciation characteristics of each trainer, and the training content and difficulty are usually fixed, which cannot be finely optimized according to the specific pronunciation quality difference of the trainer, so it may not meet the needs of different trainers; and the traditional method relies on artificial evaluation or simplified algorithm to judge the pronunciation quality, lacks sufficient data analysis and automatic diagnosis capability, and the pronunciation problems of the trainer may not be found and positioned in time and accurately, which may lead to problems being ignored or misdiagnosed, for example, the specific links of pronunciation cannot be analyzed in depth, and potential physiological or organic pronunciation disorders are ignored; and the traditional training method usually lacks continuous evaluation and real-time feedback, and the training progress and difficulty adjustment are relatively fixed, when the trainer makes progress in a certain stage, the training difficulty will not be automatically increased, which may lead to the progress of the training to stagnate, in addition, the adjustment of the training scheme lacks pertinence, and it is difficult to realize truly personalized training; and the traditional method is often difficult to effectively suppress environmental noise, so that the purity of the data cannot be guaranteed, external noise and other interference factors may affect the quality of the voice data, thereby reducing the training effect, the traditional method usually handles noise and reverberation simply, and cannot achieve the level of optimizing the training data; and the traditional method cannot analyze the key nodes in the pronunciation through an automatic way, lacks pronunciation abnormality analysis and organic disorder detection, and the trainer may not receive an accurate pronunciation disorder diagnosis report in time, and there is no targeted training scheme to correct these problems. SUMMARY

[0003] To achieve the above-mentioned purpose, the present application provides the following technical scheme: a voice auditory language training method based on artificial intelligence, comprising:

[0004] obtaining a preset language training scheme of a trainer, controlling a voice collection device to pre-train based on the preset training scheme, and obtaining actual voice data generated by the trainer in the pre-training process;

[0005] determining a preset voice model that the trainer should reach in the training process, evaluating the pronunciation quality of the trainer according to the actual voice data and the preset voice model, and generating a first evaluation result or a second evaluation result;

[0006] If it is the first evaluation result, it means that the pronunciation of the trainer meets the preset standard, and the next stage of intensive training can be entered;

[0007] If the second evaluation result is obtained, a correlation degree analysis is performed on the pronunciation segments of the trainer, and a key pronunciation node causing the pronunciation quality to be substandard is located;

[0008] An abnormal pronunciation analysis is performed on the key pronunciation node, and if a posterior probability of the key pronunciation node having a physical pronunciation disorder is greater than a preset probability threshold, a pronunciation disorder diagnosis report is generated;

[0009] If the posterior probability of the key pronunciation node having a physical pronunciation disorder is not greater than the preset probability threshold, an individualized training adjustment scheme is generated, and the individualized training adjustment scheme is sent to a terminal device of the trainer.

[0010] Preferably, the pronunciation quality of the trainer is evaluated according to the actual voice data and the preset voice model, and a first evaluation result or a second evaluation result is generated, including:

[0011] The actual voice data and the preset voice model are mapped into the vector space to obtain a voice feature vector space;

[0012] Voiceprint feature parameters of the actual voice data and reference feature parameters of the preset voice model are extracted according to the voice feature vector space, and the voiceprint feature parameters include a fundamental frequency, a formant frequency, a speech rate, and an intensity;

[0013] A feature parameter difference degree of the actual voice data and the preset voice model on the same pronunciation unit is calculated according to the voiceprint feature parameters and the reference feature parameters;

[0014] The feature parameter difference degree is calculated by a dynamic time warping algorithm to calculate a pronunciation time sequence offset and by a cosine similarity algorithm to calculate a voiceprint feature similarity;

[0015] The feature difference degree is compared with a preset difference threshold, if the feature difference degree is less than the preset difference threshold, a first evaluation result is generated, and if the feature difference degree is not less than the preset difference threshold, a second evaluation result is generated.

[0016] Preferably, if the second evaluation result is obtained, a correlation degree analysis is performed on the pronunciation segments of the trainer, and a key pronunciation node causing the pronunciation quality to be substandard is located, including:

[0017] If the second evaluation result is obtained, a pronunciation unit in the actual voice data whose feature difference degree exceeds a preset difference threshold is extracted, and the pronunciation unit is marked as an abnormal pronunciation segment;

[0018] Physiological parameters of the trainer in a pronunciation process are obtained, and the physiological parameters include tongue position coordinates, lip opening degrees, and airflow pressures.

[0019] Preferably, if the second evaluation result, the articulation of the trainer is analyzed, the key articulation node is located to trigger the quality of pronunciation not up to standard, and further comprises:

[0020] The abnormal pronunciation segment is time-aligned with the physiological parameters, and the abnormal value of the physiological parameters corresponding to the abnormal pronunciation segment is analyzed;

[0021] The physiological parameter abnormal value is input into the vocal organ movement model to determine the deviation degree of each vocal organ movement trajectory and the standard trajectory;

[0022] The vocal organ movement node with the deviation degree greater than the preset deviation threshold is marked as the key articulation node triggering the quality of pronunciation not up to standard.

[0023] Preferably, the key articulation node is analyzed for pronunciation abnormalities, and if the posterior probability of the key articulation node existing organic speech disorder is greater than a preset probability threshold, a speech disorder diagnosis report is generated, comprising:

[0024] Obtain the historical pronunciation data of the trainer; extract the abnormal pronunciation mode features in the historical pronunciation data; wherein the abnormal pronunciation mode features include deviation degree and feature difference degree;

[0025] A speech disorder diagnosis model based on neural network is constructed, and the input parameters of the speech disorder diagnosis model include abnormal pronunciation mode features;

[0026] The abnormal pronunciation mode features corresponding to the key articulation node are input into the speech disorder diagnosis model, and the posterior probability of the key articulation node existing organic speech disorder is output;

[0027] If the posterior probability is greater than a preset probability threshold, a speech disorder diagnosis report is generated; wherein the speech disorder diagnosis report includes the articulatory organs corresponding to the key articulation node.

[0028] Preferably, if the posterior probability of the key articulation node existing organic speech disorder is not greater than a preset probability threshold, a personalized training adjustment scheme is generated, and the personalized training adjustment scheme is sent to the terminal device of the trainer, comprising:

[0029] A knowledge base including a pronunciation correction case library and a training parameter optimization rule is constructed, and the case library includes correction strategies corresponding to different pronunciation defect types;

[0030] Extract the voiceprint feature abnormal mode features of the key articulation node, and generate a feature coding label according to the voiceprint feature abnormal mode features;

[0031] retrieve a matched correction strategy and a training parameter adjustment rule in the knowledge base based on the feature coding label.

[0032] Preferably, if the posterior probability of the key pronunciation node existing a physical pronunciation disorder is not greater than a preset probability threshold, a personalized training adjustment scheme is generated, and the personalized training adjustment scheme is sent to a terminal device of the trainer, and the method further comprises:

[0033] A personalized training adjustment scheme including pronunciation action guidance, repetition times, and training duration suggestions is generated according to the correction strategy and the training parameter adjustment rule;

[0034] The personalized training adjustment scheme is pushed to the trainer through a mobile terminal application and is synchronously updated to a control module of a speech training device.

[0035] Preferably, the speech training device is controlled based on the preset training scheme to perform pre-training, and actual speech data generated by the trainer in the pre-training process is acquired, including:

[0036] An original speech signal of the trainer in the pre-training process is collected through a speech collection device, and the original speech signal is pre-emphasized, framed, and windowed;

[0037] A deep learning denoising model is used to suppress background noise and eliminate reverberation of the processed original speech signal to obtain a pure speech signal;

[0038] The pure speech signal is subjected to endpoint detection and speech activity detection to segment out an effective pronunciation segment;

[0039] The effective pronunciation segment is stored as a structured speech data file, and a spectrogram and a voiceprint waveform graph of the actual speech data are generated through a speech visualization tool.

[0040] Preferably, the method further comprises:

[0041] The speech training device is controlled based on the personalized training adjustment scheme to perform a dynamic training mode;

[0042] When the first evaluation result appears for three times in succession, a training phase promotion mechanism is triggered, and a higher difficulty training module is entered.

[0043] An artificial intelligence-based sound auditory language training system, which is applicable to the artificial intelligence-based sound auditory language training method described above, comprises:

[0044] A speech collection unit is configured to acquire a preset language training scheme of a trainer, control a speech collection device to perform pre-training based on the preset training scheme, and acquire actual speech data generated by the trainer in the pre-training process.

[0045] A pronunciation evaluation unit is configured to determine a preset voice model that the trainer should achieve in a training process, evaluate the pronunciation quality of the trainer according to the actual voice data and the preset voice model, and generate a first evaluation result or a second evaluation result.

[0046] A reinforcement training unit is configured to, if the first evaluation result is obtained, indicate that the pronunciation of the trainer meets a preset standard, and the trainer can enter a next stage of reinforcement training.

[0047] A pronunciation positioning unit is configured to, if the second evaluation result is obtained, perform correlation analysis on the pronunciation links of the trainer, and locate key pronunciation nodes that cause the pronunciation quality to be substandard.

[0048] A disorder diagnosis unit is configured to perform pronunciation abnormality analysis on the key pronunciation nodes, and if a posterior probability of the key pronunciation nodes existing in organic pronunciation disorders is greater than a preset probability threshold, generate a pronunciation disorder diagnosis report.

[0049] A training adjustment unit is configured to, if the posterior probability of the key pronunciation nodes existing in organic pronunciation disorders is not greater than the preset probability threshold, generate an individualized training adjustment scheme, and send the individualized training adjustment scheme to a terminal device of the trainer.

[0050] Compared with the prior art, the present application has the following advantages:

[0051] (1) The present application can generate an individualized training adjustment scheme for each trainer by comparing and analyzing the actual voice data of the trainer with the preset voice model, which not only considers the pronunciation characteristics of each trainer, but also can accurately optimize according to the differences in pronunciation quality, thereby enhancing the targeting of the training. If an abnormality is found in the pronunciation quality evaluation, the pronunciation abnormality analysis can be further performed to locate specific key pronunciation nodes, and the possible pronunciation disorders can be analyzed according to the physiological parameters. With the assistance of a neural network model, it can automatically diagnose whether there is an organic pronunciation disorder, and provide a professional pronunciation disorder diagnosis report for the trainer.

[0052] (2) The present application combines voice acquisition, noise suppression, dynamic time warping and cosine similarity, etc. to make the processing of voice data more accurate, to better avoid environmental noise and other interference factors, to ensure the purity of data during the training process, and to improve the training quality. If the trainer reaches the preset standard at a certain stage, he can enter a higher difficulty training module to gradually improve his pronunciation ability. For unqualified pronunciation, the problem can be accurately found through the correlation analysis of the pronunciation link and the pronunciation organ movement model, and targeted training can be implemented.

[0053] (3) The present application can dynamically adjust the training scheme according to the progress of the trainer through continuous evaluation and feedback. For example, if the trainer performs excellently at a certain stage, the training difficulty will be automatically increased to ensure that the training does not stagnate. Through the mobile terminal application, the personalized training scheme can be pushed to the trainer's device in real time, greatly improving the convenience and real-time performance of the training, and enhancing the participation and sense of achievement of the trainer. BRIEF DESCRIPTION OF DRAWINGS

[0054] Figure 1 It is a step flow structure schematic diagram of the overall method in an embodiment of the present application.

[0055] Figure 2 It is a system architecture structure schematic diagram of the overall system in an embodiment of the present application.

[0056] In the figure: 1, voice acquisition unit; 2, pronunciation evaluation unit; 3, reinforcement training unit; 4, pronunciation positioning unit; 5, obstacle diagnosis unit; 6, training adjustment unit. DETAILED DESCRIPTION

[0057] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0058] Embodiment one, please refer to Figure 1 The present application provides a technical solution: a voice auditory language training method based on artificial intelligence, comprising:

[0059] S1, obtaining a preset language training scheme of a trainer, controlling a voice acquisition device to pre-train based on the preset training scheme, and obtaining actual voice data generated by the trainer during the pre-training process;

[0060] S2, determine the preset voice model that the trainer should reach in the training process, evaluate the pronunciation quality of the trainer according to the actual voice data and the preset voice model, and generate a first evaluation result or a second evaluation result;

[0061] S3, if the first evaluation result, the pronunciation of the trainer meets the preset standard, and the next stage of intensive training can be entered;

[0062] S4, if the second evaluation result, the correlation degree of the pronunciation of the trainer is analyzed, and the key pronunciation node causing the substandard pronunciation quality is located;

[0063] S5, the key pronunciation node is analyzed for pronunciation abnormality, if the posterior probability of the key pronunciation node existing organic speech disorder is greater than the preset probability threshold, a speech disorder diagnosis report is generated;

[0064] S6, if the posterior probability of the key pronunciation node existing organic speech disorder is not greater than the preset probability threshold, a personalized training adjustment scheme is generated, and the personalized training adjustment scheme is sent to the terminal device of the trainer.

[0065] It is necessary to explain that the trainer is conducting voice training, the purpose is to improve his pronunciation, especially the possible problems on some specific syllables; suppose the trainer's goal is to correctly pronounce the sh and ch sounds; S1: obtain the preset language training plan of the trainer and pre-train, in this stage, the trainer first sets his training goal with the system, such as learning how to correctly pronounce sh and ch; according to this goal, a voice collection device (such as a microphone) is controlled to record his voice; in the pre-training stage, the trainer starts to pronounce, and the system will collect the actual voice data generated when he pronounces; S2: determine the preset voice model to be reached and evaluate the pronunciation quality, the system has a preset standard model, which defines how to correctly pronounce sh and ch (for example, the coordination of lips, tongue and breathing); in this stage, the system compares the trainer's actual pronunciation with the preset standard and evaluates his pronunciation quality; the system will give two evaluation results: the first evaluation result: the trainer's pronunciation is consistent with the standard; the second evaluation result: the trainer's pronunciation is different from the standard; S3: if the first evaluation result; if the system's evaluation result is the first evaluation result, it means that the trainer's pronunciation has met the standard, he can enter the next stage of intensive training to continue to practice and consolidate the correct pronunciation skills; S4: if the second evaluation result; if the system's evaluation result is the second evaluation result, it means that the trainer's pronunciation still has problems; the system will analyze which pronunciation links are wrong (such as unclear sh pronunciation); it will find out the specific key pronunciation nodes through correlation analysis, such as the trainer's tongue position may be wrong when pronouncing sh; S5: pronunciation abnormality analysis and pronunciation disorder diagnosis; next, the system will make a deeper analysis of these key pronunciation nodes; for example, the system will detect whether there is a organic pronunciation disorder (such as tongue movement problem or oral structure problem); if the analysis result shows that the trainer's pronunciation problem is caused by some physiological reasons, and the probability of such problem exceeds the preset threshold (for example, 70% probability), the system will generate a pronunciation disorder diagnosis report, telling the trainer that he may have a pronunciation disorder, and suggesting further medical examination or treatment; S6: personalized training adjustment scheme; if the analysis result shows that the trainer's pronunciation problem does not meet the standard of organic disorder, the system will generate a personalized training adjustment scheme according to the specific problem; for example, the system finds that the trainer's tongue position is wrong when pronouncing sh, the system may recommend him to pay attention to the action of the tongue when pronouncing, or provide specific pronunciation training exercises (such as strengthening the control of tongue position by imitating correct pronunciation); this personalized scheme will be sent to the trainer through a terminal device (such as a mobile phone, a computer, etc.), helping him to improve his pronunciation in training.

[0066] In an optional embodiment, the actual voice data and the preset voice model are mapped into a vector space to obtain a voice feature vector space.

[0067] In an optional embodiment, the actual voice data and the preset voice model are mapped into a vector space to obtain a voice feature vector space.

[0068] In an optional embodiment, the actual voice data and the preset voice model are mapped into a vector space to obtain a voice feature vector space.

[0069] In an optional embodiment, the actual voice data and the preset voice model are mapped into a vector space to obtain a voice feature vector space.

[0070] In an optional embodiment, the actual voice data and the preset voice model are mapped into a vector space to obtain a voice feature vector space.

[0071] In an optional embodiment, the actual voice data and the preset voice model are mapped into a vector space to obtain a voice feature vector space.

[0072] It should be noted that the actual voice data and the information of the preset voice model are mapped into a vector space; in this vector space, each voice signal is converted into a vector, representing different features of the voice; this vector space will include various parameters related to sound characteristics, which will help the system compare the similarity between voices; from the voice data mapped into the vector space, the system will extract a set of key voiceprint feature parameters and benchmark feature parameters of the preset voice model; the following are the specific contents of these features: pitch: refers to the basic frequency of the sound, which determines the pitch of the voice; formant frequencies: this is an important audio component in the voice, which is related to the voice quality and accent; speech rate: refers to the speed of pronunciation; loudness: refers to the loudness of the voice; compare the feature parameter difference between the actual voice data and the preset voice model on the same pronunciation unit; the pronunciation unit here can be understood as a syllable or a word in language; the purpose of calculating these differences is to evaluate the deviation between the pronunciation of the trainer and the standard pronunciation; dynamic time warping algorithm is used to compare the similarity between two time series, which is commonly used in fields such as speech and handwriting recognition; here, the DTW algorithm will be used to calculate the time offset of pronunciation, that is, the time difference between the trainer's pronunciation and the standard pronunciation; for example, the trainer's pronunciation may be slightly faster or slower than the standard pronunciation, and the DTW algorithm can adjust the time alignment of the two sequences to calculate the time deviation; the cosine similarity algorithm is used to calculate the similarity between two vectors, and the cosine similarity will be used to calculate the similarity of voiceprint features, that is, to compare the similarity between the voiceprint features of the trainer and the standard voiceprint features; by calculating the angle between the feature vectors, the similarity of the two can be obtained, and the smaller the angle, the higher the similarity; compare the calculated feature difference with the preset difference threshold to determine whether the trainer's pronunciation meets the standard: if the feature difference is less than the preset difference threshold, it means that the trainer's pronunciation is very close to the standard pronunciation, and the system will generate the first evaluation result, indicating that the trainer's pronunciation meets the standard and can enter the next stage of training; if the feature difference is not less than the preset difference threshold, it means that the trainer's pronunciation has a large gap, and the system will generate the second evaluation result, meaning that the trainer's pronunciation needs to be further improved, and the system may provide personalized adjustment solutions.

[0073] In an optional embodiment, if the second evaluation result is obtained, a correlation analysis is performed on the pronunciation segment of the trainer to locate the key pronunciation node causing the substandard pronunciation quality, including:

[0074] If the second evaluation result is obtained, the pronunciation unit in the actual voice data whose feature difference exceeds the preset difference threshold is extracted, and the pronunciation unit is marked as an abnormal pronunciation segment;

[0075] acquiring physiological parameters of the trainer during pronunciation, the physiological parameters including tongue position coordinates, lip opening degree and airflow pressure.

[0076] It should be noted that.

[0077] In an optional embodiment, if the second evaluation result is obtained, the pronunciation segments of the trainer are subjected to correlation analysis, and key pronunciation nodes causing the substandard pronunciation quality are located, and the method further comprises the following steps of:

[0078] aligning the abnormal pronunciation segment with the physiological parameters in time sequence, and analyzing abnormal values of the physiological parameters corresponding to the abnormal pronunciation segment;

[0079] constructing a pronunciation organ movement model, inputting the abnormal values of the physiological parameters into the pronunciation organ movement model, and determining deviation degrees of movement trajectories of each pronunciation organ from standard trajectories;

[0080] marking the pronunciation organ movement nodes with the deviation degrees greater than a preset deviation threshold as the key pronunciation nodes causing the substandard pronunciation quality.

[0081] It should be noted that when the feature difference degree is calculated and the second evaluation result is obtained, it indicates that the pronunciation of the trainer is significantly different from the standard pronunciation; at this time, these pronunciation units will be further analyzed, and the pronunciation units with feature difference degree exceeding the preset threshold will be extracted and marked as abnormal pronunciation segments; these segments are parts of the trainer's pronunciation that have obvious problems, which may be due to inaccurate pronunciation, inconsistent tone or speed factors; physiological parameters of the trainer during pronunciation are recorded and obtained, which are used to further analyze the potential causes of pronunciation quality; specific physiological parameters include: tongue position coordinates: the position of the tongue during pronunciation; the position of the tongue is crucial to the clarity and accuracy of pronunciation, and tongue position errors can cause some pronunciation errors; lip opening degree: the degree of opening and closing of the lips, which affects the clarity of pronunciation; for example, if the lips are not fully closed during pronunciation, it may cause pronunciation to be unclear or unclear; airflow pressure: the size and pressure of airflow, which usually affects the vibration of the vocal cords and the intensity of pronunciation; unstable or uneven airflow can cause intermittent or non-standard pronunciation; time alignment of abnormal pronunciation segments and corresponding physiological parameters, i.e. synchronizing these physiological parameters with specific pronunciation time to analyze the relationship between them; time alignment ensures that the physiological changes of each pronunciation unit during pronunciation can be matched with specific speech features, so as to analyze the source of pronunciation quality problems; for example, the feature difference degree of a certain pronunciation unit is large, which may correspond to physiological abnormalities such as tongue position deviation, insufficient lip opening degree or unstable airflow pressure; through time alignment, it can accurately identify which physiological parameters are abnormal; after time alignment, the abnormal values of the physiological parameters corresponding to the abnormal pronunciation segments will be specially analyzed; these abnormal values represent the deviation of the trainer's physiological parameters during pronunciation from the standard state of normal pronunciation; for example, tongue position deviation, insufficient lip closure or abnormal airflow pressure will cause pronunciation distortion or unclear;

[0082] To further analyze the specific influence of physiological parameter abnormalities on pronunciation quality, a pronunciation organ movement model is constructed; this model usually includes a description of the movement trajectory of the pronunciation organs (such as the tongue, lips, vocal cords, etc.); by simulating and analyzing these movement trajectories, it can be determined whether the trainee's pronunciation organs move according to the standard trajectory; the standard trajectory refers to the ideal movement path of the pronunciation organs (such as the tongue, lips, etc.) when pronouncing correctly; after the movement model of the pronunciation organs is constructed, the abnormal values of the physiological parameters are input into the model to analyze the deviation degree of the movement trajectory of each pronunciation organ from the standard trajectory; the deviation degree refers to the difference between the movement path of each pronunciation organ of the trainee during pronunciation and the standard path; the movement of the pronunciation organ with a large deviation degree may be the root cause of inaccurate pronunciation; the larger the deviation degree, the more unstandard the movement of the trainee's pronunciation organs, which may lead to a decrease in pronunciation quality; finally, the pronunciation organ movement nodes with a deviation degree greater than a preset deviation threshold are marked as key pronunciation nodes; these key pronunciation nodes are the most important moments in the pronunciation process, and the deviation of the pronunciation organs is most serious at these nodes, thereby directly affecting the quality of the entire pronunciation; if the deviation degree of a pronunciation organ at a pronunciation node is greater than the threshold, it is considered that this node is a key node that leads to substandard pronunciation; these nodes need special attention, and the trainee needs to adjust the pronunciation method at these nodes.

[0083] In an optional embodiment, the key pronunciation node is subjected to pronunciation abnormality analysis, and if the posterior probability of the key pronunciation node existing organic speech disorder is greater than a preset probability threshold, a speech disorder diagnosis report is generated, including:

[0084] The historical pronunciation data of the trainee is obtained; the abnormal pronunciation pattern features in the historical pronunciation data are extracted; wherein the abnormal pronunciation pattern features include the deviation degree and the feature difference degree;

[0085] A speech disorder diagnosis model based on a neural network is constructed, and the input parameters of the speech disorder diagnosis model include the abnormal pronunciation pattern features;

[0086] The abnormal pronunciation pattern features corresponding to the key pronunciation node are input into the speech disorder diagnosis model, and the posterior probability of the key pronunciation node existing organic speech disorder is output;

[0087] If the posterior probability is greater than a preset probability threshold, a speech disorder diagnosis report is generated; wherein the speech disorder diagnosis report includes the pronunciation organs corresponding to the key pronunciation node.

[0088] It should be noted that the past pronunciation data of the trainer is collected; these data may come from the trainer's multiple pronunciation recordings, pronunciation tests or other pronunciation evaluation processes; historical pronunciation data provides the trainer's pronunciation characteristics over a long period of time and serves as the basis for subsequent analysis; in the historical pronunciation data, abnormal pronunciation pattern features are extracted, which reflect the problems that may exist in the trainer's pronunciation process, and are usually identified by analyzing the deviation or change of pronunciation; the abnormal pronunciation pattern features include the following two kinds: deviation: refers to the degree of deviation between the trainer's pronunciation organs (such as tongue position, lip opening degree, airflow, etc.) and the standard pronunciation organ movement trajectory; the greater the deviation, the higher the degree of movement of the pronunciation organs deviating from the standard path in the pronunciation process, which may indicate inaccurate or impaired pronunciation; feature difference: this refers to the difference between the trainer's pronunciation and the standard pronunciation, including differences in voice quality, tone, timbre, etc.; if these differences exceed the normal fluctuation range, it may indicate abnormal pronunciation; after the abnormal pronunciation pattern features are extracted, the system will build a neural network model based on these features for pronunciation disorder diagnosis; the input of this model is the abnormal pronunciation pattern features, and through learning these features, the neural network model can train the ability to identify pronunciation disorders; the model may consider the following aspects: input features: abnormal pronunciation features such as deviation and feature difference; network structure: the neural network may include multiple levels, such as convolutional neural network (CNN), long short-term memory network (LSTM), etc., for identifying complex time series patterns; target output: the goal of the model is to diagnose pronunciation disorders and determine whether the trainer's pronunciation has any disorders and whether it belongs to organic pronunciation disorders; then the abnormal pronunciation pattern features corresponding to the key pronunciation nodes (i.e. the nodes that are significantly abnormal in the pronunciation process) are input into the trained pronunciation disorder diagnosis model; the key pronunciation nodes refer to the moments in the pronunciation process when the movement or features of the pronunciation organs are severely abnormal, and these nodes are crucial for diagnosing pronunciation disorders; by inputting the features of these nodes into the model, the system can further analyze whether the node shows signs of organic pronunciation disorders; through the processing of the neural network model, the system outputs a posterior probability for each key pronunciation node, i.e. the probability of the key pronunciation node having organic pronunciation disorders; the posterior probability is based on the model's training of historical data and the analysis results of the current pronunciation pattern features to estimate the likelihood of the node having organic pronunciation disorders; if this probability is higher than the set threshold (usually a standard value derived from expert experience or data analysis), it can be considered that the node has possible pronunciation disorders; if the posterior probability is lower than the pre-set threshold, it is considered that the pronunciation performance of the node is normal; if the posterior probability of the key pronunciation node exceeds the set probability threshold, i.e. the model diagnoses that there may be organic pronunciation disorders, the system will generate a pronunciation disorder diagnosis report;The report includes the following: key articulatory nodes: the report will indicate which articulatory nodes are abnormal, and how these nodes deviate from the normal articulatory trajectory; articulatory organs: the report will also list the articulatory organs involved, i.e. which articulatory organs (such as the tongue, lips, vocal cords, etc.) have problems at these nodes; this helps to understand which part of the articulatory organs deviates from the normal trajectory, thereby causing the articulatory quality problem.

[0089] In an optional embodiment, if the posterior probability of the key articulatory node existing a physical speech disorder is not greater than a preset probability threshold, a personalized training adjustment scheme is generated, and the personalized training adjustment scheme is sent to the terminal device of the trainer, including:

[0090] A knowledge base including a speech correction case library and a training parameter optimization rule is constructed, and the case library includes correction strategies corresponding to different speech defect types;

[0091] An acoustic feature abnormal pattern feature of the key articulatory node is extracted, and a feature coding label is generated according to the acoustic feature abnormal pattern feature;

[0092] Based on the feature coding label, a matched correction strategy and training parameter adjustment rule are retrieved in the knowledge base.

[0093] It should be noted that the pronunciation correction case library stores different types of pronunciation defects and their corresponding correction strategies; these defects may be tongue position, pronunciation position error, airflow control deficiency, oral cavity structure problem, etc.; for each pronunciation defect, the knowledge base pre-stores effective correction strategies, such as: tongue position error: can be corrected by tongue training, imitation training, etc.; pitch too high or too low: adjust the pitch and tone of the voice; oral cavity structure problem: if the pronunciation problem is caused by structural barriers, physical correction or surgery of the pronunciation organ can be performed; these correction strategies are usually based on professional linguistic research, clinical experience and the latest pronunciation correction technology, aiming to provide targeted treatment for each pronunciation defect; the training parameter optimization rules store parameter adjustment rules related to the training and treatment process; these rules are used to optimize pronunciation practice parameters during training; for example: training intensity (such as pronunciation training frequency, duration, volume); attention to sound characteristics (such as pitch, tone, pronunciation position, etc.); simulation feedback mode (such as voice recognition system feedback frequency, difficulty setting, etc.); During pronunciation training or evaluation, the system will analyze the voiceprint features of the trainee; voiceprint features refer to the pronunciation sound wave features of each person, including timbre, pitch, tone, speed, voice duration, and other multi-dimensional information; key pronunciation nodes refer to time nodes that show significant abnormalities during pronunciation; for example, when pronouncing a certain syllable, the pronunciation organ deviates from the standard action, or the sound changes beyond the normal range; these nodes usually reflect the main problems of the trainee during pronunciation; voiceprint feature abnormal pattern features refer to the abnormal voiceprint features of these nodes; for example: pronunciation pitch is too high or too low; pronunciation timbre distortion; pronunciation speed is not uniform; by analyzing the voiceprint features of pronunciation, the abnormal pronunciation patterns of the trainee at these key pronunciation nodes can be identified, and the corresponding features can be extracted; once the abnormal patterns of voiceprint features are extracted, the next step is to convert these features into a feature encoding label; the label is a symbol or numerical representation that can encode these abnormal patterns for subsequent processing; the generation of feature encoding labels is a feature compression process that converts the original complex voiceprint features into a more concise form; these labels usually reflect the type, severity, frequency, etc. of pronunciation defects; for example, a trainee may generate the following labels: Label 1: pitch too low, duration too long; Label 2: timbre distortion, pronunciation organ position deviation; Label 3: airflow deficiency, pronunciation unclear; these labels can quantify the abnormal features in the pronunciation process, facilitating subsequent diagnosis and treatment; using the generated feature encoding labels, match in the previously constructed pronunciation correction case library; according to the label, the system will retrieve the correction strategies and training parameter optimization rules related to it; according to the label, the corresponding pronunciation defect type is matched, and the corresponding correction strategy is retrieved from the case library;For example, if the label shows that the tone is too low, the system can recommend adjusting the frequency range of pronunciation, or provide specific pronunciation exercises to help the trainer adjust the pitch; adjust the training parameters according to the pronunciation problems described in the label; if the label shows that the airflow is insufficient, the system can automatically increase the intensity of pronunciation training, or adjust the volume and syllable duration, etc.

[0094] In an optional embodiment, if the posterior probability of the key pronunciation node existing organic pronunciation disorder is not greater than the preset probability threshold, a personalized training adjustment scheme is generated, and the personalized training adjustment scheme is sent to the terminal device of the trainer, and the method further comprises:

[0095] According to the correction strategy and the training parameter adjustment rule, a personalized training adjustment scheme including pronunciation action guidance, repetition times and training duration suggestion is generated;

[0096] The personalized training adjustment scheme is pushed to the trainer through a mobile terminal application, and is synchronously updated to the control module of the speech training device.

[0097] It should be noted that,

[0098] In an optional embodiment, the speech acquisition device is controlled based on a preset training scheme for pre-training, and actual speech data generated by the trainer in the pre-training process is acquired, comprising:

[0099] The original speech signal of the trainer in the pre-training process is collected through the speech acquisition device, and the original speech signal is pre-emphasized, framed and windowed;

[0100] The processed original speech signal is subjected to background noise suppression and reverberation elimination by using a deep learning noise reduction model to obtain a pure speech signal;

[0101] The pure speech signal is subjected to endpoint detection and speech activity detection to segment out effective pronunciation segments;

[0102] The effective pronunciation segments are stored as structured speech data files, and a spectrogram and a voiceprint waveform graph of the actual speech data are generated by using a speech visualization tool.

[0103] It should be noted that the pronunciation action guidance refers to detailed pronunciation action training guidance; for example, the system can provide specific suggestions on how to adjust the tongue position, airflow control, mouth shape, and pronunciation method of syllables; for example, if the trainer has a tongue position deviation when pronouncing, the system may provide the following guidance: slightly raise the tongue, touch the gum position of the upper palate, or keep the tongue tip from touching the hard palate when pronouncing, which are the details of pronunciation correction; the number of repetitions refers to how many times each training action or pronunciation exercise needs to be performed; based on the severity of pronunciation defects and individual differences, the system can automatically set the number of repetitions for each action or exercise; for example, for a difficult syllable in pronunciation, the system may recommend repeating it 100 times a day to deepen muscle memory and coordination of oral movements; the training duration recommendation refers to the length of time for each training session; the personalized training plan will set a reasonable training duration based on the trainer's ability and needs, ensuring that the training is both effective and not overly fatiguing; for example, if the trainer's speech problem is mild, the system may recommend training for 15 minutes; if the problem is more severe, the system may recommend training for 30 minutes to 1 hour in multiple sessions; after generating the personalized training adjustment plan, the next step is to push it to the trainer's device through the mobile terminal application; the key point of this process is to push the personalized training adjustment plan to the trainer through mobile devices such as smartphones, tablets, or smartwatches; when the personalized training adjustment plan is generated, the system will immediately push this information to the trainer, ensuring that the trainer can perform the training on time; the application will display the training action guidance through a graphical interface and can provide voice, text, or video feedback to ensure that the trainer can clearly understand each action and operation step; the application program can set training reminders to remind the trainer to train at regular intervals, ensuring that the training process does not interrupt and that the training plan is fully executed; for example, the application program may prompt: today's pronunciation training begins, repeat the'shi' pronunciation 10 times for 10 minutes; in addition to mobile terminals, the training system also includes voice training devices, which may include voice training tools with microphones, speakers, or smart earphones, voice feedback devices, etc.; the control module of the voice training device refers to the core hardware and software part of the device, which controls the operation, feedback, and adjustment of the device; the system will synchronize the personalized training adjustment plan to the control module of the device, including: training plan synchronization: transmit the personalized pronunciation training plan (such as pronunciation action, number of repetitions, and training duration) to the voice training device; the control module sets the working mode of the voice training device according to these plans; for example, the device may set the frequency, tone, etc. of each pronunciation according to the synchronized plan; the control module of the training device can provide real-time feedback based on the trainer's pronunciation performance; if the trainer's pronunciation does not meet expectations, the device will provide correction suggestions through acoustic feedback or image feedback (such as a display screen or voice prompt);For example, the device can display the tongue position deviation when pronouncing through the display, or slightly raise the tone through audio prompts; the control module will record the training progress and feedback to the system and the trainer; for example, the system will track the duration of each training, the completion of the number of action repetitions, etc., and continuously optimize the training plan based on these data; the device adjusts the training intensity according to real-time feedback to ensure that the individualized characteristics of the training program are implemented; for example, if the trainer performs well on a certain action, the device may reduce the number of repetitions of that action and increase other more challenging exercises.

[0104] In an optional embodiment, the method further comprises:

[0105] Controlling the speech training device to perform a dynamic training mode based on the individualized training adjustment scheme;

[0106] When the first evaluation result appears for three consecutive times, a training phase promotion mechanism is triggered, and a higher difficulty training module is entered.

[0107] Embodiment two, please refer to Figure 2 The present application provides a technical solution: an artificial intelligence-based sound hearing language training system, which is applicable to the above-mentioned artificial intelligence-based sound hearing language training method, comprising:

[0108] The speech collection unit 1 is used to obtain a preset language training scheme of the trainer, control the speech collection device to perform pre-training based on the preset training scheme, and obtain actual speech data generated by the trainer during the pre-training process;

[0109] The pronunciation evaluation unit 2 is used to determine a preset speech model that the trainer should achieve during the training process, evaluate the pronunciation quality of the trainer according to the actual speech data and the preset speech model, and generate a first evaluation result or a second evaluation result;

[0110] The intensive training unit 3 is used to, if the first evaluation result is obtained, indicate that the pronunciation of the trainer meets the preset standard, and enter the next stage of intensive training;

[0111] The pronunciation positioning unit 4 is used to, if the second evaluation result is obtained, perform correlation analysis on the pronunciation links of the trainer, and locate the key pronunciation nodes that cause the substandard pronunciation quality;

[0112] The obstacle diagnosis unit 5 is used to perform pronunciation abnormality analysis on the key pronunciation nodes, and if the posterior probability of the key pronunciation nodes existing organic pronunciation obstacles is greater than a preset probability threshold, a pronunciation obstacle diagnosis report is generated;

[0113] The training adjustment unit 6 is configured to generate a personalized training adjustment scheme if the posterior probability of the key pronunciation node existing the organic speech disorder is not greater than a preset probability threshold, and send the personalized training adjustment scheme to a terminal device of the trainee.

[0114] The embodiments of the present application are described in detail above with reference to the drawings, but the present application is not limited thereto, and various changes can be made within the knowledge of those skilled in the art without departing from the spirit of the present application.

Claims

1. A voice auditory language training method based on artificial intelligence, characterized by , comprising: acquiring a preset language training scheme of a trainee, controlling a voice collection device to pre-train based on the preset language training scheme, and acquiring actual voice data generated by the trainee in the pre-training process; determining a preset voice model that the trainee should achieve in the training process, evaluating the pronunciation quality of the trainee according to the actual voice data and the preset voice model, and generating a first evaluation result or a second evaluation result; if it is the first evaluation result, it means that the pronunciation of the trainee meets the preset standard, and the next stage of intensive training can be entered; if it is the second evaluation result, the correlation analysis of the pronunciation links of the trainee is performed to locate the key pronunciation node that causes the substandard pronunciation quality; performing pronunciation abnormality analysis on the key pronunciation node, and if the posterior probability of the key pronunciation node existing organic speech disorder is greater than a preset probability threshold, a speech disorder diagnosis report is generated; if the posterior probability of the key pronunciation node existing organic speech disorder is not greater than the preset probability threshold, an individualized training adjustment scheme is generated and sent to the terminal device of the trainee; wherein, if it is the second evaluation result, the correlation analysis of the pronunciation links of the trainee is performed to locate the key pronunciation node that causes the substandard pronunciation quality, comprising: if it is the second evaluation result, the pronunciation unit with a feature parameter difference exceeding a preset difference threshold in the actual voice data is extracted, and the pronunciation unit is marked as an abnormal pronunciation segment; acquiring physiological parameters of the trainee in the pronunciation process, the physiological parameters including tongue position coordinates, lip opening degree, and airflow pressure; wherein, performing pronunciation abnormality analysis on the key pronunciation node, and if the posterior probability of the key pronunciation node existing organic speech disorder is greater than a preset probability threshold, a speech disorder diagnosis report is generated, comprising: acquiring historical pronunciation data of the trainee; extracting abnormal pronunciation mode features in the historical pronunciation data; wherein the abnormal pronunciation mode features include deviation and feature parameter difference; constructing a speech disorder diagnosis model based on a neural network, the input parameters of the speech disorder diagnosis model including abnormal pronunciation mode features; inputting the abnormal pronunciation mode features corresponding to the key pronunciation node into the speech disorder diagnosis model to output the posterior probability of the key pronunciation node existing organic speech disorder; if the posterior probability is greater than a preset probability threshold, a speech disorder diagnosis report is generated; wherein the speech disorder diagnosis report includes the pronunciation organ corresponding to the key pronunciation node. 2.The voice auditory language training method based on artificial intelligence according to claim 1, wherein, evaluating the pronunciation quality of the trainee according to the actual voice data and the preset voice model to generate a first evaluation result or a second evaluation result, comprising: mapping the actual voice data and the preset voice model into a vector space to obtain a voice feature vector space; extracting voiceprint feature parameters of the actual voice data and reference feature parameters of the preset voice model according to the voice feature vector space, the voiceprint feature parameters including fundamental frequency, formant frequency, speech rate, and sound intensity; Calculate a feature parameter difference degree of the actual speech data and the preset speech model on the same pronunciation unit according to the voiceprint feature parameter and the reference feature parameter; The feature parameter difference degree is calculated by a dynamic time warping algorithm to calculate a pronunciation time sequence offset and by a cosine similarity algorithm to calculate a voiceprint feature similarity; Compare the feature parameter difference degree with a preset difference threshold, if the feature parameter difference degree is less than the preset difference threshold, a first evaluation result is generated, and if the feature parameter difference degree is not less than the preset difference threshold, a second evaluation result is generated. 3.The voice auditory language training method based on artificial intelligence according to claim 2, characterized in that, If the second evaluation result is generated, a correlation degree analysis is performed on a pronunciation link of the trainer, and a key pronunciation node causing the substandard pronunciation quality is located, and the method further includes: Align the abnormal pronunciation segment and a physiological parameter in time sequence, analyze an abnormal value of the physiological parameter corresponding to the abnormal pronunciation segment; Build a pronunciation organ movement model, input the abnormal value of the physiological parameter into the pronunciation organ movement model, and determine a deviation degree of each pronunciation organ movement trajectory from a standard trajectory; Mark a pronunciation organ movement node with a deviation degree greater than a preset deviation threshold as the key pronunciation node causing the substandard pronunciation quality. 4.The voice auditory language training method based on artificial intelligence according to claim 3, characterized in that: If a posterior probability of the key pronunciation node having a physical pronunciation disorder is not greater than a preset probability threshold, an individualized training adjustment scheme is generated, and the individualized training adjustment scheme is sent to a terminal device of the trainer, and the method further includes: Build a knowledge base including a pronunciation correction case library and a training parameter optimization rule, and the case library includes correction strategies corresponding to different pronunciation defect types; Extract a voiceprint feature abnormal mode feature of the key pronunciation node, and generate a feature coding label according to the voiceprint feature abnormal mode feature; Search for matched correction strategies and training parameter adjustment rules in the knowledge base based on the feature coding label. 5.The voice auditory language training method based on artificial intelligence according to claim 4, wherein, If the posterior probability of the key pronunciation node having a physical pronunciation disorder is not greater than the preset probability threshold, the individualized training adjustment scheme is generated, and the individualized training adjustment scheme is sent to the terminal device of the trainer, and the method further includes: Generate an individualized training adjustment scheme including pronunciation action guidance, repetition times and training duration suggestions according to the correction strategies and the training parameter adjustment rules; Push the individualized training adjustment scheme to the trainer through a mobile terminal application, and synchronously update the individualized training adjustment scheme to a control module of a speech training device. 6.The voice auditory language training method based on artificial intelligence according to claim 5, wherein, Control a speech collection device to pre-train based on the preset language training scheme, and obtain actual speech data generated by the trainer in the pre-training process, and the method includes: Collect original speech signals of the trainer in the pre-training process through the speech collection device, pre-emphasize, frame and window the original speech signals; Use a deep learning noise reduction model to suppress background noise and eliminate reverberation of the processed original speech signals to obtain pure speech signals; Perform endpoint detection and speech activity detection on the pure speech signals to segment out effective pronunciation segments; Store the effective pronunciation segments as structured speech data files, and generate a spectrogram and a voiceprint waveform graph of the actual speech data through a speech visualization tool. 7.The voice auditory language training method based on artificial intelligence according to claim 6, wherein, The method further comprises: controlling the speech training device to perform a dynamic training mode based on the personalized training adjustment scheme; when the first evaluation result appears for three times in succession, triggering a training phase promotion mechanism to enter a higher difficulty training module.

8. The voice auditory language training system based on artificial intelligence, which is suitable for the voice auditory language training method based on artificial intelligence according to any one of claims 1-7, characterized in that, Comprise: a speech collection unit (1) configured to obtain a preset language training scheme of a trainee, control a speech collection device to perform pre-training based on the preset language training scheme, and obtain actual speech data generated by the trainee in the pre-training process; a pronunciation evaluation unit (2) configured to determine a preset speech model that the trainee should achieve in a training process, evaluate the pronunciation quality of the trainee according to the actual speech data and the preset speech model, and generate a first evaluation result or a second evaluation result; a reinforcement training unit (3) configured to, if the first evaluation result is obtained, indicate that the pronunciation of the trainee meets a preset standard, and enter a next stage of reinforcement training; a pronunciation positioning unit (4) configured to, if the second evaluation result is obtained, perform correlation analysis on a pronunciation link of the trainee, and locate a key pronunciation node that causes the pronunciation quality to be substandard; a disorder diagnosis unit (5) configured to perform pronunciation abnormality analysis on the key pronunciation node, and if a posterior probability of the key pronunciation node having a physical pronunciation disorder is greater than a preset probability threshold, generate a pronunciation disorder diagnosis report; a training adjustment unit (6) configured to, if the posterior probability of the key pronunciation node having the physical pronunciation disorder is not greater than the preset probability threshold, generate a personalized training adjustment scheme, and send the personalized training adjustment scheme to a terminal device of the trainee.

Citation Information

Patent Citations

  • Phoneme-level low-power consumption spoken language assessment and defect diagnosis method

    CN103985392A

  • Personalized spoken foreign language learning system and method

    CN105654785A