Speech recognition and training method for speech disorder rehabilitation patient
By constructing standard word clusters and pronunciation word clusters, assessing the degree and difficulty of patients' pronunciation disorders, and formulating personalized training plans, the problem of the existing technology being unable to accurately identify pronunciation disorders in neurosurgery patients is solved, and the training effect is improved.
Patent Information
- Application Number
- CN202510931372.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-07-07
AI Technical Summary
Existing speech training methods for neurosurgical patients with speech disorders cannot accurately identify patients' individualized pronunciation disorders, resulting in poor training results.
By obtaining speech data from neurosurgery patients, using pre-trained neural networks to screen out problem words, constructing feature vectors, forming standard word clusters and pronunciation word clusters, combining the patient's text scores and the proportion of problem words, assessing the degree of pronunciation disorder and difficulty, and formulating personalized training plans.
Accurately identifying the patient's individualized pronunciation disorder improves the pertinence and effectiveness of training and accelerates the rehabilitation process.
Smart Images

Figure CN120612926A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech recognition, and in particular to a speech recognition and training method for speech disorder rehabilitation patients. Background Art
[0002] For neurosurgical patients, due to the variety of language disorders, each type has different causes, pathological mechanisms, and effects on pronunciation, which in turn lead to different effects on pronunciation.
[0003] When conducting speech training for neurosurgical patients with speech disorders, since the brain damage of neurosurgical patients varies, if only traditional standardized processes are used to train various pronunciations word by word, the pronunciation characteristics and speech error patterns of different patients will be ignored, which will be difficult to meet the actual needs of patients, and will not be conducive to effective intervention in differentiated pronunciation disorders, which will greatly reduce the training effect. Summary of the Invention
[0004] In order to solve the technical problem that conventional training methods cannot accurately identify patients' personalized pronunciation disorders, which affects the rehabilitation process, the purpose of the present invention is to provide a speech recognition and training method for patients recovering from speech disorders. The technical solutions adopted are as follows:
[0005] Acquire speech data of neurosurgery patients responding to predetermined texts and standard speech data, and split them into separate words; use a pre-trained neural network to filter out problematic words pronounced by the patient and obtain a standard pronunciation score for each word;
[0006] Constructing a feature vector for each word based on the number of phonemes, speech duration, and pronunciation intensity of each word in the speech data; classifying the words to obtain standard word clusters based on the differences between the feature vectors of different words in standard speech; and classifying the patient's words to obtain pronunciation word clusters based on the differences between the feature vectors of each word of the patient and the feature vectors of words in the standard word clusters;
[0007] Based on the text differences and feature vector differences between the pronunciation character cluster and the corresponding standard character cluster, combined with the standard score of the patient's text and the proportion of the problem characters, the pronunciation disorder degree of each problem character of the patient is obtained; based on the distribution of the problem characters in the sentences to which they belong, combined with the pronunciation disorder degree and frequency of occurrence of the problem characters, the pronunciation difficulty of each problem character is obtained.
[0008] Furthermore, the method for obtaining the degree of pronunciation disorder includes:
[0009] Obtaining the classification accuracy of the patient's pronunciation character cluster based on the difference in the number and type of characters between the pronunciation character cluster and the corresponding standard character cluster;
[0010] Select any of the patient's problem words as a target word; obtain the pronunciation abnormality of the target word based on the difference between the feature vector of the target word and the feature vector of each word in the corresponding standard word cluster;
[0011] The pronunciation disorder degree of the target word of the patient is obtained based on the classification accuracy of the pronunciation word cluster to which the target word belongs, the proportion of the problem words and the standard scores of all words, combined with the pronunciation abnormality degree of the target word; the classification accuracy is negatively correlated with the pronunciation disorder degree.
[0012] Furthermore, the method for obtaining the classification accuracy includes:
[0013] The classification accuracy of the patient's characters is obtained based on the absolute value of the difference between the number of characters in the pronounced character cluster and the number of characters in the corresponding standard character cluster, combined with the intersection-over-union ratio of the pronounced character cluster and the corresponding standard character cluster.
[0014] Furthermore, the method for obtaining the pronunciation difficulty includes:
[0015] Obtain the continuous pronunciation difficulty of each sentence based on the distribution of the problem words in each sentence;
[0016] Combining the frequency of occurrence of each problem word in the predetermined text and the pronunciation disorder degree to obtain a difficulty weight of each problem word;
[0017] For each sentence containing any of the problem words, the pronunciation sub-difficulty corresponding to each sentence is obtained based on the continuous pronunciation difficulty and the difficulty weights of all the problem words in the sentence; and the average of the pronunciation sub-difficulties of all the sentences containing the problem word is used as the pronunciation difficulty corresponding to the problem word.
[0018] Furthermore, the method for obtaining the difficulty of continuous pronunciation includes:
[0019] The continuous pronunciation difficulty of each sentence is obtained according to the number of the problem words in each sentence and the maximum consecutive number of the problem words.
[0020] Furthermore, the method for obtaining the standard character cluster includes:
[0021] The characters are clustered according to the Euclidean distance between the feature vectors of different characters to obtain standard character clusters.
[0022] Furthermore, the method for obtaining the pronunciation character clusters includes:
[0023] Obtaining the mean of the feature vectors of all characters in each standard character cluster as a representative feature vector;
[0024] The standard character cluster whose representative feature vector is most similar to the feature vector of the patient's text is used as the standard character cluster matched by the patient's corresponding text; all the patient's texts matched by the same standard character cluster are divided into one category to obtain a pronunciation character cluster.
[0025] Furthermore, the mean of the amplitude values in the speech signal of the text is used as the pronunciation intensity of the corresponding text.
[0026] Furthermore, the text splitting method includes: using Wav2Vec to convert the voice signal in the voice data into a text sequence, and splitting the text sequence into separate characters.
[0027] Furthermore, after obtaining the pronunciation difficulty of each of the problem words, the method further includes:
[0028] According to the pronunciation disorder degree and the pronunciation difficulty of the problem word, the preset training times of the problem word are adjusted to obtain the corrected training times.
[0029] The present invention has the following beneficial effects:
[0030] The present invention first extracts text and problem words from speech data, which facilitates analysis of the pronunciation disorder manifestations of different texts of the patient, obtains a standard score for the text to quantify the degree of pronunciation standardization, and provides a basis for subsequent analysis of pronunciation disorder; further, a feature vector for each text is constructed based on the number of phonemes, speech duration, and pronunciation intensity to characterize the pronunciation characteristics of the text, providing a basis for text classification; further, texts with standard pronunciation are classified to obtain standard word clusters, and texts pronounced by the patient are classified to obtain pronunciation word clusters, which facilitates analysis of the patient's pronunciation deviation for a class of similar-pronunciation texts and accurately identifies the type and pattern of the patient's pronunciation deviation; further, based on the text differences and feature vector differences between the pronunciation word clusters and the corresponding standard word clusters, as well as the proportion of problem words, the actual pronunciation deviation of the patient is reflected from the overall perspective of a class of texts, and the pronunciation deviation of the patient is reflected from the local perspective of a single word based on the standard score of the patient's text, and the pronunciation disorder degree is comprehensively obtained to comprehensively evaluate the patient's pronunciation disorder; finally, the pronunciation disorder degree is supplemented by combining the distribution of problem words in the sentence and the frequency of occurrence of the problem words themselves, and the pronunciation difficulty of each problem word is obtained, thereby accurately identifying the patient's personalized pronunciation disorder, providing support for relevant personnel to formulate personalized training plans and improve intervention effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0032] Figure 1 A flowchart of a method for speech recognition and training for speech disorder rehabilitation patients provided by one embodiment of the present invention;
[0033] Figure 2 A flowchart of a method for obtaining a degree of articulation disorder provided by one embodiment of the present invention;
[0034] Figure 3 This is a flowchart of a method for obtaining pronunciation difficulty provided by one embodiment of the present invention. DETAILED DESCRIPTION
[0035] In order to further illustrate the technical means and effects adopted by the present invention to achieve the predetermined purpose of the invention, the following, in conjunction with the accompanying drawings and preferred embodiments, describes in detail the specific implementation, structure, features and effects of a speech recognition and training method for speech disorder rehabilitation patients proposed by the present invention. In the following description, different "one embodiment" or "another embodiment" does not necessarily refer to the same embodiment. In addition, specific features, structures or characteristics of one or more embodiments may be combined in any suitable form.
[0036] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0037] The following describes in detail a specific solution of a speech recognition and training method for speech disorder rehabilitation patients provided by the present invention in conjunction with the accompanying drawings.
[0038] See also Figure 1 , which shows a flow chart of a method for speech recognition and training for speech disorder rehabilitation patients provided by one embodiment of the present invention, specifically including:
[0039] Step S1: Acquire speech data of a neurosurgery patient for a predetermined text and speech data of a standard speech, and split them into separate words; use a pre-trained neural network to filter out problematic words pronounced by the patient and obtain a standard score for the pronunciation of each word.
[0040] Due to brain nerve damage, the causes and manifestations of speech disorders in neurosurgery patients are more complex and diverse. Different patients have different degrees and types of disorders in the pronunciation of different words due to differences in brain damaged areas and the degree of damage. Therefore, it is necessary to compare with standard speech to screen out problematic words with large pronunciation deviations. This will help locate the patient's specific pronunciation defects caused by nerve damage, accurately identify the patient's personalized pronunciation disorder, provide a reliable basis for relevant personnel, and facilitate the development of targeted training plans.
[0041] For example, if it is determined that the patient has a common pronunciation problem with retroflex consonants, retroflex consonants should be the focus of attention and training in the subsequent training process.
[0042] In one embodiment of the present invention, a predetermined text and voice data of a standard speech of the predetermined text are prepared;
[0043] When collecting voice samples, you need to choose a quiet room with good sound insulation as the collection location to reduce external noise interference and ensure that the collected voice samples are clear and intelligible;
[0044] Selection of voice collection equipment: You can use voice recording devices such as voice recorders and professional microphones to ensure that the collected voice has high sound quality and clarity. When using, the microphone should be placed in a suitable position, generally about 20-40 cm away from the patient's mouth to avoid the phenomenon of microphone splashing due to being too close or the collected sound being too small due to being too far away;
[0045] For the voice data collected from neurosurgery patients, professional audio editing software is used to remove background noise from the collected voice, while preserving the patient's voice characteristics, minimizing noise interference and improving voice clarity and recognizability. The silent parts in the voice are then identified and deleted, and the voice is organized and stored according to the patient's name, collection time, etc.
[0046] Through analog-to-digital conversion, voice data is converted into digital signals.
[0047] Preferably, in one embodiment of the present invention, Wav2Vec is used to convert the voice signal in the voice data into a text sequence, and the text sequence is split into separate characters; at the same time, the voice data of the predetermined text is also split to obtain the corresponding characters;
[0048] Furthermore, the Mel-Frequency Cepstral Coefficients (MFCC) are used to extract the acoustic feature vectors of each word pronunciation of the patient and the predetermined text; and the patient's pronunciation of each word is matched with the standard pronunciation of the predetermined text using a pre-trained neural network:
[0049] The input of the model is the acoustic feature vector of each word of the patient and the acoustic feature vector of the standard pronunciation of the given text. The output of the model is the judgment result ("yes" or "no") on whether the pronunciation of each word of the patient matches the standard pronunciation, as well as the standard score corresponding to each word, and the standard score is normalized to the range of 0-1; the pre-trained neural network is a convolutional neural network; 0 indicates a complete mismatch and a completely non-standard pronunciation, and 1 indicates a complete match and a completely standard pronunciation; the text with a no match result is a problem word pronounced by the patient.
[0050] It should be noted that Wav2Vec, Mel-frequency cepstral coefficients and methods of training neural network comparison vectors are all technical means well known to those skilled in the art. The choice of predetermined text can be selected by the implementer based on factors such as the age and educational level of the neurosurgery patient, without any limitation.
[0051] Step S2: Construct a feature vector for each word based on the number of phonemes, speech duration, and pronunciation intensity of each word in the speech data; classify the words to obtain standard word clusters based on the differences between the feature vectors of different words in the standard speech; classify the patient's words to obtain pronunciation word clusters based on the differences between the feature vectors of each word of the patient and the feature vectors of the words in the standard word clusters.
[0052] Considering that the acoustic features, tones, and phonemes of each word in the predetermined text are not exactly the same, but the pronunciations of different words may be similar, the words are first classified based on the pronunciation characteristics of the words in the standard speech. This allows the patient's pronunciation deviation for a class of similar-sounding words to be analyzed and the type and pattern of the patient's pronunciation deviation to be accurately identified.
[0053] Considering that phonemes are the smallest units that constitute syllables and represent the complexity of speech structure; speech duration can reflect the duration of text pronunciation, and pronunciation intensity can reflect the volume control when pronouncing text, so we first construct a feature vector for each text based on the number of phonemes, speech duration and pronunciation intensity of each text in the speech data to represent the pronunciation characteristics of the text and provide a basis for text classification; then, according to the differences between the feature vectors of different texts in standard speech, the text is classified to obtain standard character clusters.
[0054] In one embodiment of the present invention, considering that the Euclidean distance can measure the difference between the feature vectors of different characters, the Euclidean distance is used as the clustering distance, and the characters are clustered according to the Euclidean distance between the feature vectors of different characters to obtain standard character clusters.
[0055] It should be noted that before calculating the Euclidean distance between eigenvectors, the values of each dimension of the eigenvector can be normalized, such as linear normalization, to avoid the impact of differences in the value ranges of data in different dimensions, so that each dimension has the same impact on the clustering distance; when clustering, the K-means clustering algorithm can be used for clustering. It and the Euclidean distance are both existing technologies and will not be repeated here.
[0056] It should be noted that, in one embodiment of the present invention, when dividing the text, the voice signal is also divided, and the average of the amplitude values in the voice signal of the text is used as the pronunciation intensity of the corresponding text.
[0057] After classifying the text with normal pronunciation, the text of neurosurgery patients can be compared and matched with the standard word clusters. The patient's pronounced text can be classified with the help of the normal similar pronunciation classification results, which is convenient for subsequent identification of the deviation type of the patient's actual pronunciation and quantification of the pronunciation difficulty of the text; considering that the difference between the feature vector of the patient's pronounced text and the feature vector of the text in the standard word cluster represents the degree of matching between the two, the patient's text is classified according to the difference between the feature vector of each patient's text and the feature vector of the text in the standard word cluster to obtain the pronunciation word cluster.
[0058] Preferably, in one embodiment of the present invention, considering that the mean of all feature vectors in a cluster represents the overall pronunciation characteristics of all characters, the mean of the feature vectors of all characters in each standard character cluster is obtained as a representative feature vector to provide a matching basis for the patient's pronunciation character classification;
[0059] Taking into account the existence of multiple standard character clusters and multiple representative feature vectors, the higher the similarity between the feature vector of the patient's pronounced text and a certain representative feature vector, the more closely the actual pronunciation of the text matches the pronunciation of this type of standard text. Therefore, the standard character cluster with the representative feature vector that is most similar to the feature vector of the patient's text is used as the standard character cluster that matches the patient's corresponding text; all the patient's text that matches the same standard character cluster is divided into one category to obtain a pronunciation character cluster.
[0060] As an example, the similarity between the feature vector and the representative feature vector is measured by Euclidean distance. The smaller the Euclidean distance, the higher the similarity, which indicates the difference between the feature vector of each word of the patient and the feature vector of the words in the standard word cluster. The standard word cluster with the smallest Euclidean distance is selected as the standard word cluster that matches the corresponding word of the patient, thereby completing the classification of the patient's actual pronounced words. At this time, the patient's pronunciation word cluster corresponds one-to-one with the standard word cluster.
[0061] In other embodiments of the present invention, the implementer may also choose to match the feature vector of the patient's pronunciation text with the feature vector of the standard pronunciation of each text, and select the standard character cluster containing the standard pronunciation text with the smallest Euclidean distance as the standard character cluster matching the patient's corresponding text.
[0062] Step S3: Based on the text differences and feature vector differences between the pronunciation character cluster and the corresponding standard character cluster, combined with the standard score of the patient's text and the proportion of problem words, the pronunciation disorder degree of each problem word of the patient is obtained; based on the distribution of the problem words in the sentence to which they belong, combined with the pronunciation disorder degree and frequency of occurrence of the problem words, the pronunciation difficulty of each problem word is obtained.
[0063] Taking into account the differences in characters and character feature vectors between the patient's actual pronunciation character clusters and the corresponding standard character clusters, the deviation of the classification results of the patient's actual pronunciation characters compared with the classification results of the standard pronunciation is reflected from the overall perspective of the patient's actual pronunciation deviation.
[0064] The standard score of the patient's text represents the standard score of the neural network for the patient's text pronunciation, and reflects the patient's pronunciation deviation from the local perspective of a single word. Therefore, based on the text differences and feature vector differences between the pronunciation word cluster and the corresponding standard word cluster, combined with the standard score of the patient's text and the proportion of problem words, the pronunciation disorder degree of each problem word of the patient is obtained, and the patient's pronunciation disorder is comprehensively evaluated, which improves the accuracy and interpretability of the identification of pronunciation problems in patients with neurosurgery language disorders.
[0065] Preferably, in one embodiment of the present invention, see Figure 2 , which shows a flow chart of a method for obtaining a degree of pronunciation disorder provided by one embodiment of the present invention, specifically comprising:
[0066] Step S301: obtaining the classification accuracy of the patient's pronunciation character cluster according to the difference in the number and type of characters included in the pronunciation character cluster and the corresponding standard character cluster.
[0067] Considering that the patient's pronunciation is completely standard, the feature vector of the actual pronunciation text is highly similar to the feature vector of the standard pronunciation text, and the classification result is close to the ideal complete consistency, which is reflected in the complete consistency of the number and type of text. On the contrary, the greater the difference in the number and type of text, the lower the accuracy of the patient's text classification based on the feature vector and the standard word cluster, indicating that the patient's pronunciation deviation is greater and the patient's pronunciation disorder is greater, providing a basis for obtaining the degree of pronunciation disorder.
[0068] In one embodiment of the present invention, the classification accuracy of the patient's text is obtained based on the absolute value of the difference between the number of characters in the pronunciation character cluster and the number of characters in the corresponding standard character cluster, combined with the intersection-over-union ratio of the pronunciation character cluster and the corresponding standard character cluster.
[0069] As an example: the sum of the absolute value of the difference between the number of characters in the pronunciation character cluster and the number of characters in the corresponding standard character cluster and the preset positive division parameter 0.1 is used as the first denominator; the intersection and union of the characters contained in the pronunciation character cluster and the corresponding standard character cluster are obtained, and the ratio of the number of characters in the intersection to the number of characters in the union is used as the intersection-union ratio, the intersection-union ratio is used as the first numerator, and the ratio of the fractions corresponding to the first denominator and the first numerator is used as the classification accuracy.
[0070] The absolute value of the difference is used to express the difference in the number of characters contained in the pronunciation character cluster and the corresponding standard character cluster; the difference in the type of characters is expressed with the help of the intersection-union ratio, and the classification accuracy is obtained by combining the two to express the difference in characters between the pronunciation character cluster and the corresponding standard character cluster.
[0071] In other embodiments of the present invention, the implementer may also normalize the reciprocals of the first numerator and the first denominator respectively, and fuse the weighted sum of the normalized results to obtain the classification accuracy.
[0072] Step S302: Select any problem word of the patient as a target word; obtain the pronunciation abnormality of the target word based on the difference between the feature vector of the target word and the feature vector of each word in the corresponding standard word cluster.
[0073] First, any problem word of the patient is selected as the target word so that it can be analyzed one by one; considering that the greater the difference between the feature vector of the target word and the feature vector of each word in the standard word cluster, the more the target word deviates from the standard word cluster, and at the same time, the corresponding standard word cluster is the most matching cluster, the greater the degree of deviation, the more it indicates that the pronunciation of the target word is less standard and the greater the pronunciation abnormality.
[0074] In one embodiment of the present invention, the difference between feature vectors is quantified by Euclidean distance, and the average value of the Euclidean distance between the feature vector of the target word and the feature vector of each character in the corresponding standard character cluster is used as the pronunciation abnormality of the target word.
[0075] Step S303: Based on the classification accuracy of the pronunciation character cluster where the target character is located, the proportion of problem characters and the standard scores of all characters, combined with the pronunciation abnormality of the target character, the pronunciation disorder degree of the patient's target character is obtained; the classification accuracy is negatively correlated with the pronunciation disorder degree.
[0076] Considering that the lower the standard score of the target word itself is, the less standard the pronunciation is; the greater the pronunciation abnormality is, the less standard the pronunciation of the target word is and the greater the pronunciation disorder is;
[0077] Considering that when the proportion of problematic characters in the pronunciation character cluster where the target character is located is higher, it indicates that the patient has obvious pronunciation difficulties for this type of characters. For example, if the patient has difficulty accurately pronouncing the voiceless alveolar sibilant, there will be a large number of problematic characters in the corresponding pronunciation character cluster. At this time, the degree of pronunciation disorder is higher; at the same time, the lower the classification accuracy, the greater the deviation of the patient's pronunciation, and the greater the pronunciation disorder of the patient.
[0078] As an example, take the number of problematic characters in the pronunciation character cluster where the target character is located as the numerator, the number of all characters as the denominator, and the ratio as the proportion of problematic characters; after linearly normalizing the product of the reciprocal of the classification accuracy, the proportion of problematic characters, and the degree of pronunciation abnormality, the normalized result is used as the pronunciation disorder degree of the target character.
[0079] In another embodiment of the present invention, a weighted summation method can also be used to fuse the reciprocal of the classification accuracy, the proportion of problematic characters, and the degree of pronunciation abnormality to obtain the pronunciation disorder degree of the target character.
[0080] It should be noted that when the problematic character does not exist in the standard character cluster corresponding to it, it indicates that the pronunciation accuracy of this problematic character is relatively low. It may be a character in other clustering clusters that is classified into this clustering cluster due to a large pronunciation error. Therefore, this problematic character is recorded as a high-difficulty character. Since the pronunciation disorder degree of the problematic character is normalized to the range of 0-1 above, the pronunciation disorder degree of the high-difficulty character is directly recorded as 1 here. When the problematic character exists in multiple pronunciation character clusters, since each pronunciation character cluster corresponds to a pronunciation disorder degree, the maximum value among them is finally taken as the pronunciation disorder degree of the problematic character.
[0081] It should be noted that considering that some characters are polyphonic characters, such as "ci": "serve" and "wait for an opportunity"; "can": "participate", "irregular", and "ginseng"; due to the large differences in the different pronunciations of polyphonic characters, the implementer can regard each pronunciation of the polyphonic character as a separate character.
[0082] For a given training text, in order to fully reflect the language disorders of the patient in language expression scenarios such as daily communication and reading aloud, it is usually not the content of each character alone, but some simple phrases or sentences. The characters in different sentences are different, and the structure and context of the sentences will also affect the pronunciation of each character. Even for the same character, in different sentences, due to factors such as the connection of the preceding and following characters and the distribution of sentence stresses, the pronunciation accuracy may be different. Therefore, the distribution characteristics of problematic characters within a sentence will also affect the patient's pronunciation;
[0083] At the same time, considering that the frequency of occurrence of problem words also reflects the difficulty of pronouncing the problem words, the pronunciation difficulty of each problem word is obtained according to the distribution of problem words in the sentences to which they belong, combined with the pronunciation disorder and frequency of occurrence of problem words, and the pronunciation difficulty of each problem word is obtained. The patient's pronunciation difficulty of different problem words is carefully quantified, and the patient's personalized pronunciation disorder is accurately identified, providing a basis for relevant personnel to formulate personalized and targeted training plans, so as to improve the training effect.
[0084] Preferably, in one embodiment of the present invention, see Figure 3 , which shows a flow chart of a method for obtaining pronunciation difficulty provided by one embodiment of the present invention, specifically comprising:
[0085] Step S311: Obtain the continuous pronunciation difficulty of each sentence based on the distribution of problem words in each sentence.
[0086] Considering that the more problem words there are in a sentence and the larger the maximum consecutive number of problem words is, the more difficult it is for the patient to pronounce the sentence continuously and standardly, the continuous pronunciation difficulty of each sentence is obtained based on the number of problem words and the maximum consecutive number of problem words in each sentence.
[0087] As an example, the product of the number of problem words in each sentence and the maximum consecutive number of problem words is used as the continuous pronunciation difficulty of the corresponding sentence.
[0088] It should be noted that a single sentence is considered a single sentence. In one embodiment of the present invention, periods, question marks, exclamation marks, and semicolons are used as the basis for segmenting consecutive sentences, and the text between these symbols is considered a sentence. In other embodiments of the present invention, colons, commas, dashes, and ellipsis can also be used as the basis for segmentation because they can cause pauses in pronunciation.
[0089] Step S312: The frequency of occurrence of each problem word in the predetermined text and the degree of pronunciation disorder are integrated to obtain the difficulty weight of each problem word.
[0090] Since problem words may appear in different sentences and the distribution of problem words in different sentences is different, we first analyze the difficulty weight of each problem word to quantify the pronunciation difficulty of the sentence from the perspective of the problem word, which facilitates the subsequent analysis of the pronunciation difficulty of the problem word across all sentences in which the problem word appears.
[0091] Considering that the same problem word may appear in different sentences, the higher the frequency of its appearance in the predetermined text, the more difficult it is for the patient to pronounce the problem word; the greater the pronunciation difficulty, the more obstacles the patient has in pronunciation and the higher the pronunciation difficulty;
[0092] Based on this, as an example, the product of the frequency of occurrence of the problem character in the text and the pronunciation disorder degree is used as its own difficulty weight.
[0093] As another example, considering that the higher the probability of the character corresponding to the problem character being the problem character, the higher the difficulty of the problem character and the more difficult it is for the patient to pronounce accurately; it is also possible to use the proportion of the problem character in the characters corresponding to the problem character as the basis for calculating the difficulty weight, and use the product of the proportion of the problem character, the frequency of occurrence in the text, and the pronunciation disorder degree as its own difficulty weight.
[0094] It is also possible to fuse the frequency and the pronunciation disorder degree by means of addition or weighted summation.
[0095] It should be noted that not all the characters corresponding to the problem characters are necessarily pronounced inaccurately. For example, the character "程" corresponding to the problem character appears 10 times in total, and 5 of them are problem characters, so the proportion is 1 / 2; if the text has 400 characters in total, then the frequency is 5 / 400.
[0096] Step S313: For each sentence where any problem character is located, obtain the pronunciation sub-difficulty corresponding to each sentence according to the continuous pronunciation difficulty and the difficulty weight of all the problem characters in the sentence; use the average value of the pronunciation sub-difficulties of all the sentences where the problem character is located as the pronunciation difficulty corresponding to the problem character.
[0097] Since the problem character may appear in multiple sentences, each sentence where the problem character is located is analyzed one by one;
[0098] Considering that in a certain sentence where the target character is located, the greater the continuous pronunciation difficulty of the sentence, the greater the deviation between the sentence and the standard pronunciation in actual pronunciation, indicating that the pronunciation difficulty of the target character is greater from the overall perspective of the sentence; the greater the difficulty weight of all the problem characters in the sentence, the less standard the pronunciation of the sentence where the target character is located from the perspective of the characters, so the pronunciation sub-difficulty is obtained in this way;
[0099] When the pronunciation sub-difficulties of all the sentences where the target character is located are greater, it indicates that the pronunciation difficulty presented by the sentence is more likely to be caused by the target character, and the pronunciation difficulty of the target character is greater. Therefore, use the average value of the pronunciation sub-difficulties of all the sentences where the problem character is located as the pronunciation difficulty corresponding to the problem character.
[0100] As an example, use the product of the sum of the difficulty weights of all the problem characters in the sentence and the continuous pronunciation difficulty of the sentence itself as the pronunciation sub-difficulty of the sentence, and use the average value of the pronunciation sub-difficulties of all the sentences where the problem character is located as the pronunciation difficulty corresponding to the problem character.
[0101] At this point, the pronunciation difficulty of problem words is comprehensively assessed from multiple angles, including the scores of problem words themselves, the differences in classification clustering between actual and standard, the characteristic vector deviation of problem words, the distribution of problem words in sentences, and the frequency in texts. This helps accurately identify the patient's personalized pronunciation disorder and provides support for relevant personnel to develop personalized training plans and improve intervention effectiveness.
[0102] In another embodiment of the present invention, after obtaining the pronunciation difficulty of each problem word, the method further includes:
[0103] Taking into account that the greater the pronunciation obstacle and pronunciation difficulty of the problem word, the more emphasis is needed on training and the more training times are needed, so according to the pronunciation obstacle and pronunciation difficulty of the problem word, the preset training times of the problem word are adjusted to obtain the corrected training times.
[0104] As an example, the preset number of training times is 5, and the calculation formula for correcting the number of training times includes:
[0105]
[0106] Among them C i Indicates the number of correction trainings for the i-th problem word; norm() represents the linear normalization function; Z i N represents the pronunciation barrier degree of the i-th problem word; i Indicates the pronunciation difficulty of the i-th problem word; Indicates rounding up.
[0107] The calculation formula for the number of corrected training times integrates the pronunciation disorder degree and pronunciation difficulty by multiplication, and then limits the correction range by normalization. The greater the pronunciation disorder and the higher the pronunciation difficulty of the problem word, the higher the number of corrected training times will be, providing a reference for relevant personnel to formulate training plans.
[0108] For words with greater pronunciation difficulties, adjust the number of training sessions appropriately and practice multiple times to allow the pronunciation organs to gradually become familiar with the correct pronunciation method and position, thereby pronouncing the target sound more accurately. At the same time, repeated training of the pronunciation of specific words can enhance muscle memory and reduce incorrect pronunciations.
[0109] For words that do not have pronunciation problems, the number of training sessions can be reduced to avoid excessive fatigue. Repeated practice of content that has already been mastered may make patients feel boring or repulsive, reducing their enthusiasm and concentration for training, and reducing their attention to pronunciation that does not have problems. More energy can be focused on words that have pronunciation problems to improve overall training efficiency.
[0110] In summary, in order to solve the technical problem that conventional training methods are difficult to accurately identify the patient's personalized pronunciation disorder and affect the rehabilitation process, the present invention proposes a speech recognition and training method for speech disorder rehabilitation patients. The present invention first obtains voice data, extracts the text therein, and obtains the problem words and standard scores of the words pronounced by the patient; further, based on the number of phonemes, voice duration and pronunciation intensity, a feature vector of each word is constructed, and the text with standard pronunciation is classified to obtain standard word clusters; further, according to the difference in feature vectors between each word of the patient and the text in the standard word cluster, a pronunciation word cluster is obtained; further, according to the text difference and feature vector difference between the pronunciation word cluster and the corresponding standard word cluster, combined with the standard score of the patient's text and the proportion of problem words, the pronunciation disorder degree of each problem word of the patient is obtained; finally, according to the distribution of the problem words in the sentence to which they belong, combined with the pronunciation disorder degree and frequency of the problem words, the pronunciation difficulty of each problem word is obtained, and the patient's personalized pronunciation disorder is accurately identified, providing support for relevant personnel to formulate personalized training plans and improve intervention effects.
[0111] It should be noted that the order in which the embodiments of the present invention are described above is for illustrative purposes only and does not necessarily represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0112] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.
Claims
1. A speech recognition and training method for speech disorder rehabilitation patients, characterized in that: The method comprises: Acquire speech data of neurosurgery patients responding to predetermined texts and standard speech data, and split them into separate words; use a pre-trained neural network to filter out problematic words pronounced by the patient and obtain a standard pronunciation score for each word; Constructing a feature vector for each word based on the number of phonemes, speech duration, and pronunciation intensity of each word in the speech data; classifying the words to obtain standard word clusters based on the differences between the feature vectors of different words in standard speech; and classifying the patient's words to obtain pronunciation word clusters based on the differences between the feature vectors of each word of the patient and the feature vectors of words in the standard word clusters; Based on the text differences and feature vector differences between the pronunciation character cluster and the corresponding standard character cluster, combined with the standard score of the patient's text and the proportion of the problem characters, the pronunciation disorder degree of each problem character of the patient is obtained; based on the distribution of the problem characters in the sentences to which they belong, combined with the pronunciation disorder degree and frequency of occurrence of the problem characters, the pronunciation difficulty of each problem character is obtained.
2. The method for speech recognition and training of speech disorder rehabilitation patients according to claim 1, characterized in that: The method for obtaining the pronunciation disorder degree includes: Obtaining the classification accuracy of the patient's pronunciation character cluster based on the difference in the number and type of characters between the pronunciation character cluster and the corresponding standard character cluster; Select any of the patient's problem words as a target word; obtain the pronunciation abnormality of the target word based on the difference between the feature vector of the target word and the feature vector of each word in the corresponding standard word cluster; The pronunciation disorder degree of the target word of the patient is obtained based on the classification accuracy of the pronunciation word cluster to which the target word belongs, the proportion of the problem words and the standard scores of all words, combined with the pronunciation abnormality degree of the target word; the classification accuracy is negatively correlated with the pronunciation disorder degree.
3. The method for speech recognition and training of speech disorder rehabilitation patients according to claim 2, characterized in that: The method for obtaining the classification accuracy includes: The classification accuracy of the patient's characters is obtained based on the absolute value of the difference between the number of characters in the pronounced character cluster and the number of characters in the corresponding standard character cluster, combined with the intersection-over-union ratio of the pronounced character cluster and the corresponding standard character cluster.
4. The speech recognition and training method for speech disorder rehabilitation patients according to claim 1, characterized in that: The method for obtaining the pronunciation difficulty includes: Obtain the continuous pronunciation difficulty of each sentence based on the distribution of the problem words in each sentence; Combining the frequency of occurrence of each problem word in the predetermined text and the pronunciation disorder degree to obtain a difficulty weight of each problem word; For each sentence containing any of the problem words, the pronunciation sub-difficulty corresponding to each sentence is obtained based on the continuous pronunciation difficulty and the difficulty weights of all the problem words in the sentence; and the average of the pronunciation sub-difficulties of all the sentences containing the problem word is used as the pronunciation difficulty corresponding to the problem word.
5. The method for speech recognition and training of speech disorder rehabilitation patients according to claim 4, characterized in that: The method for obtaining the difficulty of continuous pronunciation comprises: The continuous pronunciation difficulty of each sentence is obtained according to the number of the problem words in each sentence and the maximum consecutive number of the problem words.
6. The method for speech recognition and training of speech disorder rehabilitation patients according to claim 1, characterized in that: The method for obtaining the standard character cluster includes: The characters are clustered according to the Euclidean distance between the feature vectors of different characters to obtain standard character clusters.
7. The method for speech recognition and training for speech disorder rehabilitation patients according to claim 1, characterized in that: The method for obtaining the pronunciation character clusters includes: Obtaining the mean of the feature vectors of all characters in each standard character cluster as a representative feature vector; The standard character cluster whose representative feature vector is most similar to the feature vector of the patient's text is used as the standard character cluster matched by the patient's corresponding text; all the patient's texts matched by the same standard character cluster are divided into one category to obtain a pronunciation character cluster.
8. The method for speech recognition and training for speech disorder rehabilitation patients according to claim 1, characterized in that: The mean of the amplitude values in the speech signal of the text is used as the pronunciation intensity of the corresponding text.
9. The method for speech recognition and training for speech disorder rehabilitation patients according to claim 1, characterized in that: The text splitting method includes: using Wav2Vec to convert the voice signal in the voice data into a text sequence, and splitting the text sequence into separate characters.
10. The method for speech recognition and training of speech disorder rehabilitation patients according to claim 1, characterized in that: After obtaining the pronunciation difficulty of each of the problem words, the method further includes: According to the pronunciation disorder degree and the pronunciation difficulty of the problem word, the preset training times of the problem word are adjusted to obtain the corrected training times.
Citation Information
Patent Citations
Language barrier analysis method and system based on voice information, medium and equipment
CN119380760A
Voice interaction-based dysarthria children education and rehabilitation training system
CN119741949A
Multi-logic interaction system for speech disorder rehabilitation training
CN120089285A
Pattern recognition dictionary production device and pattern recognizer
JP1998011543A
System and program for supporting foreign language learning
JP2010169973A