Voice rehabilitation training method and system and computer equipment
By constructing a voice rehabilitation training case library and performing fine-grained acoustic analysis, drawing the voice portraits of the subjects and selecting matching cases for training, the problem of difficulty in personalized implementation in voice rehabilitation training for autistic children is solved, and efficient personalized rehabilitation training results are achieved.
Patent Information
- Application Number
- CN202510694353.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-28
AI Technical Summary
The prior art has problems such as mechanized content, difficulty in personalized implementation and low family participation in speech rehabilitation training for autistic children, and it is difficult to meet the individual rehabilitation training needs of each subject.
By constructing a voice rehabilitation training case library, obtain phoneme annotations and fine-grained acoustic features of voice rehabilitation training data, perform acoustic analysis of different granularities, draw the voice portrait of the subjects, and select matching cases from the case library for the subjects to train, and iterate the correction training process.
The selection of personalized voice rehabilitation training programs was realized, which improved the rehabilitation training effect of the subjects, adapted to changes in individual rehabilitation trajectory, and enhanced family participation.
Smart Images

Figure CN120220965A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent rehabilitation training, and particularly to a speech rehabilitation training method, system and computer device. Background Art
[0002] In recent years, the incidence of autism has been continuously increasing, mostly occurring in childhood and accompanying throughout life. Language disorder is one of the core disorders of autism. Effectively carrying out language rehabilitation training and evaluation for autistic children and improving their language ability are the keys to promoting the development of autistic children and inclusive education for autism. Generally speaking, language disorders are divided into speech disorders, vocabulary and grammar disorders, communication disorders, etc.
[0003] For autistic patients (hereinafter referred to as subjects for short), due to reasons such as language cognitive development or social communication disorders, most of them lack sufficient guidance and training during the critical period of speech learning, resulting in lag or disorder in speech development. Only some subjects may have structural or functional abnormalities in organs such as hearing, oral muscles, larynx or tongue.
[0004] At present, there are still problems in autism speech rehabilitation training, such as mechanical content, difficult personalized implementation and low family participation. First of all, the individual differences of subjects are large, especially the language development varies from person to person, resulting in very difficult intervention. The traditional manual intervention method requires intervention personnel to analyze the specific situation of individuals and make a detailed assessment of their language ability status before formulating an intervention plan. This intervention method is labor-consuming, and because the training content has a high repetition frequency and is mechanical, it may make the subjects easily anxious, so the participation degree is not high.
[0005] Existing computer-aided intervention methods mostly use technologies such as learning game software, augmented reality technology, robots, etc., trying to provide a relaxed and simple test and learning environment for subjects to reduce their anxiety. However, this method is difficult to implement personalized training. It only trains by selecting pre-set fixed modes and is difficult to accurately depict the characteristics of different types of speech disorders, and cannot better meet the individual rehabilitation training needs of each subject. Summary of the Invention
[0006] The purpose of the present invention is to provide a speech rehabilitation training method, system and computer device for the deficiencies of the above-mentioned existing technologies to solve the problems in the existing technologies.
[0007] The present invention specifically provides the following technical solutions: A speech rehabilitation training method, comprising: Constructing a speech rehabilitation training case library, and obtaining phoneme annotations and fine-grained acoustic features of speech rehabilitation training data in the speech rehabilitation training case library; wherein the fine-grained is a local analysis dimension; Obtain the speech feedback data of the subject, and perform acoustic analyses at different granularities on the speech feedback data of the subject through the phoneme annotation and fine-grained acoustic features of the speech rehabilitation training data to obtain acoustic analysis results at different granularities; Statistically analyze the acoustic analysis results at different granularities, draw the speech portrait of the subject through the statistical results, and in combination with the speech portrait of the subject, select matching speech rehabilitation training cases from the speech rehabilitation training case library for the subject to train, and perform an iterative process of acoustic analysis on the speech feedback data of the subject obtained from each training and selecting matching speech rehabilitation training cases to correct the training of the subject.
[0008] Preferably, the construction of the speech rehabilitation training case library includes: Collect unannotated speech rehabilitation training data; Perform phoneme annotation on the unannotated speech rehabilitation training data, compare different annotation results, integrate to obtain speech rehabilitation training data with phonemes, and perform phoneme annotation on each data in the speech rehabilitation training data with phonemes; Perform fine-grained acoustic feature annotation on the speech rehabilitation training data with phonemes to obtain a fine-grained speech case library, and construct a speech rehabilitation training case library based on the fine-grained speech case library.
[0009] Preferably, the performing acoustic analyses at different granularities on the speech feedback data of the subject through the phoneme annotation and fine-grained acoustic features of the speech rehabilitation training data to obtain acoustic analysis results at different granularities includes: Obtain the speech phoneme annotation corresponding to the speech feedback data of the subject; each speech phoneme annotation includes syllable types and several syllables, and each syllable includes initial consonant annotation, final consonant annotation, and tone annotation; Perform acoustic pronunciation feature annotation on each syllable in the speech phoneme annotation respectively; After the acoustic pronunciation feature annotation, compare and evaluate the speech phoneme annotation and the phoneme annotation of the corresponding speech rehabilitation training data, and record the coarse-grained voiceprint analysis results; After the acoustic pronunciation feature annotation, compare and evaluate the fine-grained acoustic features of the speech feedback data of the subject and the fine-grained acoustic features of the corresponding speech rehabilitation training data to obtain the fine-grained acoustic feature differences.
[0010] Preferably, the performing acoustic pronunciation feature annotation on each syllable in the speech phoneme annotation specifically is: Obtain the initial consonant annotation of the syllable, and automatically annotate the initial consonant pronunciation features according to the determined initial consonant pronunciation types; Obtain the final consonant annotation of the syllable, and automatically annotate the final consonant pronunciation features according to the constructed final consonant pronunciation types; Obtain the tone annotation of the syllable, and automatically annotate the tone pronunciation features according to the constructed tone pronunciation types. Construct the fine-grained acoustic features of the syllable based on the annotated initial consonant pronunciation features, final pronunciation features, and tone pronunciation features.
[0011] Preferably, compare and evaluate the phoneme annotation of the speech and the phoneme annotation of the corresponding speech rehabilitation training data, and record the coarse-grained voiceprint analysis results, including: Construct a hash table with the name of the speech disorder type as the keyword, where the value of the hash table is the number of speech disorder errors. Extract the syllable types of the phoneme annotation in the speech rehabilitation training data, and the initial consonant annotation, final annotation, and tone annotation of each syllable. Extract the syllable types of the phoneme annotation of the speech, and the initial consonant annotation, final annotation, and tone annotation of each syllable. Compare the syllable types, initial consonant annotations, final annotations, and tone annotations corresponding to the phoneme annotation in the speech rehabilitation training data and the phoneme annotation of the speech respectively. If they are different, add 1 to the value of the corresponding keyword in the hash table as the coarse-grained voiceprint analysis result.
[0012] Preferably, compare and evaluate the fine-grained acoustic features of the subject's speech feedback data and the fine-grained acoustic features of the corresponding speech rehabilitation training data to obtain the fine-grained acoustic feature differences, including: Construct a hash table according to the initial fine-grained acoustic feature differences corresponding to the subject's speech feedback data. The keyword of the hash table is the initial consonant error or tone error, and the value of the hash table is the number of acoustic feature misarticulation times. Extract each syllable of the fine-grained acoustic features in the speech rehabilitation training data, and extract each syllable of the subject's fine-grained acoustic features. Compare the syllables corresponding to the fine-grained acoustic features in the speech rehabilitation training data and the subject's fine-grained acoustic features respectively. When they are different, generate a fine-grained pronunciation feature keyword by combining the initial consonant, final, and tone annotations, and look up this keyword in the hash table. If the keyword does not exist, create a new keyword in the hash table and set its value to 1 as the updated fine-grained acoustic feature difference.
[0013] Preferably, perform statistics on the acoustic analysis results of different granularities, and draw the subject's speech portrait through the statistical results, including: Perform frequency statistics on the coarse-grained voiceprint analysis results to obtain the coarse-grained portrait of the subject's speech disorder. Perform frequency statistics on the updated fine-grained acoustic feature differences to obtain the fine-grained portrait of the subject's speech disorder. Construct the subject's speech portrait jointly through the coarse-grained portrait and the fine-grained portrait.
[0014] Preferably, in combination with the subject's voice portrait, a matching voice rehabilitation training case is selected from the voice rehabilitation training case library for the subject to train, including: Taking different voice disorders in the coarse-grained portrait as the training case grouping unit, and sequentially selecting voice rehabilitation training cases that match the fine-grained portrait features in the voice rehabilitation training case library; According to the voice disorder types of the selected voice rehabilitation training cases, the voice rehabilitation training cases are assigned to different groups, and the process of repeatedly selecting voice rehabilitation training cases that match the fine-grained portrait features is repeated until the number of recommended matching voice rehabilitation training cases in each group reaches a preset number; Use the recommended matching voice rehabilitation training cases in each group for the subject to train.
[0015] The present invention provides a voice rehabilitation training system, including: A construction module for constructing a voice rehabilitation training case library and obtaining the phoneme annotation and fine-grained acoustic features of the voice rehabilitation training data in the voice rehabilitation training case library; An analysis module for obtaining the subject's voice feedback data and performing acoustic analysis of different granularities on the subject's voice feedback data through the phoneme annotation and fine-grained acoustic features of the voice rehabilitation training data to obtain acoustic analysis results of different granularities; A recommendation module for statistically analyzing the acoustic analysis results of different granularities, drawing the subject's voice portrait through the statistical results, selecting matching voice rehabilitation training cases from the voice rehabilitation training case library in combination with the subject's voice portrait for the subject to train, and performing an iterative process of acoustic analysis and selecting matching voice rehabilitation training cases on the subject's voice feedback data obtained from each training to correct the subject's training.
[0016] The present invention provides a computer device, including a memory and a processor. When a program stored in the memory is executed by the processor, the processor executes the steps of the above-mentioned voice rehabilitation training method.
[0017] Compared with the prior art, the present invention has the following remarkable advantages: The present invention performs acoustic analysis based on the obtained voice feedback data of the subject. Through phoneme annotation and fine-grained acoustic features of the voice rehabilitation training data, acoustic analysis of different granularities is performed on the voice feedback data of the subject, and acoustic analysis results of different granularities are obtained, providing richer indicators for depicting the personalized voice portrait of the subject. And by comparing and selecting with the voice rehabilitation training case library annotated with fine-grained acoustic features, an iterative process of acoustic analysis of the voice feedback data of the subject obtained in each training and selecting matching voice rehabilitation training cases is carried out to correct the training of the subject. A more targeted personalized rehabilitation training plan can be selected from the voice rehabilitation training case library, realizing automatic correction and update of the subject portrait after each training, and adapting to the change of the individual rehabilitation trajectory by correcting the training process, thereby improving the rehabilitation training effect of the subject. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is the framework of the voice rehabilitation training construction process in the present invention; Figure 2 is the flowchart of a voice rehabilitation training method in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0020] As Figure 1 and Figure 2 shown, the present embodiment provides a voice rehabilitation training method, including the following steps: S1. Establish an autism speech disorder classification system SDCatelog (Speech Disorder Catelog), construct a voice rehabilitation training case library STCases (Speech Training Cases), and obtain phoneme annotation and fine-grained acoustic features of the voice rehabilitation training data in the voice rehabilitation training case library, where the fine-grained is the local analysis dimension; the specific steps include: S11. Collect unlabeled voice rehabilitation training data (cases) USTData (Unlabeled Speech Training Data) from existing publicly available autism language assessment scales and relevant research results.
[0021] S12. Analyze Chinese speech features and construct a fine-grained speech disorder classification system. The specific steps include: S121. Analyze and determine the top-level classification of speech, mainly including initials, finals, tones, and syllables.
[0022] S122. Analyze and determine the initial sound production types InitialTypes, mainly including bilabial consonants, palatal consonants, velar consonants, and alveolar consonants before and after the tongue tip, etc. For example, bilabial consonants include "b", "p", and "m", etc., and velar consonants include "g", "k", and "h".
[0023] S123. Analyze and determine the final sound production types , mainly including closed vowels begun with i, closed vowels begun with u, closed vowels begun with ü, and open vowels, etc. For example, closed vowels begun with i include "i", "ia", and "ing", etc., and closed vowels begun with u include "u", "ua", and "uo".
[0024] S124. Analyze and construct the tone production types , mainly including high level, rising tone, falling-rising tone, falling tone, and neutral tone.
[0025] S125. Combine different sound production characteristics to establish a speech disorder classification system SDCatelog, mainly including initial deletion, initial conversion, initial addition, final deletion, final conversion, and denasalization, etc.
[0026] The present invention combines the acoustic sound production characteristics of Chinese and labels each speech rehabilitation training case with fine-grained labels. The labeled results are essentially labels, providing richer indicators for depicting the personalized speech portraits of autistic subjects.
[0027] S13. In the fine-grained speech disorder classification system, perform phoneme labeling on the unlabeled speech rehabilitation training data (speech cases) USTData, including syllable types, , syllables, initials, finals, and tones, etc., that is, perform personalized labeling of speech rehabilitation training corpus, compare different labeling results, and integrate to obtain speech rehabilitation training data with phonemes (Phoneme Speech Training Data), and perform phoneme labeling on each speech rehabilitation training data in the speech rehabilitation training data with phonemes The specific steps include: S131. Design phoneme labeling specifications, including tone symbols and labeling position rules, etc. For example: syllable types ("1" represents monosyllable, "2" represents disyllable, "0" represents polysyllable), the tone of the compound final "uo" is labeled on the last vowel "o", use "#" to mark the rising tone, and "^" to mark the falling-rising tone, etc.; S132. Recruit multiple annotators and use cross-annotation to perform phoneme annotation. The specific steps include: S1321. Evenly distribute the speech rehabilitation training data according to the number of annotators, ensuring that at least 30% of each case set after distribution overlaps with others.
[0028] S1322. Have the annotators perform annotation separately according to the phoneme annotation specifications.
[0029] S1323. Have the annotators perform cross-annotation.
[0030] S133. Compare different annotation results and integrate them to obtain the speech rehabilitation training data with phonemes , where the inconsistent annotation results are determined through centralized discussion. The phoneme annotation of each speech rehabilitation training data is (each phoneme annotation includes the syllable type , and corresponding several syllable annotations . Each syllable annotation includes an initial consonant annotation , a final consonant annotation and a tone annotation each).
[0031] For example, given the speech rehabilitation training data (training word "apple"), its phoneme annotation is "2ping#guo^". The syllable type "2" represents a disyllable and contains two syllable annotations "ping#" and "guo^". Among them, the first syllable annotation is "ping~" (the initial consonant annotation is "p", the final consonant annotation is "ing", and the tone annotation is "#"), and the second syllable annotation is "guo^" (the initial consonant annotation is "g", the final consonant annotation is "uo", and the tone annotation is "^").
[0032] S14. Automatically annotate the fine-grained acoustic features of the speech rehabilitation training data with phonemes to obtain a fine-grained speech case library , and its specific steps include: S141. Sequentially scan each piece of speech rehabilitation training data in the speech rehabilitation training data with phonemes to obtain the phoneme annotation 。
[0033] S142. Sequentially scan each syllable annotation in , and perform automatic annotation of acoustic pronunciation features on . The specific steps are as follows: S1421. Obtain the initial consonant annotation of the syllable , and automatically annotate the initial consonant pronunciation feature according to the determined initial consonant pronunciation type . 。
[0034] S1422. Obtain the final consonant annotation of the syllable , and automatically annotate the final consonant pronunciation feature according to the determined final consonant pronunciation type . 。
[0035] S1423. Obtain the tone annotation of the syllable , and automatically annotate the tone pronunciation feature according to the determined tone pronunciation type . 。
[0036] S1424. The fine-grained acoustic features of the syllable are composed of the above automatic annotation results .
[0037] For example, given speech rehabilitation training data (training word "apple"), its phoneme annotation is "2ping#guo^", and its fine-grained acoustic feature annotation is "2p1ing1#2g2uo2^3". Specifically: The annotation of the first syllable is "ping#" (the initial consonant annotation is "p", so the initial consonant pronunciation feature is "bilabial consonant", denoted as 1, the final consonant annotation is "ing", so the final consonant pronunciation feature is "front vowel", denoted as 1, the tone annotation is "#", so the tone pronunciation feature is "rising tone", denoted as 2. Thus, the fine-grained acoustic feature corresponding to this syllable is annotated as "p1ing1#2".
[0038] The annotation of the second syllable is "guo^" (the initial consonant annotation is "g", so the initial consonant pronunciation feature is "palatal consonant", denoted as 2, the final consonant annotation is "uo", so the final consonant pronunciation feature It is "rounded aperture", denoted as 2, tone marking is "^", so the tone pronunciation feature is "rising tone", denoted as 3. Thus, the fine-grained acoustic feature corresponding to this syllable is marked as "g2uo2^3".
[0039] S143. Repeat the above operations until all the speech rehabilitation training data in the speech rehabilitation training data are automatically marked completely.
[0040] S15. Speech rehabilitation training data with phonemes After fine-grained acoustic feature marking, use the fine-grained speech case library as the basis or jointly construct a speech rehabilitation training case library suitable for personalized training .
[0041] S2. Obtain the speech feedback data of the subject, and conduct acoustic analysis of different granularities on the speech feedback data of the subject through the phoneme marking and fine-grained acoustic features of the speech rehabilitation training data to obtain acoustic analysis results of different granularities.
[0042] Obtain the speech feedback data of the subject and the corresponding speech phoneme markings, conduct acoustic analysis on each speech feedback data of the subject, compare and evaluate the speech phoneme markings corresponding to the speech feedback data of the subject and the phoneme markings of the corresponding speech rehabilitation training data, where the corresponding speech rehabilitation training data is the training corpus, to obtain the coarse-grained voiceprint analysis results and the fine-grained acoustic feature differences; among them, the coarse-grained is the global analysis dimension, obtained when classifying the speech feedback data of the subject, and the fine-grained is the local analysis dimension, obtained when identifying the speech feedback data of the subject.
[0043] Conduct acoustic analysis on the speech feedback data of the subject, compare and evaluate the speech phoneme markings corresponding to the speech feedback data of the subject and the phoneme markings of the corresponding speech rehabilitation training data, to obtain the coarse-grained voiceprint analysis results and the fine-grained acoustic feature differences. The specific steps include: S21. Collect or obtain the speech feedback data of the subject from the subject's rehabilitation training behavior library .
[0044] S22. For each speech feedback data of the subject conduct acoustic analysis. The specific steps include: S221. Use a phoneme-based acoustic model (such as a model based on GMM-HMM, DNN-HMM, or Transformer) to automatically obtain the speech phoneme markings corresponding to the speech feedback data of the subject . Each speech phoneme marking includes the syllable type and several syllables , each syllable includes an initial consonant annotation , a final sound annotation and a tone annotation each. For example, given the training word "apple", the subject's voice feedback data is , and after model learning, the voice phoneme annotation is "2pi#puo^". In this embodiment, wav2vec2 is selected as the phoneme recognition model to complete the prediction of voice phoneme annotation.
[0045] S222. Automatically annotate the acoustic pronunciation features of each syllable in the voice phoneme annotation respectively. The specific steps include: S2221. Obtain the initial consonant annotation of the syllable , and automatically annotate the initial consonant pronunciation feature according to the determined initial consonant pronunciation type .
[0046] S2222. Obtain the final sound annotation of the syllable , and automatically annotate the final sound pronunciation feature according to the constructed final sound pronunciation type .
[0047] S2223. Obtain the tone annotation of the syllable , and automatically annotate the tone pronunciation feature according to the constructed tone pronunciation type .
[0048] S2224. Construct the fine-grained acoustic features of the syllable from the above automatically annotated initial consonant pronunciation features, final sound pronunciation features and tone pronunciation features.
[0049] For example, given the training word "apple", the subject's voice feedback data is , its phoneme annotation is "2pi#puo^", and its fine-grained acoustic feature annotation is "2p1i1#2 p1uo2^3". Respectively: The first syllable is "pi#" (the initial consonant annotation is "p", so the initial consonant pronunciation feature is "bilabial consonant", denoted as 1, the final sound annotation is "i", so the final sound pronunciation feature is "front tooth sound", denoted as 1, the tone annotation is "#", so the tone pronunciation feature is "rising tone", denoted as 2. Thus, the fine-grained acoustic feature corresponding to this syllable is labeled as "p1i1#2".
[0050] The second syllable is labeled as "puo^" (the initial consonant is labeled as "p", so the initial consonant pronunciation feature is "bilabial consonant", denoted as 1, and the final is labeled as "uo", so the final pronunciation feature is "rounded-mouth vowel", denoted as 2, and the tone is labeled as "^", so the tone pronunciation feature is "falling-rising tone", denoted as 3. Thus, the fine-grained acoustic feature corresponding to this syllable is labeled as "p1uo2^3".
[0051] Based on the learned speech factor annotations, perform voiceprint analysis on different granularities of the subjects' feedback.
[0052] S223. After acoustic pronunciation feature annotation, compare and evaluate the voice feedback data of the subjects with the corresponding phoneme annotations and the phoneme annotations of the corresponding voice rehabilitation training data , and record the results of the coarse-grained voiceprint analysis , for example, "initial consonant error", "initial consonant omission", and "tone error", etc. The specific steps include; S2231. Based on different types in the speech disorder classification system , construct a hash table with the keyword being the name of the speech disorder type (such as "initial consonant error" and "initial consonant omission", etc.), and the value being the number of speech disorder errors (initial value is 0).
[0053] S2232. Extract the syllable types and the initial consonant annotations of each syllable in the voice rehabilitation training data , final annotations and tone annotations .
[0054] S2233. Extract the syllable types and the initial consonant annotations of each syllable in the voice phoneme annotation , final annotations and tone annotations .
[0055] S2234, respectively compare the syllable types and initial consonant annotations, final vowel annotations and final vowel annotations corresponding to the phoneme annotations in the speech rehabilitation training data and the speech phoneme annotations, and the tone annotations and tone annotations. If they are different, the hash table The value of the corresponding keyword in is added by 1, which is used as the coarse-grained voiceprint analysis result.
[0056] For example, given speech rehabilitation training data (training word "apple"), its phonemic annotation is "2ping#guo^", the subject's voice feedback data Phonetic notation is "2pi#puo^", then the coarse-grained voiceprint analysis result The corresponding number of "initial consonant errors" is 1, and the number of "final vowel errors" is 1.
[0057] S224. After the acoustic features are annotated, compare and evaluate the subject's voice feedback data Fine-grained acoustic features and the fine-grained acoustic features of the corresponding speech rehabilitation training data , and obtain fine-grained acoustic feature differences For example, "initial consonant error" needs to be further refined into "initial consonant error: bilabial consonant b changes to bilabial consonant d", and "tone error" needs to be further refined into "tone error: rising tone changes to falling tone" or "tone error: rising tone changes to yangping tone", etc. The specific steps include: S2241. Based on the voice feedback data of the subjects Corresponding to the initial fine-grained acoustic feature difference Construct a hash table, the key of which is "initial consonant error: bilabial consonant b changes to bilabial consonant d" or "tone error: rising tone changes to falling tone", etc., and the value of the hash table is the number of acoustic feature errors (the initial value is 0).
[0058] S2242. Extracting fine-grained acoustic features from speech rehabilitation training data Each syllable in (Including initial consonant markings , vowel marking Harmony and Tone Marking , initial consonant pronunciation characteristics , the pronunciation characteristics of the finals Vocal characteristics of harmony wait).
[0059] S2243. Extract fine-grained acoustic features of subjects Each syllable of (Including initial consonant markings , vowel marking , Tone Marking , Initial sound pronunciation features , Final sound pronunciation features and Tone pronunciation features etc.).
[0060] S2244. For the syllables corresponding to the fine-grained acoustic features and the subject's fine-grained acoustic features in the speech rehabilitation training data (initial sound pronunciation features, final sound pronunciation features, and tone pronunciation features), if they are different, generate fine-grained pronunciation feature keywords by combining the annotation of the initial sound, final sound, and tone, such as "Initial sound error: bilabial consonant b changes to bilabial consonant d", and look up this keyword in the hash table . If it exists, increment the value corresponding to this keyword by 1. If this keyword does not exist, create a new keyword in the hash table and set its value to 1 as the fine-grained acoustic feature difference .
[0061] For example, given the speech rehabilitation training data (training word "apple"), its fine-grained acoustic feature annotation is "2p1ing1#2g2uo2^3", and the fine-grained acoustic feature annotation of the subject's speech feedback data is "2p1i1#2p1uo2^3". In the fine-grained acoustic feature difference , the coarse-grained "initial sound error" is refined to "Initial sound error: velar consonant g changes to bilabial consonant p", and the "final sound error" is refined to "Final sound error: front vowel ing changes to front vowel i", and the number of speech errors is 1 for both .
[0062] S23. Repeat step S22 until the analysis of the subject's speech feedback data is completed, so as to obtain the coarse-grained voiceprint analysis result of the subject and the fine-grained acoustic feature difference .
[0063] S3. Statistically analyze the acoustic analysis results of different granularities, and draw the subject's speech portrait through the statistical results. The specific steps include: S31. Conduct frequency statistics on the coarse-grained voiceprint analysis result to obtain the coarse-grained portrait of the subject's speech disorder , and the specific steps include: S311. Sequentially scan the coarse-grained voiceprint analysis result for the coarse-grained voiceprint analysis result corresponding to each speech feedback data of the subject .
[0064] S312. Accumulate the voiceprint analysis results for each voice disorder type (such as "initial consonant error" and "initial consonant omission") by category, and count the number of errors.
[0065] S313. Sort according to the number of errors of the voice disorder.
[0066] S314. Select the top K (the preset K is 3) types of voice disorders with the highest number of occurrences as the rough-grained portrait of the subject .
[0067] It should be noted that in this embodiment, the sorting is based on the number of errors, and the selection of voice disorders adopts the top K method. However, it is not limited to this. The sorting can also be implemented according to the error probability, and the selection of voice disorder types can also be implemented according to a preset threshold.
[0068] S32. Conduct frequency statistics on the fine-grained acoustic feature differences to obtain the fine-grained portrait of the subject's voice disorder. The specific steps include: S321. Scan the fine-grained acoustic feature differences corresponding to each voice feedback data of the subject in sequence. .
[0069] S322. Accumulate the number of errors of different fine-grained acoustic features by category according to the fine-grained vocalization features (such as "initial consonant error: bilabial consonant b becomes bilabial consonant d" and "tone error: rising tone becomes falling tone", etc.).
[0070] S323. Sort according to the number of errors of the fine-grained acoustic features.
[0071] S324. Select the top K (the preset K is 3) types of acoustic features with the highest number of occurrences as the fine-grained portrait of the subject .
[0072] It should be noted that in this embodiment, the sorting is based on the number of errors, and the selection of fine-grained acoustic features adopts the top K method. However, it is not limited to this. The sorting can also be implemented according to the error probability, and the selection of fine-grained acoustic features can also be implemented according to a preset threshold.
[0073] S33. The rough-grained portrait and the fine-grained portrait obtained from the above step S3 together constitute the voice portrait of the subject.
[0074] S4. Combine the voice portrait of the subject and select a matching voice rehabilitation training case from the voice rehabilitation training case library for the subject to train, that is, perform personalized recommendation of voice rehabilitation training cases. The specific steps include: S41. Use the different speech disorders in the subject's coarse-grained portrait as the grouping unit for training cases, such as "initial consonant error" or "initial consonant omission", etc.
[0075] S42. Sequentially select speech rehabilitation training cases that match (conform to) the fine-grained portrait features in the personalized speech rehabilitation training case library. For example, if there is "initial consonant error: bilabial consonant b changes to bilabial consonant d", then select cases with the initial consonant marked as "bilabial consonant b" in the phoneme annotation in .
[0076] S43. Allocate the speech rehabilitation training cases to different groups according to the speech disorder types of the selected speech rehabilitation training cases.
[0077] S434. Repeat the above process (the process of selecting cases that match the fine-grained portrait features) until the number of recommended matching speech rehabilitation training cases in each group reaches the preset number.
[0078] S5. Implement personalized speech rehabilitation training, and perform acoustic analysis on the speech feedback data of the subject obtained in each training and an iterative process of selecting matching speech rehabilitation training cases to correct the subject's training. The specific steps include: S51. Use the speech disorders in the subject's coarse-grained portrait as the training tasks respectively.
[0079] S52. Display the pictures corresponding to the training cases (accompanied by pronunciation is also possible) to guide the subject to implement speech rehabilitation training.
[0080] S53. Collect the subject's speech data from the personalized speech rehabilitation training and store it in the subject's rehabilitation training behavior library for voiceprint analysis.
[0081] S54. Return to S52 to iteratively carry out personalized speech rehabilitation training.
[0082] Use the recommended matching speech rehabilitation training cases in each group for the subject's training.
[0083] It should be noted that if the voice rehabilitation training is implemented for the subject for the first time, the cold start method (referring to the language assessment scale, such as the Chinese Communication Development Scale) can be used to select cases in combination with the estimated situation of the subject, and the training ends when the subject makes frequent repeated errors or reaches the assessment target. This method is different from the previous models. It can not only reduce the professional requirements of training assessors and is easy to promote in the family scenario, but also accurately depict the voice disorder characteristics of the subject so as to carry out more targeted personalized rehabilitation training. It is an effective and easy-to-promote training method.
[0084] The present invention provides a voice rehabilitation training system, including: a construction module, an analysis module, and a recommendation module.
[0085] Among them, the construction module is used to construct a voice rehabilitation training case library and obtain the phoneme annotation and fine-grained acoustic features of the voice rehabilitation training data in the voice rehabilitation training case library; the analysis module is used to obtain the voice feedback data of the subject, and perform acoustic analysis of different granularities on the voice feedback data of the subject through the phoneme annotation and fine-grained acoustic features of the voice rehabilitation training data to obtain acoustic analysis results of different granularities; the recommendation module is used to perform statistics on the acoustic analysis results of different granularities, draw the voice portrait of the subject through the statistical results, select the matching voice rehabilitation training cases from the voice rehabilitation training case library for the subject to train, and perform an iterative process of acoustic analysis and selection of matching voice rehabilitation training cases on the voice feedback data of the subject obtained from each training to correct the training of the subject.
[0086] The present invention also provides a computer device, including a memory and a processor. When a program stored in the memory is executed by the processor, the processor executes the steps of a voice rehabilitation training method.
[0087] According to the disclosed embodiments, the computer device can communicate with one or more external devices (such as a keyboard, a pointing device, Bluetooth communication, etc.), or communicate with any device (such as a router, a demodulator, etc.) that enables the computing device to communicate with one or more other computing devices.
[0088] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. For those skilled in the art of the present invention, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should all be regarded as belonging to the protection scope of the present invention.
Claims
1. A speech rehabilitation training method, characterized in that, Including: Construct a speech rehabilitation training case library, and obtain the phoneme annotation and fine-grained acoustic features of the speech rehabilitation training data in the speech rehabilitation training case library; wherein the fine-grained is the local analysis dimension; Obtain the speech feedback data of the subject, and perform acoustic analysis of different granularities on the speech feedback data of the subject through the phoneme annotation and fine-grained acoustic features of the speech rehabilitation training data to obtain acoustic analysis results of different granularities; Statistically analyze the acoustic analysis results of different granularities, draw a speech portrait of the subject through the statistical results, combine the speech portrait of the subject, select a matching speech rehabilitation training case from the speech rehabilitation training case library for the subject to train, and perform an iterative process of acoustic analysis and selecting a matching speech rehabilitation training case on the speech feedback data obtained from each training to correct the subject's training.
2. The voice rehabilitation training method according to claim 1, wherein The construction of the speech rehabilitation training case library includes: Collect unannotated speech rehabilitation training data; Perform phoneme annotation on the unannotated speech rehabilitation training data, compare different annotation results, integrate to obtain speech rehabilitation training data with phonemes, and perform phoneme annotation on each data in the speech rehabilitation training data with phonemes; Perform fine-grained acoustic feature annotation on the speech rehabilitation training data with phonemes to obtain a fine-grained speech case library, and construct a speech rehabilitation training case library based on the fine-grained speech case library.
3. The voice rehabilitation training method according to claim 1, wherein The performing acoustic analysis of different granularities on the speech feedback data of the subject through the phoneme annotation and fine-grained acoustic features of the speech rehabilitation training data to obtain acoustic analysis results of different granularities includes: Obtain the speech phoneme annotation corresponding to the speech feedback data of the subject; wherein each speech phoneme annotation includes a syllable type and several syllables, and each syllable includes an initial consonant annotation, a final consonant annotation, and a tone annotation; Perform acoustic pronunciation feature annotation on each syllable in the speech phoneme annotation respectively; After the acoustic pronunciation feature annotation, compare and evaluate the speech phoneme annotation and the phoneme annotation of the corresponding speech rehabilitation training data, and record the coarse-grained voiceprint analysis result; After the acoustic pronunciation feature annotation, compare and evaluate the fine-grained acoustic features of the speech feedback data of the subject and the fine-grained acoustic features of the corresponding speech rehabilitation training data to obtain the fine-grained acoustic feature difference.
4. The voice rehabilitation training method according to claim 3, characterized in that The performing acoustic pronunciation feature annotation on each syllable in the speech phoneme annotation specifically is: Obtain the initial consonant annotation of the syllable, and automatically annotate the initial consonant pronunciation feature according to the determined initial consonant pronunciation type; Obtain the final consonant annotation of the syllable, and automatically annotate the final consonant pronunciation feature according to the constructed final consonant pronunciation type; Obtain the tone annotation of the syllable, and automatically annotate the tone pronunciation feature according to the constructed tone pronunciation type; Construct the fine-grained acoustic features of the syllable according to the annotated initial consonant pronunciation feature, final consonant pronunciation feature, and tone pronunciation feature.
5. The voice rehabilitation training method according to claim 3, wherein The comparing and evaluating the speech phoneme annotation and the phoneme annotation of the corresponding speech rehabilitation training data, and recording the coarse-grained voiceprint analysis result includes: Construct a hash table with the name of the speech disorder type as the keyword, wherein the value of the hash table is the number of speech disorder errors; Extract the syllable type and the initial consonant annotation, final consonant annotation, and tone annotation of each syllable of the phoneme annotation in the speech rehabilitation training data; Extract the syllable types marked by phonemes in the speech, as well as the initial consonant markings, final vowel markings, and tone markings for each syllable. Compare the syllable types, initial consonant markings, final vowel markings, and tone markings corresponding to the phoneme markings in the speech rehabilitation training data and the phoneme markings in the speech respectively. If they are different, increment the value corresponding to the keyword in the hash table by 1 as the coarse-grained voiceprint analysis result.
6. The voice rehabilitation training method according to claim 5, characterized in that, The above-mentioned comparison evaluates the fine-grained acoustic features of the subject's speech feedback data and the fine-grained acoustic features of the corresponding speech rehabilitation training data to obtain the fine-grained acoustic feature differences, including: Construct a hash table based on the initial fine-grained acoustic feature differences corresponding to the subject's speech feedback data. The keyword of the hash table is the initial consonant error or tone error, and the value of the hash table is the number of acoustic feature mispronunciation times. Extract each syllable of the fine-grained acoustic features in the speech rehabilitation training data, and extract each syllable of the subject's fine-grained acoustic features. Compare the syllables corresponding to the fine-grained acoustic features in the speech rehabilitation training data and the subject's fine-grained acoustic features respectively. When they are different, generate a fine-grained pronunciation feature keyword by combining the initial consonant, final vowel, and tone markings, and look up this keyword in the hash table. If this keyword does not exist, create a new keyword in the hash table and set its value to 1 as the updated fine-grained acoustic feature difference.
7. The voice rehabilitation training method according to claim 6, wherein, The above-mentioned statistic on the acoustic analysis results of different granularities, and draw the subject's speech portrait through the statistical results, including: Conduct frequency statistics on the coarse-grained voiceprint analysis result to obtain the coarse-grained portrait of the subject's speech disorder. Conduct frequency statistics on the updated fine-grained acoustic feature differences to obtain the fine-grained portrait of the subject's speech disorder. Jointly construct the subject's speech portrait through the coarse-grained portrait and the fine-grained portrait.
8. The voice rehabilitation training method according to claim 7, characterized in that, The above-mentioned combines the subject's speech portrait and selects matching speech rehabilitation training cases from the speech rehabilitation training case library for the subject to train, including: Use the different speech disorders in the coarse-grained portrait as the training case grouping unit, and sequentially select speech rehabilitation training cases that match the fine-grained portrait features in the speech rehabilitation training case library. According to the speech disorder types of the selected speech rehabilitation training cases, allocate these speech rehabilitation training cases to different groups, and repeat the process of selecting speech rehabilitation training cases that match the fine-grained portrait features until the number of recommended matching speech rehabilitation training cases in each group reaches the preset number. Use the recommended matching speech rehabilitation training cases in each group for the subject to train.
9. A speech rehabilitation training system, characterized in that, Including: A construction module for constructing a speech rehabilitation training case library and obtaining the phoneme markings and fine-grained acoustic features of the speech rehabilitation training data in the speech rehabilitation training case library. An analysis module for obtaining the subject's speech feedback data, and performing acoustic analysis of different granularities on the subject's speech feedback data through the phoneme markings and fine-grained acoustic features of the speech rehabilitation training data to obtain acoustic analysis results of different granularities. A recommendation module, which is used to statistically analyze the acoustic analysis results of different granularities, draw a voice portrait of the subject through the statistical results, select a matching voice rehabilitation training case from the voice rehabilitation training case library for the subject to train in combination with the voice portrait of the subject, and perform an iterative process of acoustic analysis on the voice feedback data of the subject obtained from each training and selecting a matching voice rehabilitation training case to correct the training of the subject.
10. A computer device, characterized in that, It includes a memory and a processor. A program is stored in the memory. When the program is executed by the processor, the processor is caused to execute the steps of a voice rehabilitation training method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Speech recognition model training method and device, equipment and storage medium
CN113393841A
Amphasia patient auxiliary rehabilitation training method and device based on fusion speech recognition
CN114617769A
Language training method and system based on deep learning
CN114758647A
Voice rehabilitation analysis method and system based on acoustic analysis algorithm
CN117976141A
Pinch detecting device for vehicle with improved detection performance
KR1020250132253A