Speech rehabilitation training method, system and computer equipment

By constructing a voice rehabilitation training case library and performing phoneme annotation and fine-grained acoustic analysis, personalized voice portraits are drawn, and the problem of difficulty in personalized implementation in autistic voice rehabilitation training is solved, and the precise selection and effect improvement of personalized training plans are achieved.

CN120220965BActive Publication Date: 2025-08-12LESHAN NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510694353.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-08-12
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

In the prior art, autism voice rehabilitation training has problems such as difficulty in implementing personalized, low family participation and mechanized training content. The traditional manual intervention method consumes manpower and is difficult to meet the individual rehabilitation training needs of each subject.

Method used

A voice rehabilitation training case library was constructed, and the subjects' voice feedback data were analyzed through phoneme annotation and fine-grained acoustic characteristics, personalized voice portraits were drawn, matching rehabilitation training cases were selected and iterated correction training plans were realized to achieve personalized voice rehabilitation training.

Benefits of technology

It realizes accurate characterization of the subject's individual speech disorder characteristics, provides personalized rehabilitation training plans, improves training results and reduces the requirements of professionals, making it convenient for home applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220965B_ABST
    Figure CN120220965B_ABST
Patent Text Reader

Abstract

The present invention discloses a speech rehabilitation training method, system, and computer equipment, relating to the field of intelligent rehabilitation training technology. The method comprises the following steps: constructing a speech rehabilitation training case library and obtaining phoneme annotations and fine-grained acoustic features of speech rehabilitation training data in the speech rehabilitation training case library; performing acoustic analysis of subject speech feedback data at different granularities using the phoneme annotations and fine-grained acoustic features of the speech rehabilitation training data to obtain acoustic analysis results at different granularities; statistically analyzing the acoustic analysis results at different granularities to draw a speech profile of the subject; and selecting an optimal speech rehabilitation training case from the speech rehabilitation training case library based on the subject's speech profile for the subject to train. The present invention selects a more targeted, personalized rehabilitation training program from the speech rehabilitation training case library based on the acoustic analysis of the subject's speech feedback data, thereby improving the rehabilitation training effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent rehabilitation training, and in particular to a speech rehabilitation training method, system and computer equipment. Background Art

[0002] In recent years, the incidence of autism has continued to rise, with the condition often beginning in childhood and lasting throughout life. Language impairment is one of the core barriers to autism. Effective language rehabilitation training and assessment, as well as improving language skills in children with autism, are crucial for promoting their development and autism-inclusive education. Generally speaking, language impairments are categorized into speech, vocabulary, and grammar, as well as communication barriers.

[0003] For individuals with autism (hereinafter referred to as subjects), most lack adequate guidance and training during the critical period of speech learning due to language cognitive development or social communication difficulties, resulting in delayed or impaired speech development. Only some subjects may have structural or functional abnormalities in organs such as hearing, oral muscles, larynx, or tongue.

[0004] Currently, speech rehabilitation training for autism still faces problems such as mechanized content, difficulty in personalized implementation, and low family participation. First, there are large individual differences among the subjects, especially language development, which varies from person to person, making intervention very difficult. Traditional manual intervention methods require intervention personnel to analyze the specific circumstances of the individual and make a detailed assessment of the subject's language ability before formulating an intervention plan. This intervention method is labor-intensive, and because the training content is repetitive and mechanical, it may cause anxiety among the subjects, resulting in low participation.

[0005] Existing computer-assisted intervention methods often utilize learning game software, augmented reality technology, and robotics, attempting to provide participants with a relaxed and easy-to-use testing and learning environment and alleviate their anxiety. However, this approach is difficult to implement individually. Training is based solely on pre-defined fixed patterns, making it difficult to accurately characterize the characteristics of different speech disorders and unable to effectively meet the individual rehabilitation training needs of each participant. Summary of the Invention

[0006] The purpose of the present invention is to provide a speech rehabilitation training method, system and computer equipment to solve the problems in the prior art.

[0007] The present invention specifically provides the following technical solutions:

[0008] A speech rehabilitation training method comprising:

[0009] Construct a speech rehabilitation training case library and obtain the phoneme annotations and fine-grained acoustic features of the speech rehabilitation training data in the speech rehabilitation training case library; the fine-grainedness refers to the local analysis dimension;

[0010] Obtaining the subject's speech feedback data, performing acoustic analysis of different granularities on the subject's speech feedback data using the phoneme annotation and fine-grained acoustic features of the speech rehabilitation training data, and obtaining acoustic analysis results of different granularities;

[0011] The acoustic analysis results of different granularities are statistically analyzed, and the subject's voice portrait is drawn based on the statistical results. Combined with the subject's voice portrait, matching speech rehabilitation training cases are selected from the speech rehabilitation training case library for the subject's training. The subject's voice feedback data obtained from each training session is subjected to acoustic analysis and an iterative process of selecting matching speech rehabilitation training cases to correct the subject's training.

[0012] Preferably, the construction of a speech rehabilitation training case library includes:

[0013] Collect unlabeled speech rehabilitation training data;

[0014] Perform phoneme annotation on the unlabeled speech rehabilitation training data, compare different annotation results, integrate to obtain speech rehabilitation training data with phonemes, and perform phoneme annotation on each data in the speech rehabilitation training data with phonemes;

[0015] Fine-grained acoustic features are annotated on speech rehabilitation training data with phonemes to obtain a fine-grained speech case library, and a speech rehabilitation training case library is constructed based on the fine-grained speech case library.

[0016] Preferably, performing acoustic analysis of different granularities on the subject's speech feedback data through phoneme annotation and fine-grained acoustic features of the speech rehabilitation training data to obtain acoustic analysis results of different granularities includes:

[0017] Obtaining the speech phoneme annotations corresponding to the subject's speech feedback data; each speech phoneme annotation includes a syllable type and a number of syllables, and each syllable includes an initial consonant annotation, a final vowel annotation, and a tone annotation;

[0018] Annotate each syllable in the phonetic phoneme annotation with its acoustic pronunciation features;

[0019] After the acoustic vocal features are annotated, the phoneme annotations of the speech are compared and evaluated with the phoneme annotations of the corresponding speech rehabilitation training data, and the coarse-grained voiceprint analysis results are recorded;

[0020] After the acoustic pronunciation features are labeled, the fine-grained acoustic features of the subject's speech feedback data and the fine-grained acoustic features of the corresponding speech rehabilitation training data are compared and evaluated to obtain the fine-grained acoustic feature differences.

[0021] Preferably, the step of marking each syllable in the speech phoneme annotation with its acoustic pronunciation features is as follows:

[0022] Obtaining the initial consonant annotation of the syllable, and automatically annotating the initial consonant pronunciation characteristics according to the determined initial consonant pronunciation type;

[0023] Obtaining the final vowel annotation of the syllable, and automatically annotating the final vowel pronunciation characteristics according to the constructed final vowel pronunciation type;

[0024] Obtaining the tone annotation of the syllable and automatically annotating the tone pronunciation features according to the established tone pronunciation type;

[0025] The fine-grained acoustic features of the syllable are constructed based on the marked initial consonant pronunciation features, final vowel pronunciation features and tone pronunciation features.

[0026] Preferably, the comparative evaluation of the speech phoneme annotations and the phoneme annotations of the corresponding speech rehabilitation training data, and recording of the coarse-grained voiceprint analysis results, includes:

[0027] Construct a hash table whose key is the name of the speech disorder type, where the value of the hash table is the number of speech disorder errors;

[0028] Extracting the syllable type of the phoneme annotation and the initial consonant annotation, final vowel annotation and tone annotation of each syllable in the speech rehabilitation training data;

[0029] Extract the syllable type of the phoneme annotation and the initial consonant annotation, final vowel annotation and tone annotation of each syllable;

[0030] Compare the syllable types, initial consonant annotations, final consonant annotations, and tone annotations corresponding to the phoneme annotations and speech phoneme annotations in the speech rehabilitation training data respectively. If they are different, add 1 to the value of the corresponding keyword in the hash table as the coarse-grained voiceprint analysis result.

[0031] Preferably, the comparative evaluation of the fine-grained acoustic features of the subject's speech feedback data and the fine-grained acoustic features of the corresponding speech rehabilitation training data to obtain fine-grained acoustic feature differences includes:

[0032] A hash table is constructed based on the initial fine-grained acoustic feature differences corresponding to the subject's speech feedback data. The key of the hash table is the initial consonant error or tone error, and the value of the hash table is the number of acoustic feature errors.

[0033] Extracting each syllable of fine-grained acoustic features from speech rehabilitation training data and each syllable of fine-grained acoustic features from the subject;

[0034] The syllables corresponding to the fine-grained acoustic features in the speech rehabilitation training data and the fine-grained acoustic features of the subjects are compared respectively. The fine-grained pronunciation feature keywords are generated by combining the initial consonants, finals and tone annotations at different times. The keywords are searched in the hash table. If the keywords do not exist, new keywords are created in the hash table and their value is set to 1 as the difference in the fine-grained acoustic features after the update.

[0035] Preferably, the process of performing statistics on acoustic analysis results of different granularities and drawing a voice portrait of the subject based on the statistical results includes:

[0036] Perform frequency statistics on the coarse-grained voiceprint analysis results to obtain a coarse-grained portrait of the subject's speech disorder;

[0037] Perform frequency statistics on the updated fine-grained acoustic feature differences to obtain a fine-grained portrait of the subject's speech disorder;

[0038] The subject's voice profile is constructed through coarse-grained and fine-grained portraits.

[0039] Preferably, the step of combining the subject's voice portrait and selecting matching voice rehabilitation training cases from a voice rehabilitation training case library for the subject to train comprises:

[0040] The different speech disorders in the coarse-grained portrait are used as training case grouping units, and speech rehabilitation training cases that match the fine-grained portrait features are selected in sequence from the speech rehabilitation training case library;

[0041] Allocating the selected speech rehabilitation training cases to different groups according to their speech disorder types, and repeating the process of selecting speech rehabilitation training cases that match the fine-grained portrait features until the number of recommended matching speech rehabilitation training cases in each group reaches a preset number;

[0042] Use the matching speech rehabilitation training cases recommended in each group for subjects to train.

[0043] The present invention provides a speech rehabilitation training system, comprising:

[0044] A construction module is used to construct a speech rehabilitation training case library and obtain phoneme annotations and fine-grained acoustic features of speech rehabilitation training data in the speech rehabilitation training case library;

[0045] An analysis module is used to obtain the subject's speech feedback data, perform acoustic analysis of the subject's speech feedback data at different granularities through phoneme annotation and fine-grained acoustic features of the speech rehabilitation training data, and obtain acoustic analysis results at different granularities;

[0046] The recommendation module is used to perform statistics on the acoustic analysis results of different granularities, draw the subject's voice portrait based on the statistical results, and select matching voice rehabilitation training cases from the voice rehabilitation training case library for the subject's training based on the subject's voice portrait. The subject's voice feedback data obtained from each training session is acoustically analyzed and matching voice rehabilitation training cases are selected in an iterative process to correct the subject's training.

[0047] The present invention provides a computer device, comprising a memory and a processor, wherein a program is stored in the memory, and when the program is executed by the processor, the processor executes the steps of the above-mentioned speech rehabilitation training method.

[0048] Compared with the prior art, the present invention has the following significant advantages:

[0049] The present invention performs acoustic analysis based on the obtained voice feedback data of the subjects, performs acoustic analysis of different granularities on the voice feedback data of the subjects through the phoneme annotation and fine-grained acoustic features of the voice rehabilitation training data, obtains acoustic analysis results of different granularities, provides richer indicators for characterizing the personalized voice portrait of the subjects, and compares and selects with the voice rehabilitation training case library annotated with fine-grained acoustic features, performs acoustic analysis on the voice feedback data of the subjects obtained in each training and selects matching voice rehabilitation training cases in an iterative process, corrects the training of the subjects, and can select more targeted personalized rehabilitation training plans from the voice rehabilitation training case library, realizes automatic correction and updating of the subject portrait after each training, adapts to the changes in individual rehabilitation trajectory by correcting the training process, thereby improving the rehabilitation training effect of the subjects. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 A framework for constructing the speech rehabilitation training process in the present invention;

[0051] Figure 2 The figure is a flow chart of a speech rehabilitation training method in the present invention. DETAILED DESCRIPTION

[0052] The following is a clear and complete description of the technical solutions of the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0053] like Figure 1 and Figure 2 As shown, this embodiment provides a speech rehabilitation training method, comprising the following steps:

[0054] S1. Establish the Speech Disorder Catelog (SDCatelog) classification system for autism speech disorders, construct the speech rehabilitation training case library STCases (STCases), and obtain the phoneme annotations and fine-grained acoustic features of the speech rehabilitation training data in the speech rehabilitation training case library. Fine-grainedness refers to the local analysis dimension. The specific steps include:

[0055] S11. Collect unlabeled speech rehabilitation training data (cases) USTData (Unlabeled Speech Training Data) from existing public autism language assessment scales and related research results.

[0056] S12. Analyze Chinese speech characteristics and construct a fine-grained speech disorder classification system. The specific steps include:

[0057] S121. Analyze and determine the top-level classification of speech, mainly including initials, finals, tones and syllables.

[0058] S122. Analyze and determine the initial consonant pronunciation types (InitialTypes), which mainly include bilabial consonants, lingual consonants, lingual root consonants, and front and back tongue tip consonants. For example, bilabial consonants include "b," "p," and "m," and lingual root consonants include "g," "k," and "h."

[0059] S123. Analyze and determine the pronunciation type of the final vowel , mainly including close-toothed pronunciation, closed-mouthed pronunciation, pursed-lipped pronunciation and open-mouthed pronunciation. For example, close-toothed pronunciation includes "i", "ia" and "ing", and closed-mouthed pronunciation includes "u", "ua" and "uo".

[0060] S124. Analyze and construct tone pronunciation types , mainly including yinping, yangping, shangsheng, qusheng and light tone.

[0061] S125. Based on different vocal characteristics, a speech disorder classification system SDCatelog was established, which mainly includes initial consonant loss, initial consonant transformation, initial consonant addition, final consonant loss, final consonant transformation and denasalization.

[0062] The present invention combines the acoustic pronunciation characteristics of Chinese and annotates each speech rehabilitation training case with fine-grained annotations. The annotation results are essentially labels, which provide richer indicators for portraying the personalized speech portrait of autistic subjects.

[0063] S13. In the fine-grained speech disorder classification system, the unlabeled speech rehabilitation training data (speech cases) USTData is annotated with phonemes, including syllable type, , syllable, initial, rhyme and tone, etc., that is, to perform personalized speech rehabilitation training corpus annotation, and compare different annotation results to integrate speech rehabilitation training data with phonemes (Phoneme Speech Training Data), and each speech rehabilitation training data in the speech rehabilitation training data with phonemes Phonemic annotation , the specific steps include:

[0064] S131. Design phoneme notation standards, including tone symbols and placement rules, such as syllable type ("1" for monosyllabic, "2" for disyllabic, "0" for polysyllabic), marking the tone of the complex vowel "uo" on the last vowel "o", using "#" to mark yangping (high-pitched tone), and "^" to mark shengsheng (rising-level tone).

[0065] S132. Recruit multiple annotators to perform phoneme annotation using a cross-annotation approach. The specific steps include:

[0066] S1321. Speech rehabilitation training data according to the number of annotators The distribution was performed evenly, ensuring that at least 30% of each case set was repeated with other persons.

[0067] S1322. The annotators shall annotate the words according to the phoneme annotation standards.

[0068] S1323. The marking personnel shall perform cross-marking.

[0069] S133. Compare different annotation results and integrate them to obtain speech rehabilitation training data with phonemes , where inconsistent annotation results are determined by centralized discussion, each speech rehabilitation training data The phoneme notation is (Each phoneme annotation includes syllable type , and the corresponding syllable annotations , each syllable is marked Contains initial consonant annotations , vowel marking Harmony and tone marking one each).

[0070] For example, given speech rehabilitation training data (training word "apple"), its phonemic annotation "2ping#guo^", syllable type "2" represents a disyllabic word, which contains two syllables marked "ping#" and "guo^", of which the first syllable is marked "ping~" (initial consonant mark For "p", the final is marked For "ing", tone marking The second syllable is marked For "guo^" (initial consonant mark For "g", the final is marked "uo", tone mark is “^”).

[0071] S14. Speech rehabilitation training data with phonemes Automatically annotate fine-grained acoustic features to obtain a fine-grained speech case library , the specific steps include:

[0072] S141, sequentially scanning speech rehabilitation training data with phonemes Each speech rehabilitation training data in , get phoneme annotations .

[0073] S142, scan phoneme annotations sequentially Each syllable in ,right Automatically labeling acoustic sound features includes the following steps:

[0074] S1421. Obtain the initial consonant label of the syllable , according to the determined initial sound type Automatically label initial consonant pronunciation features .

[0075] S1422, obtain the final vowel mark of the syllable , according to the determined final pronunciation type Automatically label the pronunciation characteristics of finals .

[0076] S1423. Obtain the tone mark of the syllable , according to the determined tone phonation type Automatically label tone pronunciation features .

[0077] S1424: syllables are formed from the above automatic annotation results. Fine-grained acoustic features.

[0078] For example, given speech rehabilitation training data (training word "apple"), its phonemic annotation is “2ping#guo^”, and its fine-grained acoustic feature annotation is "2p1ing1#2g2uo2^3", respectively:

[0079] First syllable mark "ping#" (initial consonant mark It is "p", so the pronunciation characteristics of the initial consonant It is a "bilabial consonant", recorded as 1, and the final is marked It is "ing", so the pronunciation characteristics of the final vowel are It is "qi zhi hu", recorded as 1, and the tone is marked The tone pronunciation feature is “#”, which is recorded as 2. Therefore, the fine-grained acoustic feature corresponding to the syllable is annotated as “p1ing1#2”.

[0080] Second syllable mark For "guo^" (initial consonant mark It is "g", so the pronunciation characteristics of the initial consonant It is a "tongue consonant", recorded as 2, and the final is marked It is "uo", so the pronunciation characteristics of the final vowel are It is a "closed mouth call", recorded as 2, and the tone is marked The tone pronunciation feature is “^”, which is recorded as 3. The fine-grained acoustic feature corresponding to the syllable is annotated as “g2uo2^3”.

[0081] S143, repeat the above operation until the speech rehabilitation training data is Until all speech rehabilitation training data are automatically labeled.

[0082] S15, speech rehabilitation training data with phonemes After fine-grained acoustic feature annotation, a fine-grained speech case library As a basis or jointly build a speech rehabilitation training case library suitable for personalized training .

[0083] S2. Obtain the subject's speech feedback data, perform acoustic analysis of different granularities on the subject's speech feedback data through phoneme annotation and fine-grained acoustic features of the speech rehabilitation training data, and obtain acoustic analysis results of different granularities.

[0084] The subject's voice feedback data and corresponding speech phoneme annotations are obtained, and acoustic analysis is performed on each piece of subject's voice feedback data. The speech phoneme annotations corresponding to the subject's voice feedback data and the phoneme annotations of the corresponding speech rehabilitation training data are compared and evaluated. The corresponding speech rehabilitation training data is used as training corpus to obtain coarse-grained voiceprint analysis results and fine-grained acoustic feature differences; the coarse-grainedness is the global analysis dimension, which is obtained when classifying the subject's voice feedback data, and the fine-grainedness is the local analysis dimension, which is obtained when recognizing the subject's voice feedback data.

[0085] Perform acoustic analysis on the subject's speech feedback data, compare and evaluate the phoneme annotations corresponding to the subject's speech feedback data and the phoneme annotations of the corresponding speech rehabilitation training data, and obtain coarse-grained voiceprint analysis results and fine-grained acoustic feature differences. The specific steps include:

[0086] S21. Collect or obtain the subject's voice feedback data from the subject's rehabilitation training behavior library .

[0087] S22. Feedback data for each subject’s voice The specific steps of conducting acoustic analysis include:

[0088] S221. Use phoneme-based acoustic models (such as GMM-HMM, DNN-HMM, or Transformer-based models) to automatically obtain subject speech feedback data. Corresponding phonetic phoneme annotations . Each phoneme is annotated Including syllable types and several syllables , each syllable Including initial consonant marking , vowel marking Harmony and tone marking For example, given the training word "apple", the subject's speech feedback data is , after model learning, the speech phoneme annotation is obtained This embodiment selects wav2vec2 as the phoneme recognition model to complete the prediction of speech phoneme annotation.

[0089] S222. Annotate phonemes Each syllable in Automatically label the acoustic sound features separately, and the specific steps include:

[0090] S2221. Get the initial consonant label of the syllable , according to the determined initial sound type Automatically label initial consonant pronunciation features .

[0091] S2222. Get the final vowel mark of the syllable , according to the constructed final pronunciation type Automatically label the pronunciation characteristics of finals .

[0092] S2223. Obtain the tone mark of the syllable , according to the constructed tone voicing type Automatically label tone pronunciation features .

[0093] S2224: Constructing syllables based on the automatically labeled initial consonant pronunciation features, final vowel pronunciation features, and tone pronunciation features Fine-grained acoustic features.

[0094] For example, given the training word "apple", the subject's speech feedback data is , whose phoneme annotation is “2pi#puo^”, and its fine-grained acoustic feature annotation is "2p1i1#2 p1uo2^3", respectively:

[0095] First syllable For "pi#" (initial consonant mark It is "p", so the pronunciation characteristics of the initial consonant It is a "bilabial consonant", recorded as 1, and the final is marked It is "i", so the pronunciation characteristics of the final vowel are It is "qi zhi hu", recorded as 1, and the tone is marked is “#”, so the tone pronunciation characteristics It is “yangping”, recorded as 2, so the fine-grained acoustic feature corresponding to this syllable is annotated as “p1i1#2”.

[0096] Second syllable mark For "puo^" (initial consonant mark It is "p", so the pronunciation characteristics of the initial consonant It is a "bilabial consonant", recorded as 1, and the final is marked It is "uo", so the pronunciation characteristics of the final vowel are It is a "closed mouth call", recorded as 2, and the tone is marked is "^", so the tone pronunciation characteristics is "shangsheng", recorded as 3, and the fine-grained acoustic feature annotation corresponding to this syllable is "p1uo2^3".

[0097] Based on the learned speech factor annotations, voiceprint analysis of different granularities is performed on the subject's feedback.

[0098] S223. After the acoustic voice features are marked, compare and evaluate the subject's voice feedback data Corresponding phonetic phoneme annotations and phoneme annotation of corresponding speech rehabilitation training data , record the coarse-grained voiceprint analysis results , for example, “initial consonant error”, “initial consonant loss” and “tone error”, etc. The specific steps include;

[0099] S2231, Speech Disorder Classification System Based on the different types in the , a hash table is constructed with the keywords as the name of the speech disorder type (such as "initial consonant error" and "initial consonant loss", etc.) , the value is the number of speech barrier errors (initial value is 0).

[0100] S2232. Extract phoneme annotations from speech rehabilitation training data Syllable type and each syllable Initial consonant marking , vowel marking Harmony and tone marking .

[0101] S2233. Extracting speech phoneme annotations Syllable type and each syllable Initial consonant marking , vowel marking Harmony and tone marking .

[0102] S2234, respectively compare the syllable type and initial consonant annotation, final vowel annotation and final vowel annotation, and tone annotation and tone annotation corresponding to the phoneme annotation and the speech phoneme annotation in the speech rehabilitation training data. If they are different, the hash table The value of the corresponding keyword in is added by 1, which is used as the coarse-grained voiceprint analysis result.

[0103] For example, given speech rehabilitation training data (training word "apple"), its phonemic annotation The subject’s voice feedback data is “2ping#guo^” Phonemic annotations If it is "2pi#puo^", the coarse-grained voiceprint analysis result The number of “initial consonant errors” is 1, and the number of “final vowel errors” is 1.

[0104] S224. After the acoustic voice features are marked, compare and evaluate the subject's voice feedback data Fine-grained acoustic features and the fine-grained acoustic features of the corresponding speech rehabilitation training data , and obtain fine-grained acoustic feature differences For example, "initial consonant error" needs to be further refined into "initial consonant error: bilabial consonant b changes to bilabial consonant d", and "tone error" needs to be further refined into "tone error: rising tone changes to falling tone" or "tone error: rising tone changes to yangping tone", etc. The specific steps include:

[0105] S2241. Based on the subject's voice feedback data Corresponding to the initial fine-grained acoustic feature differences Construct a hash table, where the key of the hash table is "initial consonant error: bilabial consonant b changes to bilabial consonant d" or "tone error: rising tone changes to falling tone", etc. The value of the hash table is the number of acoustic feature errors (the initial value is 0).

[0106] S2242. Extracting fine-grained acoustic features from speech rehabilitation training data Each syllable in (Including initial consonant markings , vowel marking Harmony and tone marking , initial consonant pronunciation characteristics 、The pronunciation characteristics of the finals Vocal characteristics of harmonic tones wait).

[0107] S2243. Extract fine-grained acoustic features of subjects Each syllable of (Including initial consonant markings , vowel marking , tone marking , initial consonant pronunciation characteristics , the pronunciation characteristics of the finals Vocal characteristics of harmonic tones wait).

[0108] S2244. If the syllables (initial pronunciation features, final pronunciation features, and tone pronunciation features) corresponding to the fine-grained acoustic features in the speech rehabilitation training data and the subject's fine-grained acoustic features are different, the fine-grained pronunciation feature keywords are generated by combining the initial consonant, final vowel, and tone annotations, such as "Initial consonant error: bilabial consonant b changes to bilabial consonant d". If the keyword exists, the value corresponding to the keyword is increased by 1. If the keyword does not exist, a new keyword is created in the hash table and its value is set to 1 as a fine-grained acoustic feature difference. .

[0109] For example, given speech rehabilitation training data (training word "apple"), its fine-grained acoustic feature annotation For "2p1ing1#2g2uo2^3", the subject's voice feedback data Fine-grained acoustic feature annotation "2p1i1#2p1uo2^3", fine-grained acoustic feature differences The medium-coarse-grained "initial error" was refined into "initial error: the root consonant g changed to the bilabial consonant p", and the "final error" was refined into "final error: the dental ing changed to the dental i", and the number of errors was 1.

[0110] S23, repeat step S22 until the subject's voice feedback data Until the end of the analysis, the coarse-grained voiceprint analysis results of the subjects are obtained. and fine-grained acoustic feature differences .

[0111] S3. Statistically analyze the acoustic analysis results at different granularities and draw a voice profile of the subject based on the statistical results. The specific steps include:

[0112] S31. Results of coarse-grained voiceprint analysis By performing frequency statistics, we can obtain a coarse-grained portrait of the subject's speech disorder. , the specific steps include:

[0113] S311. Scan the coarse-grained voiceprint analysis results in sequence The coarse-grained voiceprint analysis results corresponding to each voice feedback data of the subjects .

[0114] S312. Categorize and accumulate the voiceprint analysis results by the type of speech disorder (such as "initial consonant error" and "initial consonant loss"). The number of errors.

[0115] S313. Sort by the number of speech disorder errors.

[0116] S314. Select the top K (K is preset to 3) speech disorders with the highest frequency as the coarse-grained profile of the subject. .

[0117] It is worth noting that the sorting in this embodiment is based on the number of errors, and the speech disorder selection adopts the topK method. However, this is not limited to this. The sorting can also be implemented based on the error probability, and the speech disorder type can also be selected based on a preset threshold.

[0118] S32. Differences in fine-grained acoustic features By performing frequency statistics, we can obtain a fine-grained portrait of the subject's speech disorder. , the specific steps include:

[0119] S321, sequentially scan fine-grained acoustic feature differences The fine-grained acoustic feature differences corresponding to each speech feedback data of the subjects .

[0120] S322. Accumulate the number of errors of different fine-grained acoustic features by category (such as "initial consonant error: bilabial consonant b changes to bilabial consonant d" and "tone error: rising tone changes to falling tone", etc.).

[0121] S323. Sort by the number of errors of fine-grained acoustic features.

[0122] S324. Select the top K (K is preset to 3) acoustic features with the highest number of times as the fine-grained portrait of the subject .

[0123] It is worth noting that the sorting in this embodiment is based on the number of errors, and the selection of fine-grained acoustic features adopts the topK method. However, it is not limited to this. The sorting can also be implemented based on the error probability, and the selection of fine-grained acoustic features can also be implemented based on a preset threshold.

[0124] S33, the coarse-grained image obtained from the above step S3 and fine-grained portraits Together they constitute the voice portrait of the subject.

[0125] S4. Based on the subject's voice profile, matching voice rehabilitation training cases are selected from the voice rehabilitation training case library for the subject to train, i.e., personalized voice rehabilitation training case recommendations are made. The specific steps include:

[0126] S41. Coarse-grained portrait of the subjects Different speech disorders in the training case grouping units, such as "initial consonant error" or "initial consonant loss", etc.

[0127] S42. In the personalized speech rehabilitation training case library Select the fine-grained portrait that matches Characteristic speech rehabilitation training case, for example, if There is "Initial error: bilabial consonant b changes to bilabial consonant d", then Select the case where the initial consonant in the phoneme annotation is marked with "bilabial consonant b".

[0128] S43. Allocate the selected speech rehabilitation training cases to different groups according to the speech disorder types of the selected speech rehabilitation training cases.

[0129] S434, repeat the above process (select matching fine-grained portrait The process of case selection based on the characteristics of the features is repeated until the number of recommended matching speech rehabilitation training cases in each group reaches the preset number.

[0130] S5. Implement personalized speech rehabilitation training, perform acoustic analysis on the subject's speech feedback data obtained during each training session, select matching speech rehabilitation training cases, and modify the subject's training. The specific steps include:

[0131] S51, respectively, using coarse-grained portraits of the subjects The speech impairment in the training task.

[0132] S52. Show the pictures corresponding to the training cases (with accompanying pronunciation is acceptable) and guide the subjects to carry out speech rehabilitation training.

[0133] S53. Collect the subject's voice data from the personalized voice rehabilitation training and store it in the subject's rehabilitation training behavior library for voiceprint analysis.

[0134] S54. Return to S52 to iteratively carry out personalized speech rehabilitation training.

[0135] Use the matching speech rehabilitation training cases recommended in each group for subjects to train.

[0136] It should be noted that if speech rehabilitation training is being implemented for a subject for the first time, a cold start approach can be used (referring to a language assessment scale, such as the Chinese Communication Development Scale). Case selection can be based on the subject's expected performance, and the training can be terminated if the subject frequently makes repeated errors or reaches the assessment goal. This approach differs from previous models in that it reduces the professional requirements for training and assessment personnel and facilitates the promotion of home use. It also allows for a more accurate characterization of the subject's speech impairment, allowing for more targeted, personalized rehabilitation training. It is an effective and easily scalable training method.

[0137] The present invention proposes a speech rehabilitation training system, which includes a construction module, an analysis module and a recommendation module.

[0138] Among them, the construction module is used to construct a speech rehabilitation training case library and obtain the phoneme annotations and fine-grained acoustic features of the speech rehabilitation training data in the speech rehabilitation training case library; the analysis module is used to obtain the subject's speech feedback data, and perform acoustic analysis of different granularities on the subject's speech feedback data through the phoneme annotations and fine-grained acoustic features of the speech rehabilitation training data to obtain acoustic analysis results of different granularities; the recommendation module is used to perform statistics on the acoustic analysis results of different granularities, draw the subject's speech portrait through the statistical results, and select matching speech rehabilitation training cases from the speech rehabilitation training case library for the subject's training based on the subject's speech portrait, and perform an iterative process of acoustic analysis and selection of matching speech rehabilitation training cases on the subject's speech feedback data obtained from each training to correct the subject's training.

[0139] The present invention also provides a computer device, comprising a memory and a processor. The memory stores a program, and when the program is executed by the processor, the processor executes the steps of a speech rehabilitation training method.

[0140] According to the disclosed embodiments, a computing device may communicate with one or more external devices (e.g., keyboards, pointing devices, Bluetooth communications, etc.), or with any device that enables a computing device to communicate with one or more other computing devices (e.g., routers, modems, etc.).

[0141] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. For those skilled in the art to which the present invention belongs, several simple deductions or replacements can be made without departing from the concept of the present invention, which should be regarded as falling within the scope of protection of the present invention.

Claims

1. A speech rehabilitation training method, characterized in that: include: Build a speech rehabilitation training case library and obtain the phoneme annotations and fine-grained acoustic features of the speech rehabilitation training data in the speech rehabilitation training case library; Among them, fine-grainedness is the local analysis dimension; Obtaining the subject's speech feedback data, performing acoustic analysis of different granularities on the subject's speech feedback data using the phoneme annotation and fine-grained acoustic features of the speech rehabilitation training data, and obtaining acoustic analysis results of different granularities; The acoustic analysis results of different granularities are statistically analyzed, and a voice profile of the subject is drawn based on the statistical results. Based on the subject's voice profile, matching voice rehabilitation training cases are selected from the voice rehabilitation training case library for the subject to train with. The subject's voice feedback data obtained during each training session is subjected to an iterative process of acoustic analysis and matching voice rehabilitation training cases, and the subject's training is corrected. The construction of the speech rehabilitation training case library includes: Collect unlabeled speech rehabilitation training data; Analyze Chinese phonetic features, perform phoneme annotation on unlabeled speech rehabilitation training data, compare different annotation results, integrate speech rehabilitation training data with phonemes, and perform phoneme annotation on each data in the speech rehabilitation training data with phonemes; The speech rehabilitation training data with phonemes is annotated with fine-grained acoustic features to obtain a fine-grained speech case library, and a speech rehabilitation training case library is constructed based on the fine-grained speech case library. By using the phoneme annotation and fine-grained acoustic features of speech rehabilitation training data, we conduct acoustic analysis of the subject's speech feedback data at different granularities, obtaining acoustic analysis results at different granularities, including: Obtaining the speech phoneme annotations corresponding to the subject's speech feedback data; each speech phoneme annotation includes a syllable type and a number of syllables, and each syllable includes an initial consonant annotation, a final vowel annotation, and a tone annotation; Annotate each syllable in the phonetic phoneme annotation with its acoustic pronunciation features; After the acoustic vocal features are annotated, the phoneme annotations of the speech are compared and evaluated with the phoneme annotations of the corresponding speech rehabilitation training data, and the coarse-grained voiceprint analysis results are recorded; After the acoustic utterance features are annotated, the fine-grained acoustic features of the subject's speech feedback data and the fine-grained acoustic features of the corresponding speech rehabilitation training data are compared and evaluated to obtain the fine-grained acoustic feature differences; The acoustic analysis results of different granularities are statistically analyzed and a voice profile of the subject is drawn based on the statistical results, including: Perform frequency statistics on the coarse-grained voiceprint analysis results to obtain a coarse-grained portrait of the subject's speech disorder; Perform frequency statistics on the updated fine-grained acoustic feature differences to obtain a fine-grained portrait of the subject's speech disorder; The subject's voice profile is constructed through coarse-grained and fine-grained portraits; The method of combining the subject's voice profile and selecting matching voice rehabilitation training cases from a voice rehabilitation training case library for the subject to train includes: The different speech disorders in the coarse-grained portrait are used as training case grouping units, and speech rehabilitation training cases that match the fine-grained portrait features are selected in sequence from the speech rehabilitation training case library; Allocating the selected speech rehabilitation training cases to different groups according to their speech disorder types, and repeating the process of selecting speech rehabilitation training cases that match the fine-grained portrait features until the number of recommended matching speech rehabilitation training cases in each group reaches a preset number; Use the matching speech rehabilitation training cases recommended in each group for subjects to train.

2. A speech rehabilitation training method according to claim 1, characterized in that: The acoustic pronunciation features of each syllable in the speech phoneme annotation are respectively marked as follows: Obtaining the initial consonant annotation of the syllable, and automatically annotating the initial consonant pronunciation characteristics according to the determined initial consonant pronunciation type; Obtaining the final vowel annotation of the syllable, and automatically annotating the final vowel pronunciation characteristics according to the constructed final vowel pronunciation type; Obtaining the tone annotation of the syllable and automatically annotating the tone pronunciation features according to the established tone pronunciation type; The fine-grained acoustic features of the syllable are constructed based on the marked initial consonant pronunciation features, final vowel pronunciation features and tone pronunciation features.

3. A speech rehabilitation training method according to claim 1, characterized in that: The comparative evaluation of the speech phoneme annotations and the phoneme annotations of the corresponding speech rehabilitation training data, and recording of the coarse-grained voiceprint analysis results, include: Construct a hash table whose key is the name of the speech disorder type, where the value of the hash table is the number of speech disorder errors; Extracting the syllable type of the phoneme annotation and the initial consonant annotation, final vowel annotation and tone annotation of each syllable in the speech rehabilitation training data; Extract the syllable type of the phoneme annotation and the initial consonant annotation, final vowel annotation and tone annotation of each syllable; Compare the syllable types, initial consonant annotations, final consonant annotations, and tone annotations corresponding to the phoneme annotations and speech phoneme annotations in the speech rehabilitation training data respectively. If they are different, add 1 to the value of the corresponding keyword in the hash table as the coarse-grained voiceprint analysis result.

4. A speech rehabilitation training method according to claim 3, characterized in that: The comparison and evaluation of the fine-grained acoustic features of the subject's speech feedback data and the fine-grained acoustic features of the corresponding speech rehabilitation training data to obtain fine-grained acoustic feature differences includes: A hash table is constructed based on the initial fine-grained acoustic feature differences corresponding to the subject's speech feedback data. The key of the hash table is the initial consonant error or tone error, and the value of the hash table is the number of acoustic feature errors. Extracting each syllable of fine-grained acoustic features from speech rehabilitation training data and each syllable of fine-grained acoustic features from the subject; The syllables corresponding to the fine-grained acoustic features in the speech rehabilitation training data and the fine-grained acoustic features of the subjects are compared respectively. The fine-grained pronunciation feature keywords are generated by combining the initial consonants, finals and tone annotations at different times. The keywords are searched in the hash table. If the keywords do not exist, new keywords are created in the hash table and their value is set to 1 as the difference in the fine-grained acoustic features after the update.

5. A system for the speech rehabilitation training method according to any one of claims 1 to 4, characterized in that: include: A construction module is used to construct a speech rehabilitation training case library and obtain phoneme annotations and fine-grained acoustic features of speech rehabilitation training data in the speech rehabilitation training case library; An analysis module is used to obtain the subject's speech feedback data, perform acoustic analysis of the subject's speech feedback data at different granularities through phoneme annotation and fine-grained acoustic features of the speech rehabilitation training data, and obtain acoustic analysis results at different granularities; The recommendation module is used to perform statistics on the acoustic analysis results of different granularities, draw the subject's voice portrait based on the statistical results, and select matching voice rehabilitation training cases from the voice rehabilitation training case library for the subject's training based on the subject's voice portrait. The subject's voice feedback data obtained from each training session is acoustically analyzed and matching voice rehabilitation training cases are selected in an iterative process to correct the subject's training.

6. A computer device, characterized in that: The method comprises a memory and a processor, wherein a program is stored in the memory, and when the program is executed by the processor, the processor executes the steps of a speech rehabilitation training method as claimed in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Speech recognition model training method and device, equipment and storage medium

    CN113393841A

  • Amphasia patient auxiliary rehabilitation training method and device based on fusion speech recognition

    CN114617769A