Education and rehabilitation training system for children with articulation disorders based on voice interaction
Through a system based on voice interaction, the voice data of children with dysarthria are extracted and analyzed, and the problem of phoneme alignment errors in the prior art is solved, which improves the accuracy of syllable pronunciation error judgment and the effect of rehabilitation training.
Patent Information
- Application Number
- CN202510260050.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-03-06
AI Technical Summary
In the prior art, in the judgment of syllable pronunciation in children with dysarthria, there is a problem of low accuracy, mainly due to the pronunciation deviation of the children, the phoneme alignment errors.
Using a system based on voice interaction, the voice data of the child is obtained through the data acquisition module. The abnormal phoneme extraction module divides the data into normal syllables and abnormal syllables, and extracts abnormal phonemes. Phoneme word formation abnormal module screens out word formation disorder syllables based on the number of occurrences of abnormal syllables and the abnormal word formation indicators, and is used for education and rehabilitation training.
It effectively avoids incorrect alignment, improves the accuracy of erroneous judgments of syllable pronunciation, and enhances the effect of education and rehabilitation training for children with dysarthria.
Smart Images

Figure CN119741949B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of dysarthria speech recognition, and in particular to a dysarthria children education and rehabilitation training system based on speech interaction. Background Art
[0002] Dysarthria is manifested by unclear speech, unfluent speech, inaccurate pronunciation, abnormal volume and rhythm, etc. Dysarthria education is education and training for children with abnormal pronunciation, phonation, resonance, breathing and rhythm. Doctors usually confirm whether there is dysarthria and the degree of pathology through examination of the pronunciation organs and speech evaluation. For preschool children, especially children with autism, the above symptoms can be improved and cured through rehabilitation training or language training.
[0003] Existing methods align audio and text to obtain phoneme sequences and phoneme boundaries, obtain the actual pronounced phoneme sequence, and use dynamic time warping to compare the two phoneme sequences to determine whether the pronunciation is wrong. However, since the pronunciation of children with dysarthria is often irregular, unclear or inaccurate, there may be a large deviation between the pronunciation of phonemes and the standard pronunciation. The incorrect pronunciation of the children may be incorrectly aligned with the standard phonemes, affecting the accuracy of the judgment of syllable pronunciation errors in children with dysarthria, thereby resulting in poor results in the education and rehabilitation training of children with dysarthria. Summary of the invention
[0004] In order to solve the technical problem that the pronunciation of children with dysarthria deviates from the standard pronunciation and reduce the low accuracy of misjudgment of syllable pronunciation of children with dysarthria, the purpose of the present invention is to provide an education and rehabilitation training system for children with dysarthria based on voice interaction. The technical solution adopted is as follows:
[0005] The present invention proposes a speech interaction-based education and rehabilitation training system for children with articulation disorders, the system comprising:
[0006] The data collection module is used to obtain the speech data of different children under guided speech, and the speech data is composed of syllables;
[0007] An abnormal phoneme extraction module is used to divide each speech data into normal syllables and abnormal syllables based on the difference between the guiding speech and the speech data of each patient, and extract abnormal phonemes in the abnormal syllables;
[0008] The phoneme word formation abnormality module is used to obtain the word formation abnormality index of each abnormal syllable in the speech data of all children according to the number of occurrences of each abnormal syllable in the speech data of all children, the number of occurrences of normal syllables containing the abnormal phonemes in each abnormal syllable, and the number of occurrences of abnormal syllables containing the abnormal phonemes in each abnormal syllable;
[0009] The word formation disorder analysis module is used to screen out word formation disorder syllables from the abnormal syllables in the speech data of all children according to the number of times the phoneme of each abnormal syllable appears separately in the speech data of each child and the word formation disorder index of each abnormal syllable, and provide articulation disorder education and rehabilitation training to the children based on the word formation disorder syllables.
[0010] Furthermore, based on the difference between the guiding speech and the speech data of each child, each speech data is divided into normal syllables and abnormal syllables, and abnormal phonemes in the abnormal syllables are extracted, including:
[0011] The speech data of all the children and the guided speech are recorded as the analysis speech; the analysis speech is segmented, and the segmented syllables are arranged in order to obtain a syllable sequence; each syllable in the syllable sequence is decomposed into phonemes, and the decomposed phonemes are arranged in order to obtain a phoneme sequence of the corresponding syllable;
[0012] For the speech data of all children, determine whether each syllable in the syllable sequence of each speech data is equal to the phoneme sequence of the syllable with the same subscript in the syllable sequence of the guide speech, if so, record each syllable in the syllable sequence of each speech data as a normal syllable; if not, record each syllable in the syllable sequence of each speech data as an abnormal syllable;
[0013] A standard syllable of each abnormal syllable is selected from the syllable sequence of the guiding speech; each abnormal syllable has the same subscript as its standard syllable; if each abnormal syllable has a different subscript phoneme in the phoneme sequence of its standard syllable, each subscript phoneme in the phoneme sequence of the standard syllable is recorded as the abnormal phoneme in each abnormal syllable.
[0014] Furthermore, the step of obtaining the abnormal word formation index of each abnormal syllable in the speech data of all children includes:
[0015] For the abnormal phonemes of abnormal syllables in the speech data of all children, the independent abnormal index of each abnormal phoneme was obtained according to the number of occurrences of normal syllables containing each abnormal phoneme and the number of occurrences of abnormal syllables containing each abnormal phoneme;
[0016] For the abnormal syllables in the speech data of all the children, the cumulative sum of the independent abnormal indicators of all the abnormal phonemes in each abnormal syllable is taken as the comprehensive abnormal value of each abnormal syllable;
[0017] According to the number of occurrences of each abnormal syllable in the speech data of all children and the comprehensive abnormal value, the word formation abnormality index of each abnormal syllable is obtained.
[0018] Furthermore, the obtaining of an independent abnormality indicator for each abnormal phoneme includes:
[0019] Selecting one abnormal phoneme from the speech data of all the children as an example phoneme, and selecting an abnormal target syllable and a normal target syllable of the example phoneme from the syllable sequence of the speech data of all the children;
[0020] The abnormal target syllable is an abnormal syllable, and the phoneme sequence of the standard syllable of the abnormal target syllable has an example phoneme; the normal target syllable is a normal syllable, and the phoneme sequence of the normal target syllable has an example phoneme;
[0021] The ratio of the number of occurrences of the deviant target syllables of the example phoneme to the number of occurrences of the normal target syllables was used as an independent deviant indicator of the example phoneme.
[0022] Furthermore, the step of screening out dysmorphic syllables from abnormal syllables in the speech data of all children includes:
[0023] According to the number of times the phoneme of each abnormal syllable in the speech data of all children appears alone in the speech data of each child, the comprehensive misreading index of each abnormal syllable in the speech data of all children is obtained;
[0024] According to the comprehensive misreading index and word formation abnormality index, obtaining the articulation disorder index of each abnormal syllable in the speech data of all children;
[0025] Among the abnormal syllables in the speech data of all children, a combined syllable of the example phonemes is selected; the example phonemes exist in the phoneme sequence of the standard syllable of the combined syllable; if the articulation disorder index of the combined syllable of the example phonemes is greater than a preset disorder threshold, the combined syllable of the example phonemes is used as a dysmorphic syllable.
[0026] Furthermore, the method of obtaining a comprehensive misreading index of each abnormal syllable in the speech data of all children includes:
[0027] A child is selected as the analyzed child, an abnormal phoneme is selected from the speech data of the analyzed child as the target phoneme, a mispronounced syllable of the target phoneme is selected from the syllable sequence of the speech data of the analyzed child, and the syllable sequence of the standard syllable of the mispronounced syllable contains an abnormal phoneme of the target phoneme; the number of occurrences of the mispronounced syllable of the target phoneme in the syllable sequence of the speech data of the analyzed child is recorded as the mispronunciation value of the target phoneme;
[0028] The cumulative sum of the misreading values of the same abnormal phoneme in the speech data of all children is used as the initial misreading index of each abnormal phoneme in the speech data of all children; the product of the initial misreading indexes of all phonemes in the phoneme sequence of each abnormal syllable in the speech data of all children is normalized to obtain the comprehensive misreading index of each abnormal syllable in the speech data of all children.
[0029] Furthermore, the method of obtaining the word formation abnormality index of each abnormal syllable according to the number of occurrences of each abnormal syllable in the speech data of all children and the comprehensive abnormal value includes:
[0030] The ratio of the number of occurrences of each abnormal syllable in the speech data of all the children to the comprehensive abnormal value is normalized to obtain the word formation abnormality index of each abnormal syllable in the speech data of all the children.
[0031] Furthermore, the method of obtaining the dysarthria index of each abnormal syllable in the speech data of all children according to the comprehensive misreading index and the abnormal word formation index includes:
[0032] The product of the comprehensive misreading index and the abnormal word formation index of each abnormal syllable in the speech data of all children is normalized to obtain the articulation disorder index of the corresponding abnormal syllable.
[0033] Furthermore, the method for performing speech segmentation on the analyzed speech is a hidden Markov model.
[0034] Furthermore, the preset obstacle threshold is 0.7.
[0035] The present invention has the following beneficial effects:
[0036] In the embodiment of the present invention, the guiding speech is the standard pronunciation, and the abnormal phonemes are determined based on the difference between the guiding signal and the speech data, so as to effectively avoid the incorrect alignment of the child's incorrect pronunciation with the standard phonemes; when the child with articulation disorder forms words, there may be a single phoneme that cannot be read accurately, resulting in syllable misreading, or there may be a simple phoneme that can be correctly spelled, but the phoneme combination increases the pronunciation difficulty, resulting in incorrect spelling of the combined syllables. The word formation abnormality index reflects the incorrect spelling when the phonemes are combined into syllables, and the number of times the phoneme of each abnormal syllable in the speech data of all children appears separately in the speech data of each child reflects the syllable misreading. The above two situations will lead to inaccurate pronunciation of the syllable, thereby reflecting the degree of articulation disorder. Through the phoneme error probability and the error probability of the phoneme combination, the pronunciation errors of syllables or words can be more comprehensively evaluated, the mutual influence between phonemes is taken into account, and it has higher fault tolerance and syllable-level modeling capabilities, so that the screened word formation disorder syllables are more accurate, the accuracy of judging the pronunciation errors of children's syllables is increased, and the effect of education and rehabilitation training for children with articulation disorder is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0038] Figure 1 A system structure diagram of a speech-interaction-based education and rehabilitation training system for children with articulation disorders provided by one embodiment of the present invention;
[0039] Figure 2 A structural diagram of a phoneme word formation anomaly module provided by an embodiment of the present invention;
[0040] Figure 3 A structural diagram of a word formation disorder analysis module provided by an embodiment of the present invention;
[0041] Figure 4 A schematic diagram of a computer device for an education and rehabilitation training device for children with articulation disorders based on voice interaction provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0042] In order to further explain the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following is a detailed description of the specific implementation method, structure, features and effects of a speech-interactive education and rehabilitation training system for children with articulation disorders proposed by the present invention in combination with the accompanying drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" does not necessarily refer to the same embodiment. In addition, specific features, structures or characteristics in one or more embodiments may be combined in any suitable form.
[0043] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0044] The following is a detailed description of a specific scheme of a speech-interaction-based education and rehabilitation training system for children with articulation disorders provided by the present invention in conjunction with the accompanying drawings.
[0045] Embodiment 1:
[0046] See also Figure 1, which shows a system block diagram of a speech-interaction-based education and rehabilitation training system for children with articulation disorders provided by an embodiment of the present invention. The system includes: a data acquisition module 110, an abnormal phoneme extraction module 120, a phoneme word formation abnormality module 130, and a word formation disorder analysis module 140.
[0047] The data collection module 110 is used to obtain the speech data of different children under the guided speech, and the speech data is composed of syllables.
[0048] Specifically, in a quiet environment without obvious noise, a digital audio player is used to play the pre-recorded guiding voice, and the parents or therapists guide the children to repeat the content of the guiding voice. A recording device is used to collect the voice signal of each child and record it as the voice data of each child under the guiding voice.
[0049] It should be noted that the guided speech should cover various pronunciation situations as much as possible so as to comprehensively evaluate the child's articulation disorder; the guided speech is the speech data of standard pronunciation, that is, the pronunciation of each word is correct; the speech data is composed of multiple syllables, one syllable corresponds to one word, and one syllable is composed of multiple phonemes.
[0050] The abnormal phoneme extraction module 120 is used to divide each speech data into normal syllables and abnormal syllables based on the difference between the guiding speech and the speech data of each child, and extract abnormal phonemes in the abnormal syllables.
[0051] The guided speech reflects the pronunciation of the child, and the guided speech is the standard pronunciation. By comparing the speech data of the child with the guided speech, the speech data is divided into normal syllables and abnormal syllables. The specific division method is as follows:
[0052] The speech data of all children and the guiding speech are recorded as analysis speech; the analysis speech is segmented, and the segmented syllables are arranged in order to obtain a syllable sequence; each syllable in the syllable sequence is decomposed into phonemes, and the decomposed phonemes are arranged in order to obtain the phoneme sequence of the corresponding syllable; for the speech data of all children, it is determined whether the phoneme sequence of each syllable in the syllable sequence of each speech data is equal to the phoneme sequence of the syllable with the same subscript in the syllable sequence of the guiding speech, if so, each syllable in the syllable sequence of each speech data is recorded as a normal syllable; if not, each syllable in the syllable sequence of each speech data is recorded as an abnormal syllable.
[0053] In an embodiment of the present invention, a Hidden Markov Model (HMM) is selected to segment the analyzed speech into different syllables, and the syllables are arranged in the order in which they appear in the analyzed speech to obtain a syllable sequence; a deep neural network is used to decompose the syllables into phonemes, and the phonemes are arranged in the order in which they appear in the syllables to obtain a phoneme sequence of the syllables. Among them, the speech segmentation of the speech signal by the Hidden Markov Model and the phoneme decomposition of the syllables by the deep neural network are both well-known techniques to those skilled in the art and will not be elaborated here.
[0054] The syllables with the same subscript in the syllable sequence of the speech data and the guiding speech represent the pronunciation of the same character. If the phoneme sequences of the syllables with the same subscript are equal, it indicates that the child pronounces the syllable correctly, and then the syllables with each subscript in the syllable sequence of the speech data are recorded as normal syllables; otherwise, it indicates that the child pronounces the syllable incorrectly, and they are recorded as abnormal syllables. It should be noted that two sequences are equal means that the lengths of the two sequences are equal and the elements with the same subscript are the same.
[0055] The method for extracting abnormal phonemes is as follows: select the standard syllables of each abnormal syllable from the syllable sequence of the guiding speech; each abnormal syllable has the same subscript as its standard syllable; if the phonemes with each subscript in the phoneme sequence of each abnormal syllable are different from those in the phoneme sequence of its standard syllable, then the phonemes with each subscript in the phoneme sequence of the standard syllable are recorded as the abnormal phonemes in each abnormal syllable.
[0056] The standard syllable of an abnormal syllable is the correct pronunciation of the abnormal syllable. As an example, if a patient misreads "guo" as "duo" when reading, then "duo" is the abnormal syllable and "guo" is the standard syllable; the phoneme sequence of the abnormal syllable "duo" is and the phoneme sequence of the standard syllable "guo" is Since the phoneme "g" is pronounced incorrectly, the phoneme "g" is recorded as the abnormal phoneme in the abnormal syllable "duo".
[0057] The phoneme word-formation abnormality module 130 is configured to obtain the word-formation abnormality index of each abnormal syllable in the speech data of all children according to the occurrence times of each abnormal syllable in the speech data of all children, and the occurrence times of the normal syllables containing the abnormal phonemes in each abnormal syllable and the occurrence times of the abnormal syllables containing the abnormal phonemes in each abnormal syllable.
[0058] When children with speech disorders form words, there may be problems with single phonemes that they do not recognize or cannot accurately pronounce, or there may be word-formation problems where single phonemes can be accurately pronounced but phonemes are misread when multiple phonemes are combined to form words.
[0059] For each abnormal syllable of the child, if the number of abnormal phonemes in the abnormal syllable appears less in the abnormal syllable of the child, and the number of abnormal phonemes appears more in the normal syllable, the child is more likely to read the abnormal phoneme accurately, the possibility of the abnormal syllable having phoneme problems is smaller, and the possibility of word formation problems is greater. The number of times each abnormal syllable appears in the child's speech data reflects the possibility of abnormal pronunciation after the phonemes constitute the abnormal syllable. Combining the above two factors, the word formation abnormality index of the abnormal syllable is obtained, which is used to measure the possibility of abnormal syllables causing pronunciation errors due to word formation of multiple phonemes.
[0060] See also Figure 2 , which shows a structural diagram of a phoneme word formation abnormality module provided by an embodiment of the present invention, and the phoneme word formation abnormality module includes: an independent abnormality analysis unit 131, a comprehensive abnormality analysis unit 132, and a phoneme word formation analysis unit 133.
[0061] The independent abnormality analysis unit 131 obtains an independent abnormality index of each abnormal phoneme for the abnormal phonemes of the abnormal syllables in the speech data of all children according to the number of occurrences of normal syllables containing each abnormal phoneme and the number of occurrences of the abnormal syllables containing each abnormal phoneme.
[0062] Preferably, in some possible implementation modes of the embodiments of the present invention, the method for obtaining an independent abnormality indicator includes: randomly selecting an abnormal phoneme in the speech data of all children as an example phoneme, and selecting an abnormal target syllable and a normal target syllable of the example phoneme from the syllable sequence of the speech data of all children; the abnormal target syllable is an abnormal syllable, and the example phoneme exists in the phoneme sequence of the standard syllable of the abnormal target syllable; the normal target syllable is a normal syllable, and the example phoneme exists in the phoneme sequence of the normal target syllable; and the ratio of the number of occurrences of the abnormal target syllable of the example phoneme to the number of occurrences of the normal target syllable is used as the independent abnormality indicator of the example phoneme.
[0063] For abnormal phonemes in the speech data of the children, if an abnormal phoneme appears less in abnormal syllables but more in normal syllables, the possibility that the children can accurately read the abnormal phoneme is greater, the possibility that the abnormal phoneme has phoneme problems is smaller, and the possibility that the abnormal phoneme has word formation problems when forming words with other phonemes is greater. At the same time, the abnormal target syllable of the example phoneme is the abnormal syllable containing the example phoneme, and the normal target syllable is the normal syllable containing the example phoneme.
[0064] Therefore, the number of occurrences of the abnormal target syllables of the example phoneme is positively correlated with the independent abnormal index, and the number of occurrences of the normal target syllables is negatively correlated with the independent abnormal index. The larger the independent abnormal index, the greater the possibility of pronunciation errors when the abnormal phoneme is combined with other phonemes to form words.
[0065] In the embodiment of the present invention, the correlation between the number of occurrences of abnormal target syllables of the example phonemes and the number of occurrences of normal target syllables and the independent abnormality index can also be constructed through other basic mathematical operations, which will not be limited or elaborated here.
[0066] It should be noted that, since the syllable containing the abnormal phoneme must contain the abnormal phoneme, the minimum value of the number of occurrences of the abnormal target syllable of the abnormal phoneme is a constant of 1, and the independent abnormal index is a constant greater than 0. The independent abnormal index of all abnormal phonemes and sample phonemes in the speech data of all children is obtained in the same way.
[0067] The comprehensive abnormality analysis unit 132 takes the cumulative sum of the independent abnormality indicators of all abnormal phonemes in each abnormal syllable as the comprehensive abnormality value of each abnormal syllable for the abnormal syllable in the speech data of all children.
[0068] Considering the possibility of incorrect pronunciation when all abnormal phonemes in the abnormal syllable are combined with other phonemes to form words, the possibility that the abnormal syllable is incorrectly pronounced due to phoneme word formation is obtained, and the comprehensive abnormal value is obtained.
[0069] It should be noted that since the independent anomaly index is a constant greater than zero, the comprehensive anomaly value is also a constant greater than zero.
[0070] The phoneme word formation analysis unit 133 obtains the word formation abnormality index of each abnormal syllable according to the number of occurrences and the comprehensive abnormality value of each abnormal syllable in the speech data of all children.
[0071] The number of occurrences of each abnormal syllable in the speech data of all children reflects the possibility of abnormal pronunciation after the phonemes form abnormal syllables; if the number of occurrences of each abnormal syllable in the speech data of all children and the comprehensive abnormality value are larger, the possibility that the abnormal syllable is a pronunciation error caused by phoneme word formation is greater. Therefore, the number of occurrences of each abnormal syllable in the speech data of all children is positively correlated with the word formation abnormality index, and the comprehensive abnormality value is negatively correlated with the word formation abnormality index.
[0072] In the embodiment of the present invention, the ratio of the number of occurrences of each abnormal syllable in the speech data of all children to the comprehensive abnormal value is normalized to obtain the word formation abnormality index of each abnormal syllable in the speech data of all children. The abnormal syllable with a larger word formation abnormality index is more likely to be mispronounced due to word formation of multiple phonemes.
[0073] In the embodiment of the present invention, other basic mathematical operations can also be used to construct the correlation between the number of occurrences of each abnormal syllable in the speech data of all children and the comprehensive abnormal value and the abnormal word formation index, which is not limited or elaborated here.
[0074] It should be noted that in the embodiment of the present invention, the Sigmoid function is used for normalization processing, and other normalization methods may also be selected, such as function conversion, maximum and minimum normalization, etc., which are not limited here. The normalization method of other positions in this scheme is the same as the normalization method of this unit.
[0075] The word formation disorder analysis module 140 is used to screen word formation disorder syllables from the abnormal syllables in the speech data of all children according to the number of times the phoneme of each abnormal syllable appears separately in the speech data of each child and the word formation disorder index of each abnormal syllable, and provide articulation disorder education and rehabilitation training for the children based on the word formation disorder syllables.
[0076] The number of times the phoneme of each abnormal syllable appears alone in the speech data of each child reflects the possibility that the child misreads a single phoneme; the abnormal word formation index reflects the possibility that the abnormal syllable is mispronounced due to phoneme word formation. Both situations can lead to incorrect pronunciation of abnormal syllables and are used to measure the child's word formation disorder.
[0077] See also Figure 3 , which shows a structural diagram of a word formation disorder analysis module provided by an embodiment of the present invention, the word formation disorder analysis module includes: a syllable misreading analysis unit 141, an articulation disorder analysis unit 142, and a word formation disorder syllable screening unit 143.
[0078] The syllable mispronunciation analysis unit 141 is used to obtain a comprehensive mispronunciation index of each abnormal syllable in the speech data of all children according to the number of times the phoneme of each abnormal syllable in the speech data of each child appears alone.
[0079] In a specific implementation of the embodiment of the present invention, the comprehensive misreading index is expressed by a formula: any one of the patients is recorded as the analyzed patient, any one of the abnormal phonemes in the speech data of the analyzed patient is recorded as the target phoneme, the misread syllable of the target phoneme is selected from the syllable sequence of the speech data of the analyzed patient, and there is an abnormal phoneme of the target phoneme in the syllable sequence of the standard syllable of the misread syllable; the number of occurrences of the misread syllable of the target phoneme in the syllable sequence of the speech data of the analyzed child is recorded as the misreading value of the target phoneme; the cumulative sum of the misreading values of the same abnormal phoneme in the speech data of all the patients is used as the initial misreading index of each abnormal phoneme in the speech data of all the children; the product of the initial misreading indexes of all phonemes in the phoneme sequence of each abnormal syllable in the speech data of all the children is normalized to obtain the comprehensive misreading index of each abnormal syllable in the speech data of all the children.
[0080] If there is only one abnormal phoneme in the abnormal syllable, and the abnormal syllable is mispronounced due to the misreading of the abnormal phoneme, then the abnormal syllable is the misread syllable of the abnormal phoneme. The misreading value of the target phoneme reflects the possibility of the target phoneme being misread in the speech data of each child. In order to improve the universality of the misreading of the target phoneme, the pronunciation of the target phoneme by all children is considered to obtain the initial misreading index of each abnormal phoneme. The larger the initial misreading index, the more likely it is that the abnormal phoneme will be misread by the child with articulation disorder.
[0081] In order to analyze the possibility of mispronunciation of the phonemes in the abnormal syllables, the initial mispronunciation indexes of all the phonemes in the phoneme sequence of each abnormal syllable are multiplied to obtain the comprehensive mispronunciation index. The larger the comprehensive mispronunciation index, the greater the possibility of mispronunciation of the abnormal syllable.
[0082] As an example, if the number of times the phonemes "g" and "e" are mispronounced is greater, the possibility that the syllable "ge" is mispronounced is greater. It should be noted that the initial mispronunciation index of the non-abnormal phonemes in the phoneme sequence of the abnormal syllable is set to zero.
[0083] The articulation disorder analysis unit 142 is used to obtain the articulation disorder index of each abnormal syllable in the speech data of all children according to the comprehensive misreading index and the abnormal word formation index.
[0084] When children with articulation disorders form words, they may not recognize or accurately read a single phoneme, which may lead to syllable misreading. Alternatively, they may be able to correctly spell a single phoneme or a simple phoneme. The combination of phonemes increases the difficulty of pronunciation, which may lead to incorrect spelling of the combined syllables. The comprehensive misreading index reflects the incorrect spelling caused by the lack of recognition of phonemes. The abnormal word formation index reflects the incorrect spelling caused by the combination of phonemes. The two situations are combined to analyze the syllable misreading and are used to analyze articulation disorders.
[0085] The product of the comprehensive misreading index and the word formation abnormality index of each abnormal syllable in the speech data of all children was normalized to obtain the articulation disorder index of the corresponding abnormal syllable. The larger the articulation disorder index, the greater the possibility of misreading of the abnormal syllable, which makes the articulation disorder greater.
[0086] The dysmorphic syllable screening unit 143 is used to select a combination syllable of example phonemes from the abnormal syllables in the speech data of all children; the example phonemes exist in the phoneme sequence of the standard syllable of the combination syllable; if the articulation disorder index of the combination syllable of the example phonemes is greater than the preset disorder threshold, the combination syllable of the example phonemes is used as a dysmorphic syllable.
[0087] The sample syllable participates in the combination of phonemes in its combined syllables; the greater the dysarthria index, the more obvious the degree of pronunciation error of the abnormal syllable, that is, the degree of dysarthria. If the dysarthria index of the combined syllables of the sample phonemes is greater than the preset obstacle threshold, it means that the child with dysarthria has made an error in pronouncing the sample phonemes, the pronunciation error of the combined syllables of the sample phonemes is more serious, and the child has dysarthria in the combined syllables of the sample phonemes; parents or therapists teach the pronunciation of the combined syllables of the sample phonemes in a targeted manner, guide the dysarthric children to repeat the pronunciation of these syllables many times, and parents correct the location of the pronunciation errors.
[0088] It should be noted that the preset obstacle threshold in the embodiment of the present invention is an empirical value of 0.7, and the implementer can set it according to the specific situation. The method for judging whether the abnormal syllables and the example syllables of the speech data of all children are dysmorphic syllables is the same.
[0089] So far, the present invention is completed.
[0090] Embodiment 2:
[0091] Figure 4 A computer device schematic diagram of a speech-interactive education and rehabilitation training device for children with articulation disorders provided by an embodiment of the present invention. Figure 4 As shown, the computer device includes: a memory 201, a processor 202, and a computer program 203 stored in the memory 201 and running on the processor 202, wherein when the processor 202 executes the computer program 203, the computer device can execute any one of the speech interaction-based education and rehabilitation training systems for children with articulation disorders introduced above.
[0092] In addition, an embodiment of the present application also protects a device, which may include a memory and a processor, wherein the memory stores executable program code, and the processor is used to call and execute the executable program code to execute a voice interaction-based education and rehabilitation training system for children with articulation disorders provided in an embodiment of the present application.
[0093] In this embodiment, the functional modules of the device can be divided according to the above method example. For example, each functional module can be corresponded, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware. It should be noted that the division of modules in this embodiment is schematic and is only a logical function division. There may be other division methods in actual implementation.
[0094] It should be understood that the device provided in this embodiment is used to execute the above-mentioned education and rehabilitation training system for children with articulation disorders based on voice interaction, and thus can achieve the same effect as the above-mentioned implementation method.
[0095] In the case of an integrated unit, the device may include a processing module and a storage module. When the device is applied to a device, the processing module may be used to control and manage the actions of the device. The storage module may be used to support the device to execute mutual program codes, etc.
[0096] The processing module may be a processor or a controller, which may implement or execute various exemplary logic blocks, modules and circuits disclosed in the present application. The processor may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of digital signal processing (DSP) and a microprocessor, etc. The storage module may be a memory.
[0097] Embodiment 3:
[0098] This embodiment also provides a computer-readable storage medium, which stores computer program code. When the computer program code runs on a computer, the computer executes the above-mentioned related method steps to implement a speech interaction-based education and rehabilitation training system for children with articulation disorders provided in the above embodiment.
[0099] Among them, the device and computer-readable storage medium provided in this embodiment are used to execute the corresponding system provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding system provided above, and will not be repeated here.
[0100] It should be noted that the sequence of the above embodiments of the present invention is only for description and does not represent the advantages and disadvantages of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0101] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from other embodiments.
[0102] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A speech-interaction-based education and rehabilitation training system for children with articulation disorders, characterized in that: The system includes: The data collection module is used to obtain the speech data of different children under guided speech, and the speech data is composed of syllables; An abnormal phoneme extraction module is used to divide each speech data into normal syllables and abnormal syllables based on the difference between the guiding speech and the speech data of each patient, and extract abnormal phonemes in the abnormal syllables; The phoneme word formation abnormality module is used to obtain the word formation abnormality index of each abnormal syllable in the speech data of all children according to the number of occurrences of each abnormal syllable in the speech data of all children, the number of occurrences of normal syllables containing the abnormal phonemes in each abnormal syllable, and the number of occurrences of abnormal syllables containing the abnormal phonemes in each abnormal syllable; The word formation disorder analysis module is used to screen out word formation disorder syllables from the abnormal syllables in the speech data of all children according to the number of times the phoneme of each abnormal syllable appears separately in the speech data of each child and the word formation disorder index of each abnormal syllable, and provide articulation disorder education and rehabilitation training to the children based on the word formation disorder syllables.
2. The speech interaction-based education and rehabilitation training system for children with dysarthria according to claim 1, characterized in that: The method of dividing each speech data into normal syllables and abnormal syllables based on the difference between the guiding speech and the speech data of each patient, and extracting abnormal phonemes in the abnormal syllables, comprises: The speech data of all the children and the guided speech are recorded as the analysis speech; the analysis speech is segmented, and the segmented syllables are arranged in order to obtain a syllable sequence; each syllable in the syllable sequence is decomposed into phonemes, and the decomposed phonemes are arranged in order to obtain a phoneme sequence of the corresponding syllable; For the speech data of all children, determine whether each syllable in the syllable sequence of each speech data is equal to the phoneme sequence of the syllable with the same subscript in the syllable sequence of the guide speech, if so, record each syllable in the syllable sequence of each speech data as a normal syllable; if not, record each syllable in the syllable sequence of each speech data as an abnormal syllable; A standard syllable of each abnormal syllable is selected from the syllable sequence of the guiding speech; each abnormal syllable has the same subscript as its standard syllable; if each abnormal syllable has a different subscript phoneme in the phoneme sequence of its standard syllable, each subscript phoneme in the phoneme sequence of the standard syllable is recorded as the abnormal phoneme in each abnormal syllable.
3. The speech interaction-based education and rehabilitation training system for children with articulation disorders according to claim 2, characterized in that: The step of obtaining the abnormal word formation index of each abnormal syllable in the speech data of all the children includes: For the abnormal phonemes of abnormal syllables in the speech data of all children, the independent abnormal index of each abnormal phoneme was obtained according to the number of occurrences of normal syllables containing each abnormal phoneme and the number of occurrences of abnormal syllables containing each abnormal phoneme; For the abnormal syllables in the speech data of all the children, the cumulative sum of the independent abnormal indicators of all the abnormal phonemes in each abnormal syllable is taken as the comprehensive abnormal value of each abnormal syllable; According to the number of occurrences of each abnormal syllable in the speech data of all children and the comprehensive abnormal value, the word formation abnormality index of each abnormal syllable is obtained.
4. The speech interaction-based education and rehabilitation training system for children with articulation disorders according to claim 3, characterized in that: The step of obtaining an independent abnormality indicator for each abnormal phoneme includes: Selecting one abnormal phoneme from the speech data of all the children as an example phoneme, and selecting an abnormal target syllable and a normal target syllable of the example phoneme from the syllable sequence of the speech data of all the children; The abnormal target syllable is an abnormal syllable, and the phoneme sequence of the standard syllable of the abnormal target syllable has an example phoneme; the normal target syllable is a normal syllable, and the phoneme sequence of the normal target syllable has an example phoneme; The ratio of the number of occurrences of the deviant target syllables of the example phoneme to the number of occurrences of the normal target syllables was used as an independent deviant indicator of the example phoneme.
5. The speech interaction-based education and rehabilitation training system for children with articulation disorders according to claim 4, characterized in that: The method of screening dysmorphic syllables from abnormal syllables in the speech data of all children includes: According to the number of times the phoneme of each abnormal syllable in the speech data of all children appears alone in the speech data of each child, the comprehensive misreading index of each abnormal syllable in the speech data of all children is obtained; According to the comprehensive misreading index and word formation abnormality index, obtaining the articulation disorder index of each abnormal syllable in the speech data of all children; Among the abnormal syllables in the speech data of all children, a combined syllable of the example phonemes is selected; the example phonemes exist in the phoneme sequence of the standard syllable of the combined syllable; if the articulation disorder index of the combined syllable of the example phonemes is greater than a preset disorder threshold, the combined syllable of the example phonemes is used as a dysmorphic syllable.
6. The speech interaction-based education and rehabilitation training system for children with dysarthria according to claim 5, characterized in that: The method of obtaining the comprehensive misreading index of each abnormal syllable in the speech data of all children includes: A child is selected as the analyzed child, an abnormal phoneme is selected from the speech data of the analyzed child as the target phoneme, a mispronounced syllable of the target phoneme is selected from the syllable sequence of the speech data of the analyzed child, and the syllable sequence of the standard syllable of the mispronounced syllable contains an abnormal phoneme of the target phoneme; the number of occurrences of the mispronounced syllable of the target phoneme in the syllable sequence of the speech data of the analyzed child is recorded as the mispronunciation value of the target phoneme; The cumulative sum of the misreading values of the same abnormal phoneme in the speech data of all children is used as the initial misreading index of each abnormal phoneme in the speech data of all children; the product of the initial misreading indexes of all phonemes in the phoneme sequence of each abnormal syllable in the speech data of all children is normalized to obtain the comprehensive misreading index of each abnormal syllable in the speech data of all children.
7. The speech interaction-based education and rehabilitation training system for children with articulation disorders according to claim 3, characterized in that: The method of obtaining the word formation abnormality index of each abnormal syllable according to the number of occurrences of each abnormal syllable in the speech data of all children and the comprehensive abnormal value includes: The ratio of the number of occurrences of each abnormal syllable in the speech data of all the children to the comprehensive abnormal value is normalized to obtain the word formation abnormality index of each abnormal syllable in the speech data of all the children.
8. The speech interaction-based education and rehabilitation training system for children with articulation disorders according to claim 5, characterized in that: The method of obtaining the dysarthria index of each abnormal syllable in the speech data of all children according to the comprehensive misreading index and the abnormal word formation index includes: The product of the comprehensive misreading index and the abnormal word formation index of each abnormal syllable in the speech data of all children is normalized to obtain the articulation disorder index of the corresponding abnormal syllable.
9. The speech interaction-based education and rehabilitation training system for children with articulation disorders according to claim 2, characterized in that: Hidden Markov model is used to perform speech segmentation on the analyzed speech.
10. The speech interaction-based education and rehabilitation training system for children with dysarthria according to claim 5, characterized in that: The preset obstacle threshold is 0.7.
Citation Information
Patent Citations
Self-sensing error tone pronunciation learning method and system
CN101661675A
Intelligent child language disorder correction treatment robot
CN114093206A