Speech training system for persons with speech disorders, speech training method for persons with speech disorders, communication support system for persons with speech disorders, communication support method for persons with speech disorders, analysis system for persons with speech disorders, analysis method for persons with speech disorders, program and recording medium
The speech training system uses speech rate conversion to allow individuals with speech disorders to practice articulation movements at their own pace, addressing location and time constraints, and enhances articulation and communication skills through independent training and feedback.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- THE UNIV OF TOKYO
- Filing Date
- 2024-11-12
- Publication Date
- 2026-05-22
AI Technical Summary
Current speech therapy for speech disorders, especially in children, is limited by location and time constraints, and patients struggle to adjust speech speed to their articulation ability, receive appropriate feedback, and feel the training effect independently.
A speech training system using speech rate conversion technology to present model audio at a slower speed, allowing individuals to imitate and practice articulation movements at their own pace, with feedback provided through high-speed conversion to the original speed.
Enables independent and continuous speech training at home, improving articulation accuracy and communication ability, and facilitates easy analysis of speech disorders through objective evaluation.
Smart Images

Figure 2026084736000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a speech training system for speech disorder patients, a speech training method for speech disorder patients, a communication support system for speech disorder patients, a communication support method for speech disorder patients, an analysis system for speech disorder patients, an analysis method for speech disorder patients, a program, and a recording medium, and is suitable for application to the training (rehabilitation) of speech disorder patients, communication support with others, or analysis of the degree of speech disorder in speech disorder patients.
Background Art
[0002] Speech disorder is a disorder that makes it difficult to speak due to the dysfunction of organs (such as the mouth, tongue, and throat) that produce sounds and words due to aging, illness, sequelae of cerebrovascular disorders, injuries, etc. When the muscles around the mouth become weak or the strength of the tongue becomes weak, it may be difficult to pronounce clearly and speak. Also, when the strength related to breathing and the vocal cords in the throat becomes weak, a voice disorder may occur simultaneously, making the voice thin and difficult to hear.
[0003] Speech therapy for speech disorder patients is handled by speech therapists (ST) at hospital facilities. However, while continuous long-term training is necessary even after discharge and during outpatient visits, the response to speech disorder patients, especially speech disorder children, has not been fully established. Current training mainly focuses on the patient repeating the pronunciation made by the ST. Since this training is basically carried out one-on-one in the hospital, there are restrictions on location and time. For training, it is necessary to adjust the speaking speed according to the patient's articulation ability, appropriately feedback the current articulation state to the patient, and let the patient feel the training effect, but it was difficult for the patient to do these alone. Especially in patients with slow movement even though the movement trajectories of the speech organs (such as the tongue and facial muscles) are correct, the pronunciation becomes slow, so there was a problem that with ordinary recording and playback feedback, the movement trajectory becomes ambiguous when trying to speak quickly and is repeated.
[0004] A known articulation disorder detection device comprises: a first line generation unit that generates a first line by averaging audio data obtained by having a subject repeatedly pronounce a voice module containing voiced plosives using a first window length set to be less than or equal to the standard pronunciation time of the voiced plosive; a second line generation unit that generates a second line by averaging the audio data using a second window length set to be greater than or equal to the standard time of the voice module and less than or equal to twice the standard time; an interval detection unit that detects intervals in which the value of the first line is greater than the value of the second line multiplied by a predetermined positive real number; and a determination unit that determines articulation disorder based on the interval detection unit (see Patent Document 1).
[0005] Changes in articulation movements associated with adjusting speech rate in healthy individuals have been reported (see Non-Patent Document 1). Non-Patent Document 1 describes that slower speech rates make it easier to achieve accurate movements, that slower speech rates improve clarity, that reducing speech rates allows for the use of a sufficient range of motion to move the articulatory organs, and that the tongue can reach the target articulation point more accurately.
[0006] The temporal changes in dysarthria associated with amyotrophic lateral sclerosis (ALS) have been reported (see Non-Patent Documents 2 and 3). Non-Patent Document 2 describes that formant transition rate decreases in ALS patients and that there is a strong correlation between decreased formant transition rate and speech clarity. Non-Patent Document 3 describes that articulation speed decreases due to muscle weakness of the articulatory organs as the disease progresses, and that syllable repetition speed decreases from an early stage.
[0007] The development of televisions and radios equipped with speech rate conversion technology has been reported (see Non-Patent Literature 4). In Non-Patent Literature 4, a note at the bottom of page 3 states that (in the case of slowing down by general speech rate conversion) "vowels become drawn out, and the boundaries between sounds become unclear. This is described as resembling the speech of an intoxicated person." This indicates that the formant transition speed that constitutes the boundary between a phoneme (the smallest unit of "sound" that can distinguish the meanings of two different words) and the phoneme is reduced and slurred.
[0008] A language training device is known that has a language training function that allows aphasic patients to practice language using pictures, sounds, and letters, and is used by being built into a dedicated computer (see Patent Documents 2 and 3). The display unit consists of a touch panel, and images for operating the language training function are displayed on the display unit. [Prior art documents] [Patent Documents]
[0009] [Patent Document 1] Patent No. 4721941 [Patent Document 2] Design Registration No. 1349389 Gazette [Patent Document 3] Design Registration No. 1349390 Gazette [Non-patent literature]
[0010] [Non-Patent Document 1] Miho Uchiyama, Yuri Fujiwara, Chieko Kojima: Changes in articulatory movements associated with adjustment of speech rate in healthy individuals. Speech and Language Medicine 57:382-390, 2016. [Non-Patent Document 2] Masaki Nishio, Seiji Niimi: Time-series changes in dysarthria associated with amyotrophic lateral sclerosis - Part 1: Examination of changes in articulation function. Speech and Language Medicine 39:410-420, 1998. [Non-Patent Document 3] Masaki Nishio, Seiji Niimi: Time-series changes in dysarthria associated with amyotrophic lateral sclerosis - Part 2: A study mainly on changes in speech rate and syllable repetition rate. Speech and Language Medicine 40: 8-16, 1999. [Non-Patent Document 4] Tomono Miki, Atsushi Tsukita, Yaichi Aoshima: Development of radio and television equipped with speech rate conversion technology, NHK Science & Technology Research Laboratories, NHK Engineering Services, and Victor Company of Japan, April 2010, Hitotsubashi University GCOE Program "Innovation in Japanese Companies - Educational and Research Center for Empirical Management Studies" Okouchi Prize Case Study Project [Non-Patent Document 5] Peter B. Denes and Elliot N. Pinson,The Speech Chain: The Physics and Biology of Spoken Language,1963, PICKLE PAERTNERS PUBLISHING [Overview of the project] [Problems that the invention aims to solve]
[0011] As mentioned above, while continuous training is necessary for individuals with speech disorders even after discharge and during outpatient visits, support for individuals with speech disorders, especially children, is not yet well-established. Currently, training is basically conducted one-on-one within the hospital, which limits the location and time. Furthermore, training requires adjusting the speech speed to match the patient's articulation ability, providing appropriate feedback to the patient on their current articulation state, and allowing the patient to feel the effects of the training, all of which have been difficult for patients to do on their own.
[0012] Therefore, the problem that this invention aims to solve is to provide a speech training system for speech disorders, including children with disabilities, that allows speech disorders to easily and continuously continue their training independently at home or elsewhere, without restrictions on time or place, even after discharge from the hospital or during outpatient visits. The invention also provides a speech training method for speech disorders, a program for the speech training method for speech disorders, and a recording medium on which this program is stored.
[0013] Another problem that this invention aims to solve is to provide a language training system for people with speech disorders, including children with disabilities, a language training method for people with speech disorders, a program for the language training method for people with speech disorders, and a recording medium on which this program is recorded. This system supports training that slows down not only the rhythm of speech but also the speed of the articulation movement itself, as slow speech tailored to the articulation ability of the person with a speech disorder, in one-on-one training with a doctor or speech-language pathologist.
[0014] Another problem that this invention aims to solve is to provide a communication support system for persons with speech impairments, including children with disabilities, that can present speech that is easy for listeners to understand, even in cases where the speech impairment is severe and unlikely to recover. This system also provides a communication support method for persons with speech impairments, a program for the communication support method for persons with speech impairments, and a recording medium on which this program is stored.
[0015] Another problem that this invention aims to solve is to provide a speech disorder analysis system, a speech disorder analysis method, a program for the speech disorder analysis method, and a recording medium on which this program is stored, which can easily analyze the disability status of speech disorders in children with speech disorders. [Means for solving the problem]
[0016] To solve the above problems, this invention provides: Using speech rate conversion technology, a model audio file converted to a slower speed is presented to individuals with speech disorders who are the target of language training support. The speech produced by the person with the speech impairment who imitates the slow-speed converted example speech is then converted at high speed using speech rate conversion technology to return it to the same speed as the original example speech before the slow-speed conversion. This is a speech training support system for persons with speech impairments, configured to present the above-mentioned high-speed converted audio to the person with the speech impairment or to the listener.
[0017] Speech rate conversion technology is a technique that allows for the free adjustment of speech speed while maintaining the fundamental frequency and frequency spectrum of speech. Generally, slowing down speech speed through speech rate conversion also slows down the formant transition speed, resulting in a slurred voice. However, by treating this as a reproduction of slow articulation movements and imitating them, training that slows down articulation movements can be achieved.
[0018] The reference voice can be created by several methods, and the creation method is selected as needed. The reference voice is created, for example, by a speech therapist (ST) or another person (typically a person who has received predetermined training) recording the voice using the recording function of a speech training support system for a person with speech disorder, or by registering a pre-prepared voice file in the speech training support system for a person with speech disorder. This voice file is created by a speech therapist or another person recording it, or by editing the voice created by speech synthesis. Typically, the reference voice and the characters to be displayed on the display corresponding to this reference voice are registered in advance as a set in the speech training support system. Then, when the reference voice is played back and when the person with speech disorder speaks, the position of the speech is presented as an animation on the display. Typically, it is configured to present the waveform of the reference voice converted to a low speed and the characters superimposed on the waveform to the person with speech disorder. The speech training support system for a person with speech disorder can arbitrarily set the magnification of the speech speed conversion based on the reference voice. Alternatively, the number of moras per unit time of the reference voice can be automatically recognized, and the speech speed can be set by arbitrarily setting the number of moras per unit time during playback.
[0019] The reference voice is the voice uttered by a speech therapist (ST) or another person at a normal speed for the training of a person with speech disorder. This voice is recorded, low-speed converted using speech speed conversion technology, and then presented to the person with speech disorder. Since it is difficult for a person with speech disorder to speak at the same speed as an ordinary healthy person, the voice is low-speed converted and then presented to the person with speech disorder. The speech speed conversion magnification at this time is selected as needed, but typically it is the fastest speed at which the person with speech disorder can speak with reasonable accuracy and correct articulation. The speed of the low-speed converted reference voice is 1 / 2 times or more and 1 times or less the speed of the reference voice before the above-mentioned low-speed conversion. Here, when the speech speed conversion magnification is 1 time, it indicates that no speech speed conversion is performed. The low-speed converted reference voice is recorded and played from the speaker towards the person with speech disorder, or is played into the ear of the person with speech disorder from headphones or earphones worn by the person with speech disorder.
[0020] By imitating the sample voice slowly converted by a person with speech disorder and speaking slowly, the person with speech disorder can conduct training that focuses on the accuracy of articulation movements rather than the speed of articulation movements, and can utter the sample voice with slow but accurate articulation movements. Then, the voice slowly spoken by the person with speech disorder is converted at high speed using speech rate conversion technology and restored to the same speed as the sample voice before the slow conversion, and the voice converted at high speed is presented to, or in other words, fed back to the person with speech disorder himself / herself or the listener. By doing so, the person with speech disorder can listen to the sample voice at the normal speed, so that the improvement of articulation movements can be achieved and the effect of training can be obtained. In addition, by listening to the voice converted at high speed, the person with speech disorder can speak slowly while imitating the sample voice slowly converted with confidence.
[0021] A speech disorder training system for persons with speech disorder typically includes a speech rate conversion device, and a speaker, headphones or earphones for presenting the sample voice slowly converted by this speech rate conversion device to the person with speech disorder and presenting the voice converted at high speed by this speech rate conversion device to the person with speech disorder himself / herself or the listener.
[0022] The types of disorders of persons with speech disorder include organic speech disorder, dyskinetic speech disorder, auditory speech disorder, functional speech disorder, etc., and persons with speech disorder have one or more of these disorders. Organic speech disorder refers to a state in which there are abnormalities in the form of organs such as the lips and tongue, which are organs for pronunciation, and proper pronunciation cannot be made, and it is seen when there are abnormalities in these organs from birth, such as after oral cancer surgery, trauma, or cleft palate. Dyskinetic speech disorder refers to a state in which, due to diseases of the brain or nerves, the commands to move muscles such as the lips and tongue properly when pronouncing do not work properly, resulting in pronunciation disorders, and it is seen in brain injuries caused by stroke or traffic accidents, and neurological diseases such as Parkinson's disease and cerebral palsy. Auditory speech disorder refers to a state in which, due to so-called hearing impairment, the correct pronunciation as a model and one's own pronunciation cannot be heard, so that correct pronunciation cannot be learned and pronunciation disorders occur. Functional speech disorder refers to a state in which, although there is no such cause as described above, there are sounds that cannot be pronounced correctly.
[0023] This speech therapy support system for individuals with speech disorders is built using a computer with a predetermined program. The computer can be a desktop computer, laptop computer, tablet computer, smartphone, or microcontroller-embedded device, and the appropriate type will be selected as needed.
[0024] Furthermore, this invention, The stage involves presenting a model audio, converted to a slower speed using speech rate conversion technology, to individuals with articulation disorders who are the target of language training support, and The process involves the following steps: first, the person with the speech impairment imitates the slow-speed converted example speech and speaks at a slow speed; then, using speech speed conversion technology, the speech is converted at a high speed to return it to the same speed as the original example speech before the slow-speed conversion; The stage of presenting the above-mentioned high-speed converted audio to the person with the speech impairment or to the listener, This is a method of supporting speech training for individuals with speech disorders.
[0025] The process from presenting a model voice converted to a slower speed using speech rate conversion technology to a person with a speech disorder who is the target of language training support can be rephrased in accordance with the so-called Speech Chain (see Non-Patent Document 5) as follows: First, the speech-language pathologist speaks the model voice at a normal speaking speed. Using speech rate conversion technology, the model voice is slowed down and presented to the person with a speech disorder at a slow speaking speed. This voice reaches the ears of the person with a speech disorder, and motor commands reach the articulatory organs via the brain and nerves. The person with a speech disorder speaks at a slower speed appropriate to their impaired function, using correct articulation. Next, the process from listening to a feedback voice in which the person with a speech disorder imitates the slower-converted model voice spoken at a slower speed, converted back to the same speed as the original model voice using speech rate conversion technology, until learning occurs, can also be rephrased in accordance with the Speech Chain as follows: The speaking speed of the person with a speech disorder's own voice is increased using speech rate conversion technology. The person with speech impairment is presented with their own accelerated speech. The accelerated speech reaches the person's ears and sounds like normal speech to them. The sound that reaches the ears travels to the brain and nerves. The person with speech impairment feels that they are speaking correctly. The sound that reaches the ears travels to the articulatory organs via the brain and nerves. The person with speech impairment learns that they were able to speak correctly at a speed appropriate to their impaired function.
[0026] With regard to this invention of a method for supporting speech training for persons with speech disorders, the above-described explanation in relation to the speech training system for persons with speech disorders is valid.
[0027] A method for supporting speech therapy for individuals with speech disorders can be easily implemented using a computer with a predetermined program. The computer is as described above in relation to the speech therapy system for individuals with speech disorders.
[0028] Furthermore, this invention, This communication support system for individuals with speech disorders is configured to present to the listener an audiographer at a speed equivalent to the original speech, using speech rate conversion technology. The audiographer speaks slowly and carefully, focusing solely on the accuracy of their articulation, and the audiographer then focuses on the accuracy of their articulation movements.
[0029] A typical example is a communication support system. A microphone for recording the voices of people with speech impairments, A wearable unit equipped with a voice conversion device, to be worn by people with speech impairments, It has, The system works by having a person with a speech impediment wear a wearable unit, recording their voice via a microphone, and then converting and outputting the voice using a speech converter.
[0030] In another example, a communication support system is: Using a smartphone with a recording function and a speaker capable of converting speech that can be carried by a person with a speech impediment, This system converts the speech of a person with a speech impediment into audio and outputs it through the speaker.
[0031] Furthermore, this invention, This communication support method for individuals with speech disorders involves a stage in which the individual, who is the target of communication support, focuses solely on the accuracy of their articulation movements and speaks slowly and carefully at a speed slower than normal for them. This speech is then converted at a high speed using speech rate conversion technology to return it to the same speed as the original speech before the conversion, and presented to the listener.
[0032] Communication support methods for individuals with speech disorders can be easily implemented using a computer with a predetermined program. The computer is as described above in relation to the speech training system for individuals with speech disorders.
[0033] Furthermore, this invention, For individuals with speech disorders, we non-invasively analyze the movements of the tongue and facial muscles, which are essential for vocalization and pronunciation, and obtain movement analysis data. This is an analysis system for individuals with speech disorders, configured to analyze the degree of speech disorder by correlating the motor analysis data with the evaluation results of a speech-language pathologist regarding the aforementioned speech disorder.
[0034] Tongue movement analysis is typically performed by examining the anterior-posterior, lateral, up-and-down movements of the tongue. Facial muscle movement analysis is performed using subjective evaluation methods such as the Yanagihara method (40-point system), House-Brackmann method, and Sunnybrook method, but objective evaluation methods can also be used. The Yanagihara method evaluates 10 items of 9 facial movements, taking into account resting asymmetry and each branch of the facial nerve, on a 3-point scale, and the total score is used for evaluation. The House-Brackmann method comprehensively assesses facial movements of the entire face on a 6-point scale and was designed for the purpose of evaluating paralysis after acoustic neuroma surgery, taking into account sequelae after paralysis recovery, such as pathological synkinesis and facial contracture. The Sunnybrook method consists of three elements: recovery of voluntary movement, resting asymmetry, and pathological synkinesis. It evaluates with a composite score obtained by subtracting the scores for resting asymmetry and pathological synkinesis from the score for voluntary movement recovery, and fully considers sequelae after paralysis recovery. Objective evaluation methods include the marker method, which involves attaching markers to major facial features and determining the amount of marker movement; the method of identifying and extracting characteristic points and shapes such as the mouth circumference, eyelids, and eyebrows to determine the amount of change; the moiré method, which projects moiré patterns onto the face and processes the moiré patterns during facial movement; the subtraction method, which binarizes images of rest and maximum movement from video images and subtracts the pixel values of the rest and maximum movement images; and the method of measuring the three-dimensional shape using a laser rangefinder.
[0035] Furthermore, this invention, The process involves non-invasively analyzing the movements of the tongue and facial muscles, which are essential for vocalization and pronunciation, in individuals with speech disorders, and obtaining movement analysis data. The process involves analyzing the degree of articulation impairment in the person with articulation impairment by correlating the motor analysis data with the evaluation results of the speech-language pathologist regarding the person with articulation impairment, This is an analysis method for individuals with speech disorders.
[0036] The analysis method for individuals with speech disorders can be easily performed using a computer with a predetermined program. The computer is as described above in relation to the training system for individuals with speech disorders. [Effects of the Invention]
[0037] The inventions for a speech training system and a speech training method for persons with speech disorders allow for easy speech training at home without the need to go to a hospital. In addition, in one-on-one training sessions with doctors or speech-language pathologists for persons with speech disorders, including children with disabilities, it is possible to support training that slows down not only the rhythm of pronunciation but also the speed of the articulation movement itself, as slow speech tailored to the articulation ability of the person with a speech disorder. Furthermore, the inventions for a communication support system and a communication support method for persons with speech disorders can improve the communication ability of persons with speech disorders even if they still have a disability after receiving speech training. Moreover, the inventions for an analysis system and a method for analyzing persons with speech disorders allow for easy and objective analysis of the disability of persons with speech disorders by correlating the results of analysis of the movement of the tongue and facial muscles with the results of evaluation by a speech-language pathologist. [Brief explanation of the drawing]
[0038] [Figure 1] This is a schematic diagram showing the processing flow of a speech training system for persons with speech disorders according to the first embodiment of this invention. [Figure 2] This is a schematic diagram illustrating an example of a speech training system for persons with articulation disorders according to the first embodiment of this invention, constructed using a tablet computer. [Figure 3A] This is a schematic diagram illustrating the principle of a speech training system for persons with speech disorders according to the first embodiment of this invention. [Figure 3B]This is a schematic diagram illustrating the principle of a speech training system for persons with speech disorders according to the first embodiment of this invention. [Figure 3C] This is a schematic diagram illustrating the principle of a speech training system for persons with speech disorders according to the first embodiment of this invention. [Figure 4] This is a schematic diagram showing the main screen displayed when the language training system for people with articulation disorders according to the first embodiment of this invention is launched and installed as an application on a tablet computer. [Figure 5] This schematic diagram shows the results of the first language training session on the first day of training for a person with a speech impediment (Person 1) who underwent language training without a model voice using an application of the language training system for persons with speech impediments according to the first embodiment of this invention, which was installed on a tablet computer. [Figure 6] This schematic diagram shows the results of the second language training session on the first day of training for a person with a speech impediment who underwent language training without a model voice using an application of the first embodiment of the speech impediment training system for persons with speech impediments, which was installed on a tablet computer. [Figure 7] This schematic diagram shows the results of the language training on the final day of training for a person with a speech impediment (Person 1) who underwent language training without a model voice using an application of the language training system for persons with speech impediments according to the first embodiment of this invention, which was installed on a tablet computer. [Figure 8] This schematic diagram shows the results of language training on the first day of training for a person with a speech impediment who underwent language training with a model voice using an application of the language training system for persons with speech impediments according to the first embodiment of this invention, which was installed on a tablet computer. [Figure 9] This schematic diagram shows the results of the language training on the final day for a person with a speech impediment (Person 1) who underwent language training with example audio using an application installed on a tablet computer, based on the first embodiment of the language training system for persons with speech impediments of this invention. [Figure 10]This schematic diagram shows the results of language training for a person with a speech impediment (Person 2) who underwent language training with a model voice using an application of the first embodiment of the speech impediment language training system for persons with speech impediments, which was installed on a tablet computer. [Figure 11] This schematic diagram shows the results of language training for a person with a speech impediment (Person 2) who underwent language training with a model voice using an application of the first embodiment of the speech impediment language training system for persons with speech impediments, which was installed on a tablet computer. [Figure 12] This is a schematic diagram showing the processing flow of a speech training system for persons with articulation disorders according to a second embodiment of the present invention. [Figure 13] This is a schematic diagram illustrating the principle of a communication support system for persons with speech impairments according to a second embodiment of the present invention. [Figure 14] This is a schematic diagram showing an example configuration of a communication support system for persons with speech impairments according to a second embodiment of the present invention. [Figure 15] This is a schematic diagram showing an example configuration of a communication support system for persons with speech impairments according to a second embodiment of the present invention. [Figure 16] This is a schematic diagram showing the processing flow of the analysis system for persons with speech disorders according to the third embodiment of this invention. [Figure 17] This is a schematic diagram illustrating the method of analyzing tongue movement in the analysis system for persons with speech disorders according to the third embodiment of this invention. [Figure 18] This is a schematic diagram illustrating the method of analyzing tongue movement in the analysis system for persons with speech disorders according to the third embodiment of this invention. [Figure 19] This is a schematic diagram illustrating the method of analyzing tongue movement in the analysis system for persons with speech disorders according to the third embodiment of this invention. [Figure 20] This is a schematic diagram illustrating a method for analyzing buccinator muscle movement in an analysis system for persons with speech disorders according to a third embodiment of the present invention. [Figure 21]This is a schematic diagram illustrating a method for analyzing the movement of the orbicularis oris muscle in an analysis system for persons with speech disorders according to a third embodiment of this invention. [Figure 22] This is a schematic diagram illustrating a method for training facial muscles in an analysis system for persons with speech disorders according to a third embodiment of the present invention. [Figure 23] This is a schematic diagram illustrating a method for training facial muscles in an analysis system for persons with speech disorders according to a third embodiment of the present invention. [Figure 24] This is a schematic diagram illustrating a method for training facial muscles in an analysis system for persons with speech disorders according to a third embodiment of the present invention. [Figure 25] This is a schematic diagram illustrating a method for training facial muscles in an analysis system for persons with speech disorders according to a third embodiment of the present invention. [Figure 26] This is a schematic diagram illustrating a method for training facial muscles in an analysis system for persons with speech disorders according to a third embodiment of the present invention. [Figure 27] This is a schematic diagram illustrating a method for training facial muscles in an analysis system for persons with speech disorders according to a third embodiment of the present invention. [Figure 28] This is a schematic diagram illustrating a method for training facial muscles in an analysis system for persons with speech disorders according to a third embodiment of the present invention. [Figure 29] This is a schematic diagram illustrating a method for training facial muscles in an analysis system for persons with speech disorders according to a third embodiment of the present invention. [Figure 30] This is a schematic diagram illustrating a method for training facial muscles in an analysis system for persons with speech disorders according to a third embodiment of the present invention. [Figure 31] This is a schematic diagram showing an image of the interface of a facial tracking app. [Figure 32] This is a schematic diagram showing an image of the interface of a facial tracking app. [Modes for carrying out the invention]
[0039] The embodiments for carrying out the invention (hereinafter referred to as "embodiments") will be described below with reference to the drawings.
[0040] <First Embodiment> [Speech therapy system for people with articulation disorders] Figure 1 shows the processing flow of a speech training system for individuals with articulation disorders according to the first embodiment. A program following this processing flow is created and executed by a computer. A speech rate conversion device is connected to the computer, or a speech rate conversion program is installed.
[0041] As shown in Figure 1, in this speech training system for individuals with speech disorders, in stage ST1, a model speech sample converted to a slower speed using speech rate conversion technology is presented to the individual with the speech disorder who is the target of the speech training support. The rate conversion ratio at this time is between 1 / 2 and 1, and multiple ratios can be set within this range.
[0042] Next, in stage ST2, the speech produced by the person with a speech impairment, who imitates the slow-speed converted model speech as described above, is converted at high speed using speech speed conversion technology to return it to the same speed as the model speech before the slow-speed conversion.
[0043] Next, in stage ST3, the rapidly converted speech is presented to the person with the speech impairment or to the listener.
[0044] The computer is equipped with speakers, headphones, or earphones connected or built-in to present a model voice converted at a slow speed by a speech rate conversion device or program to a person with a speech impairment, and to present the high-speed converted voice to the person with the speech impairment or a listener. A microphone for recording the voice of the person with the speech impairment may also be connected or built-in as needed. As an example, Figure 2 shows a tablet computer 10 with a microphone 20 connected. A headset microphone may be used as the microphone 20. The speaker for presenting the high-speed converted voice to the person with the speech impairment or a listener can be one built into the tablet computer 10.
[0045] [Examples of speech training systems for individuals with speech disorders] This paper describes an example of a speech training system for individuals with speech disorders.
[0046] I created application software (hereinafter simply referred to as "the app") for a speech training system for people with speech disorders. This app is designed to assist patients who have difficulty controlling their articulation organs due to brain and nervous system disorders, enabling them to learn the appropriate movements of their articulation organs while listening to their own voice.
[0047] Correct speech requires the proper movement of the articulatory organs. In particular, for consonants, three elements are crucial: articulation pattern, tongue position, and the distinction between voiced and voiceless sounds. This app aims to make learning the correct articulation position easier by making users aware of it through variations in speech speed. In terms of speech analysis, this corresponds to reproducing the correct formant transitions.
[0048] Generally, when speaking slowly, the speed of formant transitions does not slow down; only the vowels are lengthened. In this case, the speed at which the articulatory organs responsible for pronouncing consonants move does not slow down.
[0049] On the other hand, speech slowed down by speech rate conversion often sounds like someone is drunk. This is because it slows down the speed at which the articulatory organs move when speaking.
[0050] This app was created with the idea that by having patients imitate the latter, slower-speed converted speech audio, the speed of speech in the articulatory organs will be slowed down, allowing them to practice focusing on the appropriate articulation position while reducing the motor load according to the symptoms of impaired motor function.
[0051] Because the speech produced in this manner gives the impression of being intoxicated, patients may self-evaluate their speech as incorrect if they listen to it as is. Therefore, the slow speech is corrected by speeding up the speech rate, and by having the patient listen to this corrected version, they can confirm that they are reproducing the correct articulation and become aware of their level of achievement.
[0052] The key feature of this app is that by combining slow and fast speech, it allows users to confidently and faithfully imitate slow speech that sounds like someone is drunk, while also encouraging them to focus on proper articulation placement during practice.
[0053] Figure 3A shows how this app provides voice feedback during articulation training. As shown in Figure 3A, the app plays a training voice for the speech, and the user with articulation impairment visually checks their speaking speed by comparing it to the training voice. It then provides immediate, high-speed feedback that speaking slowly will result in the correct articulation movement (trajectory). Figure 3B shows the relationship between the speech-language pathologist's model voice and the SpeechChain. Figure 3C shows the correction of the SpeechChain for the voice of the user with articulation impairment.
[0054] I created this app and installed it on a tablet computer.
[0055] First, turn on the tablet computer and start it up, then launch the app.
[0056] Figure 4 shows the main screen of the launched application. As shown in Figure 4, the main screen has three main frames A, B, and C, and various buttons below them for checking settings and recorded audio.
[0057] Frame A is the "Teacher's Voice" frame. Here, the teacher is a speech-language pathologist. In the upper left of frame A is a pull-down menu A1 for setting the speed reduction ratio of the practice. Three multipliers are available, for example, 1.0, 1.5, and 2.0. In Figure 4, "×2.0" is selected as the speed reduction ratio, meaning that the length is doubled (speaking speed is halved). In the upper right of frame A are buttons A2 and A3 for selecting practice sentences. For example, five practice sentences are provided. Pressing button A2 selects the previous audio, and pressing button A3 selects the next audio. Practice sentences are selected by operating buttons A2 and A3. In Figure 4, the practice sentence "Dad and Mom and everyone threw beans together" is displayed in the "Teacher's Voice" frame. The waveform of the audio is displayed directly below the practice sentence. The practice sentence displays an audio waveform that serves as a model for practice, and people with speech impairments practice by imitating this audio. In the lower left of frame A are buttons A4 and A5 for playing the audio for practice. Pressing button A4 will play the video at normal speed, and pressing button A5 will play it at a slower speed. In the lower right corner of frame A is a display area A6 showing the elapsed practice time.
[0058] On the left side of frame B are buttons for recording your voice for practice, specifically button B1 for "Record after the teacher" and button B2 for "Record alone". Recording is done by capturing your voice using the tablet computer's built-in microphone or an external microphone, and the audio data is saved to the tablet computer's internal storage. The speaking speed when recording is set to the speed reduction ratio set in the pull-down A1 at the top left of frame A. Set the speed appropriately according to your progress in voice training, and then press button B1 or B2 to record. Pressing button B1 for "Record after the teacher" will play the teacher's voice first, and then recording will begin. Pressing button B2 for "Record alone" will start recording immediately without playing the teacher's voice. The waveform of the recorded voice is displayed in frame C, corresponding to the practice sentence. Record while watching the display on this screen and matching the timing. On the right side of frame B is button B3 for "Stop recording".
[0059] The "Your Voice" section in frame C is a screen where you can check your recorded voice immediately. When playing back a recorded voice, there are "Play Fast" button C1 and "Play Normal" button C2 in the lower left of frame C. The speech speed conversion speed, which reproduces the original speed, corresponds to the speed set in the pull-down A1 in the upper left of frame A. Therefore, when you press the record button to record, the value of the speed set in the pull-down A1 when you press "Play Fast" will be the same. With "Play Fast," the voice that was spoken slowly will be played back at the original speed. You can check your voice by pressing these buttons C1 and C2. There is a "Delete This Recording" button C3 in the lower center of frame C. Press this button C3 if you do not want to keep the recorded voice. If you do not press it, all recorded voices will be saved. If you play back the voice and there is a voice you want to delete, press this button C3 to delete it from the tablet computer. There is a display field C4 in the lower right of frame C that shows the elapsed time of practice.
[0060] Below frame C is frame D1, labeled "Advanced Settings." In the lower center of frame D1 are buttons D2, labeled "Reset Screen," and D3, labeled "Check Saved Audio." Pressing button D2 resets the screen. Pressing button D3 displays a list of saved audio files. The list of saved data can be viewed on the audio list screen.
[0061] Tapping frame D1 in "Advanced Settings" will bring up the settings menu screen, but this is not normally used. At the bottom of the settings menu is a button to check the update page. This is provided for updates in case of bugs.
[0062] [How to use apps to provide speech therapy for people with speech disorders] This section describes one example of how individuals with speech disorders can use this app for speech therapy.
[0063] First, in the screen shown in Figure 4, the speed multiplier (reduction ratio) for the practice is set to 2.0 by operating the pull-down A1 in the upper left of frame A. Next, the practice sentence "Dad and Mom and everyone threw beans together" (example audio) is selected by operating buttons A2 and A3 in the upper right of frame A. Here, assuming that the practice sentence will be played normally, button A4 is pressed to select "Play normally". The audio of the practice sentence is played from the speaker built into the tablet computer and can be heard by the person with articulation impairment. If button B1 for "Record after the teacher" in frame B is pressed, the teacher's voice will be played first, and then the recording will begin. If button B2 for "Record alone" in frame B is pressed, the teacher's voice will not be played and the recording will begin immediately. Recording is done using a microphone. The audio and waveform when the person with articulation impairment imitates the practice sentence are displayed in frame C. These processes correspond to stage ST1 in Figure 1.
[0064] The recorded audio is immediately converted and played back according to the setting in the pull-down menu A1 at the top left of frame A. If the setting in pull-down menu A1 is ×2.0, the conversion speed of the example audio recorded at a speed multiplier of 1 / 2 will be converted to 1x, i.e., the same speed as the original example audio. This process corresponds to stage ST2 in Figure 1.
[0065] Next, the example audio, converted back to its original speed as described above, is played to the person with the speech impairment through a speaker. This process corresponds to stage ST3 in Figure 1. In practice, stages ST2 and ST3 are performed almost simultaneously.
[0066] [Results of speech therapy for individuals with speech disorders using an app] Next, I will explain the results of using this app in speech therapy sessions with two individuals who have speech disorders.
[0067] (1) Speech disorder 1 (59 years old, male, 100 days sick, cerebellar hemorrhage) Subject 1 has ataxic dysarthria and a speech intelligibility score of 2.
[0068] (2) Speech disorder 2 (76 years old, female, 140 days sick, acute left cerebral thrombosis) Subject 3 had flaccid dysarthria in addition to apraxia of speech, and had a speech intelligibility score of 2.
[0069] In language training using the app, there are two modes: one where you press the "Record Alone" button to record audio, meaning you record by only looking at the text without playing a model audio; and another where you press the "Record After Teacher" button to listen to a recorded audio, slowed down to 1.5 times the normal speed, before starting to record.
[0070] (Speech impairment 1) Figure 5 shows the results of the first recording without a model voice on the day training began for patient 1 with articulation disorder, and Figure 6 shows the results of the second recording without a model voice on the same day. Figure 7 shows the results of the recording without a model voice 18 days after the start of training (the final day of training). Comparing Figure 6 and Figure 7, in Figure 6, both the first formant F1 and the second formant F2 / a / are dragged along with "mo", and the F2 of the vowel "me" is not raised, whereas in Figure 7, the dragging along with "mo" in "papamo" has slightly improved, and the F2 of "me" in "mamemaki" is clearly pronounced. From this, it can be considered that the anterior-posterior movement of the tongue has improved. Thus, the speech has slightly improved by the final day of training compared to the start day of training. Comparing Figure 5 and Figure 6, the second recording was slightly worse. This suggests that if you record twice in a row using the "Record Alone" button, which doesn't play the slowed-down example audio, you might mistake your own voice, sped up, for a teaching tone, leading to the counterproductive effect of trying to speak faster.
[0071] Figure 8 shows the results of recording a patient with articulation disorder 1 after playing a slow example audio on the first day of training, and Figure 9 shows the results of recording a patient with articulation disorder 1 after playing a slow example audio on the last day of training. Comparing Figures 5 and 6 with Figures 8 and 9, it can be seen that the audio recorded after playing the slow example audio shows a slight improvement. Furthermore, comparing Figure 8 with Figure 9, it can be seen that the audio on the last day of training shows an improvement.
[0072] (Speech impairment 2) For person 2 with articulation disorder, four recordings were made in one day. For the first and second recordings, the person recorded while only looking at the text of the example audio, and then received high-speed feedback. The results of the first recording are shown in Figure 10. For the third recording, the person recorded after listening to a slow example audio. For the fourth recording, the person recorded while only looking at the text of the example audio, and then received high-speed feedback. The results of the fourth recording are shown in Figure 11. Comparing Figure 10 and Figure 11, in Figure 11, compared to Figure 10, the / a / F2 in "papa mo" is not dragged down by "mo", the "nade" in "minna de" is clearer, and the F1F2 in "mamemaki" is clearer. This suggests that the way the voice is produced has changed, the volume has increased and the F2 is easier to read, and there has been a change in awareness of trying to speak clearly without rushing. Therefore, it can be said that recording after listening to a slow audio allows the person to speak more slowly and improves their voice.
[0073] As described above, according to the first embodiment, by having individuals with speech disorders undergo language training in the flow shown in Figure 1, language training can be easily conducted at home without having to go to a hospital. In addition, in one-on-one training sessions with doctors or speech-language pathologists for individuals with speech disorders, including children with disabilities, it is possible to support training that not only slows down the rhythm of pronunciation but also slows down the speed of the articulation movement itself, as slow speech tailored to the articulation ability of the individual with a speech disorder.
[0074] <Second Embodiment> [Communication support system for people with speech disorders] Figure 12 shows the processing flow of a communication support system for persons with speech impairments according to a second embodiment. A program is created according to this processing flow and executed by a computer. A speech rate conversion device is connected to the computer, or a speech rate conversion program is installed. This communication support system for persons with speech impairments provides communication support, for example, for persons with speech impairments whose speech impairment is so severe that recovery is difficult.
[0075] As shown in Figure 12, in this communication support system for people with speech disorders, in stage ST11, the speech disordered person who is the target of communication support focuses solely on the accuracy of their articulation movements and speaks slowly and carefully, slower than usual for them. This speech is then converted at high speed using speech rate conversion technology to return it to the same speed as the speech before the slow conversion, and presented to the listener.
[0076] The principle of this communication support is shown in Figure 13. Figure 13 shows the correction of the speech chain of the voice of the person with articulation disorder.
[0077] Figure 14 shows an example configuration of this communication support system for a person with a speech impediment. As shown in Figure 14, in this example, a headset microphone 40 is attached to the head 31 of the person with a speech impediment 30 to input their voice. In addition, the person with a speech impediment 30 wears a small speaker-equipped speech rate converter 50 on a collar 60. The voice recorded by the headset microphone 40 is converted into an electrical signal and sent via a cord 41 to the speech rate converter 50, where the speech rate is converted, and the converted voice is emitted to the listener from a speaker (not shown).
[0078] Figure 15 shows an example configuration of this communication support system for a person with a speech impediment (DIS). As shown in Figure 15, in this example, the person with a speech impediment (30) holds a smartphone (70) in their hand and speaks into it. The smartphone (70) has a speech rate conversion function and can convert the speaking speed of the voice spoken into it. The person with a speech impediment (30) attaches a speaker (80) that can wirelessly communicate with the smartphone (70) to a collar (not shown) or holds it in their other hand. The speaker (80) then emits the voice of the person with a speech impediment (30), whose speech rate has been converted by the smartphone (70), to the listener.
[0079] According to the second embodiment, even if speech impairment persists despite language training for individuals with speech impairments, it is possible to improve their communication abilities.
[0080] <Third Embodiment> [Analysis system for individuals with speech disorders] Figure 16 shows the processing flow of the analysis system for individuals with speech disorders according to the third embodiment. A program is created according to this processing flow and executed by a computer. A speech rate conversion device is connected to the computer, or a speech rate conversion program is installed. This analysis system for individuals with speech disorders is used, for example, to analyze the condition of a person with a speech disorder.
[0081] As shown in Figure 16, in this analysis system for individuals with speech disorders, in stage ST21, non-invasive motor analysis of the tongue and facial muscles, which are necessary functions for vocalization and pronunciation, is performed to acquire motor analysis data for individuals with speech disorders.
[0082] Next, in stage ST22, the degree of articulation impairment in individuals with articulation disorders is analyzed by correlating the motor analysis data with the evaluation results of speech-language pathologists regarding those individuals with articulation disorders.
[0083] Tongue movement analysis is performed by examining the forward, backward, left, right, up, and down movements of the tongue. Facial muscle movement analysis is performed using subjective or objective evaluation methods such as the Yanagihara method (40-point method), House-Brackmann method, and Sunnybrook method.
[0084] Let's look at an example of tongue movement analysis. It's common to examine the forward, backward, left, right, up, and down movements of the tongue. Here are a few examples.
[0085] Figure 17 shows a tongue protrusion test. The test is performed in a seated position, with the tongue tip protruding forward in an open-mouth position. The evaluation criteria are: 0 - immobile, 1 - tongue tip can be extended to the mandibular anterior teeth, 2 - tongue tip can be extended to the lower lip, 3 - tongue tip can be extended forward without deviation from the lower lip. The test should be performed twice. If the tongue protrusion is deviationd, it should be considered an error in the range from the target point, and the score should be lowered by one level. If compensatory forward movement of the mandible is clearly observed, it should be suppressed with the fingers.
[0086] Figure 18 shows the lateral movement test of the tongue. The test is performed in a seated position, with the mouth open, and the tip of the tongue is moved to the right corner of the mouth. The evaluation criteria are: 0 - no movement, 1 - the distance the tongue tip moves is less than half the distance between the midline and the corner of the mouth, 2 - the tongue tip does not reach the corner of the mouth, but the distance moved is more than half the distance between the midline and the corner of the mouth, 3 - the tongue tip reaches the corner of the mouth. The test should be performed twice. It is acceptable if the tongue reaches the corner of the mouth slightly off-center. If compensatory lateral movement of the mandible is clearly observed, it should be suppressed with the fingers.
[0087] Figure 19 shows the lateral movement test of the tongue. The test is performed in a seated position, with the mouth open, and the tip of the tongue is moved to the left corner of the mouth. The evaluation criteria are: 0 - no movement, 1 - the distance the tongue tip moves is less than half the distance between the midline and the corner of the mouth, 2 - the tongue tip does not reach the corner of the mouth, but the distance moved is more than half the distance between the midline and the corner of the mouth, 3 - the tongue tip reaches the corner of the mouth. The test should be performed twice. It is acceptable if the tongue reaches the corner of the mouth slightly off-center. If compensatory lateral movement of the mandible is clearly observed, it should be suppressed with the fingers.
[0088] This section will provide a specific example of facial muscle movement analysis. Several examples will be given.
[0089] Figure 20 shows a facial muscle movement test. This test focuses on the buccinator muscle. The test is performed while seated, and the upper and lower lips are pulled as clearly to the sides as possible. While saying "ee," the lips are pulled as far to the sides as possible. The evaluation criteria are: 0 - no movement, 1 - significantly less movement, 2 - slightly less movement, 3 - clearly lower movement possible. The test should be performed twice. Paying attention to left-right symmetry, the degree of movement on the affected side is determined by comparing it with the range of motion on the healthy side.
[0090] Figure 21 shows a lip movement test. This test focuses on the orbicularis oris muscle. The test is performed in a seated position, with the upper and lower lips protruding as clearly as possible forward. While saying "Ooo," the lips are pushed forward as far as possible. The evaluation criteria are: 0 - immobility, 1 - significantly reduced protrusion, 2 - slightly reduced protrusion, 3 - clearly protruding. The test should be performed twice. Paying attention to left-right symmetry, the degree of movement on the affected side should be determined by comparing it with the range of motion on the healthy side.
[0091] As an example of a method for analyzing facial muscle movement, I created an application using the Yanagihara method. I installed this application on a desktop computer.
[0092] First, turn on your desktop computer and start it up, then launch the application.
[0093] Figure 22 shows the main screen of the launched application. As shown in Figure 22, the left side of the main screen displays 10 facial expressions used for evaluating the Yanagihara method. The 10 expressions are: rest, facial wrinkles, slight closure, strong closure, one eye closed, nostril movement, cheek puffing, whistling, showing teeth and saying "ee," and frowning. These expressions can be selected by pressing the measurement button to the left of each expression. Before detecting an expression, the system first detects facial feature points. These feature points are the eyes and the corners of the mouth. On the right side of the main screen, an image of the face captured by a video camera placed in front of the person with a speech impediment is displayed.
[0094] Press the load button at the top of the main screen to load the user's (speech-impaired) data file. The user's data file should be placed in the "data" folder directly under the folder where the program is installed. To create a new user's data file, right-click in the "data" folder, select "New," then "Text Document," and rename the file to the user's name. The text file will record the measurement and diagnostic results. After loading, the user's name will be displayed above the log space. The file name should consist of single-byte numerals.
[0095] The measurement process involves the user with a speech impediment pressing the measurement buttons on the main screen in sequence to record the measurement results for each display. During measurement, the symmetry of the facial features is saved to the user's data file, and a series of images are saved in binary format to a folder named after the user. The measurement results are recorded in the tabi (socks) that the user clicks. Immediately after measuring each facial expression, the user changes the diagnostic value (0-4) using the slider bar below the measurement buttons on the main screen, and then presses the diagnostic record button below the slider bar. The diagnostic value is recorded in the user's data file along with the measurement results.
[0096] As described above, the evaluation obtained through the movement analysis of the tongue and facial muscles can be scored to obtain movement analysis data. On the other hand, speech-language pathologists typically qualitatively score the movement of the muscles involved in articulation in the tongue and facial muscles (for example, 2 points for sufficient movement, 1 point for insufficient movement, and 0 points for no movement) and then train the muscles involved in articulation that are insufficient. Therefore, speech-language pathologists perform evaluations on the articulation disorders they are in charge of. Thus, by correlating the movement analysis data of the tongue and facial muscles with the evaluation results of speech-language pathologists on the articulation disorders, the status of the articulation disorders of the individuals can be understood.
[0097] Next, we will explain how to conduct training (rehabilitation) for individuals with speech disorders using an application developed based on the diagnostic results obtained using the application shown in Figure 22.
[0098] Figure 23 shows the main screen of this app when it is launched. As shown in Figure 23, there are four training buttons in the upper left of the main screen, from top to bottom: "Resting," "Forehead wrinkles," "One eye closed," and "Show teeth." Before training, the app detects facial feature points. These feature points are both eyes and the corners of the mouth.
[0099] Once feature points have been detected, as shown in Figure 24, first press the "resting" training button to take a picture of the face in a resting state and measure the facial feature points (white dots).
[0100] Next, press the "forehead wrinkle" training button to measure the facial feature points when wrinkling your forehead. Figure 25 shows the case when the forehead wrinkle type is small, and Figure 26 shows the case when the forehead wrinkle type is large. In Figures 25 and 26, the arc displayed above the eyebrows is the target eyebrow line. Below the image, there is a meter that shows the degree of movement relative to the target. A higher meter indicates better movement. It can be seen that Figure 26 is closer to the target than Figure 25.
[0101] Next, press the "close one eye" training button to measure facial feature points when one eye is closed. Figure 27 shows a case where the eye is closed slightly, and Figure 28 shows a case where the eye is closed more significantly. Below the image, there is a meter that indicates the success rate of the eye-closing. A higher meter indicates a higher success rate. It can be seen that Figure 28 shows a higher success rate than Figure 27.
[0102] Next, press the "Show teeth" training button to measure the facial feature points when you show your teeth. Figure 29 shows the case where the range of motion is small when showing teeth, and Figure 30 shows the case where the range of motion is large when showing teeth. In Figures 29 and 30, the target positions at both ends of the mouth are shown as dots. Move the green dots closer to the yellow dots. Below the image, there is a meter that shows the degree of movement relative to the target. A higher meter indicates better performance. You can see that Figure 30 is closer to the target than Figure 29.
[0103] By performing the above training, it is possible to improve the motor performance of the facial muscles.
[0104] (An app for rehabilitating speech disorders using corner-of-mouth tracking) Next, we will describe the speech disorder rehabilitation app developed by the present inventor using corner-of-mouth tracking.
[0105] 1. Although there is some overlap with what has already been explained, rehabilitation for speech disorders in individuals with speech disorders mainly consists of face-to-face sessions with a speech-language pathologist and practice at home, and currently, the following challenges are considered to exist. • Limitations on the frequency of face-to-face sessions with experts • Lack of immediate feedback during home practice • Difficulty in maintaining patient motivation
[0106] The inventors developed a rehabilitation app using Apple's ARKIT® technology, particularly its facial tracking function, as a new solution to these problems.
[0107] 1.1 Real-time feedback: Precisely tracks mouth corner movements and provides immediate feedback. 1.2 Improving the quality of home-based rehabilitation: Precise movement instructions and assessment based on expert guidance 1.3 Improved Engagement: Make practice fun with AR gamification. 1.4 Data Analysis: Detailed practice records enable the development of individualized treatment plans. 1.5 Improved Accessibility: Overcoming geographical limitations and providing services to more patients
[0108] 2. Overview of ARKIT (registered trademark) and facial tracking technology ARKIT® is an augmented reality (AR) framework and a tool that makes it easy to create AR experiences on iOS devices. Its main features are as follows:
[0109] 2.1 Environmental understanding: • Uses cameras and various sensors to recognize the surrounding environment. • Provides functions such as plane detection, light source estimation, and image recognition. 2.2 Motion Tracking: • Accurately tracks device movement and displays AR containers stably. 2.3 Facial Tracking: • Available on devices with a TrueDepth camera (iPhone® / iPad®). • Tracks facial movements and expressions in real time. 2.4 Ease of Development: The details (features) of facial tracking technology are as follows: Tracking 52 facial feature points • Detects various facial expressions such as blinking, eye movements, and mouth shape. • Access to facial posture and expression data is possible through the ARFaceAnchor class. • Use blend shapes to quantify subtle changes in expression. These features make ARKIT® a suitable foundation for developing speech disorder rehabilitation apps. In particular, it can accurately track the movement of the corners of the mouth and provide real-time feedback.
[0110] 3. Analysis of ARKIT® functionality specialized for corner of the mouth tracking In speech disorder rehabilitation, the movement of the corners of the mouth is particularly important. By utilizing the facial tracking function of ARKIT®, it is possible to develop apps specifically for this area. 3.1 Related Blend Shapes and Feature Points ARKIT® tracks the movement of various parts of the face, but the key blend shapes and feature points that are particularly relevant to tracking the corners of the mouth are as follows: 3.1.1 Blend shapes related to the corners of the mouth: • mouthSmileLeft: Upward curve of the left corner of the mouth • mouthSmileRight: Upward curve of the right corner of the mouth • mouthFrownLeft: The left corner of the mouth is lowered. • mouthFrownRight: Right corner of the mouth drooping • firstLeft: Movement to the left of the mouth • mouthRight: Movement to the right of the mouth 3.1.2 Key Features • Points on the left and right corners of the mouth (index 61 and 65) • The midpoint of the upper lip (index 13) • The center point of the lower lip (index 14) By combining these data points, it is possible to precisely track the movement of the corners of the mouth and obtain the information necessary for rehabilitation. 3.2 Addition of tongue movement tracking function Furthermore, we will add a function to track tongue movements in real time. Adding this function will be extremely beneficial for the rehabilitation of speech disorders. We will implement it using the following approach. 3.2.1 Adding a new parameter: Add a new parameter related to tongue movement to the existing parameter list. • tongueTip_UpDown: Up and down movement of the tip of the tongue • tongueTip_LeftRight: Movement of the tip of the tongue from side to side. • tongueBody_UpDown: Up and down movement of the central part of the tongue • tongueBody_LeftRight: Movement of the central part of the tongue from side to side. • tongueRoot_ForwardBack: Forward and backward movement of the tongue root • tongueCurl: the degree to which the tongue is curled 3.2.2 3D Model Enhancement: The face model on the right has been improved to visualize tongue movement. • Added a semi-transparent face model option to display tongue movement inside the mouth. • Added a 3D model of the tongue, reflecting its movement in real time. • Highlight changes in tongue position and shape using color and shape. 3.2.3 Dedicated Tongue Movement Viewer: Adds a viewer specifically for tongue movements to a portion of the screen. • Simultaneously display side and front views to check the position and shape of the tongue from multiple angles. • The target tongue position is displayed as a guideline. • Visualize the difference between the actual tongue position and the target position in real time. 3.2.4 Precision analysis of tongue movement • The precision of movement of each part of the tongue (tip, middle, and base) is scored. • Calculates and displays the deviation from the target tongue movement in real time. 3.2.5 Customizable practice modes • A practice mode that targets the tongue movements necessary for pronouncing specific elements or words. • Difficulty levels can be set (e.g., only tongue tip movement, complex movements of the entire tongue, etc.) 3.2.6 Feedback System • Visual feedback: If the tongue position is correct, it will be displayed in green. If improvement is needed, it will be displayed in red. • Voice guidance: Provides voice instructions regarding tongue position and movement. 3.2.7 Data Recording and Progress Analysis • Record tongue movement data in chronological order. Integrating these functions into the existing system will create a more comprehensive and effective speech disorder rehabilitation tool. Since tongue movement plays a crucial role in many forms of speech, this addition has the potential to significantly improve the quality of speech therapy. 3.3 Evaluation of Accuracy and Response Rate The accuracy and response speed of ARKIT's (registered trademark) facial tracking function directly impact the effectiveness of rehabilitation apps. 3.3.1 Accuracy • Generally, facial movements can be tracked with an accuracy of 0.5 mm or less. Regarding the movement of the corners of the mouth, it is possible to detect an angle change of approximately 1-2 degrees. 3.3.2 Reaction Rate • Real-time tracking at 60fps (frames per second) is possible. • Delay is usually within 10-20 milliseconds These performance metrics provide sufficient accuracy and speed for many speech disorder rehabilitation applications. However, in certain pronunciation exercises that require particularly fine or rapid movements, the application of complementary algorithms or filtering techniques may be necessary. 3.4 Technical limitations and countermeasures 3.4.1 Lighting conditions • Problem: Extreme brightness or darkness affects tracking accuracy • Solution: Implement a feature within the app to guide users on appropriate lighting conditions. 3.4.2 Direction of the face Problem: It is difficult to track the orientation of the face from extreme angles. • Solution: Provide visual feedback to guide the user to the optimal face position. 3.4.3 Individual differences • Problem: Individual differences in facial shape and movement affect accuracy • Countermeasures: Establish an initial calibration stage and perform adjustments tailored to the individual. By considering these analyses and countermeasures, it becomes possible to effectively design and implement a speech disorder rehabilitation app utilizing ARKIT®.
[0111] 4. System Overview 4.1 System Configuration After sampling facial feature points using the iPhone's (registered trademark) depth sensor, the facial tracking 52 feature points are sent to a computer via Wi-Fi. The computer then uses the received data to create and display the facial shape in real time. 4.2 Screen Configuration Figures 31 and 32 show images of the facial tracking app interface. The main features of these images are as follows: 4.2.1 Screen Configuration Figure 31: List of 52 facial tracking parameters Figure 32: 3D face model updated in real time 4.2.2 Facial Tracking Parameters • Quantifying the movements of various body parts such as the nose, mouth, chin, eyes, and cheeks. Examples: mouthSmile_R (right-side smile), eyeWide_L (left eye opening), etc. Each parameter is displayed in the range of 0 to 100. 4.2.3 3D Face Model • Silver human head model • Facial expressions and movements are updated in real time according to changes in parameters. 4.2.4 Function • Record button: Records facial tracking data. • Play button: Play recorded data • All 52 parameters can be displayed and recorded in real time. This app allows you to visualize, record, and analyze detailed facial tracking data in real time.
[0112] According to the third embodiment, the condition of speech impairment in individuals with speech disorders can be easily and objectively analyzed by correlating the results of the analysis of tongue and facial muscle movements with the results of the evaluation by a speech-language pathologist.
[0113] Although embodiments of this invention have been described in detail above, this invention is not limited to the embodiments described above, and various modifications based on the technical idea of this invention are possible.
[0114] For example, the numerical values, configurations, shapes, arrangements, materials, and methods mentioned in the above-described embodiments are merely examples, and different numerical values, configurations, shapes, arrangements, materials, and methods may be used as needed. [Explanation of symbols]
[0115] 10... Tablet computer, 20... Microphone, 30... Speech impairment, 40... Headset microphone, 50... Speech rate converter, 60... Collar, 70... Smartphone, 80... Speaker
Claims
1. Using speech rate conversion technology, a model audio file converted to a slower speed is presented to individuals with speech disorders who are the target of language training support. The speech produced by the person with the speech impairment who imitates the slow-speed converted example speech is then converted at high speed using speech rate conversion technology to return it to the same speed as the original example speech before the slow-speed conversion. A speech training support system for persons with speech impairments, configured to present the above-mentioned high-speed converted audio to the person with the speech impairment or to a listener.
2. The language training system for a person with a speech impairment according to claim 1, which involves presenting the person with the speech impairment with the above-mentioned high-speed converted speech to the person with the speech impairment to provide training that focuses on the accuracy of articulation movements.
3. The language training system for persons with articulation disorders according to claim 1, wherein the speed of the model voice converted at low speed is the fastest speed at which the person with the articulation disorder can speak with accurate articulation without difficulty.
4. The speech training system for persons with articulation disorders according to claim 1, wherein the speed of the model voice converted to the slow speed is at least half the speed of the model voice before the slow speed conversion and not more than 1x the speed of the model voice before the slow speed conversion.
5. A speech rate conversion device, A speaker, headphones, or earphones is provided to present the above-mentioned slow-converted example audio to the person with the speech impairment, and to present the above-mentioned fast-converted audio to the person with the speech impairment or to the listener. A speech training system for persons with articulation disorders according to claim 1.
6. The speech training system for a person with a speech impairment according to claim 1, configured to present the waveform of the sample voice converted at a slow speed and the characters superimposed on the waveform to the person with a speech impairment.
7. The above example audio is created by a speech-language pathologist or another person recording the audio using the recording function of the above speech-language training support system for persons with speech disorders, or a pre-prepared audio file is registered in the above speech-language training support system for persons with speech disorders, as described in claim 1.
8. The speech training support system for persons with speech disorders according to claim 7, wherein the above audio file is created by editing audio recorded by a speech-language pathologist or another person, or by editing audio created by speech synthesis.
9. The speech training support system for persons with speech disorders according to claim 1, wherein the above-mentioned model voice and the characters to be displayed on the display corresponding to the above-mentioned model voice are pre-registered as a set.
10. The speech training support system for persons with speech impairments according to claim 9, wherein the position of speech is presented on the display as an animation when the above-mentioned model audio is played and when a person with a speech impairment speaks.
11. The speech training support system for persons with articulation disorders according to claim 1, wherein the ratio of speech rate conversion based on the above-mentioned model audio can be arbitrarily set.
12. The speech training support system for persons with articulation disorders according to claim 1, which automatically recognizes the number of morae per unit time of the above-mentioned example audio and sets the speaking speed by arbitrarily setting the number of morae per unit time during playback.
13. The speech training system for a person with a speech disorder according to claim 1, wherein the person with the speech disorder has organic speech disorder, motor speech disorder, auditory speech disorder, or functional speech disorder.
14. The stage involves presenting a model audio, converted to a slower speed using speech rate conversion technology, to individuals with articulation disorders who are the target of language training support, and The process involves the following steps: first, the person with the speech impairment imitates the slow-speed converted example speech and speaks at a slow speed; then, using speech speed conversion technology, the speech is converted at a high speed to return it to the same speed as the original example speech before the slow-speed conversion; The stage of presenting the above-mentioned high-speed converted audio to the person with the speech impairment or to the listener, A method for supporting speech training for individuals with speech disorders.
15. A program for causing a computer to execute the speech training support method for persons with speech disorders described in claim 14.
16. A computer-readable recording medium storing the program described in claim 15.
17. A communication support system for persons with speech disorders, which are the target of communication support, is configured to present to the listener the speech they have spoken at a slower speed than usual, focusing solely on the accuracy of their articulation, and then using speech speed conversion technology to convert the speech back to the same speed as the speech before the slow conversion.
18. A microphone for recording the voice of the person with the above-mentioned speech impediment, A wearable unit equipped with a voice conversion device, to be worn by the person with the above-mentioned speech impairment, It has, The communication support system for a person with a speech impairment according to claim 10, wherein the person with the speech impairment wears the wearable unit, and the voice of the person with the speech impairment recorded by the microphone is converted and output by the voice conversion device.
19. Using a smartphone with a recording function and a speaker capable of converting speech carried by the person with the speech impairment, The communication support system for a person with a speech impairment according to claim 18, which converts the voice of the person with the speech impairment speaking into the smartphone into voice and outputs it from the speaker.
20. A communication support method for persons with speech disorders, comprising the step of using speech rate conversion technology to convert the speech spoken by the person with a speech disorder, who is the target of communication support, at a slower speed than usual and carefully, focusing solely on the accuracy of articulation movements, back to the same speed as the speech before the slow conversion, and presenting it to the listener.
21. A program for causing a computer to execute the communication support method described in claim 20.
22. A computer-readable recording medium storing the program described in claim 21.
23. For individuals with speech disorders, we non-invasively analyze the movements of the tongue and facial muscles, which are essential for vocalization and pronunciation, and obtain movement analysis data. An analysis system for individuals with articulation disorders, configured to analyze the degree of articulation disorder in such individuals by correlating the motor analysis data with the evaluation results of a speech-language pathologist regarding the aforementioned individuals with articulation disorders.
24. The analysis system for persons with speech disorders according to claim 23, wherein tongue movement analysis is performed by examining the forward, backward, left, right, up, and down movements of the tongue, and facial muscle movement analysis is performed by the Yanagihara method, the House-Brackmann method, or the Sunnybrook method.
25. The process involves non-invasively analyzing the movements of the tongue and facial muscles, which are essential for vocalization and pronunciation, in individuals with speech disorders, and obtaining movement analysis data. The process involves analyzing the degree of articulation impairment in the person with articulation impairment by correlating the motor analysis data with the evaluation results of the speech-language pathologist regarding the person with articulation impairment, A method for analyzing individuals with speech disorders.
26. A program for causing a computer to perform the analysis method for persons with speech disorders described in claim 25.
27. A computer-readable recording medium storing the program described in claim 26.