Articulation abnormality detection method, articulation abnormality detection device, and program
The articulation abnormality detection method addresses the discomfort of facial imaging by analyzing speech patterns to detect stroke risks, enhancing detection accuracy and reducing subject burden.
Patent Information
- Application Number
- JP2021143569
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-09-02
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2041-09-02
AI Technical Summary
Existing stroke detection systems that analyze facial images impose a burden on subjects, requiring them to pose for cameras and maintain specific angles, which can be uncomfortable.
An articulation abnormality detection method that analyzes speech patterns, specifically using a detection model to identify articulation disorders through speech information, including specific sounds and phrases, to detect potential stroke indicators without requiring visual imaging.
Facilitates easy detection of articulation abnormalities and potential stroke risks by analyzing speech, reducing subject burden and improving accuracy through machine-learned models.
Smart Images

Figure 0007731102000001 
Figure 0007731102000002 
Figure 0007731102000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an articulation abnormality detection method, an articulation abnormality detection device, and a program for detecting articulation abnormalities in a subject. [Background technology]
[0002] Patent Literature 1 discloses a system for detecting a leading indicator of stroke risk. In this detection system, a video camera captures a video of the face of a subject to be evaluated for stroke risk. In this detection system, a processor analyzes processed image data associated with the video of the subject's face captured by the video camera. In this detection system, the processor determines whether the captured image data presents a leading indicator of carotid artery stenosis. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Special Publication No. 2016-522730 Summary of the Invention [Problem to be solved by the invention]
[0004] The present disclosure provides an articulation abnormality detection method, an articulation abnormality detection device, and a program that can easily detect the presence or absence of an articulation abnormality in a subject without imposing a burden on the subject. [Means for solving the problem]
[0005] A method for detecting articulation abnormalities according to one aspect of the present disclosure includes: 1. A processor-implemented method for detecting articulation abnormalities, comprising:The method includes an acquisition step and a detection step. In the acquisition step, speech information related to speech uttered by a subject is acquired. In the detection step, the presence or absence of an articulation disorder of the subject is detected based on an output result obtained by inputting the speech information acquired in the acquisition step into a detection model that has been machine-learned to input speech and output information related to the presence or absence of an articulation disorder. The audio information includes a plurality of phrases consisting of a sequence of specific sounds and plosives produced by the subject moving their tongue in a predetermined pattern, and further includes a division step of dividing the plurality of phrases from the audio information acquired in the acquisition step based on the position of the plosives, and in the detection step, each of the plurality of phrases divided in the division step is input into the detection model. [Effects of the Invention]
[0006] The present disclosure has the advantage of making it easy to detect whether or not a subject has an articulation disorder without imposing a burden on the subject. [Brief explanation of the drawings]
[0007] [Figure 1] Figure 1 is an explanatory diagram of stroke patient characteristics. [Figure 2] FIG. 2 is a diagram showing an example of a speech waveform of a healthy subject and a mel spectrogram obtained from the speech waveform. [Figure 3] FIG. 3 is a diagram showing an example of a voice waveform of a stroke patient and a melspectrogram obtained from the voice waveform. [Figure 4] FIG. 4 is a block diagram showing an example of the configuration of an articulation abnormality detection device according to an embodiment. [Figure 5] FIG. 5 is a diagram showing an example of a speech waveform of a healthy subject who uttered a plurality of phrases and a mel spectrogram obtained from the speech waveform. [Figure 6] FIG. 6 shows an example of a speech waveform of a stroke patient who uttered a number of phrases and a mel spectrogram obtained from the speech waveform. [Figure 7] FIG. 7 is a diagram showing another example of a speech waveform of a stroke patient who uttered a plurality of phrases and a mel spectrogram obtained from the speech waveform. [Figure 8] FIG. 8 shows an example of RMS envelopes obtained from the speech waveforms of a healthy subject and a stroke patient who generated multiple phrases. [Figure 9]FIG. 9 is a diagram showing an example of the learning phase of the segmented model of the articulation abnormality detection device according to the embodiment. [Figure 10] FIG. 10 is a diagram showing an example of the inference phase using the segmented model of the articulation abnormality detection device according to the embodiment. [Figure 11] FIG. 11 is a diagram illustrating an example of the learning phase of the detection model of the articulation abnormality detection device according to the embodiment. [Figure 12] FIG. 12 is a diagram illustrating an example of an inference phase using a detection model of an articulation abnormality detection device according to an embodiment. [Figure 13] FIG. 13 is a flowchart showing an example of the operation of the articulation abnormality detection device according to the embodiment. [Figure 14] FIG. 14 is a diagram illustrating an example of an outline of an articulation abnormality detection device and an articulation abnormality detection method according to an embodiment. [Figure 15] FIG. 15 is a diagram illustrating a specific example of the operation of the articulation abnormality detection device according to the embodiment. [Figure 16] FIG. 16 is a diagram showing another specific example of the operation of the articulation abnormality detection device according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0008] (Findings that led to this disclosure) Conventionally, a technique for detecting the risk of stroke by analyzing captured images of a subject's face has been known, and is disclosed, for example, in Patent Document 1. As already mentioned, the detection system disclosed in Patent Document 1 captures a video of the subject's face with a video camera. Then, this detection system analyzes processed image data associated with the video of the subject's face to determine whether the captured image data presents a leading indicator of carotid artery stenosis, which is one risk factor for stroke.
[0009] However, the detection system disclosed in Patent Document 1 has the problem that it is necessary to capture video of the subject's face using a video camera, which tends to be a heavy burden for subjects who are reluctant to be photographed by a camera or the like.
[0010] Furthermore, in the detection system disclosed in Patent Document 1, since image data of the subject's face is analyzed, it is important that the subject's face is in an appropriate position or angle in the image data. Therefore, if the subject is asked to capture their own face with a video camera, the subject must make some effort to obtain appropriate image data, which tends to be a heavy burden on the subject.
[0011] Therefore, the inventors of the present application have conducted extensive research in consideration of the above-mentioned problems and have found that it is possible to detect the presence or absence of articulation disorders in a subject from the speech of the subject, in other words, whether or not the subject can correctly pronounce the phonemes that are the elements of words when they are spoken. As will be described later, the presence or absence of articulation disorders in a subject can indicate the presence or absence of a premonition of the onset of a stroke in the subject. Therefore, the presence or absence of a premonition of the onset of a stroke in a subject can be detected simply by the subject's speech.
[0012] Therefore, according to the present disclosure, it is possible to provide an articulation abnormality detection method, an articulation abnormality detection device, and a program that can more easily detect the presence or absence of an articulation abnormality in a subject, and even the presence or absence of signs of a stroke in a subject, without placing a burden on the subject, compared to when capturing an image of the subject's face.
[0013] (Summary of the Disclosure) An outline of one aspect of the present disclosure is as follows.
[0014] A method for detecting articulation abnormalities according to one aspect of the present disclosure includes an acquisition step and a detection step. In the acquisition step, speech information related to speech uttered by a subject is acquired. In the detection step, the presence or absence of an articulation abnormality of the subject is detected based on an output result obtained by inputting the speech information acquired in the acquisition step into a detection model that has been machine-learned to input speech and output information related to the presence or absence of an articulation abnormality.
[0015] This method makes it possible to detect the presence or absence of articulation abnormalities in a subject simply by the subject uttering a voice. This has the advantage that it is easier to detect the presence or absence of articulation abnormalities in a subject without imposing a burden on the subject, compared to when the subject is asked to take a video of their own face with a video camera.
[0016] For example, in the articulation abnormality detection method according to one aspect of the present disclosure, the audio information may include a specific sound produced by the subject moving his / her tongue in a predetermined pattern.
[0017] This has the advantage that it is easier to detect the degree of tongue paralysis, which can be an indicator of whether or not a subject has an articulation disorder, compared to when the audio information does not contain specific sounds.
[0018] For example, in the articulation abnormality detection method according to an aspect of the present disclosure, the specific sound may be a popping sound.
[0019] This has the advantage that by including in the specific sounds popping sounds that are difficult to produce when the tongue is paralyzed, it becomes easier to detect whether or not the subject has an articulation disorder.
[0020] For example, in the articulation abnormality detection method according to one aspect of the present disclosure, the speech information may include a phrase in which the specific sound and a plosive sound are consecutive.
[0021] This has the advantage that by making a plosive sound, whose position is easy to identify in the speech of the subject, consecutive to a specific sound, it becomes easier to identify the position of the specific sound in the speech of the subject, making it easier to detect whether or not the subject has an articulation abnormality.
[0022] For example, in an articulation abnormality detection method according to an aspect of the present disclosure, the speech information may include a plurality of the phrases. The articulation abnormality detection method according to an aspect of the present disclosure may further include a dividing step of dividing the speech information acquired in the acquiring step into the plurality of phrases. In the detecting step, each of the plurality of phrases divided in the dividing step may be input to the detection model.
[0023] This has the advantage that it becomes easier to detect whether or not a subject has an articulation disorder compared to detecting whether or not a subject has an articulation disorder from a single phrase.
[0024] For example, in the articulation abnormality detection method according to an aspect of the present disclosure, the dividing step may divide the plurality of phrases based on a root mean square (RMS) envelope or a spectrogram as the speech information.
[0025] This has the advantage that the accuracy of distinguishing between multiple phrases can be expected to improve, since features that can distinguish between multiple phrases are likely to appear in the RMS envelope or spectrogram.
[0026] For example, in an articulation abnormality detection method according to one aspect of the present disclosure, the division step may divide the multiple phrases by inputting the speech information acquired in the acquisition step into a division model that has been machine-trained to divide the multiple phrases using speech including the multiple phrases as input.
[0027] This has the advantage that the accuracy of segmenting multiple phrases can be expected to improve compared to segmenting multiple phrases without using a segmentation model. Note that when there is a large amount of training data, a segmentation model that is a deep neural network (DNN) model can be expected to improve accuracy. Also, when there is a small amount of training data, a segmentation model that uses the RMS envelope as speech information can be expected to improve accuracy.
[0028] For example, in an articulation abnormality detection method according to an aspect of the present disclosure, the detection model may be an autoencoder that has been machine-learned to restore speech identical to that of an inputted healthy subject. Furthermore, in the detection step, the presence or absence of an articulation abnormality in the subject may be detected based on a degree of deviation between the speech information inputted to the detection model and the speech information outputted from the detection model.
[0029] This has the advantage that it is easier to prepare a large amount of training data compared to training a detection model using the speech of patients with articulation disorders, who are fewer in number than healthy individuals, making it easier to train the detection model.
[0030] For example, the articulation abnormality detection method according to one aspect of the present disclosure may further include an output step of outputting detection information regarding the presence or absence of an articulation abnormality of the subject detected in the detection step.
[0031] This has the advantage that, for example, by outputting the detection information to the subject, the subject can know whether or not he or she has an articulation disorder.
[0032] For example, the articulation abnormality detection method according to one aspect of the present disclosure may further include, before the acquisition step, a playback step of playing back to the subject a sample speech of speech uttered by the subject.
[0033] This has the advantage that the subject can attempt to speak in order to reproduce the sample voice, making it easier to acquire the subject's voice compared to when a character string is displayed to prompt the subject to speak. Also, this has the advantage that it is possible to detect the presence or absence of articulation abnormalities in the subject, including whether the subject is able to reproduce and produce the sample voice, and that this is expected to improve the accuracy of detecting the presence or absence of articulation abnormalities in the subject.
[0034] Furthermore, a program according to one aspect of the present disclosure causes one or more processors to execute the above-described articulation abnormality detection method.
[0035] This method makes it possible to detect the presence or absence of articulation abnormalities in a subject simply by the subject uttering a voice. This has the advantage that it is easier to detect the presence or absence of articulation abnormalities in a subject without imposing a burden on the subject, compared to when the subject is asked to take a video of their own face with a video camera.
[0036] An articulation abnormality detection device according to an aspect of the present disclosure includes an acquisition unit and a detection unit. The acquisition unit acquires speech information related to speech uttered by a test subject. The detection unit detects the presence or absence of an articulation abnormality of the test subject based on an output result obtained by inputting the speech information acquired by the acquisition unit into a detection model that has been machine-learned to input speech and output information related to the presence or absence of an articulation abnormality.
[0037] This method makes it possible to detect the presence or absence of articulation abnormalities in a subject simply by the subject uttering a voice. This has the advantage that it is easier to detect the presence or absence of articulation abnormalities in a subject without imposing a burden on the subject, compared to when the subject is asked to take a video of their own face with a video camera.
[0038] These comprehensive or specific aspects may be realized as a system, a method, an apparatus, an integrated circuit, a computer program, or a computer-readable recording medium such as a CD-ROM, or may be realized as any combination of a system, a method, an apparatus, an integrated circuit, a computer program, and a recording medium.
[0039] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. The numerical values, shapes, materials, components, component placement and connection configurations, steps, and step sequences shown in the following embodiments are merely examples and are not intended to limit the scope of the claims. Furthermore, among the components in the following embodiments, components that are not described in the independent claims that represent the highest concepts are described as optional components. Furthermore, each drawing is not necessarily an exact illustration. In each drawing, substantially identical components are designated by the same reference numerals, and redundant descriptions may be omitted or simplified.
[0040] (Embodiment) Hereinafter, the embodiments will be specifically described with reference to the drawings.
[0041] [1. Overview] First, before explaining an articulation abnormality detection device and an articulation abnormality detection method according to an embodiment, an overview of the finding that features that can be used to detect the presence or absence of an articulation abnormality in a subject's speech appear in the speech of the subject will be described. FIG. 1 is an explanatory diagram of the characteristics of stroke patients. Here, stroke may include, for example, cerebral infarction, such as lacunar cerebral infarction or atherothrombotic cerebral infarction, or cerebral hemorrhage. FIG. 1 shows the results of an abnormality location estimation performed by a speech-language-hearing therapist based on a total of several hundred speeches from several dozen stroke patients. In FIG. 1, the horizontal axis represents the location of oral paralysis diagnosed, and the vertical axis represents the number of subjects. As shown in FIG. 1, stroke patients often experience oral paralysis. In particular, stroke patients are thought to have significant tongue paralysis, such as anterior, middle, or posterior tongue paralysis.
[0042] In order to identify the location of paralysis in the subject's oral cavity, the subject is asked to speak a test phrase, which is then listened to by a speech-language-hearing therapist. The test phrase is, for example, "Ruri mo haru mo teraseba hikariru" (Lucuritiform glass and glass are both illuminated and shine) which is difficult for the subject to speak if their oral cavity is paralyzed.
[0043] Fig. 2 shows an example of a voice waveform of a healthy subject and a spectrogram obtained from the voice waveform, and Fig. 3 shows an example of a voice waveform of a stroke patient and a spectrogram obtained from the voice waveform.
[0044] In each of Figures 2 and 3, the upper area A1 represents the audio waveform, and the lower area A2 represents the spectrogram. The spectrogram here represents the frequency spectrum of the subject's audio over time. The audio waveforms shown in Figures 2 and 3 were obtained by having the subject utter the test phrase "lapis lazuli and bora mo teraseba hikariru" ("Ruri mo haru mo teraseba hikariru") and recording the audio.
[0045] The test phrase "lapis lazuli and bora mo teraseba hikariru" ("Ruri mo hari mo teraseba hikariru") contains a consonant from the Japanese "ra" row, which is a vowel. A vowel is a consonant produced by momentary contact between the articulatory organs in the oral cavity, for example, a sound produced by touching the tongue to the hard palate for a very short time. In other words, a vowel is a specific sound produced by the subject moving their tongue in a specific pattern. Such specific sounds are difficult to pronounce correctly if the tongue is paralyzed.
[0046] In Figures 2 and 3, the white arrows indicate the positions where the consonants in the "ra" row, i.e., the "bounce" sounds, are produced in the test phrase. As shown in Figure 2, in the mel spectrogram obtained from the speech waveform of a healthy subject, a dark linear region B1 appears vertically at the position where the "bounce" sound is produced. As shown, when the "bounce" sound is produced correctly, a power drop occurs for a very short period of time (for example, 20 ms or less).
[0047] On the other hand, as shown in Figure 3, in the spectrogram obtained from the speech waveform of a stroke patient, there is no very short-term drop in power at the position where the pop sound is produced, that is, the dark vertical line area B1 does not appear (see area C1). This inaccuracy in the sound of the pop sound at the position where it should be produced is thought to be due to the fact that the tongue of the stroke patient is paralyzed, preventing the tongue from contacting the hard palate. It can also be said that the pop sound is not being produced correctly when the drop in power is relatively weak, or when a drop in power occurs but lasts for a relatively long time.
[0048] As described above, the speech of a subject contains features that can be used to detect whether the subject has tongue paralysis, in other words, whether the subject has an articulation disorder. Therefore, by analyzing the features that appear in the speech of a subject, for example, by analyzing whether or not a pop sound is pronounced correctly, it is possible to detect whether the subject has an articulation disorder or even whether or not the subject is at risk of developing a stroke.
[0049] [2. Configuration] Next, the configuration of an articulation abnormality detection device and an articulation abnormality detection method according to an embodiment will be described in detail. FIG. 4 is a block diagram showing an example of the configuration of an articulation abnormality detection device 100 according to an embodiment. In this embodiment, the articulation abnormality detection device 100 is installed in an information terminal such as a smartphone or a tablet terminal. Of course, the articulation abnormality detection device 100 may also be installed in a desktop or laptop personal computer. The articulation abnormality detection device 100 is also called an "articulation abnormality detection system 100."
[0050] 4, the articulation abnormality detection device 100 includes an acquisition unit 11, a classification unit 12, a detection unit 13, an output unit 14, a reproduction unit 15, and a storage unit 16. The storage unit 16 also stores a classification model 17 and a detection model 18. In the embodiment, the acquisition unit 11, the classification unit 12, the detection unit 13, the output unit 14, and the reproduction unit 15 are all realized by a processor installed in an information terminal or a personal computer executing a predetermined program.
[0051] The acquisition unit 11 acquires speech information related to the speech uttered by the subject. The acquisition unit 11 is the entity that executes the acquisition step in the articulation abnormality detection method. The acquisition unit 11 acquires speech information by, for example, collecting the speech uttered by the subject using a microphone mounted on an information terminal and converting the collected speech into an electrical signal. Here, the speech information may include a speech waveform of the speech uttered by the subject, or information obtained by performing appropriate information processing on the speech waveform. As an example, the speech information may include an RMS (Root Mean Square) envelope obtained from the speech waveform, or a spectrogram (including a mel spectrogram) of the speech waveform.
[0052] In the embodiment, the subject is prompted to utter a test phrase including a plurality of phrases, and the acquisition unit 11 acquires speech information including the plurality of phrases. The phrase here is a phrase consisting of a sequence of a specific sound, such as a popping sound, that is produced by the subject moving their tongue in a predetermined pattern, followed by a plosive sound. In the embodiment, the phrase is "dere." That is, in the embodiment, the subject is prompted to utter the test phrase "derederedere..." by repeating the phrase multiple times.
[0053] Thus, in the embodiment, the audio information includes a specific sound produced by the subject moving their tongue in a predetermined pattern. Also, in the embodiment, the specific sound is a popping sound. Also, in the embodiment, the audio information includes a phrase in which the specific sound and a plosive sound are consecutive. Furthermore, in the embodiment, the audio information includes a plurality of phrases.
[0054] The reason for adopting "derederedere..." as the test phrase will be explained below. As described above, if the test phrase contains a specific sound, such as a popping sound, it is possible to detect whether the subject has an articulation disorder from the subject's speech. However, in order to analyze whether the subject correctly pronounced the specific sound, it is preferable to identify the position in the subject's speech where the specific sound should be pronounced. This is because, when a subject with an articulation disorder, such as a stroke patient, utters a test phrase, it is impossible to determine whether the subject failed to pronounce the specific sound correctly or whether the subject was not even trying to pronounce the specific sound in the first place unless the position where the specific sound should be pronounced is known.
[0055] Therefore, the inventors of the present application have found that a phrase consisting of a sequence of a plosive sound, the position of which in the speech of the test subject is relatively easy to identify, and a specific sound can be used as the test phrase. A plosive sound is a sound (consonant) that is produced when the subject suddenly releases the closure between the lips, the space between the tip of the tongue and the upper gums, or the space between the back of the tongue and the soft palate, etc., after holding the breath. Compared to a pop sound, a plosive sound is easier to produce even if the tongue is paralyzed, and its power temporarily decreases when it is produced. Therefore, the position of a plosive sound in the speech of the test subject is relatively easy to identify.
[0056] If the position of the plosive sound in the speech of the test subject can be identified, it is also possible to identify the position of the specific sound that follows the plosive sound. In this embodiment, "dere" is used as a phrase in which a plosive sound and a specific sound follow each other.
[0057] Furthermore, by using multiple phrases, such as "derederedere...," rather than the single phrase "dere" as the test phrase, we further improved the accuracy of detecting whether or not a subject has an articulation disorder. This is because, if a subject only utters the single phrase "dere," a subject with an articulation disorder, such as a stroke patient, may accidentally pronounce the specific sound correctly. In contrast, if a subject utters multiple phrases, such as "derederedere...," the probability that a subject with an articulation disorder will be unable to pronounce the specific sound correctly increases, making it easier to detect whether or not a subject has an articulation disorder. Additionally, repeating multiple phrases increases the complexity of the tongue movement requirements, making articulation disorders more likely to become apparent.
[0058] Fig. 5 shows an example of a speech waveform of a healthy subject uttering multiple phrases and a spectrogram obtained from the speech waveform. Fig. 6 shows an example of a speech waveform of a stroke patient uttering multiple phrases and a spectrogram obtained from the speech waveform. Fig. 7 shows another example of a speech waveform of a stroke patient uttering multiple phrases and a spectrogram obtained from the speech waveform.
[0059] 5 to 7, the upper area A1 represents the audio waveform, and the lower area A2 represents the spectrogram. The audio waveforms shown in each of Figures 5 to 7 were obtained by having a subject utter the test phrase "derederedere..." and recording the audio.
[0060] In each of Figures 5 to 7, the white arrow indicates the position where the "re" (i.e., the "shot") is produced in the test phrase. As shown in Figure 5, in the spectrogram obtained from the speech waveform of a healthy subject, the "shot" sound is produced correctly at the position where it should be produced, and a vertically dark linear region B2 appears, indicating a very short-term drop in power. On the other hand, as shown in Figure 6, in the spectrogram obtained from the speech waveform of a stroke patient, as shown in region C2, for example, the long vertically dark linear region indicating a very short-term drop in power does not appear at the position where the "shot" sound should be produced, and the "shot" sound is not produced correctly. Similarly, in the spectrogram obtained from the speech waveform of another stroke patient shown in Figure 7, as shown in region C3, a relatively long-term drop in power occurs at the position where the "shot" sound should be produced, and the "shot" sound is also not produced correctly.
[0061] Furthermore, features that can detect the presence or absence of articulation abnormalities can appear not only in spectrograms obtained from speech waveforms, but also in RMS envelopes obtained from speech waveforms. Figure 8 shows examples of RMS envelopes obtained from the speech waveforms of a healthy subject and a stroke patient who produced multiple phrases. Figure 8(a) shows the RMS envelope obtained from the speech waveform of the healthy subject. Meanwhile, Figure 8(b), (c), and (d) all show RMS envelopes obtained from the speech waveform of a stroke patient. The RMS envelopes in Figure 8(a), (b), (c), and (d) were all obtained by having the subject produce the test phrase "derederedere..." and then recording the speech, and then performing appropriate information processing on the speech waveform obtained.
[0062] As shown in Figure 8(a), the RMS envelope obtained from the speech waveform of a healthy subject has a consistent envelope shape for each phrase, and a slight drop in power due to the correct pronunciation of the plucking sounds in the center of each phrase. On the other hand, as shown in Figure 8(b), the RMS envelope obtained from the speech waveform of a stroke patient has an inconsistent envelope shape for each phrase, and a steep drop in power due to the incorrect pronunciation of the plucking sounds in the center of each phrase. Similarly, the RMS envelope obtained from the speech waveform of another stroke patient, as shown in Figure 8(c), also has an inconsistent envelope shape for each phrase. Similarly, the RMS envelope obtained from the speech waveform of yet another stroke patient, as shown in Figure 8(d), has an inconsistent envelope shape for each phrase, and the intervals between phrases are also inconsistent.
[0063] As described above, by using "derederedere..." as the test phrase, features that indicate whether the plucking sound is being pronounced correctly are more likely to appear in both the spectrogram and RMS envelope obtained from the audio waveform.
[0064] The classification unit 12 classifies multiple phrases from the speech information acquired by the acquisition unit 11 (acquisition step). The classification unit 12 is the entity that executes the classification step in the articulation abnormality detection method. Specifically, the test phrase uttered by the subject is a speech "derederedere..." in which the phrase "dere" is repeated multiple times as described above, and therefore includes multiple phrases. The classification unit 12 classifies the multiple phrases "derederedere..." into phrases of "dere" one by one, making it easier for the detection unit 13, described later, to handle the speech information.
[0065] In the embodiment, the division unit 12 (division step) divides multiple phrases based on an RMS envelope or a spectrogram (here, a mel spectrogram) as audio information. Also, in the embodiment, the division unit 12 (division step) divides multiple phrases by inputting audio information acquired by the acquisition unit 11 (acquisition step) into the division model 17. The division model 17 is a trained model that has been machine-learned to divide multiple phrases using audio including multiple phrases as input.
[0066] Specifically, the segmentation model 17 is, for example, a deep neural network (DNN) model, which is a sequence labeling model. The segmentation model 17 receives an RMS envelope or spectrogram obtained from a speech waveform including multiple phrases as input, and outputs label data. The label data is a collection of binary information indicating whether each frame belongs to a phrase or not. For example, if RMS envelopes or spectrograms for 100 frames are obtained from a speech waveform, the label data is a collection of binary information for 100 frames.
[0067] The segmentation unit 12 generates and outputs segment information based on the label data output from the segmentation model 17. For example, if the label data is "11...100111...", successive "1"s represent phrases, and "0"s represent boundaries between adjacent phrases. Therefore, the segmentation unit 12 generates segment information including the start and end positions of each of the multiple phrases based on the label data.
[0068] A specific example of the learning phase of the segmented model 17 will be described below with reference to Fig. 9. Fig. 9 is a diagram showing an example of the learning phase for the segmented model 17 of the articulation abnormality detection device 100 according to the embodiment. First, the acquisition unit 11 executes appropriate information processing on the collected speech waveform to acquire an RMS envelope or a mel spectrogram from the speech waveform as speech information. In the example shown in Fig. 9, an example of a mel spectrogram is illustrated.
[0069] The RMS envelope obtained from the speech waveform has a dimension of "α" (α=1) and a frame count of "p" (p is a natural number). The mel spectrogram obtained from the speech waveform has a dimension of "β" (β is a natural number, β>1) and a frame count of "p". The dimension here indicates the power resolution along the frequency axis. The frame count here indicates the number of frames obtained by cutting out the speech waveform per unit time.
[0070] Next, the speech information acquired by the acquisition unit 11 is input to a segmented model 17 for which machine learning has not yet been completed (hereinafter referred to as the "incomplete segmented model 17"). As a result, the incomplete segmented model 17 outputs label data. This label data has a dimension number of "1" and a frame number of "p".
[0071] Then, the label data output by the incomplete segmentation model 17 and the correct answer data are input into a loss function (here, a Categorical Cross Entropy Error function), and backpropagation is performed so that the output of the loss function becomes the minimum value, thereby performing machine learning on the incomplete segmentation model 17 using supervised learning. The correct answer data is label data created in advance from speech waveforms obtained by having a healthy subject speak a test phrase. The correct answer data, like the label data output by the incomplete segmentation model 17, has the number of dimensions "1" and the number of frames "p".
[0072] A specific example of the inference phase using the segmented model 17 for which machine learning has been completed will be described below with reference to FIG. 10. FIG. 10 is a diagram showing an example of the inference phase using the segmented model 17 of the articulation abnormality detection device 100 according to the embodiment. First, the acquisition unit 11 performs appropriate information processing on the collected speech waveform to acquire an RMS envelope or a mel spectrogram from the speech waveform as speech information. The example shown in FIG. 10 illustrates an example of a mel spectrogram. Note that the number of frames of the RMS envelope and the mel spectrogram are the same as in the learning phase. Furthermore, the number of dimensions of the RMS envelope and the mel spectrogram are also the same as in the learning phase.
[0073] Next, the segmentation unit 12 inputs the speech information acquired by the acquisition unit 11 to the segmentation model 17. As a result, the segmentation model 17 outputs label data. Then, the segmentation unit 12 generates segment information including the start position and end position of each of the multiple phrases based on the label data output by the segmentation model 17. The segmentation information generated by the segmentation unit 12 is used by the detection unit 13, which will be described later.
[0074] The detection unit 13 detects the presence or absence of an articulation abnormality in the subject based on the output result obtained by inputting the speech information acquired by the acquisition unit 11 (acquisition step) into the detection model 18. The detection unit 13 is the entity that executes the detection step in the articulation abnormality detection method. In the embodiment, the detection unit 13 (detection step) inputs each of the multiple phrases divided by the division unit 12 (division step) into the detection model 18. That is, in the embodiment, the speech information acquired by the acquisition unit 11 (acquisition step) is not directly input into the detection model 18, but the divided multiple phrases are indirectly input into the detection model 18 as speech information.
[0075] The detection model 18 is a model trained by machine learning to receive speech as input and output information regarding the presence or absence of articulation abnormalities. Specifically, the detection model 18 is, for example, a Convolutional Neural Network (CNN) model, which is an autoencoder model trained by machine learning to restore speech identical to the input speech of a healthy individual. For example, the detection model 18 receives as input the RMS envelope or mel spectrogram of each of the multiple phrases segmented by the segmentation unit 12, attempts to restore these, and outputs the RMS envelope or mel spectrogram corresponding to each of the multiple phrases.
[0076] Then, the detection unit 13 (detection step) detects whether or not the subject has an articulation disorder based on the degree of deviation between the speech information input to the detection model 18 and the speech information output from the detection model 18. For example, when speech information about a healthy subject is input to the detection model 18, speech information that is almost identical to the input speech information is restored and output. In this case, the degree of deviation is relatively small. On the other hand, when speech information about a subject with an articulation disorder, such as a stroke patient, is input to the detection model 18, the detection model 18 cannot restore this speech information and outputs speech information that differs from the input speech information. In this case, the degree of deviation is relatively large.
[0077] Therefore, the detection unit 13 generates detection information regarding the presence or absence of an articulation abnormality of the subject based on the degree of discrepancy between the input data input to the detection model 18 and the output data output from the detection model 18. For example, the detection unit 13 calculates the mean squared error between the input data input to the detection model 18 and the output data output from the detection model 18. If the calculated mean squared error exceeds a threshold, the detection unit 13 detects that the subject has an articulation abnormality, and if the calculated mean squared error is equal to or less than the threshold, the detection unit 13 detects that the subject does not have an articulation abnormality and is a healthy person.
[0078] Next, a specific example of the learning phase of the detection model 18 will be described with reference to FIG. 11. FIG. 11 is a diagram showing an example of the learning phase for the detection model 18 of the speech disorder detection device 100 according to the embodiment. First, the acquisition unit 11 acquires a mel spectrogram as voice information from the voice waveform by performing appropriate information processing on the collected voice waveform.
[0079] The mel spectrogram obtained from the voice waveform has a dimensionality of “γ” (γ is a natural number and β≠γ), and the number of frames is “q” (q is a natural number and q≠p).
[0080] Next, the detection unit 13 generates segmented data composed of only a plurality of frames by segmenting the voice information acquired by the acquisition unit 11 into a plurality of frames by referring to the segmentation information output by the segmentation unit 12. The segmented data has a dimensionality of “γ”, and the number of frames of the segmented data is “r” (r is a natural number and r<q). In the segmented data generated here, since the lengths of the plurality of frames are not uniform, it is hereinafter referred to as “unshaped segmented data”. Next, by resizing the plurality of frames included in the segmented data, the lengths of the plurality of frames are unified. Hereinafter, the resized segmented data is simply referred to as “segmented data”. The segmented data, similar to the unshaped segmented data, has a dimensionality of “γ” and the number of frames is “r’”.
[0081] Next, the segmented data is input into a detection model 18 (hereinafter referred to as “uncompleted detection model 18”) for which machine learning has not yet been completed. As a result, the uncompleted detection model 18 outputs restored data that attempts to restore the input segmented data. This restored data, similar to the segmented data, has a dimensionality of “γ” and the number of frames is “r’”.
[0082] Then, the segmented data and the restored data output by the uncompleted detection model 18 are input into a loss function (here, the mean squared error function), and the error backpropagation method is executed so that the output of the loss function becomes the minimum value, thereby performing machine learning on the uncompleted detection model 18 by unsupervised learning.
[0083] A specific example of the inference phase using the detection model 18 for which machine learning has been completed will be described below with reference to Fig. 12. Fig. 12 is a diagram showing an example of the inference phase using the detection model 18 of the articulation abnormality detection device 100 according to the embodiment. First, the acquisition unit 11 performs appropriate information processing on the collected speech waveform to acquire a mel spectrogram from the speech waveform as speech information. In the example shown in Fig. 12, an example of a mel spectrogram is illustrated. Note that the number of dimensions and the number of frames of the mel spectrogram are both the same as those in the learning phase.
[0084] Next, the detection unit 13 generates unformatted segmented data by segmenting the audio information acquired by the acquisition unit 11 into a plurality of phrases by referring to the segmentation information output by the segmentation unit 12. Next, the detection unit 13 generates segmented data by resizing the plurality of phrases included in the segmented data.
[0085] Next, the detection unit 13 inputs the generated segmented data to the detection model 18. As a result, the detection model 18 outputs restored data. The detection unit 13 then calculates the mean square error between the segmented data input to the detection model 18 and the restored data output by the detection model 18, and generates detection information regarding the presence or absence of an articulation abnormality in the subject by comparing the calculated mean square error with a threshold. The detection information generated by the detection unit 13 is used by the output unit 14, which will be described later.
[0086] In the embodiment, the mel spectrogram obtained from the speech waveform is used as speech information in both the learning phase of the detection model 18 and the inference phase using the detection model 18, but the RMS envelope obtained from the speech waveform may also be used as speech information.
[0087] Furthermore, the detection unit 13 may input only part of the segmented data to the detection model 18, for example, by excluding the last phrase among multiple phrases included in the segmented data, rather than inputting all of the segmented data to the detection model 18. This is because there is a possibility that the subject will not reliably utter the test phrase to the end, and in such a case, the last phrase will become noise for the detection model 18.
[0088] The output unit 14 outputs detection information regarding the presence or absence of an articulation abnormality in the subject detected by the detection unit 13 (detection step). The output unit 14 is the entity that executes the output step in the articulation abnormality detection method. The detection information may include information indicating whether or not the subject has an articulation abnormality. In the embodiment, the detection information includes information indicating the presence or absence of a sign of the onset of a stroke in the subject, which is linked to the presence or absence of the articulation abnormality in the subject. The output unit 14 outputs the detection information, for example, by displaying a character string or an image indicating the detection information on the display of an information terminal.
[0089] The playback unit 15 plays back a sample voice of the voice uttered by the subject to the subject before the acquisition unit 11 acquires the voice information (before the acquisition step). The playback unit 15 is the entity that executes the playback step in the articulation abnormality detection method. The sample voice is, for example, a mechanical voice that reads out a test phrase at a predetermined volume and a predetermined rhythm. The playback unit 15 plays back the sample voice from a speaker mounted on the information terminal, for example, when triggered by the subject performing a predetermined operation on the information terminal.
[0090] The memory unit 16 is a storage device that stores information (such as computer programs) necessary for the acquisition unit 11, classification unit 12, detection unit 13, output unit 14, and reproduction unit 15 to perform various processes. The memory unit 16 is realized by, for example, a semiconductor memory, but is not particularly limited and any known means for storing electronic information can be used. The memory unit 16 stores a classification model 17 used by the classification unit 12 and a detection model 18 used by the detection unit 13.
[0091] [3. Operation] An example of the operation of the articulation abnormality detection device 100 according to the embodiment (i.e., an articulation abnormality detection method) will be described below with reference to Figs. 13 to 15. Fig. 13 is a flowchart showing an example of the operation of the articulation abnormality detection device 100 according to the embodiment. Fig. 14 is a diagram showing an example of an overview of the articulation abnormality detection device 100 and the articulation abnormality detection method according to the embodiment. Fig. 15 is a diagram showing a specific example of the operation of the articulation abnormality detection device 100 according to the embodiment.
[0092] In the following description, it is assumed that the segmentation model 17 and the detection model 18 have both been machine-learned in advance by the methods already described, as shown in Fig. 14. Also, in the following description, it is assumed that subject 2 is a patient who has previously suffered a stroke and has at present recovered, albeit not completely, from the stroke, with a mild case. Of course, subject 2 may also be a person who has never suffered a stroke in the past.
[0093] (a) to (d) of Figure 15 all show the execution flow of an application called "Stroke Recurrence Checker" on the information terminal 3. (a) of Figure 15 shows the image displayed on the display 31 of the information terminal 3 when the application is launched. In the center of the display 31, an icon 41 containing the character string "Check with words" is displayed. When the subject 2 selects the icon 41 by touching it with his / her finger, the process moves to the flow shown in (b) of Figure 15.
[0094] 15(b), a character string M1 prompting the subject 2 to speak the test phrase, "Please speak like this," and a character string M2 indicating the test phrase, "derederederederederederederedere," are displayed on the display 31 of the information terminal 3. In addition, an icon 42 including the character string "listen to the example" and an icon 43 including the character string "start check" are displayed on the display 31 together with the character strings M1 and M2.
[0095] Here, the operation of the subject 2 selecting the icon 42 corresponds to the "playback trigger" shown in FIG. 13. That is, when the subject 2 performs the operation of selecting the icon 42, in other words, when a playback trigger is received (S1: Yes), the playback unit 15 (playback step) plays back the sample voice (S2). The timing of displaying the icon 42 on the display 31 is not limited to before acquiring the voice information, but may be after acquiring the voice information. For example, the icon 42 may be displayed on the display 31 when the test phrase of the subject 2 cannot be detected for some reason, such as when the volume of the voice uttered by the subject 2 is low. Furthermore, for example, the icon 42 may be displayed on the display 31 when the process of classifying multiple phrases from the voice information in step S4, which will be described later, cannot be executed. Furthermore, for example, the icon 42 may be displayed on the display 31 when the process of detecting the presence or absence of an articulation abnormality in step S5, which will be described later, cannot be executed.
[0096] If the subject 2 does not perform an operation to select the icon 42 (S2: No), or if the subject 2 performs an operation to select the icon 42 and then performs an operation to select the icon 43, the process proceeds to the flow shown in FIG. 15(c). Note that the icon 43 may be configured to accept an operation by the subject 2 (i.e., become active) after the subject 2 performs an operation to select the icon 42 and plays the sample audio. In this case, the process cannot proceed to the flow shown in FIG. 15(c) until the subject 2 listens to the sample audio. The icon 43 may be displayed in a manner indicating that it is inactive, for example, by being displayed in gray, until the sample audio is played, and may be displayed in a manner indicating that it is active, for example, by being displayed in white, once the sample audio is played.
[0097] As shown in (c) of Fig. 15, the display 31 of the information terminal 3 continues to display the character strings M1 and M2. Also, the display 31 displays a sub-image 5 indicating that the test phrase uttered by the subject 2 is being recorded, and an icon 44 including the character string "Judgment", together with the character strings M1 and M2. The sub-image 5 displays the character string "Recording" and a voice waveform picked up by the microphone of the information terminal 3. That is, in the flow shown in (c) of Fig. 15, the acquisition unit 11 (acquisition step) acquires voice information (S3).
[0098] Next, when the subject 2 performs an operation to select the icon 44, a series of processes for determining (detecting) the presence or absence of an articulation abnormality in the subject 2 is initiated. First, the division unit 12 (division step) divides a plurality of phrases from the speech information acquired by the acquisition unit 11 (acquisition step) (S4). Next, the detection unit 13 (detection step) inputs each of the plurality of phrases divided by the division unit 12 (division step) into the detection model 18, thereby detecting the presence or absence of an articulation abnormality in the subject 2 (S5). Then, the output unit 14 outputs detection information regarding the presence or absence of an articulation abnormality in the subject 2 detected by the detection unit 13 (detection step) (S6). Specifically, as shown in (d) of FIG. 15, the detection information is displayed on the display 31 of the information terminal 3 as a character string M3. Here, if an articulation abnormality is detected in the subject 2, in other words, if the subject 2 has a sign of the onset of a stroke, the character string M3 reading "There is a possibility that a stroke has recurred. We recommend that you consult a specialist" is displayed as detection information. If subject 2 does not have any articulation abnormalities, in other words, if subject 2 has no signs of a stroke, a string such as "There are no particular abnormalities" will be displayed on display 31.
[0099] Alternatively, the detected information may be displayed on the display 31 of the information terminal 3 in a form such as that shown in Fig. 16. Fig. 16 is a diagram showing another specific example of the operation of the articulation abnormality detection device 100 according to the embodiment.
[0100] In the example shown in (a) of FIG. 16, the detection information is displayed on the display 31 as a character string M3 and a first graph 6. The first graph 6 represents an RMS envelope obtained from the speech waveform of the subject 2, and includes a failure section 61 where the subject 2 failed to pronounce a phrase correctly (in other words, an articulation disorder was observed). By looking at the first graph 6, the subject 2 can understand which phrases he or she failed to pronounce correctly.
[0101] In the example shown in FIG. 16(b), the detection information is displayed on the display 31 as a character string M3, a first graph 6, and a character string M4 saying "The failure rate is 38%." The character string M4 indicates the proportion of the failure intervals 61 to the total intervals in which the subject 2 uttered speech (i.e., the failure rate). By looking at the character string M4, the subject 2 can understand how likely it is that the stroke has recurred.
[0102] In the example shown in FIG. 16(c), the detection information is displayed on the display 31 as a character string M3 and a second graph 7. The second graph 7 is a bar graph that shows the failure rate over time. Here, the second graph 7 shows the results of running the "Stroke Recurrence Checker" every day from August 1 to August 11. The horizontal line 71 in the second graph 7 represents a threshold value, and if the failure rate exceeds this threshold value, it indicates that there is a high possibility that stroke has recurred. By looking at the second graph 7, the subject 2 can understand the degree of possibility of stroke recurrence over time.
[0103] As described above, the articulation abnormality detection device 100 and the articulation abnormality detection method according to the embodiments make it possible to detect the presence or absence of articulation abnormalities, and even the presence or absence of signs of stroke, from the speech of the subject 2 without relying on specialists such as doctors or speech-language-hearing therapists. Therefore, by using the articulation abnormality detection device 100 and the articulation abnormality detection method according to the embodiments, if there are signs of stroke in the subject 2, it is expected that the subject 2 can be urged to seek medical attention promptly, thereby preventing the condition from becoming severe through early treatment.
[0104] [4. Effects, etc.] As described above, the articulation abnormality detection method according to the embodiment includes an acquisition step (S3) and a detection step (S5). In the acquisition step (S3), speech information about the speech uttered by the subject is acquired. In the detection step (S5), the presence or absence of an articulation abnormality of the subject is detected based on the output result obtained by inputting the speech information acquired in the acquisition step (S5) into detection model 18, which has been machine-learned to input speech and output information about the presence or absence of an articulation abnormality.
[0105] This method makes it possible to detect the presence or absence of articulation abnormalities in a subject simply by the subject uttering a voice. This has the advantage that it is easier to detect the presence or absence of articulation abnormalities in a subject without imposing a burden on the subject, compared to when the subject is asked to take a video of their own face with a video camera.
[0106] In the articulation abnormality detection method according to the embodiment, the speech information includes a specific sound that is produced when the subject moves his / her tongue in a predetermined pattern.
[0107] This has the advantage that it is easier to detect the degree of tongue paralysis, which can be an indicator of whether or not a subject has an articulation disorder, compared to when the audio information does not contain specific sounds.
[0108] In the articulation abnormality detection method according to the embodiment, the specific sound is a popping sound.
[0109] This has the advantage that by including in the specific sounds popping sounds that are difficult to produce when the tongue is paralyzed, it becomes easier to detect whether or not the subject has an articulation disorder.
[0110] In the articulation abnormality detection method according to the embodiment, the speech information includes a phrase in which a specific sound and a plosive sound are consecutively included.
[0111] This has the advantage that by making a plosive sound, whose position is easy to identify in the speech of the subject, consecutive to a specific sound, it becomes easier to identify the position of the specific sound in the speech of the subject, making it easier to detect whether or not the subject has an articulation abnormality.
[0112] In the articulation abnormality detection method according to the embodiment, the speech information includes a plurality of phrases. The articulation abnormality detection method according to the embodiment further includes a division step (S4) of dividing the speech information acquired in the acquisition step (S3) into a plurality of phrases. In the detection step (S5), each of the plurality of phrases divided in the division step (S4) is input to a detection model 18.
[0113] This has the advantage that it becomes easier to detect whether or not a subject has an articulation disorder compared to detecting whether or not a subject has an articulation disorder from a single phrase.
[0114] In the articulation abnormality detection method according to the embodiment, in the dividing step (S4), a plurality of phrases are divided based on an RMS (Root Mean Square) envelope or a spectrogram as speech information.
[0115] This has the advantage that the accuracy of distinguishing between multiple phrases can be expected to improve, since features that can distinguish between multiple phrases are likely to appear in the RMS envelope or spectrogram.
[0116] In addition, in the articulation abnormality detection method according to the embodiment, in the division step (S4), the speech information acquired in the acquisition step (S3) is input to a division model 17 that has been machine-trained to divide multiple phrases using speech containing multiple phrases as input, thereby dividing the multiple phrases.
[0117] This has the advantage that, compared to when multiple phrases are segmented without using the segmentation model 17, it is possible to expect an improvement in the accuracy of segmenting multiple phrases.
[0118] In the articulation abnormality detection method according to the embodiment, the detection model 18 is an autoencoder that has been machine-learned to restore a speech identical to the speech of a healthy individual that has been input. In the detection step (S5), the presence or absence of an articulation abnormality in the subject is detected based on the degree of deviation between the speech information input to the detection model 18 and the speech information output from the detection model 18.
[0119] This has the advantage that it is easier to prepare a large amount of training data compared to training the detection model 18 using the speech of patients with articulation disorders, who are fewer in number than healthy individuals, making it easier to train the detection model 18.
[0120] Moreover, the articulation abnormality detection method according to the embodiment further includes an output step (S6) of outputting detection information regarding the presence or absence of articulation abnormality of the subject detected in the detection step (S5).
[0121] This has the advantage that, for example, by outputting the detection information to the subject, the subject can know whether or not he or she has an articulation disorder.
[0122] Moreover, the articulation abnormality detection method according to the embodiment further includes, before the acquisition step (S3), a reproduction step (S2) of reproducing a sample speech of speech uttered by the subject to the subject.
[0123] This has the advantage that it becomes possible to detect whether or not the subject has articulation abnormalities, including whether or not the subject is able to reproduce and produce the sample voice, and that it is expected that the accuracy of detecting whether or not the subject has articulation abnormalities can be improved.
[0124] The articulation abnormality detection device 100 according to the embodiment also includes an acquisition unit 11 and a detection unit 13. The acquisition unit 11 acquires speech information related to speech uttered by the subject. The detection unit 13 detects the presence or absence of an articulation abnormality of the subject based on an output result obtained by inputting the speech information acquired by the acquisition unit 11 into a detection model 18 that has been machine-learned to input speech and output information related to the presence or absence of an articulation abnormality.
[0125] This method makes it possible to detect the presence or absence of articulation abnormalities in a subject simply by the subject uttering a voice. This has the advantage that it is easier to detect the presence or absence of articulation abnormalities in a subject without imposing a burden on the subject, compared to when the subject is asked to take a video of their own face with a video camera.
[0126] (Other embodiments) While the articulation abnormality detection method and articulation abnormality detection device 100 according to one or more aspects of the present disclosure have been described based on the above-mentioned embodiments, the present disclosure is not limited to these embodiments. As long as they do not deviate from the gist of the present disclosure, various modifications conceivable by those skilled in the art and configurations formed by combining components of different embodiments may also be included within the scope of one or more aspects of the present disclosure.
[0127] For example, in the above embodiment, the division unit 12 (division step) divides multiple phrases using the division model 17, but this is not limited to this. For example, the division unit 12 (division step) may divide multiple phrases so as to separate them at points where the power in the RMS envelope obtained from the subject's speech waveform is equal to or less than a predetermined value. In this case, the division model 17 is not necessary.
[0128] For example, in the above embodiment, a plurality of phrases are used as the test phrases to be uttered by the subject (i.e., the speech information acquired by the acquisition unit 11), but a single phrase may be used instead. In this case, the division unit 12 (division step) is not required.
[0129] Furthermore, in the above embodiment, "derederedere..." is used as the test phrase to be uttered by the subject (i.e., the audio information acquired by the acquisition unit 11), but this is not limited to this and the test phrase may be a phrase consisting of a succession of plosive sounds and popping sounds. Furthermore, the test phrase is not limited to a phrase consisting of a succession of plosive sounds and popping sounds, but may be, for example, a phrase consisting of only popping sounds. Furthermore, depending on the learning method of the detection model 18, the test phrase may not include a popping sound, and further may not include a specific sound produced by moving the tongue in a predetermined pattern.
[0130] In the above embodiment, the articulation abnormality detection device 100 is mounted on an information terminal, but this is not limiting. For example, the articulation abnormality detection device 100 may be mounted on a server device. The server device may be a cloud server or a local server. In this case, the articulation abnormality detection device 100 is realized by a processor mounted on the server device executing a predetermined program. In this case, the subject may access the server device via a network or the like using an information terminal. For example, the articulation abnormality detection device 100 may be configured such that a portion of its configuration is mounted on an information terminal and the remaining portion is mounted on a server device.
[0131] Furthermore, the articulation abnormality detection device 100 may be stored in a dedicated terminal device having an articulation abnormality detection function, rather than in a general-purpose information terminal such as a smartphone or tablet terminal. In this case, the articulation abnormality detection device 100 is realized by a processor installed in the dedicated terminal device executing a predetermined program.
[0132] For example, some or all of the components included in the articulation abnormality detection device 100 according to the above embodiment may be configured as a single system LSI (Large Scale Integration). The system LSI is an ultra-multifunctional LSI manufactured by integrating multiple components on a single chip, and specifically, is a computer system configured to include a microprocessor, a ROM (Read Only Memory), a RAM (Random Access Memory), etc. A computer program is stored in the ROM. The system LSI achieves its functions when the microprocessor operates in accordance with the computer program.
[0133] Although we refer to it as a system LSI here, it may also be called an IC (Integrated Circuit), LSI, super LSI, or ultra LSI depending on the level of integration. Furthermore, the method of integration is not limited to LSI, but may be realized using dedicated circuits or general-purpose processors. It is also possible to use FPGAs (Field Programmable Gate Arrays), which can be programmed after LSI manufacturing, or reconfigurable processors, which allow the connections and settings of circuit cells within LSI to be reconfigured.
[0134] Furthermore, if an integrated circuit technology that can replace LSI emerges due to advances in semiconductor technology or other derivative technologies, it is natural that such technology may be used to integrate functional blocks. The application of biotechnology, etc. is also a possibility.
[0135] Another aspect of the present disclosure may be a computer program that causes a computer to execute each of the characteristic steps included in the articulation abnormality detection method. Another aspect of the present disclosure may be a computer-readable non-transitory recording medium on which such a computer program is recorded. That is, the program may cause one or more processors to execute the above-mentioned articulation abnormality detection method.
[0136] This method makes it possible to detect the presence or absence of articulation abnormalities in a subject simply by the subject uttering a voice. This has the advantage that it is easier to detect the presence or absence of articulation abnormalities in a subject without imposing a burden on the subject, compared to when the subject is asked to take a video of their own face with a video camera. [Industrial Applicability]
[0137] The present disclosure is applicable to, for example, a method for determining whether or not there is a sign of the onset of stroke. [Explanation of symbols]
[0138] 100 Articulation abnormality detection device 11 Acquisition Department 12 Section 13 Detector 14 Output section 15 Playback Department 16 Memory section 17 Segment Model 18 Detection Model 2 Subjects 3. Information terminals 31 Display 41~44 Icon 5 Sub-images 6 Graph 1 61 Failed Section 7 Graph 2 71 Horizontal Line A1, A2, B1, B2, C1~C3 area M1~M4 string
Claims
1. A method for detecting articulation abnormalities executed by a processor, comprising: an acquisition step of acquiring speech information relating to speech uttered by the subject; a detection step of detecting the presence or absence of an articulation disorder of the subject based on an output result obtained by inputting the speech information acquired in the acquisition step into a detection model that has been machine-learned to input speech and output information regarding the presence or absence of an articulation disorder, the audio information includes a plurality of phrases each including a succession of specific sounds and plosive sounds uttered by the subject moving his / her tongue in a predetermined pattern; a dividing step of dividing the plurality of phrases based on the audio information acquired in the acquiring step, based on the positions of the plosives; In the detection step, each of the plurality of phrases segmented in the segmentation step is input to the detection model. Articulation abnormality detection method.
2. The specific sound is a bullet sound. The method for detecting articulation abnormalities according to claim 1 .
3. In the detection step, a root mean square (RMS) envelope or spectrogram of each of the plurality of phrases divided in the division step is input to the detection model. The method for detecting articulation abnormalities according to claim 2 .
4. In the dividing step, the plurality of phrases are divided based on an RMS (Root Mean Square) envelope or a spectrogram as the audio information. The method for detecting articulation abnormalities according to any one of claims 1 to 3.
5. In the segmentation step, the speech information acquired in the acquisition step is input to a segmentation model that has been trained by machine learning to segment the plurality of phrases using speech including the plurality of phrases as input, thereby segmenting the plurality of phrases. The method for detecting articulation abnormalities according to any one of claims 1 to 4.
6. The detection model is an autoencoder model trained by machine learning to restore a speech identical to an input speech of a healthy subject, In the detection step, presence or absence of an articulation disorder of the subject is detected based on a degree of deviation between the speech information input to the detection model and the speech information output from the detection model. The method for detecting articulation abnormalities according to any one of claims 1 to 5.
7. an output step of outputting detection information regarding the presence or absence of an articulation abnormality of the subject detected in the detection step; The method for detecting articulation abnormalities according to any one of claims 1 to 6.
8. In the output step, a failure rate indicating a proportion of a failure section in which the subject fails to utter the phrase to all sections of the speech uttered by the subject is output as the detection information. The method for detecting articulation abnormalities according to claim 7.
9. a playback step of playing back to the subject a sample voice of the voice uttered by the subject; The method for detecting articulation abnormalities according to any one of claims 1 to 8.
10. one or more processors, 10. A method for detecting articulation abnormalities according to claim 1, program.
11. an acquisition unit that acquires speech information related to speech uttered by a subject; a detection unit that detects the presence or absence of articulation abnormalities of the subject based on an output result obtained by inputting the speech information acquired by the acquisition unit into a detection model that has been machine-learned to input speech and output information regarding the presence or absence of articulation abnormalities, the audio information includes a plurality of phrases each including a succession of specific sounds and plosive sounds uttered by the subject moving his / her tongue in a predetermined pattern; a division unit that divides the plurality of phrases based on the audio information acquired by the acquisition unit at the positions of the plosives, Each of the plurality of phrases segmented by the segmentation unit is input to the detection model. Articulation abnormality detection device.
Citation Information
Patent Citations
Syllable segmentation method containing initial consonant and device thereof
CN105976811A
Multi-mode interactive speech language function disorder evaluation system and method
CN107456208A
Automatic dysarthria evaluation system and method based on speech recognition
CN112927696A
Phoneme segmentation apparatus
JP1986020099A
Voice data creation method, storage device, integrated circuit device, and voice reproduction system
JP2010048931A