Diagnostic support system, diagnostic support method, and program

The diagnostic assistance system simplifies and enhances the accuracy of dysarthria detection by analyzing speech data and maxillofacial morphology, effectively differentiating between organic and functional dysarthria types.

JP2026048377APending Publication Date: 2026-03-17HIROSHIMA CITY UNIVERSITY +5
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-05
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing methods for diagnosing dysarthria are complex and lack accuracy in detecting the onset of the disorder.

Method used

A diagnostic assistance system that acquires speech data, calculates index acoustic features using Mel-frequency cepstrum coefficients, and compares them with a normal range to determine the presence and type of dysarthria, incorporating maxillofacial morphology analysis.

Benefits of technology

Enables simple and accurate detection of dysarthria, distinguishing between morphological and functional dysarthria through acoustic and morphological assessments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026048377000001_ABST
    Figure 2026048377000001_ABST
Patent Text Reader

Abstract

To provide a diagnostic support system, diagnostic support method, and program that can detect the onset of articulation disorders more simply and accurately. [Solution] The diagnostic support system 1A generates information to assist in the diagnosis of articulation disorders. The acquisition unit 30 acquires speech data including the subject's single-syllable vocalizations. The calculation unit 31 calculates information indicating index acoustic features that show a significant difference between healthy individuals and individuals with articulation disorders, based on the speech data acquired by the acquisition unit 30. The storage unit 32 stores the normal range of the index acoustic features calculated by statistical calculation based on the speech data of healthy individuals' single-syllable vocalizations. The display unit 21 displays the information indicating the subject's index acoustic features calculated by the calculation unit 31 and the normal range stored in the storage unit 32 in a comparable manner.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] [[ID=,4]] The present invention relates to a diagnostic assistance system, a diagnostic assistance method, and a program.

Background Art

[0002] J Dysarthria is a disorder in which the ability to pronounce words accurately and clearly is impaired. Dysarthria is broadly classified into morphological dysarthria caused by organic abnormalities such as morphological abnormalities and paralysis, and functional dysarthria that occurs without organic abnormalities due to a delay in articulation itself or incorrect habits.

[0003] [[ID=,16]]In general, the diagnosis of dysarthria mainly involves evaluation by a doctor and brain function tests. In addition, a dysarthria detection device that detects the onset of dysarthria based on the voice sounds uttered by a subject has been disclosed (see, for example, Patent Document 1).

Prior Art Document

Patent Document

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] There is a demand for the emergence of a system or the like that can detect the onset of dysarthria more simply and accurately.

[0006] The present invention has been made under the above circumstances, and an object thereof is to provide a diagnostic assistance system, a diagnostic assistance method, and a program that can detect the onset of dysarthria more simply and accurately.

Means for Solving the Problems

[0007] To achieve the above object, a diagnostic assistance system according to a first aspect of the present invention is A diagnostic support system that generates information to assist in the diagnosis of speech disorders, An acquisition unit that acquires speech data including monosyllabic sounds of the subject, Based on the voice data acquired by the acquisition unit, a calculation unit calculates information indicating index acoustic features that show a significant difference between healthy individuals and individuals with speech disorders among the acoustic features of the voice data. A storage unit that stores the normal range of the index acoustic feature, calculated by statistical calculation based on the speech data of single-syllable vocalizations of healthy individuals, A display unit that displays information indicating the subject's index acoustic feature quantity calculated by the calculation unit and the normal range stored in the storage unit in a comparable manner, It is equipped with.

[0008] The acquisition unit acquires information showing the maxillofacial morphology of the subject, The display unit simultaneously displays the subject's index acoustic features and normal range, as well as information indicating the subject's maxillofacial morphology acquired by the acquisition unit. It would be acceptable to do so.

[0009] The system includes a determination unit that determines if the subject's index acoustic characteristic quantity is not within the normal range, and determines that the subject has a speech disorder. It would be acceptable to do so.

[0010] The acquisition unit acquires information showing the maxillofacial morphology of the subject, If the determination unit determines that the subject has developed a speech disorder, it will determine, based on the information indicating the subject's maxillofacial morphology acquired by the acquisition unit, whether the speech disorder the subject has developed is a morphological speech disorder accompanied by an organic abnormality or a functional speech disorder without an organic abnormality. It would be acceptable to do so.

[0011] The index acoustic features used to diagnose speech disorders include Mel-frequency cepstrum coefficients. It would be acceptable to do so.

[0012] The Mel-frequency cepstrum coefficients include at least one of MFCC(4) or MFCC(12). It would be acceptable to do so.

[0013] A diagnostic assistance method according to a second aspect of the present invention is: A diagnostic assistance method performed by an information processing device that generates information to assist in the diagnosis of speech disorders, Acquisition step to obtain speech data including monosyllabic vocalizations of the subject, Based on the voice data acquired in the acquisition step, a calculation step is performed to calculate information indicating index acoustic features that show a significant difference between healthy individuals and individuals with speech disorders among the acoustic features of the voice data. A display step that displays information showing the subject's index acoustic feature calculated in the calculation step and the normal range of the index acoustic feature calculated by statistical calculation based on the speech data of a healthy person's single-syllable vocalization, in a way that allows for comparison. Includes.

[0014] A program relating to the third aspect of the present invention is: A computer that generates information to assist in the diagnosis of speech disorders, Acquisition unit that acquires speech data including monosyllabic sounds of the subject. A calculation unit calculates information indicating index acoustic features that show a significant difference between healthy individuals and individuals with speech disorders, based on the audio data acquired by the acquisition unit. A storage unit that stores the normal range of the index acoustic feature, calculated by statistical calculation based on the speech data of single-syllable vocalizations of healthy individuals. A display unit that displays information indicating the subject's index acoustic feature quantity calculated by the calculation unit and the normal range stored in the storage unit in a comparable manner. To make it function as such. [Effects of the Invention]

[0015] According to the present invention, the onset of articulation disorders can be detected in a simple and accurate manner. [Brief explanation of the drawing]

[0016] [Figure 1] It is a block diagram showing the functional configuration of the diagnostic support system according to Embodiment 1 of the present invention. [Figure 2] It is a table showing an example of an index acoustic feature amount having a significant difference between a healthy person and a person with dysarthria. [Figure 3] It is a block diagram showing the configuration of the arithmetic unit in FIG. 1. [Figure 4] (A) and (B) are diagrams showing display examples of the index acoustic feature amount. [Figure 5] It is a table showing an example of an index acoustic feature amount having a significant difference between types of the jaw and facial morphology of a subject. [Figure 6] It is a block diagram showing the hardware configuration of the diagnostic support system in FIG. 1. [Figure 7] It is a flowchart of data registration processing. [Figure 8] It is a flowchart of data display processing. [Figure 9] It is a block diagram showing the configuration of the diagnostic support system according to Embodiment 2 of the present invention.

Embodiments for Carrying Out the Invention

[0017] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. In each drawing, the same or equivalent parts are denoted by the same reference numerals. In the following embodiments, expressions such as "having", "including", or "containing" also include the meaning of "consisting of" or "composed of".

[0018] Embodiment 1 First, Embodiment 1 of the present invention will be described. The diagnostic support system 1A according to this embodiment, shown in Figure 1, generates information to assist in the diagnosis of articulation disorders in a subject. Articulation disorders refer to a condition in which a person is unable to speak properly due to damage to parts that play an important role in producing sound, such as the mouth, tongue, and vocal cords. The subject is the person to be diagnosed with an articulation disorder, and the generated information is provided to the dentist who diagnoses the subject.

[0019] As shown in Figure 1, the diagnostic support system 1A comprises a terminal device 2 and a server device 3 as an information processing device. The terminal device 2 may be a mobile terminal or smartphone, or a tablet, wearable device, or personal computer. The server device 3 is a computer that is connected to the terminal device 2 via a communication network (not shown). The server device 3 can connect to multiple terminal devices 2 via a communication network (not shown).

[0020] [Terminal device: Voice input unit] Terminal device 2 comprises an audio input unit 20 and a display unit 21. The audio input unit 20 is, for example, a microphone and inputs audio data including the subject's voice. The subject pronounces a pre-specified type of monosyllabic sound, such as Japanese monosyllabic sounds like "o", "sa", or "so". A monosyllabic sound is a syllable consisting of one vowel as the syllable root, either alone or accompanied by one or more consonants before or after it. The audio input unit 20 inputs audio data including this monosyllabic sound.

[0021] The voice input unit 20 transmits the input voice data to the server device 3. During transmission, the voice data is supplemented with header information, which includes the identification information of the subject who spoke and information indicating the type of monosyllabic speech.

[0022] In practice, the subject is asked to speak multiple times, and the voice input unit 20 inputs voice data containing the spoken sound each time. By obtaining multiple voice data sets, the acoustic features of each voice data set can be determined, and a representative value (e.g., the mean) of those acoustic features can be generated. There is no particular limit to the number of times the subject speaks, but it can be, for example, around 20 times. It is desirable that the number of times the subject speaks be sufficient to obtain statistical information on the features of the voice data of the spoken sound.

[0023] The display unit 21 displays information regarding the acoustic features of the audio data, including the subject's voice, transmitted from the server device 3. The operation of the display unit 21 will be described later.

[0024] [Server equipment] The server device 3 comprises an acquisition unit 30, a calculation unit 31, a storage unit 32, and a data reading unit 33. The acquisition unit 30 receives input from the voice input unit 20 of the terminal device 2 and acquires voice data, including the single-syllable vocalizations of the subject transmitted from the terminal device 2.

[0025] The calculation unit 31 calculates information indicating index acoustic features that show a significant difference between healthy individuals and individuals with speech disorders, based on the speech data acquired by the acquisition unit 30. Figure 2 shows a table illustrating an example of index acoustic features that show a significant difference between healthy individuals and individuals with speech disorders. As shown in Figure 2, an example of an index acoustic feature that shows a significant difference between healthy individuals and individuals with speech disorders is the Mel-Frequency Cepstrum Coefficient (MFCC). In Figure 2, MFCC(4), MFCC(10), MFCC(12), MFCC(13), MFCC(15), and MFCC(18) are shown as examples of Mel-Frequency Cepstrum Coefficients in the column direction. Note that in Figure 2, parentheses such as MFCC(4) are omitted.

[0026] In the table shown in Figure 2, the types of monosyllables (ANOVA) are shown in the row direction. For example, "ka" represents the monosyllable sound (ka) with the vowel "a" and the consonant "k". Other monosyllables are represented in a similar manner. Also in Figure 2, * indicates that the p-value based on the null hypothesis that there is no significant difference between healthy individuals and those with articulation disorders is less than 0.05. ** indicates that the p-value is less than 0.01. *** indicates that the p-value is less than 0.001. As shown in the table in Figure 2, for example, for MFCC(4), it has been revealed that there is a significant difference between healthy individuals and those with articulation disorders for the monosyllables ki, ku, ke, se, ti, tu, te, ri, ru, and re. Similarly, for MFCC(10), MFCC(12), MFCC(13), MFCC(15), and MFCC(18), it has been revealed that there is a significant difference between healthy individuals and those with articulation disorders for specific monosyllables.

[0027] Diagnostic support system 1A utilizes the relationship between indicator acoustic features and single syllables, which show a significant difference between healthy individuals and those with articulation disorders. It has subjects pronounce such single-syllable sounds, calculates the indicator acoustic features of the resulting speech data, and detects the presence or absence of articulation disorders. Below, the indicator acoustic features are described as Mel-frequency cepstrum coefficients, but are not limited to this.

[0028] As shown in Figure 3, the calculation unit 31 includes a Fourier transform unit 10, a Mel-banding unit 11, a logarithmic unit 12, a discrete cosine transform unit 13, a normalization unit 14, and a coefficient extraction unit 15. The Fourier transform unit 10 performs a Fourier transform on the audio data acquired by the acquisition unit 30. The audio data is subjected to a Fast Fourier Transform, and the spectrum of the audio data is calculated. The Mel-banding unit 11 performs a Mel-banding on the calculated spectrum of the audio data and calculates the Mel-band spectrum. The logarithmic unit 12 logarithms the calculated Mel-band spectrum and calculates the Mel-logarithmic spectrum. The discrete cosine transform unit 13 performs a discrete cosine transform on the calculated Mel-logarithmic spectrum and calculates the Mel-frequency cepstrum. The normalization unit 14 normalizes the calculated Mel-frequency cepstrum and calculates the normalized Mel-frequency cepstrum. The coefficient extraction unit 15 extracts and outputs Mel-frequency cepstrum coefficients (MFCC(1), MFCC(2), ...) from the calculated normalized Mel-frequency cepstrum.

[0029] The subject makes multiple vocalizations, and each time, vocal data is input to the vocal input unit 20, the vocal data is acquired by the acquisition unit 30, and the index acoustic features of the vocal data are calculated by the calculation unit 31. The calculation unit 31 calculates the average value of the index acoustic features of the multiple vocal data as a representative value. The calculated representative value of the index acoustic features is stored in the storage unit 32 as information indicating the index acoustic features.

[0030] Returning to Figure 1, the memory unit 32 stores information indicating the index acoustic features of speech data including the subject's monosyllabic vocalizations, as well as the normal range of these index acoustic features in healthy individuals. This normal range is data calculated in advance by statistical calculations based on the speech data of healthy individuals' monosyllabic vocalizations. The normal range is defined by an upper limit and a lower limit.

[0031] The data reading unit 33 reads from the storage unit 32 information indicating the index acoustic feature of the speech data, including the subject's monosyllabic vocalizations, along with the normal range of the index acoustic feature. The data reading unit 33 transmits the read normal range of the index acoustic feature and the information indicating the index acoustic feature of the speech data of the subject's monosyllabic vocalizations to the terminal device 2.

[0032] [Terminal device: Display unit] The display unit 21 of the terminal device 2 displays information indicating the index acoustic feature quantity of the speech data, including the subject's monosyllabic vocalizations, transmitted from the server device 3, and the normal range of the index acoustic feature quantity stored in the storage unit 32, in a manner that allows for comparison.

[0033] Figures 4(A) and 4(B) show examples of graphs displayed by the display unit 21. The graph in Figure 4(A) displays MFCC(4), MFCC(5), MFCC(8), MFCC(10)~(13), MFCC(15), MFCC(17), and MFCC(18) as information indicating the index acoustic features of the subject's speech data, including monosyllabic vocalizations. This graph also displays the upper and lower limits of the normal range for the index acoustic features. This allows for confirmation of whether the information indicating the index acoustic features falls within the normal range. As shown in Figure 4(A), MFCC(11) and (12) are below the lower limit, while MFCC(15) and MFCC(18) are above the upper limit. A dentist who sees this display can confirm that the subject may be suffering from a speech disorder. On the other hand, in the graph shown in Figure 4(B), MFCC(8) is above the upper limit. A dentist who sees this display can confirm that the subject may be suffering from a speech impediment.

[0034] [Measuring device] Returning to Figure 1, the server device 3 is connected to the measuring device 4. The measuring device 4 is a device that measures the maxillofacial morphology of a subject, for example, by cephalogram analysis. The measuring device 4 acquires a standardized X-ray image of the subject's head. Based on this image, the measuring device 4 calculates parameters that represent the subject's maxillofacial morphology. A dentist who sees these parameters can diagnose, for example, whether the subject's maxillofacial morphology is normal, whether the molars are in contact but the front teeth are not (open bite), or whether the upper front teeth overlap the lower front teeth when biting (deep bite). The parameters representing the subject's maxillofacial morphology measured by the measuring device 4 are transmitted to the server device 3.

[0035] The acquisition unit 30 of the server device 3 acquires information indicating the maxillofacial morphology of the subject from the measurement device 4. The acquisition unit 30 stores the information indicating the maxillofacial morphology of the subject in the storage unit 32. The data reading unit 33 transmits the subject's index acoustic feature quantity, the normal range of said index acoustic feature quantity, and the information indicating the maxillofacial morphology of the subject acquired by the acquisition unit 30 to the terminal device 2.

[0036] The display unit 21 of the terminal device 2 simultaneously displays the subject's index acoustic feature quantity, the normal range of said index acoustic feature quantity, and parameters indicating the subject's maxillofacial morphology acquired by the acquisition unit 30. By displaying the index acoustic feature quantity and parameters simultaneously in this way, for example, if it is determined that the subject has a speech disorder in which the index acoustic feature quantity falls outside the normal range, and the parameters indicate a structure that causes speech disorders, the dentist can diagnose that the subject has a morphological speech disorder accompanied by an organic abnormality. An organic abnormality is an abnormality that can be confirmed by machines, testing equipment, etc. Figure 4(A) shows an example of the index acoustic feature quantity of a subject with a morphological speech disorder.

[0037] On the other hand, if the index acoustic features are outside the normal range, indicating dysarthria, and the parameters do not show a structure that would cause dysarthria, the dentist can diagnose that the subject has functional dysarthria without organic abnormalities. Figure 4(B) shows an example of index acoustic features for a subject with functional dysarthria.

[0038] The treatment methods differ between morphological and functional speech disorders. For morphological speech disorders, treatment to correct the subject's craniofacial morphology takes priority. For functional speech disorders, since there is no need to correct craniofacial morphology, vocal exercises such as changing the position of the tongue during speech are performed.

[0039] The memory unit 32 may store a table showing the correspondence between each of several types of monosyllabic vocalizations with different vowel and consonant combinations, information indicating the subject's craniofacial morphology acquired by the acquisition unit 30, and index acoustic features that show a significant difference between healthy individuals and those with articulation disorders. Figure 5 shows an example of such a table. This table shows the correspondence between monosyllabic vocalizations (ka, ki, ku, ke, ko, sa, si, su, se, so) and index acoustic features that show a significant difference between craniofacial morphology types. In Figure 5, N indicates that the subject's craniofacial morphology is normal, O indicates that the molars are in contact but the front teeth are not (open bite), and D indicates that when biting, the upper front teeth overlap the lower front teeth (deep bite). Looking at the table in Figure 5, for example, MFCC(4) is an index acoustic feature that shows a significant difference between normal and deep bite for the single syllable (ke). This table shows that MFCC(4), (10), (12), (13), (15), and (18) show significant differences between normal and deep bite, between normal and open bite, and between open bite and deep bite when pronouncing specific single syllables.

[0040] If a subject is suspected of having morphological dysarthria, having them pronounce ka~ko and sa~so as shown in Figure 5, and calculating the MFCC(4), (10), (12), (13), (15), and (18) as index acoustic features at that time, makes it easier to narrow down to some extent whether or not a dysarthria is present and the factors causing the dysarthria.

[0041] The contents of the table shown in Figure 2 or Figure 5 may be stored in the memory unit 32 and displayed on the display unit 21 as needed. This display makes it easier for the dentist to understand the relationship between the type of monosyllable, the index acoustic features, and the subject's maxillofacial morphology. If the subject is diagnosed with morphological dysarthria, treatment for the morphological dysarthria (e.g., surgical treatment) is performed, and if the subject is diagnosed with functional dysarthria, treatment for the functional dysarthria (e.g., vocal exercises) is performed.

[0042] [Hardware configuration] The diagnostic support system 1A shown in Figure 1 is realized, for example, by a terminal device 2 and a server device 3 having the hardware configuration shown in Figure 6 executing a software program. Specifically, the server device 3 includes a CPU (Central Processing Unit) 41, which is a processor that controls the entire device; main memory 42 such as RAM (Random Access Memory); external memory 43 consisting of non-volatile memory such as flash memory and hard disk; a communication interface 46 for data communication with the terminal device 2; and an internal bus 48 that connects these.

[0043] Program 49 is loaded from external memory 43 into main memory 42 and executed by CPU 41. This enables the functionality of server device 3. During the execution of program 49, CPU 41 communicates data with an external computer via communication interface 46 as needed.

[0044] The functions of server device 3 can be implemented in a computer system consisting of one or more computers, each containing one or more processors and one or more storage devices, including non-temporary storage media. Multiple computers communicate with each other via an interconnected communication network to implement the functions of server device 3. For example, some of the functions of server device 3 may be implemented on one computer, while other functions may be implemented on other computers. The functions of server device 3 may also be implemented on a cloud computer.

[0045] Similarly, the terminal device 2 shown in Figure 1 is realized by a computer having the hardware configuration shown in Figure 6 executing a software program. Specifically, the terminal device 2 includes a CPU (Central Processing Unit) 51, main memory 52, external memory 53 for storing the program 59, an operation unit 54 which is a device such as a keyboard and mouse, a display 55 which is composed of a display device such as a CRT (Cathode Ray Tube) and an LCD monitor, a communication interface 56 for data communication with other computers, a microphone 57 for capturing images, and an internal bus 58 for connecting these. The microphone 57 corresponds to the audio input unit 20 mentioned above.

[0046] Program 59 is loaded from external memory 53 into main memory 52 and executed by the CPU 51. The execution of program 59 is controlled by the operation input of the operation unit 54, and data communication with an external computer is performed via the communication interface 56 as needed, and an image is displayed on the display 55. This realizes the functions of the terminal device 2.

[0047] [Data registration process] Next, the operation of the diagnostic support system 1A will be described. First, the data registration process will be described. Prior to the data registration process, the memory unit 32 stores the normal range of the index acoustic feature in healthy individuals, which is calculated by statistical calculation based on the speech data of single-syllable vocalizations of healthy individuals (memory step).

[0048] As shown in Figure 7, first, in the terminal device 2, the voice input unit 20 inputs voice data and transmits the voice data, which includes the subject's identification information and the type of sound produced, to the server device 3 (step S1). The voice input unit 20 repeats inputting voice data (step S1) until the scheduled number of times has been completed (step S2; No).

[0049] Meanwhile, in the server device 3, the acquisition unit 30 waits until it acquires voice data (Step S11; No). When the acquisition unit 30 acquires voice data from the terminal device 2 (Step S11; Yes; Acquisition step), the calculation unit 31 calculates index acoustic features from the acoustic features of the voice data acquired in the acquisition step, which serve as indicators for diagnosing articulation disorders (Step S12; Calculation step). The calculation unit 31 determines whether the planned number of rounds has been completed (Step S13). If it has not been completed (Step S13; No), the acquisition unit 30 returns to waiting for data acquisition (Step S11; No).

[0050] When the acquisition unit 30 acquires data (step S11; Yes), the calculation unit 12 again calculates the index acoustic feature (step S12) and determines when the planned number of iterations has ended (step S13). In this way, steps S11 → S12 → S13 are repeated until the planned number of iterations has ended (step S13; No).

[0051] When the scheduled number of operations is completed (Step S13; Yes), the calculation unit 31 calculates the average value of the index acoustic features calculated so far (Step S14) and stores it (Step S15). After that, the server device 3 waits for data acquisition again (Step S11; No).

[0052] In this way, information indicating the index acoustic features of the subject's voice data is stored in the memory unit 32.

[0053] [Data display processing] Next, we will explain the data display process.

[0054] As shown in Figure 8, first, the display unit 21 in the terminal device 2 transmits the specified display content (step S21). The display content is specified by the operation input of the operation unit 54 (see Figure 6) of the terminal device 2. The subject's identification information, the type of monosyllable, etc., are specified as the display content. Next, the display unit 21 waits until it receives data (step S22; No).

[0055] Meanwhile, in the server device 3, the data reading unit 33 waits until it receives the specified display content (step S31; No). When it receives the specified display content (step S31; Yes), the data reading unit 33 reads data from the storage unit 32, such as information indicating the index acoustic feature quantity and the normal range (upper limit, lower limit), corresponding to the specified display content (step S32), and transmits the data to the terminal device 2 (step S33).

[0056] Meanwhile, in terminal device 2, when the display unit 21 receives data (step S22; Yes), it displays the received data (step S23; display step). For example, as shown in Figure 4(A) or Figure 4(B), information is displayed indicating whether the representative values ​​of MFCC (4), (5), (8), (10), (11), (12), (13), (15), (17), and (18) as index acoustic features fall between the upper and lower limits. The dentist refers to this display as auxiliary information for diagnosis and makes a diagnosis of the patient.

[0057] Embodiment 2 Next, Embodiment 2 of the present invention will be described. As shown in Figure 9, the diagnostic support system 1B according to this embodiment differs from the diagnostic support system 1A according to Embodiment 1 in that it includes a determination unit 34.

[0058] The determination unit 34 determines whether the subject has developed a speech disorder based on whether the subject's index acoustic feature quantities, calculated by the calculation unit 31 and stored in the memory unit 32, fall within the normal range of the index acoustic feature quantities stored in the memory unit 32. Furthermore, if the determination unit 34 determines that the subject has developed a speech disorder, it determines, based on the information indicating the subject's maxillofacial morphology acquired by the acquisition unit 30, whether the speech disorder the subject has developed is a morphological speech disorder accompanied by an organic abnormality or a functional speech disorder without an organic abnormality. For example, if there are no particular problems with the maxillofacial morphology of the subject, but some of the index acoustic feature quantities are outside the normal range, it will be determined to be a functional speech disorder.

[0059] As described in detail above, the diagnostic support systems 1A and 1B according to the above embodiments acquire speech data including monosyllabic sounds of the subject, and display index acoustic features from the acquired speech data that show a significant difference between healthy individuals and those with articulation disorders in a way that allows comparison with the normal range. Therefore, the onset of articulation disorders can be detected in a simple and accurate manner.

[0060] In this embodiment, Mel-frequency cepstrum coefficients are used as acoustic features. However, this is not the only option. The calculation unit 31 only needs to calculate at least one index acoustic feature from the acoustic features of the speech data that shows a significant difference between healthy individuals and those with speech disorders. Such index acoustic features are not limited to Mel-frequency cepstrum coefficients. Various parameters that are acoustic features of the speech data can be included as index acoustic features in a wide variety of analysis processes such as bandpass filtering, linear prediction analysis, cepstrum analysis, average power analysis, and formant frequency analysis.

[0061] Furthermore, in the above embodiment, index acoustic features of speech data of monosyllabic pronunciations in Japanese are used. However, index acoustic features of speech data of monosyllabic pronunciations in other languages ​​may also be used.

[0062] The hardware and software configurations of server device 3 and terminal device 2 are examples only and can be changed and modified as needed.

[0063] The core processing portion of the server device 3 and terminal device 2, which consist of CPUs 41 and 51, main memory 42 and 52, external memory 43 and 53, operation unit 54, display 55, communication interfaces 46 and 56, microphone 57, and internal buses 48 and 58, can be implemented using a normal computer system, not a dedicated system. For example, the computer program for performing the above operations may be stored on a computer-readable recording medium (flexible disk, CD-ROM, DVD-ROM, etc.) and distributed, and the server device 3 and terminal device 2 that perform the above processing may be configured by installing the computer program on a computer. Alternatively, the computer program may be stored on a storage device of a server device on a communication network such as the Internet, and the server device 3 and terminal device 2 may be configured by downloading it from a normal computer system.

[0064] When the functions of server device 3 and terminal device 2 are realized through a division of labor between the OS (operating system) and application programs, or through cooperation between the OS and application programs, only the application program portion may be stored on a recording medium or storage device.

[0065] The configurations shown in each of the embodiments described above can be partially replaced or modified, and it is also possible to combine configurations from different embodiments. The dimensions, scale, and aspect ratio in the drawings can also be changed as appropriate.

[0066] This invention allows for various embodiments and modifications without departing from the broad spirit and scope of the invention. Furthermore, the embodiments described above are for illustrative purposes only and do not limit the scope of the invention. In other words, the scope of this invention is indicated not by the embodiments, but by the claims. Various modifications made within the scope of the claims and the equivalent scope of the meaning of the invention are considered to be within the scope of this invention. [Industrial applicability]

[0067] This invention can be applied to assist in the diagnosis of speech disorders. [Explanation of symbols]

[0068] 1A, 1B Diagnostic support system, 2 Terminal device, 3 Server device, 4 Measurement device, 10 Fourier transform unit, 11 Mel-banding unit, 12 Logarithmic unit, 13 Discrete cosine transform unit, 14 Normalization unit, 15 Coefficient extraction unit, 20 Voice input unit, 21 Display unit, 30 Acquisition unit, 31 Calculation unit, 32 Memory unit, 33 Data read unit, 34 Judgment unit, 41 CPU, 42 Main memory, 43 External memory, 46 Communication interface, 48 Internal bus, 49 Program, 51 CPU, 52 Main memory, 53 External memory, 54 Operation unit, 55 Display, 56 Communication interface, 57 Microphone, 58 Internal bus, 59 Program

Claims

1. A diagnostic support system that generates information to assist in the diagnosis of speech disorders, An acquisition unit that acquires speech data including monosyllabic sounds of the subject, Based on the voice data acquired by the acquisition unit, a calculation unit calculates information indicating index acoustic features that show a significant difference between healthy individuals and individuals with speech disorders among the acoustic features of the voice data. A storage unit that stores the normal range of the index acoustic feature, calculated by statistical calculation based on the speech data of single-syllable vocalizations of healthy individuals, A display unit that displays information indicating the subject's index acoustic feature quantity calculated by the calculation unit and the normal range stored in the storage unit in a comparable manner, A diagnostic support system equipped with the following features.

2. The acquisition unit acquires information showing the maxillofacial morphology of the subject, The display unit simultaneously displays the subject's index acoustic features and normal range, as well as information indicating the subject's maxillofacial morphology acquired by the acquisition unit. The diagnostic support system according to claim 1.

3. The system includes a determination unit that determines if the subject's index acoustic characteristic quantity is not within the normal range, and determines that the subject has a speech disorder. The diagnostic support system according to claim 1.

4. The acquisition unit acquires information showing the maxillofacial morphology of the subject, If the determination unit determines that the subject has developed a speech disorder, it will determine, based on the information indicating the subject's maxillofacial morphology acquired by the acquisition unit, whether the speech disorder the subject has developed is a morphological speech disorder accompanied by an organic abnormality or a functional speech disorder without an organic abnormality. The diagnostic support system according to claim 3.

5. The index acoustic features used to diagnose speech disorders include Mel-frequency cepstrum coefficients. A diagnostic support system according to any one of claims 1 to 4.

6. The Mel-frequency cepstrum coefficients include at least one of MFCC(4) or MFCC(12). The diagnostic support system according to claim 5.

7. A diagnostic assistance method performed by an information processing device that generates information to assist in the diagnosis of speech disorders, Acquisition step to obtain speech data including monosyllabic vocalizations of the subject, Based on the voice data acquired in the acquisition step, a calculation step is performed to calculate information indicating index acoustic features that show a significant difference between healthy individuals and individuals with speech disorders among the acoustic features of the voice data. A display step that displays information showing the subject's index acoustic feature calculated in the calculation step and the normal range of the index acoustic feature calculated by statistical calculation based on the speech data of a healthy person's single-syllable vocalization, in a way that allows for comparison. Diagnostic support methods including those mentioned above.

8. A computer that generates information to assist in the diagnosis of speech disorders, Acquisition unit that acquires speech data including monosyllabic sounds of the subject. A calculation unit calculates information indicating index acoustic features that show a significant difference between healthy individuals and individuals with speech disorders, based on the audio data acquired by the acquisition unit. A storage unit that stores the normal range of the index acoustic feature, calculated by statistical calculation based on the speech data of single-syllable vocalizations of healthy individuals. A display unit that displays information indicating the subject's index acoustic feature quantity calculated by the calculation unit and the normal range stored in the storage unit in a comparable manner. A program that makes it function as such.

Citation Information

Patent Citations

  • Articulation disorder detection device and articulation disorder detection method

    JP2023146782A