Articulation abnormality examination device, articulation abnormality correction support system, articulation abnormality examination method, and program

The sound abnormality inspection device and system enhance the accuracy of dysarthria evaluations by using machine learning to analyze voice data and identify articulation abnormalities, addressing the limitations of existing technologies in supporting individuals with dysarthria.

JP2025074685APending Publication Date: 2025-05-14IBARAKI UNIVERSITY +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023185676
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-10-30
Publication Date
2025-05-14

AI Technical Summary

Technical Problem

Existing technologies for evaluating dysarthria using machine learning face challenges in achieving accurate evaluations, particularly in distinguishing distorted sounds and improving support for individuals with dysarthria, especially children who may not recognize their own errors.

Method used

A sound abnormality inspection device and system that acquire voice data, generate voice characteristic information for syllables containing consonants or vowels, and use a trained model to evaluate articulation abnormalities, thereby enhancing the accuracy of dysarthria evaluations.

Benefits of technology

The proposed solution significantly improves the accuracy of dysarthria evaluations, enabling more effective identification and correction of articulation abnormalities, which can lead to faster and more targeted support for individuals with dysarthria.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025074685000001_ABST
    Figure 2025074685000001_ABST
Patent Text Reader

Abstract

To improve accuracy of evaluation in a technique for evaluating articulation disorder using machine learning.SOLUTION: An articulation abnormality examination device includes: a voice data acquisition unit that acquires voice data of a subject; a voice feature information generation unit that generates voice feature information of a portion corresponding to syllables including at least one of consonants and vowels from the voice data acquired by the voice data acquisition unit; and an evaluation unit that evaluates the articulation abnormality of the subject on the basis of the output result acquired by inputting the voice feature information generated by the voice feature information generation unit to a learned model that has been learned so as to output evaluation information indicating the evaluation of the articulation abnormality when the voice feature information is input.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an apparatus for inspecting speech abnormalities, a system for supporting correction of speech abnormalities, a method for inspecting speech abnormalities, and a program. [Background technology]

[0002] Previously, articulation disorders were evaluated subjectively based on the examiner's hearing judgment. Therefore, it was difficult to make a correct evaluation unless the examiner had a lot of experience. Distorted sounds are particularly difficult to evaluate, and children who are unable to recognize their own incorrect sounds do not improve easily, and it takes time for the effects of support to appear.

[0003] There is known a technique for evaluating articulation disorders using machine learning. For example, a method for detecting the presence or absence of articulation disorders is known (Patent Document 1), which detects the presence or absence of articulation disorders by inputting a spectrogram or envelope obtained from a speech waveform into a detection model that has been machine-learned to use speech as input and output information on the presence or absence of articulation disorders. In addition, a method for detecting articulation disorders is known in which acoustic features are calculated from a speaker's speech data, and the speaker's features are calculated using a trained DNN (Deep Neural Network), and the speaker's features are calculated to determine the speaker's articulation disorders by calculating the similarity with the features of a healthy speaker (Patent Document 2). [Prior art documents] [Patent documents]

[0004] [Patent Document 1] JP 2023-036486 A [Patent Document 2] JP 2023-002421 A Summary of the Invention [Problem to be solved by the invention]

[0005] There is a demand for improved accuracy in technology that uses machine learning to evaluate speech disorders.

[0006] The present invention has been made in consideration of the above points, and provides an articulation abnormality inspection device, an articulation abnormality correction support system, an articulation abnormality inspection method, and a program that can improve the accuracy of evaluation in a technology for evaluating articulation disorders using machine learning. [Means for solving the problem]

[0007] The present invention has been made to solve the above-mentioned problems, and one aspect of the present invention is an articulation abnormality inspection device comprising: a voice data acquisition unit that acquires voice data of a subject; a voice feature information generation unit that generates voice feature information of a portion of the voice data acquired by the voice data acquisition unit that corresponds to a syllable including at least one of a consonant or a vowel; and an evaluation unit that evaluates the articulation abnormality of the subject based on an output result obtained by inputting the voice feature information generated by the voice feature information generation unit into a trained model that is trained to output evaluation information indicating an evaluation of an articulation abnormality when the voice feature information is input.

[0008] In addition, one aspect of the present invention is that in the above-mentioned speech abnormality inspection device, the speech feature information generation unit extracts a portion corresponding to a syllable containing at least one of a consonant or a vowel from the speech data, and generates the speech feature information from the extracted portion.

[0009] In addition, one aspect of the present invention is that in the above-mentioned speech abnormality inspection device, the speech feature information generation unit extracts a portion corresponding to a syllable containing at least one of a consonant or a vowel from the speech data based on the time and frequency range in which the syllable appears in the speech data.

[0010] In addition, one aspect of the present invention is that in the above-mentioned speech abnormality inspection device, the speech feature information generation unit extracts a portion from the speech data corresponding to a syllable containing at least one of a consonant or a vowel, based on the sound pressure or volume indicated by the speech data.

[0011] In one aspect of the present invention, in the speech abnormality inspection device, the speech feature information is a spectrogram or a spectral envelope.

[0012] In one aspect of the present invention, in the articulation abnormality inspection device, the speech feature information is speech feature information of a portion corresponding to a consonant.

[0013] In addition, one aspect of the present invention is that in the above-mentioned speech abnormality inspection device, the evaluation unit evaluates the speech abnormality of the subject based on whether or not frequency components greater than a predetermined threshold are missing from the frequency components indicated by the speech feature information.

[0014] In addition, one aspect of the present invention is that in the above-mentioned speech abnormality inspection device, the evaluation information indicates the presence or absence of speech abnormalities, and the evaluation unit determines the presence or absence of speech abnormalities of the subject based on an output result obtained by inputting the speech feature information generated by the speech feature information generation unit to the trained model.

[0015] In addition, one aspect of the present invention is that in the above-mentioned speech abnormality inspection device, the evaluation information indicates the degree of speech abnormality, and the evaluation unit judges the degree of the speech abnormality of the subject based on the output result obtained by inputting the speech feature information generated by the speech feature information generation unit to the trained model.

[0016] Another aspect of the present invention is a speech abnormality correction support system comprising the above-mentioned speech abnormality inspection device, a sound collection unit that collects the voice of the subject, and an evaluation result presentation unit that presents the evaluation results by the evaluation unit.

[0017] In one aspect of the present invention, the articulation disorder correction support system further comprises a character presenting unit that presents one or more of a monosyllable, a word, or a short sentence to be spoken by the subject.

[0018] Moreover, one aspect of the present invention is a method for testing articulation abnormalities, comprising: a voice data acquisition step of acquiring voice data of a subject; a voice feature information generation step of generating voice feature information of a portion of the voice data acquired by the voice data acquisition step that corresponds to a syllable including at least one of a consonant or a vowel; and an evaluation step of evaluating the articulation abnormalities of the subject based on an output result obtained by inputting the voice feature information generated by the voice feature information generation step into a trained model trained to output evaluation information indicating an evaluation of an articulation abnormality when the voice feature information is input.

[0019] Another aspect of the present invention is a program for causing a computer to execute the following steps: a voice data acquisition step of acquiring voice data of a subject; a voice feature information generation step of generating voice feature information of a portion of the voice data acquired by the voice data acquisition step that corresponds to a syllable including at least one of a consonant or a vowel; and an evaluation step of evaluating the subject's articulation abnormality based on an output result obtained by inputting the voice feature information generated by the voice feature information generation step into a trained model trained to output evaluation information indicating an evaluation of an articulation abnormality when the voice feature information is input. Effect of the Invention

[0020] According to the present invention, it is possible to improve the accuracy of evaluation in a technique for evaluating speech disorders using machine learning. [Brief description of the drawings]

[0021] [Figure 1] 1 is a diagram showing an example of a configuration and an overview of processing of an articulation abnormality inspection system 1 according to a first embodiment of the present invention. [Diagram 2] 1 is a diagram illustrating an example of a functional configuration of an articulation abnormality inspection device 2 according to a first embodiment of the present invention. [Diagram 3] FIG. 2 is a diagram showing an example of a flow of an articulation abnormality evaluation process according to the first embodiment of the present invention. [Figure 4]1 is a diagram showing an example of audio data A1 according to the first embodiment of the present invention, and a fast Fourier transform of the audio data A1. [Diagram 5] 1A to 1C are diagrams showing an example of a change in fundamental frequency over time according to the first embodiment of the present invention, and an example of audio data of a sound showing the change in fundamental frequency over time. [Figure 6] FIG. 11 is a diagram showing an example of waveforms of audio signals of portions corresponding to consonants and vowels extracted from audio data A1 according to a modified example of the first embodiment of the present invention. [Figure 7] FIG. 2 is a diagram showing an example of a mel spectrogram generated from the waveform of a speech signal of portions corresponding to consonants and vowels, respectively, according to the first embodiment of the present invention. [Figure 8] 2 is a diagram showing an example of various data generated from audio data A1 according to the first embodiment of the present invention. FIG. [Figure 9] FIG. 4 is a diagram showing an example of an evaluation result of an articulation abnormality according to the first embodiment of the present invention. [Figure 10] FIG. 4 is a diagram showing another example of an evaluation result of an articulation abnormality according to the first embodiment of the present invention. [Figure 11] FIG. 13 is a diagram showing an example of a spectral envelope for a person with an articulation disorder according to a modified example of the first embodiment of the present invention. [Figure 12] FIG. 13 is a diagram showing an example of a spectral envelope for a person with an articulation disorder according to a modified example of the first embodiment of the present invention. [Figure 13] FIG. 11 is a diagram showing an example of a spectral envelope for a healthy subject according to a modified example of the first embodiment of the present invention. [Figure 14] FIG. 11 is a diagram showing an example of a spectral envelope for a healthy subject according to a modified example of the first embodiment of the present invention. [Figure 15] FIG. 11 is a diagram showing an example of a spectral envelope for a healthy subject according to a modified example of the first embodiment of the present invention. [Figure 16] FIG. 11 is a diagram showing an example of the configuration of an articulation disorder correction support system 5a according to a second embodiment of the present invention. [Figure 17]FIG. 11 is a diagram showing a functional configuration of an articulation abnormality inspection device 2a according to a second embodiment of the present invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0022] (First embodiment) Hereinafter, an embodiment of the present invention will be described in detail with reference to the drawings. Fig. 1 is a diagram showing an example of the configuration and processing overview of an articulation abnormality inspection system 1 according to this embodiment. The articulation abnormality inspection system 1 is a system for inspecting an articulation abnormality of a subject. The articulation abnormality inspection system 1 includes an articulation abnormality inspection device 2, a sound collection unit 3, and an evaluation result presentation unit 4.

[0023] As an example, the subject is a child with an articulation disorder. The subject may be an adult with an articulation disorder, or a healthy child or person. The subject may also be a person with another disability in which abnormalities in articulation are observed. The other disability may be, for example, a hearing impairment or aphasia. The examiner reads out loud the character string for the test. In the example shown in FIG. 1, the character string for the test is "ringo" (apple). The examiner is unable to pronounce the "ri" sound correctly.

[0024] The examiner collects the voice uttered by the subject using the sound collection unit 3. The sound collection unit 3 is, for example, a recorder. The voice collected by the sound collection unit 3 is stored in the sound collection unit 3 as voice data A1. The voice data A1 is data indicating the waveform of the voice uttered by the subject. The articulation abnormality inspection device 2 acquires the voice data A1 from the sound collection unit 3.

[0025] The articulation abnormality inspection device 2 performs processing for evaluating articulation abnormalities from the speech data A1. Articulation abnormalities are also described as articulation disorders. The articulation abnormality inspection device 2 is, for example, a personal computer (PC). The articulation abnormality inspection device 2 may be a mobile terminal such as a smartphone or a tablet terminal. The articulation abnormality inspection device 2 may also be realized as a server.

[0026] The articulation abnormality inspection device 2 extracts, from the speech data A1, a waveform portion corresponding to a consonant (hereinafter referred to as a consonant portion) and a waveform portion corresponding to a vowel (hereinafter referred to as a vowel portion).

[0027] The articulation abnormality inspection device 2 generates a spectrogram of the extracted consonant part.The articulation abnormality inspection device 2 generates a spectrogram of the extracted vowel part.

[0028] The articulation abnormality inspection device 2 judges the presence or absence of articulation abnormality in the subject and the degree of the articulation abnormality based on machine learning from the generated spectrogram. The articulation abnormality inspection device 2 inputs the spectrograms of the consonant part and the vowel part into a trained model, and makes a judgment based on the results output from the trained model. In other words, the articulation abnormality inspection device 2 makes a judgment based on an image of an area in the spectrogram that corresponds to the phoneme of the consonant part or the vowel part. In the following description, "judgment" is also referred to as "identification" or "classification".

[0029] The learning data used in the machine learning includes articulation disorder data, severity data, and speech data of healthy individuals. The articulation disorder data is data showing the waveform of speech uttered by an articulation disorder person. The severity data is data showing the degree of articulation disorder for an articulation disorder person. The speech data of healthy individuals is data showing the waveform of speech uttered by healthy individuals. In the learning data, consonant parts and vowel parts are respectively extracted from the data showing the speech waveform, and spectrograms extracted from the extracted consonant parts and vowel parts are respectively used as explanatory variables. In addition, in the learning data, the degree of articulation disorder is used as a label.

[0030] The articulation abnormality inspection device 2 may generate an envelope spectrum instead of a spectrogram. In that case, the articulation abnormality inspection device 2 converts the extracted consonant part into an envelope spectrum by the Cepstrum method. The articulation abnormality inspection device 2 converts the extracted vowel part into an envelope spectrum by the Cepstrum method. When an envelope spectrum is generated, the envelope spectrum is used as an explanatory variable in the training data.

[0031] As a result of the machine learning, the voice data A1 is classified into two classes, namely, with articulation abnormality and without articulation abnormality. Furthermore, the class with articulation abnormality is further classified into, for example, severe, moderate, and mild classes. Therefore, the voice data A1 is classified into, for example, four classes, namely, severe, moderate, mild, and normal, by determining the presence or absence of articulation abnormality and the degree of articulation abnormality.

[0032] As another example, as a result of the machine learning determination, the speech data A1 may be classified into a class of syllables with articulation disorders and a class of syllables without articulation disorders. Furthermore, the class of syllables with articulation disorders may be classified according to the degree of articulation disorders.

[0033] The articulation abnormality inspection device 2 causes the evaluation result presentation unit 4 to present the classification result of the articulation abnormality as the evaluation result of the subject's articulation abnormality. As an example, the evaluation result presentation unit 4 presents the evaluation result by displaying it as one or more of an image such as a graph, and letters, numbers, etc. As an example, the evaluation result presentation unit 4 is a display.

[0034] 2 is a diagram showing an example of a functional configuration of the articulation abnormality inspection device 2 according to the present embodiment. The articulation abnormality inspection device 2 includes a processing unit 20 and a storage unit .

[0035] The processing unit 20 executes various processes. The processing unit 20 includes a voice data acquisition unit 200, a voice feature information generation unit 201, an evaluation unit 202, and an output unit 203. Each of these functional units is realized, for example, by a CPU loading a program read from a ROM (Read Only Memory) into a RAM (Random Access Memory) and executing processing according to the program. The ROM and RAM are included in the storage unit 21.

[0036] The voice data acquiring section 200 acquires the voice data A1 of the subject. In this embodiment, the voice data acquiring section 200 acquires the voice data A1 from the sound collecting section 3.

[0037] The voice feature information generating unit 201 generates voice feature information. The voice feature information is information indicating the features of a voice. In this embodiment, the voice feature information is, for example, a spectrogram. Note that the voice feature information may be an envelope of a spectrum.

[0038] The speech feature information generated by the speech feature information generating unit 201 includes consonant speech feature information and vowel speech feature information. The consonant speech feature information is speech feature information of a portion of the speech data A1 acquired by the speech data acquiring unit 200 that corresponds to a consonant. The vowel speech feature information is speech feature information of a portion of the speech data A1 acquired by the speech data acquiring unit 200 that corresponds to a vowel. The voice feature information may include voice feature information of a portion of the voice data A1 corresponding to a syllable including a consonant and a vowel. In other words, the voice feature information is voice feature information of a portion of the voice data A1 corresponding to a syllable including at least one of a consonant and a vowel.

[0039] When the voice data A1 includes a plurality of syllables, the voice feature information generating unit 201 generates consonant voice feature information or vowel voice feature information for each of the plurality of syllables. A syllable is a voice consisting of one or more consonants before and after a vowel, or a voice consisting of one vowel. For example, when the voice data A1 indicates a waveform of a voice when the phrase "ringo" is pronounced, the voice feature information generating unit 201 divides the voice data A1 into a part corresponding to the syllable "ri", a part corresponding to the syllable "n", and a part corresponding to the syllable "go". The voice feature information generating unit 201 generates consonant voice feature information or vowel voice feature information for each of the syllables "ri", "n", and "go".

[0040] The evaluation unit 202 evaluates the articulation abnormality of the subject based on the output result obtained by inputting the consonant speech feature information generated by the speech feature information generation unit 201 to the trained model B1. The evaluation unit 202 also evaluates the articulation abnormality of the subject based on the output result obtained by inputting the vowel speech feature information generated by the speech feature information generation unit 201 to the trained model B1.

[0041] Here, in this embodiment, the trained model B1 is trained to output evaluation information indicating an evaluation of articulation abnormality when consonant speech feature information is input, and the trained model B1 is trained to output evaluation information indicating an evaluation of articulation abnormality when vowel speech feature information is input.

[0042] As one example, the evaluation information indicates the presence or absence of an articulation abnormality. As another example, the evaluation information may indicate the degree of the articulation abnormality. Furthermore, the evaluation information may indicate both the presence or absence of an articulation abnormality and the degree of the articulation abnormality.

[0043] The degree of the articulation disorder is classified into classes indicating the degree of the articulation disorder, such as "severe," "moderate," and "mild," for example. The number of classes into which the degree of the articulation disorder is classified may be two or more. The degree of the articulation disorder may be indicated by a score. The score is indicated by, for example, a real number from 0 to 100 or an integer value. The degree of the articulation disorder may be indicated by an image showing the degree of the articulation disorder. The image is indicated by, for example, different icons according to the degree of severity. In the case of a severe case, an icon of a troubled face or an icon of a small tree is indicated. In the case of a mild case, an icon of a smiling face or an icon of an overgrown tree is indicated.

[0044] The trained model B1 is, for example, a trained Convolutional Neural Network (CNN). The trained model B1 may be a machine learning model other than CNN (for example, a Deep Neural Network (DNN)). When the information input to the trained model B1 is an image such as a spectrogram, it is preferable that the trained model B1 is a training model suitable for image recognition such as CNN.

[0045] The output unit 203 outputs the evaluation result C1 by the evaluation unit 202. In this embodiment, the output unit 203 outputs the evaluation result C1 to the evaluation result presenting unit 4.

[0046] The storage unit 21 stores various types of information. A trained model B1 is stored in advance in the storage unit 21. The storage unit 21 is configured using a storage device such as a magnetic hard disk device or a semiconductor storage device.

[0047] Next, an articulation abnormality evaluation process in which the articulation abnormality inspection device 2 evaluates articulation abnormalities of a subject will be described with reference to Figs. Fig. 3 is a diagram showing an example of the flow of the articulation abnormality evaluation process according to the present embodiment. The articulation abnormality evaluation process is executed by the processing unit 20. Fig. 4 to Fig. 8 are diagrams showing an example of the voice data A1 according to the present embodiment and various data generated from the voice data A1.

[0048] Step S10: The voice data acquiring unit 200 acquires voice data A1 of the subject. The voice data acquiring unit 200 acquires the voice data A1 from the sound collecting unit 3. An example of the voice data A1 is shown in Fig. 4. In the example shown in Fig. 4, the voice data A1 is a waveform of the voice uttered by the subject for approximately 0.5 seconds.

[0049] As an example, after the subject finishes reading aloud the character string for the test, the sound collection unit 3 stores the collected voice of the subject as voice data A1. The voice data acquisition unit 200 acquires the voice data A1 stored in the sound collection unit 3 from the sound collection unit 3. The voice data acquisition unit 200 acquires the voice data A1 from the sound collection unit 3, for example, by wireless communication between the articulation abnormality inspection device 2 and the sound collection unit 3. The sound collection unit 3 may be connected to the articulation abnormality inspection device 2 by a cable. In this case, the voice data acquisition unit 200 acquires the voice data A1 from the sound collection unit 3 via the cable. The voice data A1 stored in the sound collection unit 3 may also be stored in an external storage device, and the voice data acquisition unit 200 acquires the voice data A1 via the external storage device.

[0050] The voice data acquiring section 200 may instantly acquire the voice of the subject picked up by the sound collecting section 3 as the voice data A1 while the subject is reading aloud the character string for the test.

[0051] Step S20: The voice feature information generation unit 201 extracts a portion corresponding to a consonant from the voice data A1. The voice feature information generation unit 201 also extracts a portion corresponding to a vowel from the voice data A1. The voice feature information generation unit 201 extracts a portion corresponding to a consonant from the voice data A1 based on the time, frequency range, and sound pressure in which the consonant appears in the voice data A1. Similarly, the voice feature information generation unit 201 extracts a portion corresponding to a vowel from the voice data A1 based on the time, frequency range, and sound pressure in which the vowel appears in the voice data A1.

[0052] When the voice data A1 includes a plurality of syllables, the voice feature information generating unit 201 generates consonant voice feature information or vowel voice feature information for a portion corresponding to each of the plurality of syllables. The voice feature information generating unit 201 divides the voice data A1 into portions corresponding to each of the plurality of syllables, for example, based on the time and sound pressure in the voice data A1. When the voice data A1 includes a plurality of syllables, in the following description of step S20, the voice data A1 is a portion corresponding to each of the plurality of syllables. Moreover, each process from step S30 to step S40 is executed sequentially or in parallel for each syllable.

[0053] The audio feature information generating unit 201 converts the waveform represented by the audio data A1 into a spectrum by Fast Fourier Transform (FFT) in order to calculate the frequency of the waveform represented by the audio data A1. FIGS. 4(B) and (C) show an example of a spectrum obtained by converting the waveform represented by the audio data A1 by FFT. Here, the frequency range corresponding to consonants is approximately 2000 to 10000 Hz. The frequency range corresponding to vowels is approximately 0 to 1000 Hz. FIG. 4(B) is a spectrum for the frequency range corresponding to consonants. FIG. 4(C) is a spectrum for the frequency range corresponding to vowels. Note that the scales of the frequency axes are different between FIG. 4(B) and FIG. 4(C).

[0054] The voice feature information generating unit 201 extracts a portion corresponding to a consonant from the voice data A1 based on a frequency range corresponding to the consonant. The voice feature information generating unit 201 extracts a frequency range of 2000 Hz or more and 10000 Hz or less from the voice data A1 as a portion corresponding to a consonant. The voice feature information generating unit 201 extracts a portion corresponding to a vowel from the voice data A1 based on a frequency range corresponding to a vowel. The voice feature information generating unit 201 extracts a frequency range of 0 Hz or more and 1000 Hz or less from the voice data A1 as a portion corresponding to a vowel.

[0055] The time during which a consonant appears in the voice data A1 is the time from when the sound starts to be produced until about 0.1 seconds (or 0.2 seconds) have elapsed. The time during which a vowel appears in the voice data A1 is the time from when the sound starts to be produced until about 0.1 seconds (or 0.2 seconds) have elapsed. In FIG. 4, time T1 indicates the time during which a consonant appears. Time T2 indicates the time during which a vowel appears.

[0056] The voice feature information generating unit 201 may extract a portion corresponding to a consonant from the voice data A1 by regarding a time during which a consonant appears as a predetermined time longer than 0.1 seconds and not longer than 0.2 seconds from the start of the utterance of the sound. The voice feature information generating unit 201 may extract a portion corresponding to a vowel from the voice data A1 by regarding a time during which a vowel appears as a time from the predetermined time to the end of the utterance of the sound.

[0057] The voice feature information generating unit 201 extracts a portion corresponding to a consonant from the voice data A1 based on the sound pressure corresponding to the consonant. The sound pressure of a consonant generally tends to be smaller than the sound pressure of a vowel. The voice feature information generating unit 201 divides a time range of the voice data A1 into a portion corresponding to a consonant and a portion corresponding to a vowel, for example, based on the average sound pressure per time. The voice feature information generating unit 201 extracts the range with the smaller average sound pressure per time as the portion corresponding to a consonant. The voice feature information generating unit 201 extracts the range with the larger average sound pressure per time as the portion corresponding to a vowel. Therefore, the voice feature information generating unit 201 extracts the part corresponding to the consonant from the voice data A1 based on the sound pressure indicated by the voice data A1. The voice feature information generating unit 201 may extract the part corresponding to the consonant from the voice data A1 based on the volume indicated by the voice data A1.

[0058] In the present embodiment, the voice feature information generating unit 201 extracts a portion corresponding to a consonant from the voice data A1 based on the time, frequency range, sound pressure, and volume of the consonant in the voice data A1, but the present invention is not limited to this. The voice feature information generating unit 201 may extract a portion corresponding to a consonant from the voice data A1 based only on the time and frequency range of the consonant in the voice data A1. That is, in this case, the voice feature information generating unit 201 extracts a portion corresponding to a consonant from the voice data A1 based on the time and frequency range of the consonant in the voice data A1. In this case, the voice feature information generating unit 201 similarly extracts a portion corresponding to a vowel from the voice data A1 based on the time and frequency range of the vowel in the voice data A1.

[0059] Furthermore, the voice feature information generating unit 201 may extract a portion corresponding to a consonant or a portion corresponding to a vowel from the voice data A1 based on the fundamental frequency (pitch) in the voice data A1. FIG. 5(B) shows a time change in the fundamental frequency. The time change in the fundamental frequency shown in FIG. 5(B) is, as an example, a time change in the fundamental frequency when a person with an articulation disorder utters the consonant "ki". FIG. 5(A) shows, for reference, the voice data of the voice whose time change in the fundamental frequency is shown in FIG. 5(B). Note that FIG. 5 is a diagram for explaining the distinction between consonants and vowels in the fundamental frequency, and the data shown in FIG. 5 is different from the voice data A1 shown in FIG. 4.

[0060] In Fig. 5(B), time T3 indicates the time when a consonant appears. Time T4 indicates the time when a vowel appears. Comparing the time change of the fundamental frequency at time T3 and time T4, the time change of the fundamental frequency at time T3 is dispersed like noise, whereas the time change of the fundamental frequency at time T4 is small and stable. In other words, the time change of the fundamental frequency is dispersed during the time when a consonant appears, whereas the time change is small and stable during the time when a vowel appears.

[0061] The voice feature information generating unit 201 extracts a portion corresponding to a consonant or a portion corresponding to a vowel from the voice data A1 based on the difference in tendency between the time when a consonant appears and the time when a vowel appears in the fundamental frequency change over time. For example, the voice feature information generating unit 201 extracts a portion of the fundamental frequency change over time that changes little and is stable (a portion with small variance) as a portion corresponding to a vowel. The voice feature information generating unit 201 may also extract the remaining portion of the fundamental frequency change over time after extracting the portion corresponding to a vowel as a portion corresponding to a consonant. The voice feature information generating unit 201 may also extract a portion of the fundamental frequency change over time that has large variance as a portion corresponding to a consonant.

[0062] The voice feature information generating unit 201 may extract a portion corresponding to a consonant from the voice data A1 based on one or more of a time when a consonant appears in the voice data A1, a frequency range, a sound pressure, a volume, and a fundamental frequency. The voice feature information generating unit 201 may extract a portion corresponding to a vowel from the voice data A1 based on one or more of a time when a vowel appears in the voice data A1, a frequency range, a sound pressure, a volume, and a fundamental frequency.

[0063] Fig. 6(A) shows the waveform of a speech signal of a portion corresponding to a consonant extracted from the speech data A1 shown in Fig. 4. Fig. 6(B) shows the waveform of a speech signal of a portion corresponding to a vowel extracted from the speech data A1 shown in Fig. 4.

[0064] Step S30: The voice feature information generating unit 201 generates consonant voice feature information from the portion corresponding to the extracted consonant. The voice feature information generating unit 201 generates vowel voice feature information from the portion corresponding to the extracted vowel. In this embodiment, the voice feature information is a spectrogram. A spectrogram indicates the frequency spectrum of a voice in a time series. Therefore, the consonant voice feature information is an image indicating the frequency spectrum of the portion corresponding to the consonant extracted from the voice data A1 in a time series. The vowel voice feature information is an image indicating the frequency spectrum of the portion corresponding to the vowel extracted from the voice data A1 in a time series. Note that the voice feature information generating unit 201 uses FFT in generating the spectrogram. Note that the spectrogram includes a mel spectrogram.

[0065] Fig. 7(A) shows a mel spectrogram generated from the waveform of the speech signal of the portion corresponding to the consonant shown in Fig. 6(A). Fig. 7(B) shows a mel spectrogram generated from the waveform of the speech signal of the portion corresponding to the vowel shown in Fig. 6(B).

[0066] In the present embodiment, an example of generating consonant speech feature information and vowel speech feature information after extracting consonant parts and vowel parts from speech data A1 has been described, but the present invention is not limited to this. After generating speech feature information from speech data A1, a part corresponding to a consonant and a part corresponding to a vowel may be extracted from the speech feature information.

[0067] FIG. 8(C) shows a spectrogram generated from the voice data A1. For example, from the spectrogram shown in FIG. 8(C), a part corresponding to a consonant and a part corresponding to a vowel may be extracted based on at least the time when the consonant appears and the frequency range. Also, from the spectrogram shown in FIG. 8(C), extraction may be performed from the whole including the consonant and the vowel. From the spectrogram shown in FIG. 8(C), a part corresponding to a consonant and a part corresponding to a vowel may be extracted based on the time when the consonant appears, the frequency range, and the sound. In FIG. 8(C), the sound pressure (spectral intensity) is indicated by color or shading. The area R1 in FIG. 8(C) corresponds to the consonant sound feature information shown in FIG. 8(A). The area R2 in FIG. 8(C) corresponds to the vowel sound feature information shown in FIG. 8(B).

[0068] Fig. 8(D) shows a wavelet generated from the speech data A1. Regions R3 and R4 shown in Fig. 8(D) are parts corresponding to consonants and vowels in the speech data A1 based on the wavelet, respectively.

[0069] Step S40: The evaluation unit 202 evaluates the articulation abnormality of the subject based on the output result obtained by inputting the consonant speech feature information generated by the speech feature information generation unit 201 to the trained model B1. The evaluation unit 202 also evaluates the articulation abnormality of the subject based on the output result obtained by inputting the vowel speech feature information generated by the speech feature information generation unit 201 to the trained model B1. Therefore, in this embodiment, the evaluation of the articulation abnormality is performed in units of phonemes, that is, consonants and vowels.

[0070] Here, when the trained model B1 is trained to output evaluation information indicating the presence or absence of an articulation abnormality, the evaluation unit 202 judges the presence or absence of an articulation abnormality of the subject based on the output result obtained by inputting the consonant speech feature information generated by the speech feature information generation unit 201 to the trained model B1. In this case, the evaluation unit 202 judges the presence or absence of an articulation abnormality of the subject based on the output result obtained by inputting the vowel speech feature information generated by the speech feature information generation unit 201 to the trained model B1.

[0071] On the other hand, when the trained model B1 is trained to output evaluation information indicating the degree of articulation abnormality, the evaluation unit 202 judges the degree of articulation abnormality of the subject based on the output result obtained by inputting the consonant speech feature information generated by the speech feature information generation unit 201 to the trained model B1. In this case, the evaluation unit 202 judges the degree of articulation abnormality of the subject based on the output result obtained by inputting the vowel speech feature information generated by the speech feature information generation unit 201 to the trained model B1.

[0072] Step S50: The output unit 203 outputs the evaluation result C1 by the evaluation unit 202. The output unit 203 outputs the evaluation result C1 to the evaluation result presentation unit 4. The evaluation result presentation unit 4 presents the evaluation result C1 by displaying it as an image such as a graph. With this, the processing unit 20 ends the articulation abnormality evaluation process.

[0073] In addition, when CNN is used as a learning model, the inference process is often unclear. When CNN is used as the trained model B1 as in this embodiment, it may be necessary to interpret which feature value from the spectrogram was used by the CNN to evaluate the speech disorder. In order to interpret the evaluation result, a visualization method may be performed on the spectrogram.

[0074] For example, CNN is trained using consonant speech feature information of normal articulation as a spectrogram. From the trained CNN (i.e., trained model B1), features to be visualized for the consonant speech feature information of normal articulation are extracted based on the visualization method. For the consonant speech feature information of the evaluation target having an articulation disorder, the consonant speech feature information to be evaluated is input to the trained CNN, and features to be visualized are similarly extracted based on the visualization method. By comparing the extracted features between the consonant speech feature information of normal articulation and the consonant speech feature information of the articulation disorder being evaluated, it is possible to interpret which feature was used by the trained CNN (i.e., trained model B1) to evaluate the articulation disorder.

[0075] Next, the results of evaluation of articulation abnormalities by the articulation abnormality inspection device 2 will be described with reference to FIGS. Fig. 9 is a diagram showing an example of an evaluation result of articulation abnormality according to the present embodiment. The evaluation result shown in Fig. 9 is an evaluation result of articulation abnormality for each of distorted sounds (sounds with articulation abnormality) and normal articulation. As the speech data A1 to be evaluated, speech in the case of distorted sounds and speech in the case of normal articulation were used for five monosyllables, "gi", "nya", "nyu", "nyo", and "ri". In the evaluation result shown in Fig. 9, each syllable is classified into a class with articulation abnormality and a class without articulation abnormality.

[0076] According to the evaluation results shown in Figure 9, the accuracy of identifying monosyllables of distorted sounds was approximately 86 percent, and the accuracy of identifying monosyllables of normal speech was approximately 93 percent. The accuracy of identifying monosyllables is the percentage of cases where the five types of monosyllables were correctly classified into classes. On the other hand, the prediction accuracy for distorted sounds was about 98 percent, and for normal speech was about 100 percent. The prediction accuracy for distorted sounds is the percentage of distorted syllables that could be classified into the class with articulation abnormalities. The prediction accuracy for normal speech is the percentage of normally articulated syllables that could be classified into the class without articulation abnormalities. From the evaluation result shown in FIG. 9, it can be seen that the articulation abnormality evaluation process by the articulation abnormality inspection device 2 can classify whether or not there is an articulation abnormality with a high accuracy of 98 percent or more.

[0077] Fig. 10 is a diagram showing another example of the evaluation result of articulation abnormality according to the present embodiment. The evaluation result shown in Fig. 10 is an evaluation result of the degree of articulation abnormality for 43 types of data of distorted sounds obtained from the voice of a person with articulation abnormality as voice data A1. The true class of the degree of articulation abnormality is a class obtained by the applicants subjectively evaluating the degree of articulation abnormality. Fig. 10 also shows the result of evaluating the presence or absence of articulation abnormality for normal articulation of a syllable for which the degree of articulation abnormality was evaluated.

[0078] According to the evaluation results shown in Figure 10, the prediction accuracy of the degree of articulation disorder was about 84.6 percent. The prediction accuracy of the degree of articulation disorder is the percentage of cases in which the degree of articulation disorder was classified to match the true class (the subjective evaluation of the person with articulation disorder).

[0079] As described above, in this embodiment, the articulation abnormality is evaluated in units of phonemes, namely, consonants and vowels, or in units of syllables including consonants and vowels. The articulation abnormality may be evaluated based only on consonants among consonants and vowels. When the articulation abnormality is evaluated based only on consonants, the evaluation unit 202 evaluates the articulation abnormality of the subject based on an output result obtained by inputting consonant speech characteristic information to the trained model B1. The trained model B1 is trained to output evaluation information indicating the evaluation of the articulation abnormality when the consonant speech characteristic information is input. When the articulation abnormality is evaluated based only on consonants, the trained model B1 is not trained using vowel speech characteristic information as an input.

[0080] As described above, the articulation abnormality inspection device 2 according to this embodiment includes the speech data acquisition unit 200, the speech feature information generation unit 201, and the evaluation unit 202. The voice data acquiring section 200 acquires voice data A1 of the subject. The voice feature information generating unit 201 generates voice feature information (in this embodiment, consonant voice feature information) of a portion of the voice data A1 acquired by the voice data acquiring unit 200 that corresponds to a syllable including at least one of a consonant and a vowel. The evaluation unit 202 evaluates the articulation abnormalities of the subject based on the output result obtained by inputting the speech feature information generated by the speech feature information generation unit 201 into a trained model B1 that has been trained to output evaluation information indicating an evaluation of an articulation abnormality when speech feature information is input.

[0081] With this configuration, the speech abnormality inspection device 2 of this embodiment can evaluate speech abnormalities based on syllables that contain at least one of a consonant or a vowel, thereby improving the accuracy of evaluation in technology that uses machine learning to evaluate speech disorders.

[0082] In conventional methods for testing articulation disorders, examiners subjectively evaluate the hearing of patients. Therefore, the examiner must be highly skilled to perform articulation tests. Furthermore, there was no technology to distinguish the severity of articulation disorders. The articulation abnormality inspection device 2 according to this embodiment can objectively and quantitatively evaluate the presence or absence of articulation abnormality and which sound is distorted based on machine learning (AI) from acoustic features (speech feature information) in phoneme units. Evaluating which sound is distorted means, for example, evaluating which of the sounds "shi", "chi", and "ji" is distorted. Furthermore, the articulation abnormality inspection device 2 can encourage the examiner to become aware of the abnormality and improve the examiner's proficiency.

[0083] The evaluation unit 202 may evaluate the articulation abnormality of the test subject based on whether or not a frequency component greater than a predetermined threshold is missing from the frequency components indicated by the speech feature information of the portion corresponding to the consonant. This process is also referred to as evaluation based on missing frequency components. In this case, the evaluation based on missing frequency components by the evaluation unit 202 is performed as screening before performing the processes of steps S30 and S40 in the articulation abnormality evaluation process shown in FIG. 3, for example, as follows.

[0084] After the process of step S20, the evaluation unit 202 calculates a spectrum (also referred to as a cepstrum) based on a cepstrum analysis from a portion corresponding to a consonant extracted from the speech data A1 by the speech feature information generation unit 201. When the evaluation unit 202 determines that a frequency component larger than a predetermined threshold is missing based on the envelope of the spectrum, it determines that the subject has an articulation disorder. When the evaluation unit 202 determines that the subject has an articulation disorder, the processing unit 20 omits the processes of steps S30 and S40 and executes the process of step S50. In this case, in step S50, the output unit 203 outputs an evaluation result C1 indicating that there is an articulation disorder. On the other hand, when the evaluation unit 202 determines that no frequency components greater than the predetermined threshold are missing based on the calculated envelope, it does not determine whether the subject has an articulation disorder. After that, the processing unit 20 executes the processes of steps S30, S40, and S50 as in the above-mentioned embodiment.

[0085] Here, the predetermined threshold used by the evaluation unit 202 for the above-mentioned judgment is, for example, 4 kHz. The threshold may be any value between 3 kHz and 4 kHz. The threshold is determined by comparing the envelope of the spectrum calculated based on the cepstrum analysis between the person with an articulation disorder and the healthy person. Figures 11 and 12 each show the envelope of the spectrum calculated based on the cepstrum analysis from the speech data when the person with an articulation disorder utters the consonant "ki". Meanwhile, Figures 13, 14, and 15 each show the envelope of the spectrum calculated based on the cepstrum analysis from the speech data when the healthy person utters the consonant "ki".

[0086] Comparing the envelope curves for the speech disorder (FIGS. 11 and 12) with those for the healthy subjects (FIGS. 13, 14, and 15), it can be seen that the envelope curves for the speech disorder are missing spectral power in a frequency range greater than 3 kHz to 4 kHz compared to the envelope curves for the healthy subjects. Therefore, it can be seen that the threshold value used by the evaluation unit 202 to determine whether or not a frequency component greater than a predetermined threshold is missing based on the envelope curve is in the range of 3 kHz to 4 kHz. Therefore, if the threshold value is determined to be 4 kHz, which is the upper limit of the range of 3 kHz to 4 kHz, the probability of erroneously determining that a speech disorder is a healthy subject in screening can be reduced compared to a value less than 4 kHz and equal to or greater than 3 kHz.

[0087] The screening may be performed before the consonant portion is extracted in step S20. In that case, after step S10 is performed, the evaluation unit 202 calculates a spectrum based on a cepstrum analysis from the voice data A1 acquired by the voice data acquisition unit 200, and calculates an envelope of the spectrum. When the evaluation unit 202 determines that a frequency component larger than a predetermined threshold is missing based on the calculated envelope, it determines that the subject has an articulation disorder. When the evaluation unit 202 determines that the subject has an articulation disorder, the processing unit 20 omits the processes of steps S20, S30, and S40, and performs the process of step S50.

[0088] Furthermore, after the evaluation unit 202 performs the evaluation based on the missing frequency components, the processing unit 20 may execute the processes of steps S30, S40, and S50 regardless of the evaluation result. In this case, the output unit 203 outputs both the result of the evaluation based on the missing frequency components and the evaluation result C1 by machine learning.

[0089] Second embodiment The second embodiment of the present invention will now be described in detail with reference to the drawings. In this embodiment, a case will be described in which the articulation abnormality inspection device is used for training a person with an articulation disorder to correct his or her articulation abnormality. The articulation abnormality inspection device according to this embodiment is referred to as an articulation abnormality inspection device 2a. The same components as those in the first embodiment described above are denoted by the same reference numerals, and descriptions of the same components and operations may be omitted.

[0090] 16 is a diagram showing an example of the configuration of an articulation abnormality correction support system 5a according to this embodiment. The articulation abnormality correction support system 5a includes an articulation abnormality inspection device 2a, a sound collection unit 3a, an evaluation result presentation unit 6a, and a character presentation unit 7a. The articulation abnormality correction support system 5a is a system for carrying out speech training for correcting articulation abnormalities at home or the like, for example, by a person with an articulation disorder.

[0091] The articulation abnormality inspection device 2a performs a process for evaluating articulation abnormalities from the speech data A1. The articulation abnormality inspection device 2a is, for example, a PC.

[0092] The sound collection unit 3a collects the voice of the subject. The sound collection unit 3a is, for example, a microphone. The sound collection unit 3a outputs the collected voice of the subject to the articulation abnormality inspection device 2a in real time as voice data A1. The sound collection unit 3a may be a recorder. When the sound collection unit 3a is a recorder, the sound collection unit 3a records the collected voice of the subject and outputs the recorded voice as voice data A1 to the articulation abnormality inspection device 2a.

[0093] The evaluation result presenting unit 6a presents the evaluation result by the evaluation unit 202. The evaluation result presenting unit 6a presents the evaluation result by displaying one or more of an image such as a graph, and letters and numbers. The evaluation result presenting unit 6a may present the evaluation result by a stimulus such as sound or light. For example, the evaluation result presenting unit 6a displays the letters corresponding to a sound determined to have an articulation abnormality in a different manner from the letters corresponding to other sounds determined to have no articulation abnormality. For example, the evaluation result presenting unit 6a displays the letters corresponding to a sound determined to have an articulation abnormality in a different color from the letters corresponding to other sounds determined to have no articulation abnormality. The evaluation result presenting unit 6a may display the degree of articulation abnormality using a bar graph (gauge) or an image, which is an icon that differs according to the degree of severity, as described above.

[0094] The character presentation unit 7a presents one or more of a monosyllable, a word, or a short sentence to be spoken by the test subject. The one or more of the monosyllable, the word, or the short sentence is also referred to as a practice phrase.

[0095] In this embodiment, the evaluation result presentation unit 6a and the character presentation unit 7a are integrated into one display. Alternatively, the evaluation result presentation unit 6a and the character presentation unit 7a may be separately implemented as two displays.

[0096] In the articulation abnormality correction support system 5a, practice phrases according to the degree of the articulation abnormality are sequentially presented to the character presentation unit 7a. Based on the evaluation results by the articulation abnormality inspection device 2a, the presence or absence of articulation abnormality (distorted sound) and the degree of the articulation abnormality are presented on the screen of the evaluation result presentation unit 6a, and feedback is given to the subject. The evaluation result presentation unit 6a displays a graph based on the evaluation results by the articulation abnormality inspection device 2a. The graph shows, for example, the daily transition of the frequency of articulation abnormality and the degree of the articulation abnormality. In the articulation abnormality correction support system 5a, feedback to the subject and the display of the graph can encourage the subject to become aware of the articulation abnormality, and can improve the effectiveness of the speech training to correct the articulation abnormality.

[0097] 17 is a diagram showing the functional configuration of an articulation abnormality inspection device 2a according to this embodiment. Comparing the articulation abnormality inspection device 2a according to this embodiment (FIG. 17) with the articulation abnormality inspection device 2 according to the first embodiment (FIG. 2), the difference is that character data D1 is stored in a processing unit 20a and a storage unit 21.

[0098] The processing unit 20a includes a voice data acquisition unit 200, a voice feature information generation unit 201, an evaluation unit 202a, an output unit 203, and a character output unit 204a. The functions of the voice data acquisition unit 200, the voice feature information generation unit 201, and the output unit 203 are the same as those in the first embodiment.

[0099] The character output unit 204a outputs character data D1 to the character presentation unit 7a. The character data D1 indicates one or more of a monosyllable, a word, or a short sentence presented by the character presentation unit 7a. The character data D1 is stored in the storage unit 21 in advance.

[0100] The evaluation unit 202a has a feedback function and a graph display function in addition to the functions of the evaluation unit 202. For example, as the feedback function, the evaluation unit 202a causes the evaluation result presentation unit 6a to display characters corresponding to sounds determined to have an articulation abnormality in a manner different from characters corresponding to other sounds determined not to have an articulation abnormality. In addition, as the graph display function, the evaluation unit 202a causes the evaluation result presentation unit 6a to display, for example, a graph showing the daily transition of the frequency and degree of articulation abnormality.

[0101] The character presenting unit 7a may be omitted from the configuration of the articulation disorder correction support system 5a. In that case, the test subject utters the practice phrase while looking at a piece of paper on which the practice phrase is written, for example.

[0102] As described above, the articulation abnormality correction support system 5a according to this embodiment includes the articulation abnormality inspection device 2a, the sound collection unit 3a, and the evaluation result presentation unit 6a. The sound collection section 3a collects the voice of the subject. The evaluation result presenting unit 6a presents the evaluation results obtained by the evaluation unit 202a.

[0103] With this configuration, the articulation disorder correction support system 5a according to this embodiment can collect the voice of the subject and present the evaluation result. In the articulation disorder correction support system 5a, by including not only the presence or absence of articulation disorder symptoms but also their severity in the evaluation results, it is possible to provide feedback on changes in the severity to a supporter of the person with the articulation disorder or to the person with the articulation disorder himself / herself.

[0104] In the past, it was difficult for people with articulation disorders to notice distorted sounds, and it took time for the effects of practice to appear. The articulation disorder correction support system 5a instantly provides feedback on the presence or absence of distorted sounds and their severity, which promotes the awareness of the person with the articulation disorder and makes it easier for the effects of practice to appear. In addition, in conventional support, the supporter provides verbal feedback to the person with an articulation disorder to encourage them to become aware of distorted sounds. With the articulation disorder correction support system 5a, feedback can be obtained even without a supporter, so the person with an articulation disorder can easily practice at home, etc.

[0105] In each of the above-described embodiments, voice data included in the learning data used for machine learning may be superimposed with voice data corresponding to the background (hereinafter referred to as background data). In this case, the background data is superimposed on the speech disorder data and the healthy subject voice data, respectively, and machine learning is performed. The background data is, for example, pink noise, white noise, or environmental sound.

[0106] As the environmental sounds used for the background data, it is preferable to use environmental sounds collected in advance in an environment similar to the usage environment. The usage environment is an environment in which the articulation abnormality inspection device 2 is used to actually evaluate articulation abnormalities, or an environment in which a person with an articulation disorder practices to correct articulation abnormalities using the articulation abnormality correction support system 5a (articulation abnormality inspection device 2a).

[0107] The usage environment in which the articulation abnormality inspection device 2 is used to actually evaluate articulation abnormalities is, for example, a school or a medical examination at a clinic. For example, when the usage environment is a school, environmental sounds collected at the school are used as background data. In this case, it is preferable that environmental sounds are collected at a plurality of schools and used as background data.

[0108] The environment in which the articulation disorder correction support system 5a (articulation disorder inspection device 2a) is used to practice for correcting articulation disorders is, for example, the home of a person with an articulation disorder, or a clinic or facility for correcting articulation disorders. For example, it is preferable that environmental sounds are collected in each of a plurality of homes, a plurality of clinics, or a plurality of facilities and used as background data.

[0109] Therefore, speech data (background data) obtained by collecting environmental sounds in advance in an environment in which the articulation abnormality inspection device 2, 2a is actually used may be used for machine learning, with the speech data superimposed on the speech disorder data and the speech data of healthy subjects. In this case, the trained model is trained to output evaluation information indicating an evaluation of an articulation abnormality when speech feature information generated from a waveform (consonant part or vowel part) of a part corresponding to a syllable extracted from the speech data superimposed with the background is input. By performing machine learning with the background data superimposed on the speech data included in the training data, the accuracy of the evaluation of an articulation abnormality by the articulation abnormality inspection device 2, 2a can be further improved.

[0110] In addition, a part of the articulation abnormality inspection device 2, 2a in the above-mentioned embodiment, for example, the voice data acquisition unit 200, the voice feature information generation unit 201, the evaluation unit 202, 202a, the output unit 203, and the character output unit 204a may be realized by a computer. In that case, a program for realizing this control function may be recorded in a computer-readable recording medium, and the program recorded in the recording medium may be read into a computer system and executed to realize the control function. In addition, the "computer system" here refers to a computer system built into the articulation abnormality inspection device 2, 2a, and includes hardware such as an operating system (OS) and peripheral devices. In addition, the "computer-readable recording medium" refers to a portable medium such as a flexible disk, a magneto-optical disk, a ROM (Read Only Memory), a CD-ROM (Compact Disc-Read Only Memory), and a storage device such as a hard disk built into a computer system. Furthermore, the term "computer-readable recording medium" may include a medium that dynamically stores a program for a short period of time, such as a communication line when transmitting a program via a network such as the Internet or a communication line such as a telephone line, and a medium that stores a program for a certain period of time, such as a volatile memory inside a computer system that serves as a server or client in such a case. The above program may be one that realizes part of the above-mentioned functions, or may be one that can realize the above-mentioned functions in combination with a program already recorded in the computer system. In addition, a part or the whole of the speech abnormality inspection devices 2, 2a in the above-mentioned embodiments may be realized as an integrated circuit such as an LSI (Large Scale Integration). Each functional block of the speech abnormality inspection devices 2, 2a may be individually processed, or a part or the whole may be integrated and processed. The integrated circuit method is not limited to LSI, and may be realized by a dedicated circuit or a general-purpose processor. In addition, when an integrated circuit technology that replaces LSI appears due to the progress of semiconductor technology, an integrated circuit based on that technology may be used.

[0111] Although one embodiment of the present invention has been described in detail above with reference to the drawings, the specific configuration is not limited to the above, and various design changes, etc. are possible within the scope that does not deviate from the gist of the present invention. [Explanation of symbols]

[0112] 2, 2a... Articulation abnormality inspection device, 200... Voice data acquisition unit, 201... Voice feature information generation unit, 202, 202a... Evaluation unit, A1... Voice data, B1... Trained model

Claims

1. A voice data acquisition unit that acquires voice data of a subject; a voice feature information generating unit that generates voice feature information of a portion of the voice data acquired by the voice data acquiring unit that corresponds to a syllable including at least one of a consonant and a vowel; an evaluation unit that evaluates the articulation abnormality of the subject based on an output result obtained by inputting the speech feature information generated by the speech feature information generation unit into a trained model that has been trained to output evaluation information indicating an evaluation of an articulation abnormality when the speech feature information is input; A speech abnormality inspection device comprising:

2. The voice feature information generating unit extracts a portion corresponding to a syllable including at least one of a consonant and a vowel from the voice data, and generates the voice feature information from the extracted portion. The speech abnormality inspection device according to claim 1.

3. The voice feature information generating unit extracts a portion corresponding to a syllable including at least one of a consonant and a vowel from the voice data based on a time and a frequency range in which the syllable appears in the voice data. The speech abnormality inspection device according to claim 2.

4. The voice feature information generating unit extracts a portion corresponding to a syllable including at least one of a consonant and a vowel from the voice data based on a sound pressure or a volume indicated by the voice data. The speech abnormality inspection device according to claim 3.

5. The audio feature information is a spectrogram or a spectral envelope. The speech abnormality inspection device according to claim 1.

6. The speech feature information is speech feature information of a portion corresponding to a consonant. The speech abnormality inspection device according to claim 1.

7. The evaluation unit evaluates the articulation abnormality of the subject based on whether or not a frequency component greater than a predetermined threshold is missing from the frequency components indicated by the voice feature information. The speech abnormality inspection device according to claim 6.

8. The evaluation information indicates the presence or absence of an articulation disorder, The evaluation unit judges the presence or absence of an articulation disorder of the subject based on an output result obtained by inputting the speech feature information generated by the speech feature information generation unit to the trained model. The speech abnormality inspection device according to claim 1.

9. The evaluation information indicates a degree of articulation disorder, The evaluation unit judges the degree of articulation disorder of the subject based on an output result obtained by inputting the speech feature information generated by the speech feature information generation unit to the trained model. The speech abnormality inspection device according to claim 1.

10. The speech abnormality inspection device according to any one of claims 1 to 9, A sound collection unit that collects the voice of the subject; an evaluation result presentation unit that presents an evaluation result by the evaluation unit; Equipped A system to assist in correcting articulation disorders.

11. A character presentation unit is further provided to present one or more of a monosyllable, a word, or a short sentence to be spoken by the subject. The articulation disorder correction support system according to claim 10.

12. A voice data acquisition step of acquiring voice data of a subject; a voice feature information generating step of generating voice feature information of a portion of the voice data acquired by the voice data acquiring step, the portion corresponding to a syllable including at least one of a consonant and a vowel; an evaluation step of evaluating the articulation abnormality of the subject based on an output result obtained by inputting the speech feature information generated by the speech feature information generation step into a trained model trained to output evaluation information indicating an evaluation of an articulation abnormality when the speech feature information is input; The method for testing speech abnormalities comprises:

13. On the computer, A voice data acquisition step of acquiring voice data of a subject; a voice feature information generating step of generating voice feature information of a portion of the voice data acquired by the voice data acquiring step, the portion corresponding to a syllable including at least one of a consonant and a vowel; an evaluation step of evaluating the articulation abnormality of the subject based on an output result obtained by inputting the speech feature information generated by the speech feature information generation step into a trained model trained to output evaluation information indicating an evaluation of an articulation abnormality when the speech feature information is input; A program for executing the above.

Citation Information

Patent Citations

  • Abnormal articulation detection method, abnormal articulation detection device, and program

    JP2023002421A

  • Articulation abnormality detection method, articulation abnormality detection device, and program

    JP2023036486A