Method for detecting speech disorder, speech disorder detection device, and program
The method employs a DNN to calculate speaker identity and acoustic statistics for speech disorder detection, addressing the reliance on training data limitations and enhancing detection accuracy by leveraging pre-trained models and additional acoustic features.
Patent Information
- Application Number
- JP2021103673
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-06-22
- Publication Date
- 2025-07-30
- Estimated Expiration
- 2041-06-22
AI Technical Summary
Existing methods for detecting speech disorders rely heavily on the quantity and quality of training data, making it difficult to achieve accurate detection when sufficient data is not available, especially for pathological conditions not included in the training set.
A method using a learned Deep Neural Network (DNN) to calculate acoustic feature quantities, determine speaker identity through first and second speaker feature quantities, and assess similarity to detect speech disorders without requiring additional training data, incorporating acoustic statistics like pitch variation, waveform periodicity, and skewness for enhanced accuracy.
Enables high-accuracy detection of speech disorders by leveraging pre-trained DNNs for speaker identity recognition, reducing reliance on training data volume and utilizing acoustic statistics, thus improving detection precision.
Smart Images

Figure 0007715339000001 
Figure 0007715339000002 
Figure 0007715339000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to a method for detecting speech disorder, a device for detecting speech disorder, and a program.
Background Art
[0002] As a technique for detecting speech disorder (also called speech impairment), which is a state where normal pronunciation cannot be made, for example, Non-Patent Document 1 discloses a method of analyzing speech using deep learning.
Prior Art Documents
Non-Patent Documents
[0003]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, in such a method using learning, the detection accuracy depends on the quantity and quality of training data (also called teacher data), but it is difficult to secure a large amount of training data at the time of speech disorder. As a result, for example, it is impossible to detect a speech disorder of a pathological condition that is not included in the training data.
[0005] An object of the present disclosure is to provide a method for detecting speech disorder, a device for detecting speech disorder, and a program that can improve detection accuracy without depending on the amount of training data at the time of speech disorder.
Means for Solving the Problems
[0006] The method for detecting speech disorder according to one aspect of the present disclosure calculates acoustic feature quantities from the utterance data of a speaker, and uses a learned DNN (Deep Neural Network) to calculate, from the acoustic feature quantities, a first speaker feature quantity representing the speaker identity of the utterance data, calculates the similarity between the second speaker feature quantity, which is the speaker feature quantity of the speaker in a normal state, and the first speaker feature quantity, and determines the speech disorder of the speaker based on the similarity. and the acoustic statistic quantity and calculates, from the acoustic feature quantities, a first speaker feature quantity representing the speaker identity of the utterance data, calculates the similarity between the second speaker feature quantity, which is the speaker feature quantity of the speaker in a normal state, and the first speaker feature quantity, and determines the speech disorder of the speaker based on the similarity. and the acoustic statistic quantity Based on the similarity, the speech disorder of the speaker is determined.
Advantages of the Invention
[0007] The present disclosure can provide a method for detecting speech disorder, a device for detecting speech disorder, and a program that can improve the detection accuracy without depending on the amount of training data in the case of speech disorder.
Brief Description of the Drawings
[0008]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Modes for Carrying Out the Invention
[0009] The method for detecting speech disorder according to one aspect of the present disclosure calculates acoustic feature quantities from the utterance data of a speaker, and uses a trained DNN (Deep Neural Network) to calculate, from the acoustic feature quantities, a first speaker feature quantity representing the speaker identity of the utterance data. The method calculates the similarity between a second speaker feature quantity, which is the speaker feature quantity of the speaker in a healthy state, and the first speaker feature quantity, and determines the speech disorder of the speaker based on the similarity.
[0010] According to this, since the method for detecting speech disorder does not require training data during speech disorder, a detection process that does not depend on the amount of training data during speech disorder can be realized. Further, the method for detecting speech disorder calculates a first speaker feature quantity using a trained DNN (Deep Neural Network) for identifying speaker identity, and determines speech disorder based on the similarity between this first speaker feature quantity and the second speaker feature quantity in a healthy state, thereby enabling detection of speech disorder with high accuracy with a simple configuration.
[0011] For example, in the detection of the speech disorder, when the similarity is lower than a predetermined first threshold value, it may be determined that the speaker has a speech disorder.
[0012] For example, in the calculation of the acoustic feature quantities, a plurality of acoustic feature quantities including the acoustic feature quantities are calculated from each of a plurality of utterance data of the speaker including the utterance data. In the calculation of the first speaker feature quantity, a plurality of first speaker feature quantities including the first speaker feature quantity are calculated from the plurality of acoustic feature quantities using the trained DNN. In the calculation of the similarity, a plurality of similarities between the second speaker feature quantity including the similarity and the plurality of first speaker feature quantities are calculated. In the detection of the speech disorder of the speaker, the variance of the plurality of similarities is calculated, and when the variance is larger than a predetermined second threshold value, it may be determined that the speaker has a speech disorder.
[0013] According to this, the method for detecting speech disorder can detect speech disorder with high accuracy by utilizing the property that it becomes difficult to repeat the same phrase during speech disorder.
[0014] For example, the speech disorder detection method may further calculate an acoustic statistic from the speech data, and in detecting a speech disorder of the speaker, determine the speech disorder of the speaker based on the similarity and the acoustic statistic.
[0015] According to this, the speech disorder detection method can detect a speech disorder with high accuracy by making a determination taking into account an acoustic statistic in addition to the similarity between the first speaker feature amount and the second speaker feature amount in a healthy state.
[0016] For example, the acoustic statistic includes pitch variation, and in detecting a speech disorder of the speaker, it may be determined that the possibility of a speech disorder is higher as the pitch variation is smaller.
[0017] For example, the acoustic statistic includes waveform periodicity, and in detecting a speech disorder of the speaker, it may be determined that the possibility of a speech disorder is higher as the waveform periodicity is lower.
[0018] For example, the acoustic statistic includes skewness, and in detecting a speech disorder of the speaker, it may be determined that the possibility of a speech disorder is higher as the skewness is larger.
[0019] A speech disorder detection device according to an aspect of the present disclosure includes an acoustic feature amount calculation unit that calculates an acoustic feature amount from speech data of a speaker, a speaker feature amount calculation unit that calculates a first speaker feature amount representing the speaker property of the speech data from the acoustic feature amount using a learned DNN (Deep Neural Network), a similarity calculation unit that calculates a similarity between a second speaker feature amount that is a speaker feature amount of the speaker in a healthy state and the first speaker feature amount, and a speech disorder determination unit that determines a speech disorder of the speaker based on the similarity.
[0020] According to this, since the speech disorder detection device does not require training data during speech disorders, a detection process that does not depend on the amount of training data during speech disorders can be realized. Furthermore, the speech disorder detection device calculates a first speaker feature amount using a learned DNN for identifying speaker characteristics, and determines speech disorders based on the similarity between this first speaker feature amount and a second speaker feature amount during normal speech, enabling the detection of speech disorders with high accuracy in a simple configuration.
[0021] A program according to an aspect of the present disclosure causes a computer to execute the speech disorder detection method.
[0022] Note that these general or specific aspects may be implemented by a system, a method, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or may be implemented by any combination of a system, a method, an integrated circuit, a computer program, and a recording medium.
[0023] Hereinafter, embodiments will be specifically described with reference to the drawings. Note that the embodiments described below are all specific examples of the present disclosure. The numerical values, shapes, materials, components, arrangement positions and connection forms of the components, steps, order of steps, etc. shown in the following embodiments are merely examples and are not intended to limit the present disclosure. In addition, among the components in the following embodiments, the components not described in the independent claims are described as optional components.
[0024] (Embodiment 1) FIG. 1 is a block diagram showing the configuration of a speech disorder detection device 100 according to the present embodiment. The speech disorder detection device 100 detects a speech disorder of a speaker (user). That is, the speech disorder detection device 100 determines whether the speaker has a speech disorder (or the possibility of a speech disorder). For example, the speech disorder detection device 100 is included in a terminal device such as a smartphone or a tablet terminal. Note that the functions of the speech disorder detection device 100 may be realized by a single device or by a plurality of devices. For example, some functions of the speech disorder detection device 100 may be realized by a terminal device, and some other functions may be realized by a server or the like that can communicate with the terminal device.
[0025] As shown in FIG. 1, the speech disorder detection device includes an audio acquisition unit 101, an acoustic feature quantity calculation unit 102, a speaker feature quantity calculation unit 103, a storage unit 104, a similarity calculation unit 105, a speech disorder determination unit 106, and an output unit 107.
[0026] The audio acquisition unit 101 acquires utterance data that is audio data of the speaker's utterance. For example, the audio acquisition unit 101 is a microphone, and generates utterance data by converting the acquired audio into an audio signal. Note that the audio acquisition unit 101 may acquire utterance data generated outside the speech disorder detection device 100.
[0027] The acoustic feature quantity calculation unit 102 calculates an acoustic feature quantity for the speech of the utterance from the utterance data. For example, the acoustic feature quantity calculation unit 102 calculates MFCC (Mel Frequency Cepstral Coefficient), which is a feature quantity of the speech of the utterance, as an acoustic feature quantity from the utterance data. MFCC is a feature quantity representing the vocal tract characteristics of the speaker and is generally used in speech recognition. More specifically, MFCC is an acoustic feature quantity obtained by analyzing the frequency spectrum of speech based on human auditory characteristics. Note that the acoustic feature quantity calculation unit 102 may calculate, as an acoustic feature quantity, the result of applying a mel filter bank to the speech signal of the utterance from the utterance data, or may calculate the spectrogram of the speech signal of the utterance as an acoustic feature quantity.
[0028] The speaker feature quantity calculation unit 103 extracts a first speaker feature quantity for identifying the speaker of the utterance indicated by the utterance data from the acoustic feature quantities calculated from the utterance data. In other words, the first speaker feature quantity represents the speaker identity of the utterance data. More specifically, the speaker feature quantity calculation unit 103 extracts the first speaker feature quantity from the acoustic feature quantities using a pre-trained DNN.
[0029] For example, the speaker feature quantity calculation unit 103 extracts the first speaker feature quantity using the x-vector method. Here, the x-vector method is a method for calculating a speaker feature quantity, which is a speaker-specific feature called an x-vector. FIG. 2 is a block diagram showing a configuration example of the speaker feature quantity calculation unit 103. As shown in FIG. 2, the speaker feature quantity calculation unit 103 includes a frame connection processing unit 201 and a DNN 202.
[0030] The frame connection processing unit 201 connects a plurality of acoustic feature quantities and outputs the obtained acoustic feature quantities to the DNN 202. For example, the frame connection processing unit 201 connects a plurality of frames of MFCCs, which are acoustic feature quantities, and outputs them to the input layer of the DNN 202. For example, the frame connection processing unit 201 generates a 1200-dimensional vector by connecting 50 frames of MFCC parameters each consisting of 24-dimensional feature quantities for each frame, and outputs the generated vector to the input layer of the DNN 202.
[0031] The DNN 202 is a pre-trained machine learning model that outputs a first speaker feature quantity corresponding to the input acoustic feature quantities. In the example shown in FIG. 2, the DNN 202 is a neural network including an input layer, a plurality of intermediate layers, and an output layer. Also, the DNN 202 is pre-generated by machine learning using a plurality of training data 203. Each of the plurality of training data 203 is data that associates information for identifying a speaker with the speaker's utterance data. That is, the DNN 202 is a pre-trained model that takes utterance data as input and outputs information (speaker label) for identifying the speaker of the utterance data. However, in the present embodiment, the DNN 202 outputs a first speaker feature quantity generated as intermediate data.
[0032] Specifically, the output layer consists of nodes that output speaker labels corresponding to the number of speakers included in the training data 203. The plurality of intermediate layers consists of, for example, two to three intermediate layers, and includes an intermediate layer that calculates the first speaker feature amount. The intermediate layer that calculates the first speaker feature amount outputs the calculated first speaker feature amount as the output of the DNN 202.
[0033] The storage unit 104 is composed of a rewritable non-volatile memory such as a hard disk drive or a solid state drive, for example. The storage unit 104 stores the second speaker feature amount, which is the first speaker feature amount when the speaker is in a healthy state.
[0034] The similarity calculation unit 105 calculates the similarity between the first speaker feature amount output from the speaker feature amount calculation unit 103 and the second speaker feature amount stored in the storage unit 104. For example, the similarity calculation unit 105 calculates the cosine distance (also referred to as cosine similarity), which indicates the angle between the vectors of the first speaker feature amount and the second speaker feature amount, by calculating the cosine using the inner product in the vector space model, as the similarity. In this case, the larger the numerical value of the angle between the vectors, the lower the similarity. Note that the similarity calculation unit 105 may calculate a cosine distance that takes a value from -1 to 1 using the inner product of the vector indicating the first speaker feature amount and the vector indicating the second speaker feature amount as the similarity. In this case, the larger the numerical value indicating the cosine distance, the higher the similarity.
[0035] The articulation disorder determination unit 106 determines whether the speaker has an articulation disorder based on the similarity calculated by the similarity calculation unit 105. For example, the articulation disorder determination unit 106 determines that there is an articulation disorder when the similarity is lower than a predetermined threshold value. Note that the articulation disorder determination unit 106 may determine the possibility of an articulation disorder instead of determining whether there is an articulation disorder. For example, the articulation disorder determination unit 106 may determine that the lower the similarity, the higher the possibility of an articulation disorder. Note that the result of the determination may be a multi-stage classification such as "possible", "high possibility", "very high possibility", or a numerical value indicating the possibility.
[0036] The output unit 107 notifies the speaker of the determination result obtained by the speech disorder determination unit 106. For example, the output unit 107 is a display or a speaker provided in the terminal device, and notifies the speaker of the determination result by display or voice. Note that the output unit 107 may output the determination result to an external device.
[0037] Hereinafter, the speech disorder detection process by the speech disorder detection device 100 will be described. FIG. 3 is a flowchart of the speech disorder detection process by the speech disorder detection device 100. Here, a case where one speaker is registered in the speech disorder detection device 100 in advance will be described.
[0038] First, the speech disorder detection device 100 instructs the speaker (user) to utter a predetermined phrase (S101). For example, this instruction is given by display or voice.
[0039] Next, the voice acquisition unit 101 acquires the utterance data of the above phrase uttered by the speaker according to the instruction (S102). Next, the acoustic feature amount calculation unit 102 calculates the acoustic feature amount from the utterance data (S103). Next, the speaker feature amount calculation unit 103 calculates the first speaker feature amount from the acoustic feature amount (S104). Specifically, the speaker feature amount calculation unit 103 outputs the first speaker feature amount corresponding to the input acoustic feature amount.
[0040] Next, the similarity calculation unit 105 calculates the similarity between the first speaker feature amount output from the speaker feature amount calculation unit 103 and the second speaker feature amount stored in the storage unit 104 (S105). For example, the second speaker feature amount is the first speaker feature amount when it is determined that there is no speech disorder in the past speech disorder detection process. Note that the second speaker feature amount may be calculated from a plurality of first speaker feature amounts obtained in a plurality of past speech disorder detection processes. For example, it may be an average value or a median value of a plurality of first speaker feature amounts obtained in a plurality of past speech disorder detection processes.
[0041] The articulation disorder determination unit 106 determines the articulation disorder of the speaker based on the similarity calculated by the similarity calculation unit 105. Specifically, the articulation disorder determination unit 106 compares the similarity with a predetermined threshold (S106). When the similarity is less than the predetermined threshold (Yes in S106), the articulation disorder determination unit 106 determines that the speaker has an articulation disorder (S107). When the similarity is equal to or greater than the predetermined threshold (No in S106), the articulation disorder determination unit 106 determines that the speaker does not have an articulation disorder (normal) (S108). Note that the articulation disorder determination unit 106 may determine the possibility of an articulation disorder instead of determining whether there is an articulation disorder. For example, the articulation disorder determination unit 106 may determine that the lower the similarity, the higher the possibility of an articulation disorder.
[0042] Next, the output unit 107 outputs the determination result obtained by the articulation disorder determination unit 106 (S109). For example, the output unit 107 notifies the speaker of the determination result obtained by the articulation disorder determination unit 106.
[0043] In the above description, an example in which one speaker is registered in advance is shown, but a plurality of speakers may be registered. In this case, the second speaker feature amount for each speaker is stored in the storage unit 104. Further, information for identifying the speaker is input to the articulation disorder detection device 100, and the above processing is performed using the second speaker feature amount of the identified speaker.
[0044] As described above, the articulation disorder detection device 100 calculates an acoustic feature amount from the utterance data of the speaker. The articulation disorder detection device 100 calculates a first speaker feature amount representing the speaker identity of the utterance data from the acoustic feature amount using a learned DNN (Deep Neural Network). The articulation disorder detection device 100 calculates the similarity between the second speaker feature amount, which is the speaker feature amount when the speaker is normal, and the first speaker feature amount. The articulation disorder detection device 100 determines the articulation disorder of the speaker based on the similarity. For example, when the similarity is lower than a predetermined first threshold, the articulation disorder detection device 100 determines that the speaker has an articulation disorder.
[0045] That is, the speech disorder detection device 100 uses the first speaker feature amount obtained by the learned DNN that calculates the first speaker feature amount representing speaker identity from the utterance data, and utilizes the fact that the first speaker feature amount changes during speech disorders compared to normal speech to detect the speaker's speech disorder. In this way, by reusing the already created DNN for speaker identification, there is no need to create a new DNN for determining speech disorders. Therefore, the speech disorder detection device 100 does not require training data during speech disorders, and can realize a detection process that does not depend on the amount of training data during speech disorders.
[0046] (Embodiment 2) FIG. 4 is a block diagram of the speech disorder detection device 100A according to the present embodiment. The speech disorder detection device 100A shown in FIG. 4 mainly has a different function of the speech disorder determination unit 106A from the speech disorder determination unit 106 shown in FIG. 1.
[0047] The speech disorder detection device 100A calculates the similarity of each of a plurality of utterance data corresponding to a plurality of utterances. The speech disorder determination unit 106A calculates the variance of the calculated plurality of similarities, and determines the speech disorder of the speaker based on the calculated variance.
[0048] FIG. 5 is a flowchart of the speech disorder detection process by the speech disorder detection device 100A. Here, the case where one speaker is registered in advance in the speech disorder detection device 100A will be described.
[0049] First, the speech disorder detection device 100A instructs the speaker (user) to utter the same phrase a plurality of times as predetermined (S121). For example, this instruction is given by display or voice.
[0050] Next, the voice acquisition unit 101 acquires the utterance data of the above phrase uttered by the speaker according to the instruction (S122). Next, the acoustic feature amount calculation unit 102 calculates the acoustic feature amount from the utterance data (S123). Next, the speaker feature amount calculation unit 103 calculates the first speaker feature amount from the acoustic feature amount (S124). Next, the similarity calculation unit 105 calculates the similarity between the first speaker feature amount output from the speaker feature amount calculation unit 103 and the second speaker feature amount stored in the storage unit 104 (S125). Also, until the processing for a plurality of utterances is completed, the processing of steps S122 to S125 is repeated (S126), and a plurality of similarities corresponding to the plurality of utterances are calculated.
[0051] Next, the articulation disorder determination unit 106A calculates the variance of the calculated plurality of similarities (S127), and determines whether or not the calculated variance is equal to or greater than a predetermined first threshold value (S128). When the variance is equal to or greater than the first threshold value (Yes in S128), the articulation disorder determination unit 106A determines that the speaker has an articulation disorder (S130).
[0052] On the other hand, when the variance is less than the first threshold (No in S128), the articulation disorder determination unit 106A determines whether all of the plurality of similarities are less than the second threshold (S129). When all of the plurality of similarities are less than the second threshold (Yes in S129), the articulation disorder determination unit 106A determines that the speaker has an articulation disorder (S130). Further, when at least one of the plurality of similarities is equal to or greater than the second threshold (No in S129), the articulation disorder determination unit 106A determines that the speaker does not have an articulation disorder (is normal) (S131). Note that the articulation disorder determination unit 106A may determine whether at least one of the plurality of similarities is less than the second threshold in step S129. That is, when at least one of the plurality of similarities is less than the second threshold, the articulation disorder determination unit 106A determines that the speaker has an articulation disorder (S130), and when all of the plurality of similarities are equal to or greater than the second threshold, the articulation disorder determination unit 106A may determine that the speaker does not have an articulation disorder (S131). Alternatively, the articulation disorder determination unit 106A may determine whether a first evaluation value calculated from the plurality of similarities is less than the second threshold in step S129. That is, when the first evaluation value is less than the second threshold, the articulation disorder determination unit 106A determines that the speaker has an articulation disorder (S130), and when the first evaluation value is equal to or greater than the second threshold, the articulation disorder determination unit 106A may determine that the speaker does not have an articulation disorder (S131). This first evaluation value is, for example, the average value, median value, maximum value, or minimum value of the plurality of similarities.
[0053] Also, the order of steps S128 and S129 may be reversed. That is, the articulation disorder determination unit 106A may determine that there is an articulation disorder when at least one of the first condition that the variance is equal to or greater than the first threshold and the second condition that all of the plurality of similarities are less than the second threshold is satisfied, and determine that there is no articulation disorder when neither the first condition nor the second condition is satisfied. Note that the articulation disorder determination unit 106A may determine that there is an articulation disorder when both the first condition and the second condition are satisfied, and determine that there is no articulation disorder when at least one of the first condition and the second condition is not satisfied.
[0054] Note that the articulation disorder determination unit 106A may determine the possibility of an articulation disorder instead of determining whether there is an articulation disorder. For example, the articulation disorder determination unit 106A may calculate a second evaluation value based on the variance and the first evaluation value, and determine the possibility of an articulation disorder based on the second evaluation value. For example, the articulation disorder determination unit 106A calculates the second evaluation value by weighted addition of the reciprocal of the variance and the first evaluation value. Further, the articulation disorder determination unit 106A determines that the lower the second evaluation value, the higher the possibility of an articulation disorder. That is, the articulation disorder determination unit 106A determines that the higher the variance, the higher the possibility of an articulation disorder, and the lower the first evaluation value, the higher the possibility of an articulation disorder.
[0055] Next, the output unit 107 outputs the determination result obtained by the articulation disorder determination unit 106A (S132). For example, the output unit 107 notifies the speaker of the determination result obtained by the articulation disorder determination unit 106A.
[0056] As described above, the articulation disorder detection device 100A calculates a plurality of acoustic feature amounts from each of a plurality of utterance data of a speaker, calculates a plurality of first speaker feature amounts from the plurality of acoustic feature amounts using a learned DNN, and calculates a plurality of similarities between the second speaker feature amount and the plurality of first speaker feature amounts. The articulation disorder detection device 100A calculates the variance of the plurality of similarities, and determines that the speaker has an articulation disorder when the variance is greater than a predetermined second threshold value.
[0057] Thereby, the articulation disorder detection device 100A can detect an articulation disorder with high accuracy by utilizing the property that it becomes difficult to repeat the same phrase during an articulation disorder.
[0058] (Embodiment 3) FIG. 6 is a block diagram of an articulation disorder detection device 100B according to the present embodiment. The articulation disorder detection device 100B shown in FIG. 6 includes an acoustic statistic calculation unit 108 in addition to the configuration of the articulation disorder detection device 100 shown in FIG. 1. Further, the function of the articulation disorder determination unit 106B is different from that of the articulation disorder determination unit 106.
[0059] The acoustic statistic calculation unit 108 calculates acoustic statistics from the speech data. For example, the acoustic statistics include at least one of pitch variation (intonation), waveform periodicity, and skewness. The articulation disorder determination unit 106B determines the articulation disorder of the speaker based on the similarity and the acoustic statistics.
[0060] FIG. 7 is a flowchart of the articulation disorder detection process by the articulation disorder detection device 100B. Here, the case where one speaker is registered in advance in the articulation disorder detection device 100B will be described.
[0061] First, the articulation disorder detection device 100B instructs the speaker (user) to utter a predetermined phrase (S141). For example, this instruction is given by display or voice.
[0062] Next, the voice acquisition unit 101 acquires the speech data of the above-mentioned phrase uttered by the speaker according to the instruction (S142). Next, the acoustic feature calculation unit 102 calculates the acoustic features from the speech data (S143). Next, the speaker feature calculation unit 103 calculates the first speaker feature from the acoustic features (S144). Next, the similarity calculation unit 105 calculates the similarity between the first speaker feature output from the speaker feature calculation unit 103 and the second speaker feature stored in the storage unit 104 (S145).
[0063] Also, the acoustic statistic calculation unit 108 calculates the acoustic statistics from the speech data (S146). Next, the articulation disorder determination unit 106B determines the articulation disorder of the speaker based on the similarity and the acoustic statistics.
[0064] Specifically, the articulation disorder determination unit 106B determines whether the similarity is less than a predetermined first threshold value (S147). When the similarity is less than the first threshold value (Yes in S147), the articulation disorder determination unit 106B determines whether a first evaluation value based on acoustic statistics is less than a predetermined second threshold value (S148). When the first evaluation value is less than the second threshold value (Yes in S148), the articulation disorder determination unit 106B determines that the speaker has an articulation disorder (S149). Further, when the similarity is greater than or equal to the first threshold value (No in S147), or when the first evaluation value is greater than or equal to the second threshold value (No in S148), the articulation disorder determination unit 106B determines that the speaker does not have an articulation disorder (is healthy) (S150).
[0065] For example, the greater the pitch variation, the higher the first evaluation value, the higher the waveform periodicity, the greater the first evaluation value, and the greater the skewness, the lower the first evaluation value. For example, the articulation disorder determination unit 106B calculates the first evaluation value by weighted addition of pitch variation, waveform periodicity, and the reciprocal of skewness.
[0066] That is, the articulation disorder determination unit 106B determines that the higher the possibility of an articulation disorder, the lower the pitch variation. Further, the articulation disorder determination unit 106B determines that the higher the possibility of an articulation disorder, the lower the waveform periodicity. Note that a low waveform periodicity means a large amount of noise. Further, the articulation disorder determination unit 106B determines that the higher the possibility of an articulation disorder, the greater the skewness.
[0067] Also, the order of steps S147 and S148 may be reversed. That is, the articulation disorder determination unit 106B determines that the speaker has an articulation disorder when both a first condition that the similarity is less than the first threshold value and a second condition that a first evaluation value based on acoustic statistics is less than the second threshold value are satisfied, and determines that the speaker does not have an articulation disorder (is healthy) when at least one of the first condition and the second condition is not satisfied.
[0068] In addition, when at least one of the first condition and the second condition is satisfied, the articulation disorder determination unit 106B may determine that the speaker has an articulation disorder, and when both the first condition and the second condition are not satisfied, the articulation disorder determination unit 106B may determine that the speaker does not have an articulation disorder (is normal). Alternatively, the articulation disorder determination unit 106B may calculate a second evaluation value from the similarity and the first evaluation value, and determine that the speaker has an articulation disorder when the second evaluation value is less than a third threshold value, and determine that the speaker does not have an articulation disorder when the second evaluation value is greater than or equal to the third threshold value. For example, the articulation disorder determination unit 106B calculates the second evaluation value by weighted addition of the similarity and the first evaluation value. Further, the articulation disorder determination unit 106B may determine that the lower the second evaluation value, the higher the possibility of an articulation disorder.
[0069] Note that, instead of calculating the first evaluation value, the articulation disorder determination unit 106B may compare each of pitch variation, waveform periodicity, and skewness with a corresponding threshold value. For example, the articulation disorder determination unit 106B may determine that there is an articulation disorder when at least one of the conditions that the pitch variation is less than a fourth threshold value, the waveform periodicity is less than a fifth threshold value, and the skewness is greater than or equal to a sixth threshold value is satisfied, and determine that there is no articulation disorder otherwise.
[0070] Next, the output unit 107 outputs the determination result obtained by the articulation disorder determination unit 106B (S151). For example, the output unit 107 notifies the speaker of the determination result obtained by the articulation disorder determination unit 106B.
[0071] Here, an example of further using acoustic statistics with respect to the configuration of the first embodiment has been described, but acoustic statistics may be further used with respect to the configuration of the second embodiment.
[0072] As described above, the articulation disorder detection device 100B calculates acoustic statistics from the utterance data, and determines the articulation disorder of the speaker based on the similarity and the acoustic statistics. According to this, the articulation disorder detection device 100B can detect the articulation disorder with high accuracy by making a determination considering the acoustic statistics in addition to the similarity between the first speaker feature amount and the second speaker feature amount in the normal state.
[0073] As described above, the articulation disorder detection device according to the embodiment of the present disclosure has been explained, but the present disclosure is not limited to this embodiment.
[0074] Moreover, each processing unit included in the articulation disorder detection device according to the above embodiment is typically realized as an LSI which is an integrated circuit. These may be individually formed into one chip, or may be formed into one chip so as to include some or all of them.
[0075] Moreover, the integration into an integrated circuit is not limited to an LSI, and it may be realized by a dedicated circuit or a general-purpose processor. An FPGA (Field Programmable Gate Array) which can be programmed after manufacturing the LSI, or a reconfigurable processor which can reconfigure the connection and setting of circuit cells inside the LSI may be used.
[0076] Also, in each of the above embodiments, each component may be configured by dedicated hardware, or may be realized by executing a software program suitable for each component. Each component may be realized by a program execution unit such as a CPU or a processor reading and executing a software program recorded on a recording medium such as a hard disk or a semiconductor memory.
[0077] Further, the present disclosure may be realized as an articulation disorder detection method or the like executed by an articulation disorder detection device or the like.
[0078] Also, the division of functional blocks in the block diagram is an example, and a plurality of functional blocks may be realized as one functional block, one functional block may be divided into a plurality, or some functions may be transferred to other functional blocks. Also, the functions of a plurality of functional blocks having similar functions may be processed by a single piece of hardware or software in parallel or in time division.
[0079] Also, the order in which each step in the flowchart is executed is for illustrative purposes to specifically describe the present disclosure, and it may be in an order other than the above. Also, some of the above steps may be executed simultaneously (in parallel) with other steps.
[0080] As described above, the articulation disorder detection device and the like according to one or more aspects have been described based on the embodiments. However, the present disclosure is not limited to these embodiments. Without departing from the spirit of the present disclosure, various modifications conceived by those skilled in the art applied to these embodiments, or forms constructed by combining components in different embodiments may also be included within the scope of one or more aspects.
Industrial Applicability
[0081] The present disclosure can be applied to an articulation disorder detection device.
Explanation of Signs
[0082] 100, 100A, 100B Articulation disorder detection device 101 Voice acquisition unit 102 Acoustic feature quantity calculation unit 103 Speaker feature quantity calculation unit 104 Storage unit 105 Similarity calculation unit 106, 106A, 106B Articulation disorder determination unit 107 Output unit 108 Acoustic statistic quantity calculation unit 201 Frame connection processing unit 202 DNN 203 Training data
Claims
1. Calculate an acoustic feature quantity and an acoustic statistic quantity from the speaker's utterance data, Using a trained DNN (Deep Neural Network), calculate a first speaker feature quantity representing the speaker identity of the utterance data from the acoustic feature quantity, Calculate the similarity between the second speaker feature quantity, which is the speaker feature quantity of the speaker in a healthy state, and the first speaker feature quantity, Based on the similarity and the acoustic statistic quantity, determine the speaker's speech disorder Speech disorder detection method.
2. Calculate a plurality of acoustic feature quantities from each of a plurality of utterance data of a speaker, Using a trained DNN (Deep Neural Network), calculate a plurality of first speaker feature quantities representing the speaker identity of the plurality of utterance data from the plurality of acoustic feature quantities, Calculate a plurality of similarities between the second speaker feature quantity, which is the speaker feature quantity of the speaker in a healthy state, and the plurality of first speaker feature quantities, Calculate the variance of the plurality of similarities, When the variance is greater than a predetermined second threshold, determine that the speaker has a speech disorder Speech disorder detection method.
3. In the detection of the speech disorder, when the similarity is lower than a predetermined first threshold, determine that the speaker has a speech disorder The speech disorder detection method according to Claim 1 or 2.
4. The acoustic statistic quantity includes pitch variation, In the detection of the speaker's speech disorder, it is determined that the lower the pitch variation, the higher the possibility of speech disorder The speech disorder detection method according to Claim 1.
5. The acoustic statistic quantity includes waveform periodicity, In the detection of the speaker's speech disorder, it is determined that the lower the waveform periodicity, the higher the possibility of speech disorder The speech disorder detection method according to Claim 1.
6. The acoustic statistic quantity includes skewness, In the detection of the speaker's speech disorder, it is determined that the greater the skewness, the higher the possibility of speech disorder The speech disorder detection method according to Claim 1.
7. An acoustic feature quantity calculation unit that calculates an acoustic feature quantity from the speaker's utterance data, An acoustic statistic quantity calculation unit that calculates an acoustic statistic quantity from the utterance data, A speaker feature quantity calculation unit that calculates a first speaker feature quantity representing the speaker identity of the utterance data from the acoustic feature quantity using a trained DNN (Deep Neural Network), A similarity calculation unit that calculates the similarity between the second speaker feature quantity, which is the speaker feature quantity of the speaker in a healthy state, and the first speaker feature quantity, A speech disorder determination unit that determines the speaker's speech disorder based on the similarity and the acoustic statistic quantity, and is provided with Speech disorder detection device. [
8. ] An acoustic feature quantity calculation unit that calculates a plurality of acoustic feature quantities from each of a plurality of utterance data of a speaker, A speaker feature quantity calculation unit that calculates a plurality of first speaker feature quantities representing the speaker characteristics of the plurality of utterance data from the plurality of acoustic feature quantities using a learned DNN (Deep Neural Network); A similarity calculation unit that calculates a plurality of similarities between a second speaker feature quantity that is the speaker feature quantity of the speaker in a healthy state and the plurality of first speaker feature quantities; A dysarthria determination unit that calculates the variance of the plurality of similarities and determines that the speaker has dysarthria when the variance is greater than a predetermined second threshold value, A dysarthria detection device. [
9. ] A program for causing a computer to execute the dysarthria detection method according to claim 1 or 2.
Citation Information
Patent Citations
Training method, speaker identification method, and recording medium
JP2021033260A
Telephone pathology assessment
US20070005357A1