Threshold generation method, threshold generation device, and program
The method generates thresholds for keyword detection by analyzing noise distribution, addressing the inefficiencies of manual adjustment and enhancing performance in noisy environments.
Patent Information
- Application Number
- JP2022118134
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-07-25
- Publication Date
- 2025-11-05
- Estimated Expiration
- 2042-07-25
AI Technical Summary
Conventional detection devices require manual adjustment of thresholds for keyword detection, which is time-consuming, and perform poorly in noisy environments with high false detection or missed detection rates.
A method for generating thresholds that automatically adjusts based on noise distribution analysis, calculating keyword scores and setting thresholds to minimize false positives by considering the distribution of noise scores.
Enables accurate keyword detection without user intervention, reducing false positives and improving performance in noisy conditions.
Smart Images

Figure 0007764329000020 
Figure 0007764329000021 
Figure 0007764329000022
Abstract
Description
[Technical Field]
[0001] FIELD Embodiments of the present invention relate to a threshold generation method, a threshold generation device, and a program. [Background technology]
[0002] There are known detection devices that detect predetermined keywords contained in speech for the purpose of operating devices by speech, etc. Such detection devices calculate a score that represents the similarity between the speech contained in the speech signal and the keyword, and determine that the keyword is contained in the speech signal if the calculated score is greater than a preset threshold.
[0003] Such a detection device needs to adjust the threshold appropriately, for example, by having the user repeatedly utter a keyword and adjusting the threshold so that the keyword can be more easily detected by the detection device.
[0004] However, in conventional detection devices, the threshold value is not adjusted to an appropriate value at the start of use, and users must repeatedly speak keywords until the appropriate value is reached, which is extremely time-consuming.Furthermore, in noisy environments, such detection devices have a high probability of false detection of keywords or a high probability of not detecting keywords even when spoken by the user. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Japanese Patent Application Publication No. 2019-184633 Summary of the Invention [Problem to be solved by the invention]
[0006] The problem to be solved by the present invention is to provide a threshold generation method, a threshold generation device, and a program for generating a threshold that enables appropriate keyword detection without requiring the user to perform adjustment processing. [Means for solving the problem]
[0007] The threshold value generation method according to the embodiment generates a threshold value set for a keyword detection device. , by an information processing device The keyword detection device detects whether the voice signal contains a keyword based on a comparison result between a keyword score representing a degree of similarity between a voice included in the voice signal and a preset keyword and a threshold value. The information processing device, Multiple Reference Voices Multiple noises that are The threshold value generation method calculates the keyword score representing the degree of similarity with the keyword for each of the keywords. The information processing device, The aforementioned Multiple noises The keyword scores are calculated based on noise A parameter representing the distribution of the score set is calculated. noise Based on the parameters that represent the distribution of the score set, a value that makes the keyword score included in the noise score set smaller with a predetermined probability, The threshold as Generate. [Brief explanation of the drawings]
[0008] [Figure 1] 1 is a configuration diagram of a voice operation system according to a first embodiment. [Figure 2] 1 is an external view of a keyword detection device according to a first embodiment. [Figure 3] FIG. 10 is a diagram showing an example of the operation of the operation target device. [Figure 4] FIG. 2 is a configuration diagram of a keyword detection unit according to the first embodiment. [Figure 5] FIG. 4 is a diagram showing thresholds of a keyword detection unit according to the first embodiment. [Figure 6] FIG. 10 is a diagram showing keyword scores. [Figure 7] FIG. 7 is a diagram showing the detection results when the keyword scores of FIG. 6 are calculated. [Figure 8] FIG. 3 is a configuration diagram of a keyword score calculation unit. [Figure 9] FIG. 1 is a configuration diagram of a threshold generating device according to a first embodiment. [Figure 10] 3 is a flowchart showing the flow of processing in the first embodiment. [Figure 11] FIG. 11 is a diagram showing an example of threshold values generated in the process shown in FIG. 10. [Figure 12] FIG. 10 is a diagram showing keyword scores when uttered. [Figure 13] FIG. 13 is a diagram showing the detection results when the keyword scores of FIG. 12 are calculated. [Figure 14] FIG. 10 is a configuration diagram of a keyword detection unit according to a modified example of the first embodiment. [Figure 15] 10 is a flowchart showing the flow of processing in the second embodiment. [Figure 16] FIG. 16 is a diagram showing an example of threshold values generated in the process shown in FIG. 15. [Figure 17] 10 is a flowchart showing the flow of processing in a third embodiment. [Figure 18] FIG. 18 is a diagram showing an example of threshold values generated in the process shown in FIG. 17. [Figure 19] 10 is a flowchart showing the flow of processing in a fourth embodiment. [Figure 20] FIG. 13 is a configuration diagram of a keyword detection unit according to the fifth embodiment. [Figure 21] FIG. 13 is a configuration diagram of a keyword detection unit according to the sixth embodiment. [Figure 22] FIG. 2 is a diagram illustrating an example of a hardware configuration of a threshold generating device. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.
[0010] (First embodiment) Fig. 1 is a diagram showing the configuration of a voice operation system 10 according to the first embodiment. Fig. 2 is a diagram showing an example of the appearance of a keyword detection device 22 according to the first embodiment.
[0011] The voice operation system 10 includes an operation target device 20, a keyword detection device 22, and a threshold generation device 24.
[0012] The operation target device 20 is, for example, a household electrical appliance or an electronic device that operates in response to a user's operation. In the first embodiment, the operation target device 20 is an air conditioner. The operation target device 20 receives an operation signal from the keyword detection device 22 and operates in response to the received operation signal.
[0013] The keyword detection device 22 collects voice spoken by a user. The keyword detection device 22 determines whether the collected voice contains a preset keyword. If the collected voice contains a preset keyword, the keyword detection device 22 transmits an operation signal to the operation target device 20, causing the operation target device 20 to perform an operation corresponding to the keyword. For example, the keyword detection device 22 transmits the operation signal to the operation target device 20 by infrared rays, radio waves, or the like. The keyword detection device 22 may be incorporated in the operation target device 20 and transmit the operation signal to the operation target device 20 via a wired line.
[0014] As an example, the keyword detection device 22 includes a microphone 32, a keyword detection unit 34, and a communication unit 36, as shown in FIGS.
[0015] The microphone 32 picks up surrounding sounds and converts them into analog audio signals.
[0016] The keyword detection unit 34 receives the audio signal from the microphone 32. A plurality of keywords are set in advance in the keyword detection unit 34. The keyword detection unit 34 calculates a keyword score for each of the plurality of keywords for each frame, which is a predetermined time interval. The keyword score represents the degree of similarity between the audio included in the audio signal and the preset keyword.
[0017] The keyword detection unit 34 has a preset threshold value for each of the multiple keywords. For each of the multiple keywords, the keyword detection unit 34 detects whether the audio signal contains a corresponding keyword based on the result of comparing the calculated keyword score with the threshold value for each frame. For example, if the keyword score is greater than the threshold value, the keyword detection unit 34 detects that the audio signal contains the corresponding keyword. If the keyword detection unit 34 detects that the audio signal contains any of the multiple keywords, it outputs an operation signal instructing an operation corresponding to the included keyword. The keyword detection unit 34 is realized, for example, by an information processing circuit including a processing circuit, a memory, etc.
[0018] When the keyword detection unit 34 detects that the voice signal contains a keyword, the communication unit 36 transmits to the operation target device 20 an operation signal corresponding to the detected keyword.
[0019] The threshold generation device 24 generates thresholds corresponding to each of a plurality of keywords prior to the keyword detection operation by the keyword detection device 22. The threshold generation device 24 sets the thresholds for each of the generated keywords in the keyword detection device 22. For example, the threshold generation device 24 stores the generated thresholds in a non-volatile memory inside the keyword detection device 22.
[0020] The threshold generation device 24 is realized by, for example, an information processing device including a processing circuit, a memory, etc. executing a program. The threshold generation device 24 may be provided integrally with the keyword detection device 22. Furthermore, the threshold generation device 24 may be realized by a processing circuit, a memory, etc. shared with the keyword detection unit 34.
[0021] FIG. 3 is a diagram showing an example of the operation of the operation target device 20 when a keyword is uttered by the user.
[0022] The keyword detection device 22 assigns a keyword ID, which is identification information, to each of a plurality of preset keywords. When the keyword detection device 22 detects that the audio signal contains one of the plurality of keywords, it transmits an operation signal including the keyword ID assigned to the detected keyword to the operation target device 20. The operation target device 20 stores a table or the like that associates the keyword ID with the operation content. When the operation target device 20 receives the operation signal, it executes the operation corresponding to the keyword ID.
[0023] The keyword detection device 22 has "Danbo" set as a keyword with a keyword ID of "1." When the keyword voice "Danbo" is uttered by the user, the keyword detection device 22 causes the operation target device 20 to start heating operation.
[0024] Furthermore, the keyword detection device 22 has set "RE-BO" as a keyword having a keyword ID of "2." When the keyword voice "RE-BO" is uttered by the user, the keyword detection device 22 causes the operation target device 20 to start cooling operation.
[0025] Furthermore, the keyword detection device 22 has "power off" set as the keyword with the keyword ID of "3." When the user utters the keyword voice "power off," the keyword detection device 22 causes the operation target device 20 to stop operation.
[0026] Furthermore, the keyword detection device 22 has set "hot" as the keyword with a keyword ID of "4." When the user utters the keyword voice "hot," the keyword detection device 22 causes the operation target device 20 to lower the set temperature by 1 degree.
[0027] Furthermore, the keyword detection device 22 has set "cold" as the keyword with a keyword ID of "5." When the user utters the keyword voice "cold," the keyword detection device 22 causes the operation target device 20 to increase the set temperature by one degree.
[0028] 4 is a diagram showing the configuration of the keyword detection unit 34 according to the first embodiment. The keyword detection unit 34 includes an AD conversion unit 40, a feature generation unit 42, a keyword model storage unit 44, a keyword score calculation unit 46, a threshold storage unit 48, and a determination unit 50.
[0029] The AD conversion unit 40 samples the audio signal output from the microphone 32 and converts it into a digital audio signal. For example, the AD conversion unit 40 converts it into a 16-bit PCM digital audio signal with a sampling frequency of 16 kHz.
[0030] The feature generation unit 42 receives a digital audio signal and generates, for each frame, a feature vector that represents the features of the audio contained in the audio signal. For example, the feature generation unit 42 performs a short-time Fourier transform on the time-domain digital audio signal with a frame length of 160 samples and a window length of 512 samples. This allows the feature generation unit 42 to convert the time-domain digital audio signal into a frequency-domain audio signal. The feature generation unit 42 then generates a feature vector for each frame based on the frequency-domain audio signal. For example, the feature generation unit 42 generates a 40-dimensional Mel filter bank feature vector.
[0031] The keyword model storage unit 44 stores a score calculation model for calculating a keyword score from a feature vector for each of a plurality of keywords. In the first embodiment, the score calculation model is realized by a neural network and a directed graph search algorithm using a Viterbi algorithm or the like. The keyword model storage unit 44 stores neural network parameters, directed graphs, and the like as the score calculation model for each of a plurality of keywords.
[0032] The keyword score calculation unit 46 calculates a keyword score for each of a plurality of keywords for each frame using the corresponding score calculation model stored in the keyword model storage unit 44. In the first embodiment, the more similar the voice and the keyword, the larger the keyword score value.
[0033] The threshold storage unit 48 stores a threshold for each of the plurality of keywords. Prior to a keyword detection operation, the threshold storage unit 48 receives and stores the threshold for each of the plurality of keywords from the threshold generation device 24.
[0034] The determination unit 50 receives, for each frame, the keyword score for each of the multiple keywords from the keyword score calculation unit 46. For each frame, the determination unit 50 compares the received keyword score with the corresponding threshold stored in the threshold storage unit 48 to determine whether the audio signal contains the corresponding keyword. For example, if the received keyword score is greater than the corresponding threshold, the determination unit 50 determines that the audio signal contains the corresponding keyword. Then, the determination unit 50 provides the determination result to the communication unit 36.
[0035] Fig. 5 is a diagram showing an example of thresholds set in the keyword detection unit 34 according to the first embodiment. Fig. 6 is a diagram showing an example of keyword scores detected by the keyword detection unit 34. Fig. 7 is a diagram showing an example of detection results by the keyword detection unit 34 when the keyword scores shown in Fig. 6 are calculated.
[0036] The keyword detection unit 34 sets a threshold value for each of a plurality of keywords. In the first embodiment, the keyword detection unit 34 sets a threshold value as shown in Fig. 5 for each of the keywords having keyword IDs "1" to "5" shown in Fig. 3.
[0037] t is an integer representing the frame, and increases by 1 from a predetermined value for each frame. i (t) represents the keyword score for the keyword with keyword ID i at frame t.
[0038] The keyword detection unit 34 calculates a keyword score for each of a plurality of keywords for each frame. In the first embodiment, the keyword detection unit 34 calculates a keyword score for each keyword with a keyword ID of "1" to "5" for each frame. Then, in a frame where the calculated keyword score is greater than a set threshold, the keyword detection unit 34 outputs, as a detection result, a keyword ID that identifies the keyword whose keyword score is greater than the threshold.
[0039] In the examples of FIGS. 5 to 7, the keyword detection unit 34 calculates the keyword score for each of the frames from t=130 to t=140. For the keyword "power off" with a keyword ID of "3", the keyword detection unit 34 detects a maximum keyword score of 451 in the frame at t=136. Since the threshold for the keyword with a keyword ID of "3" is 339, the keyword detection unit 34 determines that the keyword "power off" is included in the audio signal in the frame at t=136. Then, as shown in FIG. 7, the keyword detection unit 34 outputs 3, which is the keyword ID of the keyword "power off", as the detection result in the frame at t=136. Note that in the first embodiment, the keyword detection unit 34 outputs 0 as the detection result if the keyword score of any keyword is not greater than the threshold.
[0040] 8 is a diagram showing the configuration of keyword score calculation unit 46. Keyword score calculation unit 46 includes neural network unit 52 and search unit 54. For each of a plurality of keywords, keyword score calculation unit 46 executes score calculation processing in accordance with a score calculation model using neural network unit 52 and search unit 54.
[0041] A keyword is represented by a directed graph that represents the time transition of minute elements of speech. In the first embodiment, the directed graph represents a syllable string. Each syllable included in the syllable string represented by the directed graph is modeled by a left-to-right hidden Markov model that represents three states. If the number of syllables in a keyword is n (an integer equal to or greater than 1), the directed graph that represents the keyword is made up of N states {y1, y2, ..., y N}, and each of the N states includes a self-transition and an inter-state transition from the previous state to the next state. N is 3 × n. For example, the three-syllable keyword "hot" is represented by a directed graph containing nine states.
[0042] The neural network unit 52 acquires a feature vector for each frame from the feature generator 42. For each frame, the neural network unit 52 calculates, based on the feature vector, a likelihood score indicating the likelihood that a voice will be in the corresponding state for each of a plurality of states included in the directed graph representing the keyword.
[0043] Here, in the t-th frame, the feature vector (x t ) is obtained, the qth state (y q ) is the likelihood score of score(x t ,y q ) for each of the plurality of keywords. The neural network unit 52 generates a set of N states {y1, y2, ..., y N} and calculate the likelihood score for each of them.
[0044] The neural network unit 52 executes a calculation according to the neural network for each frame. The neural network is, for example, a fully connected network. The neural network includes four hidden layers. Each layer includes 256 nodes. The neural network applies, for example, a Sigmoid function as an activation function. The output layer of the neural network includes, for example, nodes corresponding to all syllables and a node corresponding to silence. The output layer of the neural network applies, for example, a Softmax function as an activation function. Each parameter of the neural network is preset in the keyword model storage unit 44.
[0045] Then, the neural network unit 52 outputs a likelihood score obtained from the output layer of the neural network for each of the multiple keywords. In this case, the neural network unit 52 outputs a likelihood score obtained from the output layer of the neural network for N states {y1, y2, ..., y N}, and outputs likelihood scores from multiple nodes corresponding to the
[0046] The search unit 54 searches the directed graph for the best sequence for each of the multiple keywords for each frame, with the best sequence having the maximum total likelihood score. Then, the search unit 54 calculates the total likelihood score in the best sequence for each frame as the keyword score.
[0047] Specifically, the search unit 54 performs a search process for calculating the formula (1) for each frame to obtain the keyword score (S i (t)) is calculated.
number
[0048] In equation (1), S i(t) represents the keyword score of the i-th keyword in the frame to be processed. t is an integer representing the frame to be processed, and increases by 1 for each frame. b represents the initial frame corresponding to the first state among the multiple states included in the directed graph when the frame to be processed is t.
[0049] Q represents the sequence of state numbers in each of the multiple paths from the first state to the tth state in the directed graph. x τ represents the feature vector at frame τ. qτ represents the q-th state among the states included in the directed graph at frame τ. score(x τ ,y qτ ) represents the likelihood score of the qth state at frame τ.
[0050] The search unit 54 performs the following process as a search process corresponding to the calculation shown in equation (1). That is, the search unit 54 selects one best route from among a plurality of routes from the first state to the t-th state included in the directed graph, which has the largest total value of likelihood scores. The search unit 54 also changes the initial frame (b) under the condition that the initial frame (b) is smaller than t, and selects such a best route for each initial frame (b). Furthermore, the search unit 54 calculates a normalized total value by multiplying the total value of the likelihood scores of each selected best route by 1 / (t-b+1). Then, the search unit 54 calculates the largest value of the normalized total values of the selected plurality of best routes as the keyword score (S i (t)).
[0051] By performing this processing, the search unit 54 can search for the best sequence that maximizes the total likelihood score from the directed graph for each frame, and calculate the total likelihood score for the best sequence as the keyword score. The search unit 54 can solve the problem of searching for the best sequence that maximizes the total likelihood score from the directed graph, for example, by using the Viterbi algorithm.
[0052] 9 is a diagram showing the configuration of the threshold generation device 24 according to the first embodiment. Prior to the detection operation by the keyword detection device 22, the threshold generation device 24 generates a threshold for each of a plurality of keywords and sets the threshold in the keyword detection device 22.
[0053] The threshold generation device 24 includes an acquisition unit 60 , a score calculation unit 62 , a distribution calculation unit 64 , a threshold generation unit 66 , and a setting unit 68 .
[0054] The acquiring unit 60 acquires an input signal including a plurality of reference sounds collected in advance. In the first embodiment, the acquiring unit 60 acquires an input signal including a plurality of noises as a plurality of reference sounds.
[0055] The score calculation unit 62 calculates a keyword score representing the degree of similarity between each of the plurality of reference voices and the keyword. In the first embodiment, a keyword score representing the degree of similarity between each of the plurality of noises and the keyword is calculated.
[0056] The score calculation unit 62 calculates a keyword score (S i 4 without the threshold storage unit 48 and the determination unit 50. In addition, when acquiring a digitally converted input signal, the score calculation unit 62 has the same configuration as the keyword detection unit 34 without the AD conversion unit 40.
[0057] Then, the score calculation unit 62 generates a score set including a plurality of keyword scores calculated based on a plurality of reference voices for each of the plurality of keywords. In the first embodiment, the score calculation unit 62 generates a noise score set including a plurality of keyword scores calculated based on a plurality of noises as the score set for each of the plurality of keywords.
[0058] The distribution calculation unit 64 calculates parameters representing the distribution of the score set for each of the multiple keywords. In the first embodiment, the distribution calculation unit 64 calculates parameters representing the distribution of the noise score set for each of the multiple keywords. For example, the distribution calculation unit 64 assumes that the noise score set approximates a normal distribution and calculates the mean value and standard deviation as parameters representing the distribution of the noise score set.
[0059] The threshold generation unit 66 generates a threshold for each of the multiple keywords based on a parameter representing the distribution of the score set. For example, the threshold generation unit 66 generates a threshold such that the keyword score included in the score set is larger with a predetermined probability, or such that the keyword score included in the score set is larger with a predetermined probability, based on the parameter representing the distribution of the score set. In the first embodiment, the threshold generation unit 66 generates, as a threshold for each of the multiple keywords, a value such that the keyword score calculated based on noise is smaller with a predetermined probability, based on the parameter representing the distribution of the noise score set. For example, the threshold generation unit 66 generates, as a threshold for each of the multiple keywords, a value such that the majority of the keyword scores included in the noise score set are smaller, based on the mean value and standard deviation representing the distribution of the noise score set.
[0060] The setting unit 68 sets the generated threshold value in the keyword detection device 22 for each of the plurality of keywords.
[0061] 10 is a flowchart showing the flow of processing by the threshold generating device 24 according to the first embodiment. The threshold generating device 24 according to the first embodiment generates thresholds according to the flow shown in FIG.
[0062] First, in S101, the acquiring unit 60 acquires an input signal including a plurality of noises as a plurality of reference sounds.
[0063] In the first embodiment, the input signal is, for example, an audio signal collected in an environment in which the keyword detection device 22 is used or in an acoustic environment similar to that in which the keyword detection device 22 is used. In the first embodiment, the input signal is, for example, an audio signal collected inside a car when the keyword detection device 22 is used inside a car. In the first embodiment, the input signal is, for example, an audio signal collected in a living room when the keyword detection device 22 is used in a living room. In addition, the input signal may be an audio signal of a long duration, such as several hours or tens of hours. This allows the input signal to include a greater number of different types of noise.
[0064] Next, the threshold generation device 24 executes the processes from S103 to S106 for each of the multiple keywords (loop processing between S102 and S107). The threshold generation device 24 may execute the processes from S103 to S106 sequentially for each of the multiple keywords, or may execute the processes for multiple keywords in parallel.
[0065] In S103 in the loop, the score calculation unit 62 calculates a keyword score (S i Then, the score calculation unit 62 calculates the plurality of keyword scores (S i (t)) is stored as a noise score set, which is a score set for the keyword to be processed.
[0066] For example, the score calculation unit 62 may add T n If the frame contains noise, T n Let t={1,2,…,T n}. Then, the score calculation unit 62 assigns T n Keyword score (S i (t)) and calculate the calculated T n Keyword score (S iThe score set containing (t) is stored as the noise score set for the i-th keyword.
[0067] Next, in S104, the distribution calculation unit 64 calculates parameters representing the distribution of the noise score set for the keyword to be processed. For example, the distribution calculation unit 64 assumes that the noise score set approximates a normal distribution and calculates the mean value and standard deviation of the distribution of the noise score set as parameters representing the distribution of the noise score set.
[0068] For example, the distribution calculation unit 64 performs the calculation shown in equation (2) to obtain the average value (m ni ) is calculated.
number
[0069] Furthermore, for example, the distribution calculation unit 64 performs the calculation shown in equation (3) to calculate the standard deviation (σ ni ) is calculated.
number
[0070] Next, in S105, the threshold generation unit 66 generates a threshold for the keyword to be processed based on parameters representing the distribution of the noise score set. For example, the threshold generation unit 66 regards the distribution of the noise score set as a normal distribution and generates, as the threshold, a value at which the keyword scores included in the noise score set are smaller with a predetermined probability, based on the mean value and standard deviation. For example, the threshold generation unit 66 generates, for the keyword to be processed, a value at which the majority of keyword scores included in the noise score set are smaller, based on parameters representing the distribution of the noise score set.
[0071] For example, the threshold value generating unit 66 performs the calculation shown in equation (4) to obtain the threshold value (θ ni ) is calculated.
number
[0072] The threshold value generating unit 66 sets a value equal to or greater than the value of the formula (4) as the threshold value (θ ni ) may be generated as the average value (m ni ) and the standard deviation of the noise score set (σ ni ) multiplied by a predetermined first magnification (A) and added to the result (m ni +Aσ ni ) is set to a value equal to or greater than the threshold (θ ni )
[0073] The threshold value shown in equation (4) is determined by the normal distribution table, which shows that the frequency at which the calculated keyword score is larger when noise is input is 2.87 × 10 -7 In other words, the threshold value shown in formula (4) is a value that, when noise is continuously input for 24 hours, the frequency of falsely detecting noise as a keyword due to the keyword score being larger than the threshold value is about 2.5 times. As a result, the threshold value generation unit 66 can generate, for the i-th keyword, a value at which the majority of keyword scores included in the noise score set are smaller, that is, a value at which the majority of keyword scores included in the noise score set are not detected.
[0074] Furthermore, the threshold generation unit 66 generates a threshold for each of a plurality of keywords by the same calculation, thereby making it possible for the threshold generation unit 66 to maintain a constant probability of false positive for each of a plurality of keywords.
[0075] Next, in S106, the setting unit 68 sets the generated threshold value in the keyword detection device 22.
[0076] When the threshold generation device 24 has completed the processes from S103 to S106 for each of the multiple keywords, it exits the loop process between S101 and S107 and ends this flow.
[0077] FIG. 11 is a diagram showing an example of the average value, standard deviation, and threshold value generated in the process shown in FIG.
[0078] The threshold generation device 24 generates a threshold for each of the plurality of keywords by executing the process shown in Fig. 10. Each of the plurality of thresholds is set based on the keyword score (S i (t)) is a value that will be smaller with a predetermined probability. Therefore, by generating such a threshold value for each of a plurality of keywords, the threshold generation device 24 can make the probability of false positives for each keyword constant.
[0079] Fig. 12 is a diagram showing an example of keyword scores when a user utters "hot," a keyword with a keyword ID of "4," in a noisy environment. Fig. 13 is a diagram showing an example of detection results by the keyword detection unit 34 when the keyword scores shown in Fig. 12 are calculated.
[0080] The examples shown in FIGS. 12 and 13 are based on the assumption that speech is made in an environment where noise is generated by the airflow of an air conditioner or the sound of a television set.
[0081] In the frame t=38, the keyword score for keyword ID 4 is S4(38)=458, and the threshold θ n4 On the other hand, in the frame at t=37, the keyword score for keyword ID 5 is S5(37)=471, which is larger than the threshold value for keyword ID 4, S4(38)=458, but is smaller than the threshold value for keyword ID 5, θ n5= 512. If the thresholds for "hot" with keyword ID "4" and "cold" with keyword ID "5" are the same, "cold" will be mistakenly detected, and the correct answer, "hot," will not be detected.
[0082] In contrast, the keyword detection device 22 according to the first embodiment sets a threshold value for each keyword based on a noise score distribution, which is a distribution of keyword scores relative to noise, to suppress erroneous detection. Therefore, the keyword detection device 22 according to the first embodiment can accurately detect the correct answer while suppressing erroneous detection.
[0083] As described above, the threshold generation device 24 according to the first embodiment can generate a threshold that allows the keyword detection device 22 to appropriately detect keywords, without requiring the user to perform adjustment processing.
[0084] (Variation) FIG. 14 is a diagram showing the configuration of a keyword detection unit 34 according to a modified example of the first embodiment.
[0085] The keyword detection unit 34 of the keyword detection device 22 may have the configuration shown in Fig. 14 instead of the configuration shown in Fig. 4. In the keyword detection unit 34 according to this modification, the threshold value stored in the threshold storage unit 48 is provided to the keyword score calculation unit 46 instead of the determination unit 50. In the following, the same reference numerals are used to denote components having substantially the same functions and configurations as the components included in the first embodiment described with reference to Figs. 1 to 13, and differences between the modified example and the first embodiment will be described.
[0086] In this modification, the keyword detection unit 34 calculates a keyword score from which a threshold value has been subtracted in advance. Then, in this modification, the determination unit 50 compares the received keyword score with 0 for each of a plurality of keywords to detect whether the corresponding keyword is included in the audio signal. Thus, even in this modification, the determination unit 50 can detect whether the corresponding keyword is included in the audio signal based on the comparison result between the keyword score and the corresponding threshold value.
[0087] More specifically, the search unit 54 of the keyword detection unit 34 performs a search process for calculating the operation of equation (5) for each frame, thereby obtaining a keyword score (S i (t)) is calculated.
number
[0088] The search unit 54 according to the modified example performs the following process as a search process corresponding to the calculation shown in equation (5). That is, the search unit 54 selects one best route from among a plurality of routes from the first state to the Nth state included in the directed graph, which route has the largest total value of subtracted likelihood scores obtained by subtracting a threshold value from the likelihood score. Furthermore, the search unit 54 varies the initial frame (b) under the condition that the initial frame (b) is smaller than t, and selects such a best route for each initial frame (b). Then, the search unit 54 calculates the largest value of the total value of the subtracted likelihood scores of the selected plurality of best routes as the keyword score (S i (t)).
[0089] Equation (5) does not include the operation of multiplying the total likelihood score by 1 / (t-b+1). Therefore, the search unit 54 can independently and sequentially search for the best sequence regardless of the position of the initial frame (b). This allows the search unit 54 to perform the search process equivalent to the operation of equation (5) with a smaller amount of calculation than when performing the search process for the operation of equation (1).
[0090] In addition, in the process of S103, the threshold generation device 24 performs a search process corresponding to the calculation of equation (5) to obtain a keyword score (S i (t)). In this case, the threshold generation device 24 sets an initial threshold value for each of the multiple keywords at the start of the search process. The initial threshold value for each of the multiple keywords may be the same. Then, in the process of S105, the threshold generation device 24 generates a final threshold value by adding the initial value to the threshold calculated based on the distribution. This allows the threshold generation device 24 to generate a threshold value with a small amount of calculation.
[0091] Furthermore, the threshold generation device 24 according to the first embodiment calculates a keyword score (S i (t)) and generate a keyword score distribution for each of the multiple keywords. Alternatively, the threshold generation device 24 may generate a likelihood score distribution for each of the multiple states included in the directed graph representing the keyword. The threshold generation device 24 may then generate a keyword score distribution based on the likelihood score distribution for each of the multiple states. In this case, the threshold generation device 24 may generate a likelihood score distribution for each of all states obtained from the neural network and select from these distributions the likelihood score distribution for the multiple states included in the keyword. This allows the threshold generation device 24 to easily generate a threshold for a new keyword when the keyword is changed without performing a search process again.
[0092] In the first embodiment, five keywords are set in the keyword detection device 22. However, any number of keywords may be set in the keyword detection device 22 as long as the number is one or more. In the first embodiment, the keyword detection device 22 generates a Mel filter bank feature vector as the feature vector. However, the keyword detection device 22 may generate a feature vector other than the Mel filter bank feature vector.
[0093] In the first embodiment, a keyword is a directed graph representing a string of multiple syllables. A keyword may be represented by a graph representing transitions of various minute elements such as phonemes, two-phoneme sequences, three-phoneme sequences, subwords, or words. A keyword may also be represented by a unit obtained by clustering a predetermined number of these minute elements.
[0094] In the first embodiment, the keyword detection device 22 calculates the likelihood score of each state using a neural network. However, the keyword detection device 22 may calculate the likelihood score of each state using other models, such as a Gaussian mixture distribution model. In the first embodiment, the keyword detection device 22 uses a fully connected network that uses a Sigmoid function as the activation function as the neural network. However, the keyword detection device 22 may also use a convolutional neural network or a recurrent neural network. In addition, the keyword detection device 22 may use other functions, such as Tanh or ReLU, as the activation function.
[0095] The threshold generation device 24 calculates the threshold by adding five times the standard deviation to the average value in equation (4). However, the threshold generation device 24 may calculate the threshold by adding a multiple of the standard deviation other than five to the average value. The designer of the threshold generation device 24 may set an appropriate multiple in equation (4) based on constraints on keyword false detection, etc. Furthermore, the threshold generation device 24 sets the threshold by regarding the keyword score distribution as a normal distribution. However, the threshold generation device 24 may calculate the distribution parameters by regarding the keyword score distribution as a distribution other than a normal distribution. Furthermore, the threshold generation device 24 may generate a threshold by using, as the parameter of the keyword score distribution, the maximum value of the keyword scores included in the distribution or a predetermined cumulative frequency, etc.
[0096] (Second embodiment) Next, a voice-operated system 10 according to a second embodiment will be described. The voice-operated system 10 according to the second embodiment has substantially the same functions and configuration as the voice-operated system 10 according to the first embodiment, and therefore, hereinafter, substantially the same components are denoted by the same reference numerals, and detailed description thereof will be omitted except for the differences.
[0097] 15 is a flowchart showing the flow of processing by the threshold generating device 24 according to the second embodiment. The threshold generating device 24 according to the second embodiment generates thresholds according to the flow shown in FIG.
[0098] The threshold generation device 24 executes the processes from S202 to S206 for each of the multiple keywords (loop process between S201 and S207).
[0099] In S202 within the loop, the acquisition unit 60 acquires an input signal including a plurality of keyword voices, each of which is a keyword spoken by one or a plurality of speakers, as a plurality of reference voices. It is preferable that the plurality of keyword voices include a large number of speakers who have spoken the keyword. It is also preferable that the plurality of keyword voices include a large number of times each speaker has spoken the keyword. It is also preferable that the input signal is a voice signal collected by a speaker speaking a keyword in, for example, an environment in which the keyword detection device 22 is used or in an acoustic environment similar to the environment in which the keyword detection device 22 is used.
[0100] Next, in S203, the score calculation unit 62 calculates a keyword score (S i When a speaker utters a keyword voice once, the score calculation unit 62 calculates a keyword score (S i When one keyword voice is uttered, the score calculation unit 62 calculates the keyword score for each of a plurality of frames from the start to the end of the utterance. Therefore, for each utterance of one keyword voice, the score calculation unit 62 calculates the calculated plurality of keyword scores (S i(k)) i (k)).
[0101] The score calculation unit 62 calculates a plurality of keyword scores (S i (k)) as an utterance score set, which is a score set for the keyword to be processed. For example, when the input signal contains K keyword sounds, the score calculation unit 62 assigns frame numbers of the K keyword sounds as k={1, 2, ..., K}. Then, the score calculation unit 62 calculates K keyword scores (S i (k)), and a score set including the calculated K keyword scores (S(k)) is stored as a speech score set for the i-th keyword.
[0102] Next, in S204, the distribution calculation unit 64 calculates parameters representing the distribution of the speech score set for the keyword to be processed. For example, the distribution calculation unit 64 assumes that the speech score set approximates a normal distribution and calculates the mean value and standard deviation of the distribution of the speech score set as parameters representing the distribution of the speech score set.
[0103] For example, the distribution calculation unit 64 performs the calculation shown in Equation (6) to obtain the average value (m ui ) is calculated.
number
[0104] Furthermore, for example, the distribution calculation unit 64 performs the calculation shown in equation (7) to calculate the standard deviation (σ ui ) is calculated.
number
[0105] Next, in S205, the threshold generation unit 66 generates a threshold for the keyword to be processed based on parameters representing the distribution of the speech score set. For example, the threshold generation unit 66 regards the distribution of the speech score set as a normal distribution and generates, as the threshold, a value at which the keyword scores included in the speech score set will be larger with a predetermined probability, based on the mean value and standard deviation. For example, the threshold generation unit 66 generates, for the i-th keyword, a value at which the majority of keyword scores included in the speech score set will be larger.
[0106] For example, the threshold value generating unit 66 performs the calculation shown in equation (8) to obtain the threshold value (θ ui ) is calculated.
number
[0107] The threshold value generating unit 66 sets a value equal to or less than the value of the formula (8) as the threshold value (θ ui ) may be generated as the average value (m ui ) to the standard deviation of the speech score set (σ ui ) multiplied by a predetermined second magnification (B) minus the value (m ui -Bσ ui ) below the threshold (σ ui )
[0108] The threshold value shown in equation (8) is a value that, based on the normal distribution table, when a keyword speech is input, causes the frequency at which the calculated keyword score is smaller to be approximately 0.00135. In other words, the threshold value shown in equation (8) is a value that causes the frequency at which the keyword speech is not detected due to the keyword score being smaller than the threshold value to be approximately 1.4 times when the keyword is spoken 1,000 times. This allows the threshold value generation unit 66 to generate, for the i-th keyword, a value at which the majority of keyword scores included in the speech score set are larger, i.e., a value at which the majority of keyword scores included in the speech score set are detected, as the threshold value.
[0109] Furthermore, the threshold generation unit 66 generates a threshold for each of a plurality of keywords by the same calculation, thereby making it possible for the threshold generation unit 66 to make the non-detection probability for each of a plurality of keywords constant.
[0110] Subsequently, in S206, the setting unit 68 sets the generated threshold value in the keyword detection device 22.
[0111] When the threshold generation device 24 has completed the processes from S202 to S206 for each of the multiple keywords, it exits the loop process between S201 and S207 and ends this flow.
[0112] FIG. 16 is a diagram showing an example of the average value, standard deviation, and threshold value generated in the process shown in FIG.
[0113] The threshold value generating device 24 generates a threshold value for each of the plurality of keywords by executing the process shown in Fig. 15. Each of the plurality of threshold values is determined based on the keyword score (S i (k)) is a value that is larger with a predetermined probability. Therefore, the threshold generation device 24 according to the second embodiment can make the non-detection probability for each keyword constant by generating such a threshold for each of a plurality of keywords.
[0114] As described above, the threshold generation device 24 according to the second embodiment can generate a threshold that allows the keyword detection device 22 to appropriately detect keywords, without requiring the user to perform adjustment processing.
[0115] The threshold value generating device 24 calculates the threshold value (θ ui ) is calculated by subtracting three times the standard deviation from the average value to obtain the threshold value. However, the threshold generation device 24 may calculate the threshold value by subtracting a multiple of the standard deviation other than three times from the average value. The designer of the threshold generation device 24 can set an appropriate multiple in equation (8) based on constraints such as whether keywords are not detected.
[0116] Furthermore, the threshold generation device 24 according to the second embodiment collects keyword speech uttered by a user to prepare an input signal. However, the threshold generation device 24 may also prepare a large amount of speech data of arbitrary content to which syllable labels are attached, generate scores for each state constituting the keyword, calculate a distribution of the scores for each state, and generate a keyword score distribution from the score distribution for each state. Because such a threshold generation device 24 does not require collection of keyword speech, it reduces the cost of collecting keyword speech and can generate thresholds in a short time even when the keyword is changed.
[0117] (Third embodiment) Next, a voice operation system 10 according to a third embodiment will be described. The voice operation system 10 according to the third embodiment has substantially the same functions and configurations as the voice operation systems 10 according to the first and second embodiments, and therefore, hereinafter, substantially the same components are denoted by the same reference numerals, and detailed description thereof will be omitted except for the differences.
[0118] 17 is a flowchart showing the flow of processing by the threshold generating device 24 according to the third embodiment. The threshold generating device 24 according to the third embodiment generates thresholds according to the flow shown in FIG.
[0119] First, the threshold generation device 24 executes the processes of S101, S102, S103, S104, S105, and S107. The processes of S101, S102, S103, S104, S105, and S107 are the same as the processes of the first embodiment shown in Fig. 10. However, in the third embodiment, the threshold generated in S105 is called a noise threshold.
[0120] Next, the threshold generation device 24 executes the processes of S201, S202, S203, S204, S205, and S207. The processes of S201, S202, S203, S204, S205, and S207 are the same as the processes of the second embodiment shown in Fig. 15. However, in the third embodiment, the threshold generated in S205 is called the speech threshold.
[0121] Next, the threshold generation device 24 executes the processes from S302 to S304 for each of the multiple keywords (loop process between S301 and S305).
[0122] In S302 in the loop, the threshold generation unit 66 calculates the noise threshold (θ ni ) and the speech threshold (θ ui For example, the threshold generation unit 66 performs the calculation of equation (9) to generate a threshold (θ nui )
number
[0123] By performing such processing, the threshold generation unit 66 can generate thresholds that balance the frequency of false detections and the frequency of non-detections by using a noise threshold generated based on the noise score distribution and a speech threshold generated based on the speech score distribution.
[0124] Next, in S303, the threshold generation device 24 calculates the false detection probability or the false detection frequency as an evaluation value based on the threshold generated in S302 and the noise score set generated in S103. Alternatively, the threshold generation device 24 calculates the non-detection probability or the false detection frequency as an evaluation value based on the threshold generated in S302 and the speech score set generated in S203. For example, the threshold generation device 24 calculates the non-detection probability or the false detection frequency as an evaluation value based on the threshold generated in S302 and the speech score set generated in S203. nui -m ni ) / σ ni The threshold value generating device 24 may calculate the false detection probability when noise is input based on a normal distribution table from the value of (m ui -θ nui ) / σ ui The threshold generation device 24 may calculate the probability of undetection when the keyword voice is uttered based on a normal distribution table from the value of (1). Then, the threshold generation device 24 outputs at least one of the evaluation values calculated in this way to the user by displaying it on a monitor or the like.
[0125] Next, in S304, the setting unit 68 sets the generated threshold value in the keyword detection device 22.
[0126] When the threshold generation device 24 has completed the processes from S302 to S304 for each of the multiple keywords, it exits the loop process between S301 and S305 and ends this flow.
[0127] FIG. 18 is a diagram showing an example of the average value, standard deviation, threshold value, false detection frequency, and non-detection probability generated in the flow shown in FIG.
[0128] FA in Figure 18 24 is the false positive frequency per 24 hours. FR in Figure 18 is the probability (%) of a keyword not being detected.
[0129] In the example of FIG. 18, the keyword “cold” with keyword ID 5 is u5 <θ n5 Therefore, θ un5 <θ n5 and θu5 <θ un5 Therefore, the keyword "cold" with the keyword ID of 5 is ni =m ni +5θ ni The false detection probability set by ui =m ui -3θ ui The constraint on the non-detection probability set by
[0130] Therefore, the keyword with keyword ID 5, “Cold”, is 24 It is estimated that 54.1 times and FR is 27.4%. Other keywords are θ n5 <θ un5 and θ u5 <θ u5 Therefore, it is estimated that the constraints on the false positive and false negative probabilities are satisfied and the errors are further reduced to almost zero.
[0131] The threshold generation device 24 according to the third embodiment can prompt the user to reconsider the keywords by presenting such evaluation values to the user. For example, the threshold generation device 24 according to the third embodiment can prompt the user to change the keyword "cold" to another word that instructs the air conditioner to perform the same action, such as "turn up the temperature." This allows the threshold generation device 24 to improve the detection accuracy of the keyword detection device 22 and improve usability for the user.
[0132] The threshold value generating device 24 uses the false detection frequency (FA) per 24 hours as an evaluation value. 24 Although the example in which the evaluation value and the probability of non-detection of the keyword (FR) are output to the user has been shown, values other than these may be calculated and presented to the user. Furthermore, the threshold generation device 24 may convert the evaluation value into a qualitative indicator such as "high," "medium," or "low" based on a predetermined standard and output it.
[0133] (Fourth embodiment) Next, a voice-operated system 10 according to a fourth embodiment will be described. The voice-operated system 10 according to the fourth embodiment has substantially the same functions and configurations as the voice-operated systems 10 according to the first to third embodiments, and therefore, hereinafter, substantially the same components are denoted by the same reference numerals, and detailed description thereof will be omitted except for the differences.
[0134] For example, when the number of keywords set in the keyword detection device 22 is large, or when similar keyword pairs are included among the multiple keywords, there is a high possibility that a spoken keyword will be erroneously detected as another keyword. For example, "dengen off" and "dengen on" have many identical syllables and are therefore likely to be erroneously detected. The threshold generation device 24 according to the fourth embodiment sets a threshold to improve the accuracy of correct answer detection while suppressing erroneous detection due to such similar keywords.
[0135] 19 is a flowchart showing the flow of processing by the threshold generating device 24 according to the fourth embodiment. The threshold generating device 24 according to the fourth embodiment generates thresholds according to the flow shown in FIG.
[0136] In S401, the acquisition unit 60 acquires an input signal including a plurality of first keyword speeches, in which one or more speakers have spoken a first keyword, as a plurality of reference speeches. The first keyword is any one of a plurality of keywords set in the keyword detection device 22. In S401, the acquisition unit 60 executes the same process as S202 in FIG. 15 of the second embodiment for the first keyword.
[0137] In S402, the score calculation unit 62 calculates a first keyword score (S i Then, the score calculation unit 62 calculates the calculated keyword scores (S i(k)) is stored as a set of positive detection scores for the primary keywords. In S402, the score calculation unit 62 executes the same process as S203 in Fig. 15 of the second embodiment for the primary keywords.
[0138] Next, in S403, the distribution calculation unit 64 calculates parameters representing the distribution of the set of positive detection scores for the primary keyword. In S403, the distribution calculation unit 64 executes the same process as S204 in FIG. 15 of the second embodiment for the primary keyword.
[0139] Next, in S404, the threshold generation unit 66 generates a positive detection threshold for the first keyword based on parameters representing the distribution of the positive detection score set. For example, the threshold generation unit 66 regards the distribution of the positive detection score set as a normal distribution and generates, based on the mean value and standard deviation, a value at which the keyword scores included in the positive detection score set are larger with a predetermined probability as the positive detection threshold. In S404, the threshold generation unit 66 executes the same process as S205 of FIG. 15 of the second embodiment for the first keyword.
[0140] Next, the threshold generation device 24 executes the processes from S406 to S409 for each of one or more secondary keywords different from the primary keyword (loop process between S405 and S410). Each of the one or more secondary keywords is any one of the multiple keywords set in the keyword detection device 22. For example, each of the one or more secondary keywords is a keyword that is likely to be erroneously detected as a primary keyword when uttered.
[0141] In the loop, at S406, the acquiring unit 60 acquires an input signal including a plurality of second keyword speeches, in which one or more speakers have uttered the second keyword to be processed, as a plurality of reference speeches. At S406, the acquiring unit 60 executes the same process as at S202 in FIG. 15 of the second embodiment for the second keyword to be processed.
[0142] In S407, the score calculation unit 62 calculates a second keyword score (S ij Then, the score calculation unit 62 calculates a plurality of second keyword scores (S ij (k)) is stored as a false positive score set, which is a score set for the second keyword to be processed.
[0143] For example, when the input signal contains K secondary keyword voices, the score calculation unit 62 assigns frame numbers of the K keyword voices as k={1, 2, ..., K}. The score calculation unit 62 calculates K secondary keyword scores (S ij Then, the score calculation unit 62 calculates the calculated K second keyword scores (S ij The score set including (k) is stored as a false positive score set for the j-th secondary keyword.
[0144] Next, in S408, the distribution calculation unit 64 calculates parameters representing the distribution of the false positive score set for the second keyword to be processed. For example, the distribution calculation unit 64 assumes that the false positive score set approximates a normal distribution and calculates the mean value and standard deviation of the distribution of the false positive score set as parameters representing the distribution of the false positive score set.
[0145] For example, the distribution calculation unit 64 performs the calculation shown in equation (10) to obtain the average value (m uij ) is calculated.
number
[0146] Furthermore, for example, the distribution calculation unit 64 performs the calculation shown in Equation (11) to calculate the standard deviation (σ uij ) is calculated.
number
[0147] Next, in S409, the threshold generation unit 66 generates a false positive threshold for the second keyword to be processed based on parameters representing the distribution of the false positive score set. For example, the threshold generation unit 66 regards the distribution of the false positive score set as a normal distribution and generates, based on the mean value and standard deviation, a value at which the second keyword scores included in the false positive score set will be smaller with a predetermined probability, as the false positive threshold. For example, the threshold generation unit 66 generates, as the false positive threshold, a value at which the majority of second keyword scores included in the false positive score set will be smaller.
[0148] For example, the threshold value generating unit 66 performs the calculation shown in equation (12) to obtain the false detection threshold value (θ uij ) is calculated.
number
[0149] When the threshold generation device 24 has completed the processes from S406 to S409 for each of one or more secondary keywords, it exits the loop process between S405 and S410.
[0150] Next, in S411, the threshold value generating unit 66 calculates the false detection threshold value (θ uij ) the maximum false positive threshold (maxθ uij ).
[0151] Next, in S412, the threshold value generating unit 66 calculates the positive detection threshold value (θ ui ) and the maximum false positive threshold (maxθ uij ) is used as the threshold value (θ i For example, the threshold value generating unit 66 performs the calculation of the formula (13) to generate the intermediate value between the correct detection threshold value and the maximum erroneous detection threshold value as the threshold value (θ i) is calculated as
number
[0152] Next, in S413, the setting unit 68 sets the generated threshold value in the keyword detection device 22.
[0153] When the process of S413 is completed, the threshold generation device 24 ends the process of generating the threshold value of the first keyword.
[0154] Such a threshold generation device 24 can reduce the non-detection probability of a first keyword below a predetermined probability, and can reduce the erroneous detection probability of a second keyword that is most likely to be erroneously detected as the first keyword below a predetermined probability, provided that the correct detection threshold is greater than the maximum erroneous detection threshold. For example, the threshold generation device 24 can reduce the non-detection frequency to about 1.4 times or less when the first keyword (e.g., "Danbo") is uttered 1,000 times, and can reduce the erroneous detection frequency to about 1.4 times or less when the second keyword (e.g., "Rebo") that is most similar to the first keyword is uttered 1,000 times.
[0155] Furthermore, if the correct detection threshold is equal to or less than the maximum erroneous detection threshold, the threshold generation device 24 may output to the user that the target secondary keyword is likely to be erroneously detected as the primary keyword, thereby allowing the threshold generation device 24 to prompt the user to change the target secondary keyword.
[0156] As described above, the threshold generation device 24 according to the fourth embodiment can set a plurality of keywords in the keyword detection device 22 so that they do not cause erroneous detections.
[0157] The threshold generation device 24 according to the fourth embodiment collects keyword speech uttered by a user to prepare an input signal. However, the threshold generation device 24 may also prepare a large amount of speech data of arbitrary content to which syllable labels are attached, generate scores for each state constituting the keyword, calculate a distribution of the scores for each state, and generate a keyword score distribution from the score distribution for each state. Since this type of threshold generation device 24 does not require collection of keyword speech, it reduces the cost of collecting keyword speech and can generate thresholds in a short time even when the keyword is changed.
[0158] (Fifth embodiment) Next, a voice-operated system 10 according to a fifth embodiment will be described. The voice-operated system 10 according to the fifth embodiment has substantially the same functions and configuration as the voice-operated system 10 according to the first embodiment, and therefore, hereinafter, substantially the same components are denoted by the same reference numerals, and detailed description thereof will be omitted except for the differences.
[0159] The voice operation system 10 according to the fifth embodiment may not include the threshold generating device 24. When the voice operation system 10 does not include the threshold generating device 24, the keyword detection device 22 has an initial threshold value set in advance for each of the multiple keywords. Then, the keyword detection device 22 according to the fourth embodiment updates the threshold for each of the multiple keywords during the detection operation of detecting whether or not a keyword is included in a voice signal.
[0160] FIG. 20 is a diagram showing the configuration of the keyword detection unit 34 according to the fifth embodiment.
[0161] Compared to the keyword detection unit 34 according to the first embodiment shown in FIG. 9, the keyword detection unit 34 according to the fifth embodiment further includes a keyword score acquisition unit 82, a distribution calculation unit 64, a threshold generation unit 66, and an update unit 84.
[0162] During a detection operation for detecting whether or not a keyword is included in an audio signal, the keyword score obtaining unit 82 obtains, for each of a plurality of keywords, a keyword score for a frame in which noise is included in the audio signal from the keyword score calculation unit 46. That is, during the detection operation, the keyword score obtaining unit 82 obtains, for each of a plurality of keywords, a keyword score for each frame in a period in which the keyword voice is not uttered from the keyword score calculation unit 46.
[0163] For example, the keyword score acquiring unit 82 may be configured not to acquire the keyword output from the keyword detecting unit 34 in a predetermined number of frames before and after the frame in which the keyword is detected, based on the determination result of the determining unit 50. This allows the keyword score acquiring unit 82 to acquire a keyword score based on noise without being affected by the utterance of the keyword voice.
[0164] The distribution calculation unit 64 sequentially receives, for each of the multiple keywords, the keyword scores acquired by the keyword score acquisition unit 82. Then, for each of the multiple keywords, the distribution calculation unit 64 generates a parameter representing the distribution of a noise score set including multiple keyword scores in frames in which noise is included in the audio signal.
[0165] In the fifth embodiment, the distribution calculation unit 64 updates the mean value and standard deviation of the noise score set for each of the multiple keywords every time it receives a keyword score. For example, the distribution calculation unit 64 performs the calculation shown in Equation (14) to calculate the mean value (m ni (t)) is calculated.
number
[0166] In addition, m ni(t-1) represents the average value of the noise score set for the i-th keyword immediately before the t-th frame. i (t) is the keyword score for the i-th keyword obtained in the t-th frame.
[0167] In addition, α is a real number greater than 0 and less than 1. For example, α may be a real number such as 0.9. In addition, m ni (t-1) is set to an initial value before the start of the detection operation. ni The initial value of (t-1) may be 0 or another predetermined value.
[0168] Furthermore, for example, the distribution calculation unit 64 performs the calculations shown in equations (15) and (16) to calculate the standard deviation (σ ni (t)) is calculated.
number
number
[0169] V ni (t) represents the variance of the noise score set for the i-th keyword in the t-th frame. ni (t-1) represents the variance of the noise score set for the i-th keyword just before the t-th frame. V ni The initial value of (t-1) may be 0 or another predetermined value.
[0170] The distribution calculation unit 64 can calculate the average value and standard deviation by exponential moving average processing by performing calculations using equations (14) to (16).
[0171] The threshold generation unit 66 generates a new threshold for each of the multiple keywords based on parameters representing the distribution of the noise score set. For example, the threshold generation unit 66 regards the distribution of the noise score set as a normal distribution and generates, for each of the multiple keywords, a threshold value that makes the keyword score included in the noise score set smaller with a predetermined probability based on the average value and standard deviation.
[0172] For example, the threshold value generating unit 66 performs the calculation shown in equation (17) to obtain the threshold value (θ ni (t)) is calculated.
number
[0173] The update unit 84 updates the threshold used for comparison with the keyword score for each of the plurality of keywords to a new threshold generated by the threshold generation unit 66 for each predetermined period. In the fifth embodiment, the update unit 84 rewrites the threshold stored in the threshold storage unit 48 to the new threshold generated by the threshold generation unit 66. The predetermined period may be a frame or a period longer than a frame.
[0174] The keyword detection device 22 according to the fifth embodiment updates the threshold value as needed based on the noise contained in the audio signal during the detection operation to detect whether the audio signal contains a keyword, thereby enabling the keyword detection device 22 according to the fifth embodiment to set an appropriate threshold value in accordance with the actual noise environment.
[0175] In addition, the threshold generation unit 66 calculates the threshold value by adding five times the standard deviation to the average value in equation (17). However, the threshold generation unit 66 may calculate the threshold value by adding a multiple of the standard deviation other than five times to the average value. The designer of the threshold generation unit 66 may set an appropriate multiple in equation (17) based on constraints such as keyword misdetection. Furthermore, the distribution calculation unit 64 calculated the average value and standard deviation using exponential moving average processing. However, the distribution calculation unit 64 may also calculate the average value and standard deviation based on the noise score set for each block divided into blocks of a predetermined number of frames. Furthermore, the distribution calculation unit 64 may also calculate the average value and standard deviation using moving average processing within a window of a predetermined number of frames. Furthermore, the threshold generation unit 66 may set upper and lower limit values to clip the threshold value so that it does not become extremely large or small.
[0176] (Sixth embodiment) Next, a voice-operated system 10 according to a sixth embodiment will be described. The voice-operated system 10 according to the sixth embodiment has substantially the same functions and configurations as the voice-operated system 10 according to the modified example of the first embodiment and the voice-operated system 10 according to the fifth embodiment, and therefore, hereinafter, substantially the same components are denoted by the same reference numerals, and detailed description thereof will be omitted except for the differences.
[0177] FIG. 21 is a diagram showing the configuration of the keyword detection unit 34 according to the sixth embodiment.
[0178] The keyword detection unit 34 according to the sixth embodiment further includes a keyword score acquisition unit 82, a distribution calculation unit 64, a threshold generation unit 66, and an update unit 84, compared to the keyword detection unit 34 according to the modified example of the first embodiment shown in FIG. 14.
[0179] The keyword score acquisition unit 82 and the distribution calculation unit 64 have the same configuration as in the fifth embodiment.
[0180] The threshold generator 66 generates a modified threshold value for each of the multiple keywords based on a parameter representing the distribution of the noise score set. For example, the threshold generator 66 performs the calculation shown in Equation (18) to obtain the modified threshold value (δ ni (t)) is calculated.
number
[0181] The update unit 84 reads out the immediately preceding threshold value stored in the threshold value storage unit 48, updates the read-out threshold value based on the corrected value, and writes the updated threshold value back into the threshold value storage unit 48. For example, the update unit 84 performs the calculation shown in equation (19) to obtain the threshold value (θ ni Update (t).
number
[0182] In addition, θ ni (t-1) represents the threshold value of the i-th keyword immediately before the t-th frame.
[0183] The keyword detection device 22 according to the sixth embodiment updates the threshold value as needed based on the noise contained in the audio signal during the detection operation to detect whether the audio signal contains a keyword, thereby enabling the keyword detection device 22 according to the sixth embodiment to set an appropriate threshold value in accordance with the actual noise environment.
[0184] In addition, the threshold generation unit 66 calculates the correction value by adding five times the standard deviation to the average value in equation (18). However, the threshold generation unit 66 may calculate the correction value by adding a multiple of the standard deviation other than five times to the average value. The designer of the threshold generation unit 66 can set an appropriate multiple in equation (18) based on constraints on keyword detection errors, etc.
[0185] Fig. 22 is a diagram showing an example of the hardware configuration of the threshold generation device 24 according to each embodiment. The threshold generation device 24 is realized by, for example, a computer, which is an information processing device having the hardware configuration shown in Fig. 22. The threshold generation device 24 includes a CPU (Central Processing Unit) 301, a RAM (Random Access Memory) 302, a ROM (Read Only Memory) 303, an operation input device 304, a display device 305, a storage device 306, and a communication device 307. These components are connected via a bus.
[0186] The CPU 301 is a processor that executes arithmetic processing, control processing, etc. in accordance with a program. The CPU 301 uses a predetermined area of the RAM 302 as a work area and executes various processes in cooperation with programs stored in the ROM 303, the storage device 306, etc.
[0187] The RAM 302 is a memory such as an SDRAM (Synchronous Dynamic Random Access Memory), and functions as a work area for the CPU 301. The ROM 303 is a memory that stores programs and various types of information in a non-rewritable manner.
[0188] The operation input device 304 is an input device such as a mouse, a keyboard, etc. The operation input device 304 receives information input by a user as an instruction signal, and outputs the instruction signal to the CPU 301.
[0189] The display device 305 is a display device such as an LCD (Liquid Crystal Display), etc. The display device 305 displays various information based on a display signal from the CPU 301.
[0190] The storage device 306 is a device that writes and reads data to a semiconductor storage medium such as a flash memory, or a magnetically or optically recordable storage medium, etc. The storage device 306 writes and reads data to the storage medium in response to control from the CPU 301. The communication device 307 communicates with external devices via a network in response to control from the CPU 301.
[0191] The program executed by the computer has a modular configuration including an acquisition module, a score calculation module, a distribution calculation module, a threshold generation module, and a setting module.
[0192] This program is deployed on RAM 302 and executed by CPU 301 (processor), causing the computer to function as an acquisition unit 60, score calculation unit 62, distribution calculation unit 64, threshold generation unit 66, and setting unit 68. Note that some or all of the acquisition unit 60, score calculation unit 62, distribution calculation unit 64, threshold generation unit 66, and setting unit 68 may be realized by hardware circuits.
[0193] In addition, the program to be executed by a computer is provided as a file in a format that can be installed on a computer or in a format that can be executed by a computer, and is recorded on a computer-readable recording medium such as a CD-ROM, a flexible disk, a CD-R, or a DVD (Digital Versatile Disk).
[0194] This program may be stored on a computer connected to a network such as the Internet and provided by being downloaded via the network. This program may also be provided or distributed via a network such as the Internet. The program executed by the threshold generating device 24 may also be provided by being pre-installed in the ROM 303 or the like.
[0195] Although several embodiments of the present invention have been described, these embodiments are presented as examples and are not intended to limit the scope of the invention. These novel embodiments can be embodied in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their modifications are included within the scope and spirit of the invention, and are also included in the scope of the invention and its equivalents as defined in the claims. [Explanation of symbols]
[0196] 10 Voice Control System 20 Operation target device 22 Keyword detection device 24 Threshold generator 40 AD conversion section 42 Feature generation unit 44 Keyword model memory section 46 Keyword score calculation section 48 Threshold memory unit 50 Judgment section 52 Neural Network Department 54 Search Department 60 Acquisition Department 62 Score calculation section 64 Distribution calculation part 66 Threshold generation unit 68 Setting section 82 Keyword score acquisition unit 84 Update section
Claims
1. A threshold generation method for generating, by an information processing device, a threshold set for a keyword detection device that detects whether a keyword is included in a speech signal, based on a comparison result between a keyword score representing a similarity between a speech included in the speech signal and a predetermined keyword and the keyword score, the method comprising: the information processing device calculates the keyword score representing the degree of similarity between each of a plurality of noises that are a plurality of reference sounds and the keyword; the information processing device calculates a parameter representing a distribution of a noise score set including the plurality of keyword scores calculated based on the plurality of noises; The information processing device generates, as the threshold value, a value at which the keyword score included in the noise score set is smaller with a predetermined probability, based on a parameter representing a distribution of the noise score set. Threshold generation method.
2. The information processing device further sets the threshold value in the keyword detection device. The threshold generation method according to claim 1 .
3. The keyword detection device the information processing device sets the threshold for each of a plurality of preset keywords; calculating the keyword score for each of the plurality of keywords; For each of the plurality of keywords, the keyword score is compared with the threshold value to detect whether the corresponding keyword is included in the audio signal. The threshold generation method according to claim 1 .
4. The information processing device, in calculating the keyword scores, calculates the keyword scores for each of the plurality of keywords and each of the plurality of noises; the information processing device calculates, for each of the plurality of keywords, a parameter representing a distribution of the noise score set, in calculating the parameter representing the distribution; The information processing device generates the threshold for each of the plurality of keywords in generating the threshold. The threshold generation method according to claim 3 .
5. The information processing device, in calculating the keyword score, calculates a mean value and a standard deviation of the distribution of the noise score set as parameters representing the distribution of the noise score set; In generating the threshold value, the information processing device generates a value equal to or greater than a value obtained by adding the average value of the noise score set and a value obtained by multiplying the standard deviation of the noise score set by a predetermined first magnification. The threshold generation method according to claim 1 .
6. In calculating the keyword score, the information processing device calculates the keyword score representing a similarity between the keyword and each of a plurality of keyword voices uttering the keyword, the plurality of reference voices being the plurality of reference voices; the information processing device calculates, in the calculation of the parameter representing the distribution, a parameter representing a distribution of an utterance score set including the plurality of keyword scores calculated based on the plurality of keyword voices; The information processing device, in generating the threshold value, generating a noise threshold based on a parameter representing the distribution of the noise score set such that the keyword score included in the noise score set is smaller with a predetermined probability; generating an utterance threshold at which the keyword score included in the utterance score set is larger with a predetermined probability based on a parameter representing a distribution of the utterance score set; generating a value between the noise threshold and the speech threshold as the threshold; The threshold generation method according to claim 1 .
7. The information processing device, in calculating the parameters representing the distribution, calculates a mean value and a standard deviation of the distribution of the noise score set as parameters representing the distribution of the noise score set; The information processing device generates, in generating the threshold, a value obtained by adding the average value of the noise score set and a value obtained by multiplying the standard deviation of the noise score set by a predetermined first magnification as the noise threshold, the information processing device calculates a mean value and a standard deviation of the distribution of the speech score set as parameters representing the distribution of the speech score set, the information processing device, in generating the threshold, generates, as the speech threshold, a value obtained by subtracting a value obtained by multiplying the standard deviation of the distribution of the speech score set by a predetermined second magnification from the average value of the distribution of the speech score set; The information processing device generates a value between the noise threshold and the speech threshold as the threshold in generating the threshold. The threshold generation method according to claim 6 .
8. The information processing device, in generating the threshold, outputs to a user at least one of a false detection probability or frequency calculated based on the threshold and the noise score set, and a non-detection probability or frequency calculated based on the threshold and the speech score set. The threshold generation method according to claim 7 .
9. The information processing device, in generating the threshold value, generating the new threshold based on a parameter representing the distribution of the noise score set; The threshold used for comparison with the keyword score is updated to the generated new threshold for each predetermined period. The threshold generation method according to claim 1 .
10. A threshold generation device that generates a threshold to be set for a keyword detection device that detects whether a voice signal contains a keyword based on a comparison result between a keyword score representing a similarity between a voice included in the voice signal and a predetermined keyword and a threshold, the threshold generation device comprising: a score calculation unit that calculates the keyword score representing the degree of similarity between each of a plurality of noises that are a plurality of reference voices and the keyword; a distribution calculation unit that calculates a parameter representing a distribution of a noise score set including the plurality of keyword scores calculated based on the plurality of noises; a threshold value generating unit that generates, as the threshold value, a value at which the keyword score included in the noise score set is smaller with a predetermined probability based on a parameter representing a distribution of the noise score set; A threshold generating device comprising:
11. A program for causing a computer to function as a threshold generation device that generates a threshold to be set for a keyword detection device, the keyword detection device detects whether the speech signal contains the keyword based on a comparison result between a keyword score representing a degree of similarity between a speech included in the speech signal and a predetermined keyword and the threshold value; The computer a score calculation unit that calculates the keyword score representing the degree of similarity between each of a plurality of noises that are a plurality of reference voices and the keyword; a distribution calculation unit that calculates a parameter representing a distribution of a noise score set including the plurality of keyword scores calculated based on the plurality of noises; a threshold value generating unit that generates, as the threshold value, a value at which the keyword score included in the noise score set is smaller with a predetermined probability based on a parameter that represents a distribution of the noise score set; A program that makes it work.
Citation Information
Patent Citations
Pattern recognition device
JP1996254991A
Voice spotting device
JP1999190999A
Speech recognition device and speech recognition method
JP2013080015A
Voice recognition system
JP2019184633A
Information processing device, information processing method and program
JP2021033051A