Continuous utterance estimation method, continuous utterance estimation device, and program
The continuous speech estimation method addresses the issue of overlapping sounds by adapting operations based on keyword detection and continuous speech analysis, enhancing voice recognition accuracy in devices like smart speakers and in-car systems.
Patent Information
- Application Number
- JP2025134205
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2025-10-22
AI Technical Summary
Conventional voice-controlled devices face issues when users utter keywords and target sounds consecutively, leading to cut-off or overlapping sounds that are difficult to recognize, as they fail to distinguish between separate and consecutive speech methods.
A continuous speech estimation method that detects keywords and determines if continuous speech follows, controlling response sound emission and target sound intervals based on this detection to adapt operations to different usage methods.
The system effectively distinguishes between separate and consecutive speech methods, preventing sound overlap and improving recognition accuracy by adjusting response and target sound timings.
Smart Images

Figure 2025160510000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a technique for estimating whether or not a target sound is uttered consecutively after a keyword is uttered. [Background technology]
[0002] For example, devices that can be controlled by voice, such as smart speakers and in-car systems, are sometimes equipped with a function called keyword wake-up, which starts voice recognition when a trigger keyword is pronounced. Such a function requires technology to input a voice signal and detect the pronunciation of the keyword.
[0003] FIG. 1 shows the configuration of the conventional technology disclosed in Non-Patent Document 1. In the conventional technology, when a keyword detection unit 91 detects the pronunciation of a keyword from an input voice signal, a target sound output unit 99 turns on a switch and outputs the voice signal as a target sound to be subjected to voice recognition or the like. Furthermore, a response sound output unit 92 outputs a response sound when a keyword is detected to notify the user that the pronunciation of the keyword has been detected. At this time, in order to control the timing of each process, a delay unit 93 may be further provided to delay the output of the keyword detection unit 91 (see FIG. 1A) or the input voice (see FIG. 1B). [Prior art documents] [Non-patent literature]
[0004] [Non-Patent Document 1] Sensory, Inc., "TrulyHandsfreeTM," [online], [searched August 17, 2018], Internet<URL: http: / / www.sensory.co.jp / product / thf.htm> Summary of the Invention [Problem to be solved by the invention]
[0005] However, in the conventional technology, in addition to the usage method of uttering a keyword, waiting for a response sound, and then uttering the target sound, there is also a usage method of uttering a keyword and the target sound consecutively. If the start position of the target sound interval is set after the response sound, assuming a usage method of waiting for the response sound and then uttering the target sound, a problem occurs in that the beginning of the target sound is cut off when the user utters the keyword and the target sound consecutively. Furthermore, if the start position of the target sound interval is set immediately after the utterance of the keyword, assuming a usage method of uttering a keyword and the target sound consecutively, there is a problem in that the response sound overlaps in time with the utterance of the target sound, resulting in a sound that is difficult to recognize.
[0006] In view of the technical problems described above, the object of this invention is to automatically distinguish between a usage method in which a keyword is spoken, then a response sound is heard, and then a target sound is spoken, and a usage method in which a keyword and a target sound are spoken successively, and to change operation appropriately to suit each usage method. [Means for solving the problem]
[0007] In order to solve the above problem, a continuous speech estimation method according to one aspect of the present invention detects whether a keyword is included in a user's speech, controls the emission of a response sound so that if there is continuous speech after the keyword, no response sound is emitted, and if there is no continuous speech after the keyword, a response sound is emitted, and if there is no continuous speech after the keyword, the target sound after the emission of the response sound is targeted for speech recognition. [Effects of the Invention]
[0008] According to this invention, it is advantageous to wait for a response sound after uttering a keyword and then utter a target sound. Since the system can automatically distinguish between a method of using a voice recorder and a method of using a voice recorder in which a keyword and target sound are spoken consecutively, it can change its operation appropriately to suit each method of use. [Brief explanation of the drawings]
[0009] [Figure 1]FIG. 1 is a diagram illustrating the functional configuration of a conventional keyword detection device. [Figure 2] FIG. 2 is a diagram for explaining the principle of the invention. [Figure 3] FIG. 3 is a diagram illustrating the functional configuration of a continuous utterance estimation device according to the first embodiment. [Figure 4] FIG. 4 is a diagram illustrating a processing procedure of the continuous utterance estimation method according to the first embodiment. [Figure 5] FIG. 5 is a diagram illustrating the functional configuration of a continuous utterance estimation device according to the second embodiment. [Figure 6] FIG. 6 is a diagram illustrating the functional configuration of a continuous utterance estimation device according to the third embodiment. [Figure 7] FIG. 7 is a diagram illustrating the functional configuration of a continuous utterance estimation device according to the fourth embodiment. [Figure 8] FIG. 8 is a diagram illustrating the functional configuration of a continuous utterance estimation device according to the fifth embodiment. [Figure 9] FIG. 9 is a diagram illustrating the functional configuration of a continuous utterance estimation device according to the sixth embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0010] In the prior art, it was difficult to support both a usage method in which a target sound is uttered after a response sound is received after a keyword is spoken, and a usage method in which the keyword and target sound are spoken consecutively. If a response sound is emitted when a keyword is detected assuming a usage method in which a target sound is uttered after a keyword is spoken and a response sound is received before the keyword is spoken, the main problem is that the response sound and the target sound overlap when the user speaks assuming a usage method in which the keyword and target sound are spoken consecutively.
[0011] The object of this invention is to automatically distinguish between a usage method in which a target sound is uttered after a response sound is received and a usage method in which the keyword and target sound are uttered consecutively, and to change the start position of the target sound interval and whether or not a response sound is emitted based on the result of the distinction, thereby performing an operation appropriate for each usage method. Specifically, when it is determined that the usage method is one in which a target sound is uttered after a response sound is received and then the target sound is uttered, the response sound is emitted first, and the target sound interval begins after the response sound has been emitted (see FIG. 2A). On the other hand, when it is determined that the usage method is one in which a keyword and target sound are uttered consecutively, the response sound is not emitted, and the target sound interval begins immediately after the keyword has been uttered (see FIG. 2B).
[0012] Hereinafter, an embodiment of the present invention will be described in detail. In the drawings, components having the same functions are designated by the same reference numerals, and duplicated explanations will be omitted.
[0013] [First embodiment] A continuous utterance estimation device 1 of the first embodiment receives a user's voice (hereinafter referred to as "input voice") as input, and outputs a keyword detection result that determines whether the input voice contains a pronunciation of a keyword, and a continuous utterance detection result that determines whether the input voice contains a continuous utterance after the utterance of the keyword. As shown in Fig. 3, the continuous utterance estimation device 1 includes a keyword detection unit 11, a voice detection unit 12, and a continuous utterance detection unit 13. The continuous utterance estimation method S1 of the first embodiment is realized by the continuous utterance estimation device 1 performing the processing of each step shown in Fig. 4.
[0014] The continuous utterance estimation device 1 is implemented by a known or dedicated computer having, for example, a central processing unit (CPU), a main memory (RAM: Random Access Memory), etc. The continuous utterance estimation device 1 is a special device configured by loading a special program into the device. The continuous utterance estimation device 1 executes each process under the control of, for example, a central processing unit. Data input to the continuous utterance estimation device 1 and data obtained in each process are stored in, for example, a main memory device, and the data stored in the main memory device is read out to the central processing unit as needed and used for other processes. At least a part of each processing unit of the continuous utterance estimation device 1 may be configured by hardware such as an integrated circuit.
[0015] Hereinafter, with reference to FIG. 4, a continuous utterance estimation method executed by the continuous utterance estimation device of the first embodiment will be described.
[0016] In step S11, the keyword detection unit 11 detects the pronunciation of a predetermined keyword from the input speech. The keyword detection is performed, for example, by determining whether a power spectrum pattern obtained over a short period of time is similar to a keyword pattern recorded in advance using a pre-trained neural network. The keyword detection unit 11 outputs a keyword detection result indicating whether a keyword has been detected or not to the continuous speech detection unit 13.
[0017] In step S12, the speech detection unit 12 detects a speech interval from the input speech. The speech interval detection is performed, for example, as follows: First, the steady noise level N(t) is calculated from the long-term average of the input speech. Next, a threshold is set by multiplying the steady noise level N(t) by a predetermined constant α. Then, an interval in which the short-term average level P(t) exceeds the threshold is detected as a speech interval. Alternatively, the speech interval may be detected by a method that adds to the determination factors whether the shape of the spectrum or cepstrum matches the characteristics of the speech. The speech detection unit 12 outputs a speech interval detection result indicating whether a speech interval has been detected or not to the continuous speech detection unit 13.
[0018] The short-term average level P(t) is calculated by multiplying the root mean square power by a rectangular window of the average keyword utterance time T or by an exponential window. If the power at discrete time t is P(t) and the input signal is x(t), then:
[0019]
number
[0020] where α is a forgetting coefficient and is set in advance to a value between 0 and α and 1. α is set so that the time constant is the average keyword utterance time T (samples). In other words, α = 1-1 / T. Alternatively, the absolute average power multiplied by a rectangular window of the keyword utterance time T or the absolute average power multiplied by an exponential window may be calculated as shown in the following formula.
[0021]
number
[0022] In step S13, the continuous speech detection unit 13 determines that the utterance is continuous speech if the keyword detection result output by the keyword detection unit 11 indicates that a keyword has been detected and the voice activity detection result output by the voice detection unit 12 indicates that a voice activity has been detected. Because keyword detection by the keyword detection unit 11 involves a delay of several hundred milliseconds, the utterance of the keyword has already ended by the time the keyword detection process is completed. Therefore, the presence or absence of a voice activity at the time of keyword detection can determine whether or not a speech segment has been detected in the continuous speech. The continuous speech detection unit 13 outputs the continuous speech detection result indicating that continuous speech has been detected or not, together with the keyword detection result output by the keyword detection unit 11, to the continuous speech estimation device 1.
[0023] By configuring in this way, according to the first embodiment, after the keyword is uttered, Since it is possible to determine whether or not a target sound section is being spoken, it becomes possible to change the start position of the target sound section and whether or not a response sound is being emitted based on the continuous utterance detection result output by the continuous utterance estimation device 1.
[0024] [Second embodiment] Similar to the first embodiment, the continuous utterance estimation device 2 of the second embodiment receives user speech as input and outputs keyword detection results and continuous utterance detection results. As shown in Fig. 5, the continuous utterance estimation device 2 further includes a delay unit 21 in addition to the keyword detection unit 11, speech detection unit 12, and continuous utterance detection unit 13 of the first embodiment.
[0025] The delay unit 21 adds a delay to the keyword detection result output by the keyword detection unit 11. This delay is provided to make up for the shortfall in the delay of keyword detection when it is too short to determine whether or not a speech start of continuous speech exists, by adding a delay of XY to the output of the keyword detection unit 11. If the appropriate delay for determining whether or not a speech start of continuous speech exists is X, and the delay of keyword detection is Y, the delay is set to XY.
[0026] With this configuration, according to the second embodiment, it is possible to determine the presence or absence of continuous speech at an appropriate timing.
[0027] [Third embodiment] The third embodiment is configured to change whether or not a response sound is emitted based on the continuous speech detection result of the first or second embodiment. Consider emitting a response sound when a keyword is detected to notify the user that the keyword has been detected. When the target sound is pronounced consecutively with the keyword, the target sound is uttered before the response sound is emitted, so the response sound is unnecessary. Furthermore, if a response sound were to be emitted in this case, the response sound would be superimposed on the target sound, which would be inconvenient for voice recognition, etc. Therefore, in the third embodiment, if continuous speech is detected during keyword detection, the response sound is not emitted, but if continuous speech is not detected during keyword detection, the response sound is emitted.
[0028] The continuous utterance estimation device 3 of the third embodiment receives a user's voice as input, and if continuous utterance is not detected when detecting keywords from the input voice, emits a response sound. As shown in Fig. 6, the continuous utterance estimation device 3 includes a continuous utterance detection-enabled keyword detection unit 10, a switch unit 20, and a response sound output unit 30.
[0029] Specifically, the keyword detection unit with continuous speech detection 10 is configured in the same manner as the continuous speech estimation device 1 of the first embodiment or the continuous speech estimation device 2 of the second embodiment. That is, the keyword detection unit with continuous speech detection 10 includes at least a keyword detection unit 11, a speech detection unit 12, and a continuous speech detection unit 13, receives a user's speech as input, and outputs a keyword detection result and a continuous speech detection result.
[0030] The switch unit 20 controls whether or not to transmit the keyword detection result output by the continuous speech detection-equipped keyword detection unit 10 to the response sound output unit 30. If the continuous speech detection result output by the continuous speech detection-equipped keyword detection unit 10 is true (i.e., if continuous speech is detected), the keyword detection result is not transmitted to the response sound output unit 30, and if the continuous speech estimation result is false (i.e., if continuous speech is not detected), the keyword detection result is transmitted to the response sound output unit 30.
[0031] When a keyword detection result indicating that a keyword has been detected is transmitted from the switch section 20, the response sound output section 30 outputs a predetermined response sound.
[0032] By configuring in this way, according to the third embodiment, when continuous speech is made following a keyword, unnecessary response sounds are not emitted, and deterioration in accuracy of voice recognition and the like can be prevented.
[0033] [Fourth embodiment] The fourth embodiment is configured to change the start position of the target sound section based on the continuous speech detection result of the first or second embodiment. In a usage method in which a keyword and a target sound are spoken consecutively, it is assumed that the target sound starts to be spoken before the keyword is detected due to a delay in keyword detection. Therefore, when the keyword is detected, it is necessary to go back in time and extract the target sound. In a usage method in which a keyword is spoken and then a response sound is heard before the target sound is spoken, it is necessary to extract the target sound from a point in time when the length of the response sound has elapsed since the keyword was detected, in order to extract the part after the response sound as the target sound. If this is not done, the response sound will be superimposed on the target sound, causing inconvenience to speech recognition, etc.
[0034] The continuous utterance estimation device 4 of the fourth embodiment receives a user's voice as input, and if continuous speech is detected when a keyword is detected from the input voice, outputs the target sound immediately after the keyword is uttered, and if continuous speech is not detected when a keyword is detected from the input voice, outputs the target sound after the response sound has been emitted. As shown in Figure 7, the continuous utterance estimation device 4 includes delay units 41 and 43, switch units 42 and 44, and a target sound output unit 45 in addition to the continuous utterance detection-equipped keyword detection unit 10 of the third embodiment.
[0035] The delay unit 41 gives a delay equivalent to the length of the response sound to the keyword detection result output by the continuous speech detection and keyword detection unit 10.
[0036] When the delayed keyword detection result output by the delay unit 41 indicates that a keyword has been detected, the switch unit 42 turns on the switch and outputs the input voice to the target voice output unit 45. That is, the switch operates so that it turns on after the response voice has been emitted.
[0037] The delay unit 43 gives the input voice a delay equivalent to the delay of the keyword detection performed by the continuous speech detection and keyword detection unit 10 .
[0038] When the keyword detection result (i.e., the undelayed keyword detection result) output by the continuous speech detection with keyword detection unit 10 indicates that a keyword has been detected, the switch unit 44 turns on the switch and outputs the delayed input voice output by the delay unit 43 to the target voice output unit 45. In other words, the switch operates so as to be on immediately after the keyword is uttered.
[0039] The target sound output unit 45 selects either the output of the switch unit 42 or the output of the switch unit 44 and outputs it as the target sound. Specifically, if the continuous speech detection result output by the continuous speech detection-with-keyword detection unit 10 is true (i.e., if continuous speech is detected), the target sound output unit 45 selects the output of the switch unit 44 (i.e., the input voice from immediately after the keyword utterance), and if the continuous speech detection result is false (i.e., if continuous speech is not detected), the target sound output unit 45 selects the output of the switch unit 42 (i.e., the input voice from after the response sound is emitted), and outputs it as the target sound. In this way, if continuous speech is detected during keyword detection, the target sound is output immediately after the keyword is uttered, and if continuous speech is not detected during keyword detection, the target sound is output after the response sound has finished being emitted.
[0040] By configuring in this way, according to the fourth embodiment, when a continuous utterance is made following a keyword, the input voice immediately after the keyword utterance is output as the target voice, and the voice recognition is performed. In addition, when a target sound is uttered after a response sound is output following a keyword utterance, the input voice after the target sound is output is output as the target sound, thereby preventing degradation of speech recognition due to the superimposition of the response sound.
[0041] [Fifth embodiment] The fifth embodiment is a combination of the third and fourth embodiments. The continuous utterance estimation device 5 of the fifth embodiment receives a user's voice as input, and if continuous utterance is detected when a keyword is detected from the input voice, outputs a target sound immediately after the keyword is uttered. If continuous utterance is not detected when a keyword is detected from the input voice, the device outputs a response sound and then outputs the target sound after the response sound has been emitted.
[0042] 8, the continuous utterance estimation device 5 includes the continuous utterance detection-equipped keyword detection unit 10, switch unit 20, and response sound output unit 30 of the third embodiment, and delay units 41 and 43, switch units 42 and 44, and target sound output unit 45 of the fourth embodiment. The operation of each processing unit is the same as in the third and fourth embodiments.
[0043] [Sixth embodiment] The continuous utterance estimation device 6 of the sixth embodiment receives multi-channel speech as input and outputs keyword detection results and continuous utterance detection results for each channel. As shown in Fig. 9, the continuous utterance estimation device 6 includes pairs of keyword detection unit 11 and continuous utterance detection unit 14 of the first embodiment, each set corresponding to the number M (≥ 2) of channels of the input speech, and further includes a multi-input speech detection unit 62 with M-channel input / output.
[0044] The multi-input speech detector 62 receives multi-channel speech as input, and for each integer i between 1 and M, detects a speech segment from the speech signal of channel i, and outputs the detected speech segment to the continuous speech detector 14-i. The multi-input speech detector 62 can detect speech segments more accurately by exchanging speech level information between channels. The method for detecting speech segments from multi-channel input can be, for example, the method described in Reference 1 below.
[0045] [Reference 1] JP 2017-187688 A With this configuration, according to the sixth embodiment, when a multi-channel audio signal is input, it is possible to accurately detect a speech segment, thereby improving the accuracy of continuous speech estimation.
[0046] Although the embodiments of the present invention have been described above, the specific configuration is not limited to these embodiments, and it goes without saying that the present invention includes appropriate design changes and the like within the scope of the spirit of the present invention. The various processes described in the embodiments may not only be executed in chronological order in the order described, but may also be executed in parallel or individually depending on the processing capacity of the device executing the processes or as needed.
[0047] [Programs, recording media] When the various processing functions of each device described in the above embodiments are realized by a computer, the processing contents of the functions that each device should have are described by a program, and the various processing functions of each device are realized on the computer by executing this program.
[0048] The program describing the processing contents can be recorded on a computer-readable recording medium, which may be, for example, a magnetic recording device, an optical disk, a magneto-optical recording medium, a semiconductor memory, or any other suitable recording medium.
[0049] In addition, distribution of this program may be restricted to, for example, DVDs, CD-ROMs, etc. on which the program is recorded. This can be done by selling, transferring, lending, etc. the portable recording medium. Furthermore, this program can be distributed by storing it in a storage device of a server computer and transferring the program from the server computer to other computers via a network.
[0050] A computer that executes such a program, for example, first stores the program recorded on a portable recording medium or transferred from a server computer in its own storage device. Then, when executing a process, the computer reads the program stored in its own storage device and executes the process in accordance with the read program. As another execution mode of this program, the computer may read the program directly from a portable recording medium and execute the process in accordance with the program, or may execute the process in accordance with the received program each time a program is transferred from the server computer to this computer. Also, there is a so-called ASP (Application Service Provider) type service in which the server computer does not transfer the program to this computer, but instead realizes the processing function by simply issuing an execution instruction and obtaining the results. The program in this embodiment includes information used for processing by a computer and equivalent to a program (data that is not a direct command to a computer but has the properties of defining the processing of a computer, etc.).
[0051] Furthermore, in this embodiment, the device is configured by executing a predetermined program on a computer, but at least a part of the processing contents may be realized by hardware. [Explanation of symbols]
[0052] 1,2,3,4,5,6 Continuous speech estimation device 9 Keyword detector 11,91 Keyword detection section 12 Voice detection unit 13 Continuous speech detection unit 30,92 Response sound output section 21, 41, 43, 93 Delay section 20, 42, 44 Switch section 45,99 Target sound output section 62 Multi-input audio detector
Claims
1. Detect whether the user's voice contains keywords, controlling the emission of the response sound so that if there is a continuous utterance after the keyword, the response sound is not emitted, and if there is no continuous utterance after the keyword, the response sound is emitted; If there is no subsequent speech after the keyword, the target sound after the end of the response sound is used for speech recognition. A method for continuous speech estimation.
2. 2. The continuous speech estimation method according to claim 1, wherein, if there is no continuous speech following the keyword, the target sound is extracted from a point in time when a length of the response sound has elapsed since the keyword was detected.
3. 3. The continuous speech estimation method according to claim 2, wherein if there is a continuous speech following the keyword, a portion immediately after the utterance of the keyword is extracted as the target sound.
4. a keyword detection unit with continuous speech detection that detects whether a keyword is included in the user's voice and whether there is a continuous speech following the keyword; a control unit that controls the emission of the response sound so that the response sound is not emitted when there is a consecutive utterance after the keyword, and the response sound is emitted when there is no consecutive utterance after the keyword; a target sound output unit that, when there is no subsequent speech after the keyword, recognizes a target sound after the end of output of the response sound; A continuous speech estimation device including:
5. On the computer, Detect whether the user's voice contains keywords, controlling the emission of the response sound so that if there is a continuous utterance after the keyword, the response sound is not emitted, and if there is no continuous utterance after the keyword, the response sound is emitted; If there is no subsequent speech after the keyword, the target sound after the end of the response sound is used for speech recognition. A program for executing a process.