Channel selection device, channel selection method, and program
Patent Information
- Application Number
- JP2026108988
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2026-06-22
- Publication Date
- 2026-09-01
AI Technical Summary
【0008】 この発明によれば、複数チャネルの音響信号からキーワードの発音が含まれるチャネルを少ない演算量で適切に選択することができる。
Smart Images

Figure 2026139868000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a technology for selecting a channel containing pronunciation of a keyword from acoustic signals of a plurality of channels.
Background Art
[0002] For example, voice-controllable devices such as smart speakers and in-vehicle systems are sometimes equipped with a function called keyword wake-up, which starts speech recognition when a trigger keyword is pronounced. Such a function requires a technology that receives an audio signal as input and detects the pronunciation of a keyword.
[0003] Figure 1 shows the configuration of the conventional technology disclosed in Non-Patent Document 1. In the conventional technology, when a keyword detection unit 91 detects pronunciation of a keyword from an input audio signal, a target sound output unit 99 turns on a switch and outputs the audio signal as a target sound to be subjected to speech recognition or the like. When the input audio is multi-channel, as shown in Figure 1, if as many sets of the keyword detection unit 91 and the target sound output unit 99 as the number of channels are prepared, a channel containing the keyword can be selected from the plurality of channels. For example, if the above processing is performed using acoustic signals collected by a plurality of microphones installed in a room as input, it is possible to know near which microphone the keyword was pronounced, and to specify the utterance position and perform speech recognition triggered by the keyword.
Prior Art Documents
Non-Patent Documents
[0004]
Non-Patent Document 1
Summary of the Invention
[0005] However, conventional technology requires keyword detection processing for each channel, resulting in an enormous amount of computation. Furthermore, in the case of multiple microphones installed in the same room, the same keyword utterance may be picked up by multiple microphones, and the keyword may be included in multiple channels. In this case, the microphone closest to the location of the keyword utterance should be selected, but conventional technology selects all channels that detected the pronunciation of the keyword.
[0006] In view of the technical problems described above, the objective of this invention is to appropriately select the channel containing the pronunciation of a keyword from multiple channels of acoustic signals with a small amount of computation. [Means for solving the problem]
[0007] To solve the above problems, the selection device of the first embodiment of this invention includes the steps of: acquiring channels of a plurality of audio signals picked up by three or more microphones; acquiring the power of each of the plurality of channels; and selecting the audio signal of the channel with the highest power among the channels with the highest power among the channels of the plurality of audio signals, from among the channels in which a keyword is detected, as the target sound to be recognized. [Effects of the Invention]
[0008] According to this invention, it is possible to appropriately select the channel containing the pronunciation of a keyword from multiple channels of acoustic signals with a small amount of computation. [Brief explanation of the drawing]
[0009] [Figure 1] Figure 1 illustrates the functional configuration of a conventional keyword detection device. [Figure 2] Figure 2 illustrates the functional configuration of the channel selection device according to the first embodiment. [Figure 3]Figure 3 illustrates the processing procedure of the channel selection method according to the first embodiment. [Figure 4] Figure 4 is a diagram illustrating the principle of the first embodiment. [Figure 5] Figure 5 illustrates the functional configuration of the channel selection device according to the second embodiment. [Figure 6] Figure 6 illustrates the processing procedure of the channel selection method according to the second embodiment. [Figure 7] Figure 7 is a diagram illustrating the principle of the third embodiment. [Figure 8] Figure 8 illustrates the functional configuration of the channel selection device according to the third embodiment. [Figure 9] Figure 9 is a diagram illustrating the functional configuration of the channel selection device according to the fourth embodiment. [Modes for carrying out the invention]
[0010] The embodiments of this invention will be described in detail below. In the drawings, components having the same function will be numbered the same, and redundant explanations will be omitted.
[0011] [First Embodiment] The channel selection device 1 of the first embodiment takes multiple channels of audio signals (hereinafter referred to as "input audio signals") as input and selects and outputs the audio signal of the channel suitable for the target sound to be targeted for speech recognition, etc., from among the channels in which the pronunciation of a keyword is detected. As shown in Figure 2, the channel selection device 1 comprises an adder 11, a keyword detection unit 12, M power calculation units 13-1, ..., 13-M, M delay units 14-1, ..., 14-M, a maximum power detection unit 15, and a channel selection unit 16. Herein, M is the number of channels of the input audio signal and is an integer of 2 or more. The channel selection method S1 of the first embodiment is realized when this channel selection device 1 performs the processing of each step shown in Figure 3.
[0012] The channel selection apparatus 1 is, for example, a known or dedicated computer including a central processing unit (CPU: Central Processing Unit), a main memory (RAM: Random Access Memory), and the like which is a special apparatus configured by loading a special program therein. The channel selection apparatus 1 executes each process under control of the central processing unit, for example. Data input to the channel selection apparatus 1 and data obtained in each process are stored in, for example, the main memory, and the data stored in the main memory is read out to the central processing unit as needed and used for other processes. At least a part of each processing unit of the channel selection apparatus 1 may be configured by hardware such as an integrated circuit.
[0013] Hereinafter, a channel selection method executed by the channel selection apparatus of the first embodiment will be described with reference to FIG. 3.
[0014] In step S11, an adding unit 11 adds all channels of an input M-channel audio signal (hereinafter referred to as "input audio signal") to generate a 1-channel audio signal (hereinafter referred to as "synthesized audio signal"). The adding unit 11 outputs the synthesized audio signal to a keyword detecting unit 12.
[0015] In step S12, the keyword detecting unit 12 receives the synthesized audio signal output from the adding unit 11 as an input, and detects pronunciation of a predetermined keyword from the synthesized audio signal. Keyword detection is performed, for example, by determining whether a power spectrum pattern obtained in a short period is similar to a pre-recorded keyword pattern using a pre-trained neural network. Instead of using the voice of a keyword, a sound-producing action such as whistling or clapping may be used. The keyword detecting unit 12 outputs a keyword detection result indicating whether a keyword has been detected or not detected to a maximum power detecting unit 15.
[0016] In step S13, a power calculation unit 13-i (i=1, …, M) calculates the power of the i-th channel of the input audio signal (hereinafter referred to as "channel i"). The power calculation unit 13-i outputs the power of channel i to a delay unit 14-i. For power calculation, the mean-square power obtained by applying a rectangular window with an average keyword utterance duration T or the mean-square power obtained by multiplying by an exponential window is calculated. If the power of channel i at discrete time t is represented as Pi(t) and the input signal is represented as xi(t),
[0017] [Math]
[0018] where α is a forgetting coefficient, and a value satisfying 0<α<1 is set in advance. α is set such that its time constant corresponds to the average keyword utterance duration T (in samples). That is, α=1-1 / T. Alternatively, as shown in the following formula, the average absolute value power obtained by applying a rectangular window with the keyword utterance duration T or the average absolute value power obtained by multiplying by an exponential window may be calculated.
[0019] [Math]
[0020] The power calculated by the power calculation unit 13-i may be a value obtained by subtracting a noise level. The noise level can be obtained from the average value of long-term signal power or a dip-hold value. Dip-hold processing that holds the floor of the calculated power Pi(t) is performed to obtain stationary noise pow er Ni(t). This calculation can be realized, for example, by performing averaging processing with a long time constant when the power rises, and performing averaging processing with a short time constant when the power falls.
[0021] [Math]
[0022] However, β < γ, and each value can be between 0 and 1 (inclusive).
[0023] Noise level subtraction can also be performed in the frequency domain. By calculating the power and noise levels in each frequency domain and subtracting them, noise subtraction can be performed more accurately.
[0024] In step S14, the delay unit 14-i (i=1,…,M) delays the power of channel i output by the power calculation unit 13-i by time D. Time D is set to the time corresponding to the detection delay of keyword detection. The delay unit 14-i then calculates the power of channel i after the delay. Output is sent to the high-power detection unit 15.
[0025] In step S15, when the keyword detection result output by the keyword detection unit 12 indicates that a keyword has been detected, the maximum power detection unit 15 selects the channel with the highest power among the powers of each channel output by the delay units 14-1, ..., 14-M as the output channel. The maximum power detection unit 15 outputs information indicating the selected output channel to the channel selection unit 16.
[0026] In step S16, the channel selection unit 16 selects the audio signal of the output channel from the input audio signal according to the information indicating the output channel output by the maximum power detection unit 15, and outputs it as the target sound.
[0027] The channel selection device 1 of the first embodiment estimates the keyword's speech channel by calculating the power of the portion of the channel corresponding to the keyword's speech interval (see Figure 4) for each channel, based on the hypothesis that the signal power of the channel containing the keyword is greatest during the keyword's speech interval.
[0028] With this configuration, according to the first embodiment, a single keyword detection process can be used to select a channel containing a keyword utterance from multiple channels. Furthermore, if multiple channels contain audio components of a keyword utterance, such as signals from multiple microphones placed in a room, the channel with the highest signal level can be selected.
[0029] [Second Embodiment] In the first embodiment, keyword detection is performed after adding all channels of the input audio signal. Therefore, if the system includes audio signals from channels without keyword utterances in addition to the audio signals from channels with keyword utterances, the signal-to-noise ratio (SNR) of the synthesized audio signal after addition will deteriorate, and the accuracy of keyword detection will decrease. In the second embodiment, when three or more audio signals are input, the K channel audio signal with the highest power is selected from the M channel audio signals. Keyword detection processing is then performed on each of the selected K channel audio signals, and the channel with the highest power among the audio signals in which keyword detection occurred is selected as the target sound. In this way, candidate channels are selected first using only power information, and keyword detection is performed on each candidate channel. This reduces the number of keyword detection processes while avoiding the decrease in SNR caused by addition.
[0030] The channel selection device 2 of the second embodiment takes three or more audio signals as input and selects and outputs the audio signal of the channel suitable for the target sound to be targeted for speech recognition, etc., from among the channels in which the pronunciation of a keyword is detected. As shown in Figure 5, the channel selection device 2 further includes K keyword detection units 12-1, ..., 12-K, M delay units 21-1, ..., 21-M, candidate selection unit 22, and candidate channel selection unit 23, in addition to the power calculation units 13-1, ..., 13-M, delay units 14-1, ..., 14-M, maximum power detection unit 15, and channel selection unit 16 of the first embodiment. However, K is an integer between 1 and M. The channel selection device 2 performs the processing of each step shown in Figure 6 to realize the channel selection method S2 of the second embodiment.
[0031] The channel selection method performed by the channel selection device of the second embodiment will be described below with reference to Figure 6, focusing on the differences from the channel selection method of the first embodiment.
[0032] In step S21, the delay unit 21-i (i=1,…,M) delays the audio signal of channel i of the input audio signal. This is related to the processing of the power calculation unit 13-i and the candidate selection unit 22. This delay is implemented to prevent the beginning of a keyword from being cut off due to selection delay, and it provides a delay of several hundred milliseconds. The delay unit 21-i outputs the audio signal of channel i after the delay to the candidate channel selection unit 23.
[0033] In step S22, the candidate selection unit 22 selects the K channel with the highest power among the M channels of the input audio signal as a candidate channel based on the power of each channel output by the power calculation units 13-1, ..., 13-M. The candidate selection unit 22 outputs information indicating the selected candidate channel to the candidate channel selection unit 23.
[0034] In step S23, the candidate channel selection unit 23 selects the audio signal of a candidate channel from the delayed input audio signal output by the delay unit 21-i, according to the information indicating the candidate channel output by the candidate channel selection unit 22. The candidate channel selection unit 23 outputs the audio signal of the j-th candidate channel (hereinafter referred to as "candidate channel j") to the keyword detection unit 12-j.
[0035] In step S12, the keyword detection unit 12-j receives the audio signal of candidate channel j output by the candidate channel selection unit 23 as input and detects the pronunciation of a predetermined keyword from the audio signal. Keyword detection can be performed in the same manner as in the first embodiment. The keyword detection unit 12-j outputs the keyword detection result to the maximum power detection unit 15.
[0036] In step S15, when the keyword detection result output by the keyword detection unit 12-j indicates that a keyword has been detected, the maximum power detection unit 15 selects the channel with the highest power among the outputs of the delay units 14-1, ..., 14-M corresponding to the candidate channel j that indicated that the keyword was detected as the output channel. The maximum power detection unit 15 outputs information indicating the selected output channel to the channel selection unit 16.
[0037] By configuring it in this way, according to the second embodiment, it is possible to select a channel containing the utterance of a keyword from multiple channels without causing a decrease in the signal-to-noise ratio by adding the audio signals of each channel of the input audio signal.
[0038] [Third Embodiment] In the first embodiment, it was assumed that the channel containing the pronunciation of the keyword would have the highest power during the keyword utterance interval. However, this assumption is not always met. In the third embodiment, in addition to the assumption that the channel containing the pronunciation of the keyword has high power during the keyword utterance interval, an assumption is made that the speaker does not speak before the keyword utterance. Since the keyword utterance is always considered to be at the beginning of the utterance, it is assumed that there is a period of time or longer without utterance before the keyword utterance (see Figure 7). In the third embodiment, focusing on this point, a weight is assigned to make it easier to detect channels with low power in the period before the keyword utterance, and then the channel with the highest power is detected.
[0039] The channel selection device 3 of the third embodiment, similar to the first embodiment, takes audio signals from multiple channels as input and selects and outputs the audio signal of the channel suitable for the target sound to be targeted for speech recognition, etc., from among the channels in which the pronunciation of a keyword is detected. As shown in Figure 8, the channel selection device 3 further includes M power calculation units 31-1, ..., 31-M, M delay units 32-1, ..., 32-M, M weight calculation units 33-1, ..., 33-M, and a weighted maximum power detection unit 34, in addition to the summing unit 11, keyword detection unit 12, power calculation units 13-1, ..., 13-M, delay units 14-1, ..., 14-M, and channel selection unit 16 of the first embodiment.
[0040] The channel selection method performed by the channel selection device of the third embodiment will be described below, focusing on the differences from the channel selection method of the first embodiment.
[0041] The power calculation unit 31-i (i=1,…,M) calculates the power of channel i of the input audio signal. The power calculation unit 31-i outputs the power of channel i to the delay unit 32-i. The power calculation involves calculating the mean square power multiplied by a rectangular window of length A, which is assumed to be the length of the silent interval that is assumed to exist before the pre-set keyword utterance, or the mean square power multiplied by an exponential window. The detailed procedure for power calculation is the same as in the first embodiment. For example, the assumed length of the silent interval A is set in advance to 1 second.
[0042] The delay unit 32-i (i=1,…,M) delays the power of channel i output by the power calculation unit 31-i. The delay amount is the sum of the detection delay time equivalent D for keyword detection, the average keyword utterance time T, and the margin time B (see Figure 7). The delay unit 32-i outputs the delayed power of channel i to the weight calculation unit 33-i.
[0043] The weight calculation unit 33-i (i=1,…,M) calculates weights from the output of the delay unit 14-i and the output of the delay unit 32-i. The output of the delay unit 14-i and the output of the delay unit 32-i are the average power Pi(t) of the keyword utterance interval shown in Figure 7 and the assumed silence before the keyword utterance, respectively. This is the average power Qi(t) over the given interval. For keyword utterances, it is assumed that the relationship Pi(t) > Qi(t) holds. Therefore, the weights are set so that the value increases as Pi(t) becomes larger than Qi(t). For example, the ratio Zi(t) = Pi(t) / Qi(t) is found, and a monotonically increasing function f is given to this to calculate Wi(t) = f(Pi(t) / Qi(t)), and the weight Wi(t) is calculated. However, the function f is a sigmoid function, for example.
[0044] The weighted maximum power detection unit 34 multiplies the power Pi(t) output by the delay unit 14-i by the weight Wi(t) calculated by the weight calculation unit 33-i for each channel i, and selects the channel with the highest weighted power after multiplication as the output channel.
[0045] Other processing is the same as described in the first embodiment above.
[0046] In the third embodiment, a more accurate determination can be made by determining the channel containing the keyword utterance based on two assumptions: the assumption that the channel containing the pronunciation of the keyword has high power in the keyword utterance interval, and the assumption that the speaker has not uttered any words before the keyword utterance.
[0047] [Fourth Embodiment] The fourth embodiment is a channel selection device of the second embodiment, configured in the same way as the third embodiment, in which a weight is assigned to make it easier to detect channels with low power in the section preceding the keyword utterance, and then the channel with the highest power is detected.
[0048] The channel selection device 4 of the fourth embodiment, similar to the second embodiment, takes three or more audio signals as input and selects and outputs the audio signal of the channel suitable for the target sound to be targeted for speech recognition, etc., from among the channels in which the pronunciation of a keyword is detected. As shown in Figure 9, the channel selection device 4 further includes a weighted candidate selection unit 41 and M delay units 42-1, ..., 42-M, in addition to the keyword detection units 12-1, ..., 12-K, power calculation units 13-1, ..., 13-M, delay units 14-1, ..., 14-M, channel selection unit 16, delay units 21-1, ..., 21-M, and candidate channel selection unit 23 of the second embodiment, and the power calculation units 31-1, ..., 31-M, delay units 32-1, ..., 32-M, weight calculation units 33-1, ..., 33-M, and weighted maximum power detection unit 34 of the third embodiment.
[0049] The channel selection method performed by the channel selection device of the fourth embodiment will be described below, focusing on the differences from the channel selection method of the fourth embodiment.
[0050] The weighted candidate selection unit 41 multiplies the power Pi(t) output by the power calculation unit 13-i by the weight Wi(t) calculated by the weight calculation unit 33-i for each channel i, and selects the K channel with the largest weighted power after multiplication as the candidate channel. The weighted candidate selection unit 41 outputs information indicating the selected candidate channel to the candidate channel selection unit 23.
[0051] The delay unit 42-i (i=1,…,M) adjusts the weight Wi(t) output by the weight calculation unit 33-i over time. A delay of time D is introduced. Time D is set to a time corresponding to the detection delay for keyword detection. The delay unit 42-i outputs the weight Wi(t) after the delay to the weighted maximum power detection unit 34.
[0052] The weighted maximum power detection unit 34 calculates the weighted power for each channel i by multiplying the power Pi(t) output by the delay unit 14-i by the weight Wi(t) output by the delay unit 42-i. When the keyword detection result output by the keyword detection unit 12-j indicates that a keyword has been detected, the weighted maximum power detection unit 34 selects the channel with the highest weighted power among the candidate channels j that indicated that the keyword was detected as the output channel.
[0053] Other processing is the same as described in each of the embodiments above.
[0054] Although embodiments of this invention have been described above, the specific configuration is not limited to these embodiments, and it goes without saying that any modifications to the design, etc., made as appropriate without departing from the spirit of this invention will still be included. The various processes described in the embodiments may be executed not only in chronological order according to the order described, but may also be executed in parallel or individually depending on the processing capacity of the device performing the processes or as necessary.
[0055] [Programs, recording media] When the various processing functions of each device described in the above embodiment are implemented by a computer, the processing content of the functions that each device should have is described by a program. Then, by executing this program on the computer, the various processing functions of each device are implemented on the computer.
[0056] The program describing this process can be recorded on a computer-readable recording medium. Any computer-readable recording medium can be used, such as a magnetic recording device, optical disc, magneto-optical recording medium, or semiconductor memory.
[0057] Furthermore, the distribution of this program may involve, for example, DVDs, CD-ROMs, etc. containing the program. This can be done by selling, transferring, or lending portable recording media. Furthermore, the program may be stored in the storage device of a server computer and distributed by transferring the program from the server computer to other computers via a network.
[0058] A computer executing such a program, for example, first stores the program recorded on a portable storage medium or a program transferred from a server computer in its own memory. Then, when processing is to be executed, the computer reads the program stored in its memory and executes the processing according to the read program. Alternatively, as another form of execution of this program, the computer may directly read the program from the portable storage medium and execute the processing according to that program. Furthermore, each time a program is transferred to this computer from a server computer, it receives it sequentially. It is also possible to execute processing according to the program obtained. Alternatively, a so-called ASP (Application Service Provider) type service can be implemented where the processing function is realized only by issuing execution instructions and obtaining results, without transferring the program from the server computer to this computer. The above processing may be performed by S. In this embodiment, the program includes information used for processing by an electronic computer that is equivalent to a program (data that is not a direct instruction to the computer but has the property of defining the processing of the computer).
[0059] Furthermore, in this configuration, the device is configured by executing a predetermined program on a computer, but at least a part of these processes may be implemented in hardware. [Explanation of Symbols]
[0060] Channel 1, 2, 3, 4 selection device 9 Keyword detection device 11 Addition section 12, 91 Keyword detection unit 13, 31 Power Calculation Unit 14, 21, 32, 42 Delay section 15. Maximum power detection unit 16 Channel Selection Section 22 Candidate Selection Section 23 Candidate channel selection unit 33 Weight Calculation Unit 34 Weighted Maximum Power Detection Unit 41 Weighted Candidate Selection Unit 99 Target sound output section
Claims
1. The steps include acquiring channels of multiple audio signals picked up by three or more microphones, The steps include acquiring the power of each of the aforementioned multiple channels, The steps include selecting the audio signal of the channel with the highest power among the channels with the highest power among the multiple audio signal channels, from among the channels in which the keyword was detected, as the target sound for speech recognition, and A selection method that includes this.
2. The channels with high power in the aforementioned audio signal are channels selected from among the multiple audio signal channels, with a predetermined number of channels having the highest power. The selection method according to claim 1.
3. Furthermore, the selection is made taking into consideration that the speaker had not uttered any words prior to the utterance of the keyword. The selection method according to claim 2.
4. In the above selection, weights are assigned to the interval preceding the utterance of the keyword in the audio signal channel in which the keyword is detected among the multiple audio signal channels. The selection method according to claim 3.
5. The keyword detection is performed only for the audio signal of each of the channels with high power. The detection of the aforementioned keywords is performed for each channel with high power. The selection method according to claims 2 to 4.
6. A means for acquiring channels of multiple audio signals picked up by three or more microphones, Means for acquiring the power of each of the aforementioned multiple channels, A means for selecting the audio signal of the channel with the highest power among the channels with the highest power among the multiple audio signal channels, from among the channels in which a keyword is detected, as the target sound for speech recognition. A selection device including a selection device.
7. A program for causing a computer to perform each step of the selection method described in claim 1.