Voice detection device and voice detection method
The voice detection device uses frequency band division and a simple algorithm to distinguish voice from noise, addressing the need for cost-effective and efficient voice detection.
Patent Information
- Application Number
- JP2024100632
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-06-21
- Publication Date
- 2025-07-31
- Estimated Expiration
- 2044-06-21
AI Technical Summary
Existing audio detection devices require large-capacity memory and high computing power, making them expensive, and existing voice detection methods are prone to misclassifying noise as voice due to small level variations.
A voice detection device that divides input acoustic signals into frequency bands, determines the power value of the lowest frequency band, and uses a simple algorithm to distinguish voice from noise using a low-pass and high-pass filter combination.
Enables voice detection with minimal memory and CPU resources, effectively distinguishing voice from noise, reducing data communication and battery consumption, and extending recording time.
Smart Images

Figure 0007716059000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an audio detection device and an audio detection method.
Background Art
[0002] There is a device that compresses audio data for audio communication. When the communication speed is slow, real-time communication is realized by increasing the compression ratio. Wireless devices tend to have a slow communication speed and a small number of data that can be communicated in real time. Therefore, there is a device that determines voice to reduce the amount of data and effectively uses the limited communication speed by reducing the data to be transmitted during non-voice periods.
[0003] If voice detection is performed only based on the power value of the input to the communication device, it will be determined as voice even for noise. Therefore, an audio detection device can be used to distinguish between audio and noise.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0005] The audio detection device described in Patent Document 1 above converts an input acoustic signal into the frequency domain by FFT (Fast Fourier Transform), calculates a power spectrum, extracts a fundamental frequency from this spectrum, and uses it for audio detection determination. In the audio detection device described in this Patent Document 1, like the technology used for audio signal processing, information processing for performing complex audio processing is required, and a large-capacity memory and a CPU with a high computing function are required, so the device becomes expensive.
[0006] The voice detection device described in Patent Document 2 above uses the histogram of the signal level of the input acoustic signal to determine the level variation for voice detection determination. In the voice detection device described in this Patent Document 2, there is a possibility of determining it as voice even when the level variation is small even for noise.
[0007] The present invention solves the above-described problems of the prior art, and an object thereof is to provide a voice detection device and method capable of determining voice with inexpensive resources such as a small-capacity memory and an inexpensive CPU and a simple algorithm.
Means for Solving the Problems
[0008] A first aspect of the present invention is a voice detection device that determines voice included in an input acoustic signal, the device including: a band division unit that divides the input acoustic signal into a plurality of frequency bands; a power determination unit that determines that the power value of the lowest frequency band among the divided bands is the maximum and exceeds a predetermined threshold value; and a voice determination unit that determines whether it is voice based on the output of the power determination unit.
[0009] Note that the band division unit preferably divides by combining two divisions using a low-pass filter and a high-pass filter.
[0010] Another aspect of the present invention is a voice detection method for determining voice included in an input acoustic signal, the method including: a band division step of dividing the input acoustic signal into a plurality of frequency bands; a power determination step of determining that the power value of the lowest frequency band among the divided frequencies is the maximum and exceeds a predetermined threshold value; and a voice determination step of determining whether it is voice based on the output of the power determination unit.
Effects of the Invention
[0011] By taking advantage of the fact that voice has high power in the low frequency band and noise has its power dispersed across all frequency bands, and determining whether there is voice based on whether the power in the low frequency band is greater than that in other frequency bands, voice detection can be achieved with a small amount of memory, a low-cost CPU, and a simple algorithm, and can be realized with an inexpensive device.
Brief Description of the Drawings
[0012]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Embodiments for Carrying Out the Invention
[0013] Embodiments of the present invention will be described below with reference to the drawings. FIG. 1 is a diagram showing the configuration of a voice detection device 10 according to an embodiment of the present invention. The voice detection device 10 of the present embodiment includes an A / D conversion unit 11 that converts an acoustic signal from an input environment into a digital signal, a band division unit 12 that divides the frequency band of the acoustic signal converted into a digital signal, a power determination unit 13 that determines that the power value of the lowest band among the divided bands is the largest and exceeds a predetermined threshold, and a voice determination unit 14 that determines whether it is voice based on the output of the power determination unit.
[0014] FIG. 2 is a diagram for explaining the band division unit 12 of the embodiment. The band division unit 12 has a function of dividing an acoustic signal converted into a digital signal by the A / D conversion unit 11 into a plurality of frequency bands in a combination of two-band division so that the lowest frequency band becomes a predetermined band. The predetermined band is a band in which it is assumed that the voice power value is concentrated. For example, assuming that the voice power value is from 0 to 1000 Hz and the sampling frequency of the A / D conversion unit 11 is 8000 Hz, when three combinations of two-band division are made, Band1 is from 0 to 1000 Hz (however, the lowest frequency can also be the audible band), Band2 is from 1000 to 2000 Hz, Band3 is from 2000 to 3000 Hz, and Bnad4 is divided into four frequency bands from 3000 to 4000 Hz. The divided acoustic signal is output to the subsequent power determination unit 13.
[0015] FIG. 3 is a diagram for explaining the two-band division unit 121. The two-band division unit 121 has a function of dividing the frequency band into two by dividing the acoustic signal by a low-pass filter and a high-pass filter.
[0016] FIG. 4 is a diagram for explaining the power determination unit 13 of the embodiment. It has a power comparison unit 131 and a maximum value selection unit 132. From the power values of the frequency bands divided by the band division unit 12, if the power value of the lowest frequency band, in this case the frequency band from 0 Hz to 1000 Hz, exceeds the power values of the other frequency bands and exceeds a predetermined threshold A, for example, in the case of ordinary speech, when the maximum sound pressure is 0 dB and the sound pressure exceeds -40 dB, it outputs "1", and if the power value of the lowest band is lower than the power values of the other frequency bands or lower than the predetermined threshold A, it outputs "0". The power comparison unit 131 outputs "1" if the input power value exceeds the reference power value, and "0" if it is lower. The maximum value selection unit 132 outputs the maximum power value from the power values of the frequency bands divided by the band division unit 12 excluding the lowest frequency band. The power comparison unit 131 can be realized by using the comparison function of the CPU that compares the input power value with the reference power value. The maximum value selection unit 132 can be realized by using the comparison function of the CPU that compares a plurality of input power values.
[0017] Referring to FIG. 5, the voice determination principle will be described. FIG. 5(a) shows the frequency characteristics of voice, and (b) shows the frequency characteristics of noise etc., with the horizontal axis representing frequency and the vertical axis representing power. Voice has power concentrated below 1000 Hz, while sounds such as noise have power dispersed across the entire band, so it can be distinguished from voice.
[0018] FIG. 6 is a diagram for explaining the voice determination unit 14 of the embodiment. It includes 16 shift registers and a count comparison unit 141. The output of the power determination unit 13 is stored in these shift registers, and it has the function of outputting voice detection when the sum of the values in the shift registers exceeds a predetermined threshold B. The count comparison unit 141 outputs "1" when the input value exceeds the reference value, and "0" when it is below the reference value.
[0019] For example, the output of the power determination unit 13 is sampled every 10 milliseconds and stored in the shift registers. Voice is determined based on the sum of "1"s in 16, that is, 10 milliseconds × 16, 160 milliseconds.
[0020] FIG. 7 is a flowchart showing the above-described voice detection operation. If the power value in the lowest frequency band exceeds the power values in other frequency bands every 10 milliseconds, "1" is put into the shift register, and if not, "0" is put in (step S1).
[0021] During a 160 - millisecond period, it is determined by the sum of "1"s in the shift register (step S2). This determination is made by the sum of "1"s in the 16 shift registers. In step S2, if the sum of "1"s is 8 or more, it is determined that it is voice, and voice detection is output (step S3).
[0022] (Description of the Effect) The above-described embodiments are understood as follows. An audio detection device 10 that determines audio included in an input acoustic signal, including a band division unit 12 that divides the input acoustic signal into a plurality of frequency bands, a power determination unit 13 that determines that the power value of the lowest frequency band among the divided bands is the maximum and exceeds a predetermined threshold, and an audio determination unit 14 that determines whether it is audio based on the output of the power determination unit.
[0023] With this configuration, the determination of whether it is audio can be realized by the operation of a shift register and counting the data in the register, so the determination operation can be performed with a small memory amount and a small calculation amount. As a result, an audio detection device can be realized with inexpensive hardware.
[0024] Also, the band division unit 12 can be realized by Fourier transform using FFT, but since the number of bands for dividing the frequency is small, it can be realized by a combination of a low-pass filter and a high-pass filter, so there is an advantage that a high computing function is not required. Also, there is an advantage that band filters do not have to be used. The process of dividing the band into two can make the filter coefficients odd when using a FIR (Finite Impulse Response) filter, and can reduce the amount of arithmetic processing. Also, there is an advantage in the frequency separation characteristics.
[0025] Furthermore, in the embodiments of the present invention, since the determination is made by the fact that the power value of the lowest band is the maximum and exceeds a predetermined threshold, audio detection is possible with a simple algorithm and few resources.
[0026] In the present invention, it can be determined that the input acoustic signal is not audio with respect to noise.
[0027] In the present invention, voice can be determined. Therefore, by transmitting only voice with a handy type of wireless device, it is possible to suppress the data communication volume and battery consumption. Further, by combining a voice recording device with the voice detection device, only voice can be stored, and the recording time can be extended while suppressing the memory capacity.
[0028] Note that the present invention is suitable for use in a voice device that discriminates voice from noise and performs gain control in an environment where the sound source is not diverse and there is meaningless noise, such as a voice device for communication between the inside and outside of an elevator or a voice device for communication inside a house or an office. Even if a low-frequency sound source or a notification sound, etc. is determined as voice, the effect of the present invention for distinguishing and discriminating between noise and voice is not reduced.
Explanation of Reference Numerals
[0029] 10 Voice detection device 11 A / D conversion unit 12 Band division unit 121 Band 2 division unit 13 Power determination unit 131 Power comparison unit 132 Maximum value selection unit 14 Voice determination unit 141 Number comparison unit
Claims
1. An audio detection device for determining speech included in an input acoustic signal, comprising: a band division unit that divides the input acoustic signal into a predetermined frequency band; a power determination unit that determines that the power value of the lowest-frequency band among the divided frequency bands is the largest and exceeds a predetermined threshold; an audio determination unit that determines whether it is audio based on the output of the power determination unit; and having the audio determination unit has a count determination unit that determines that the output of the power determination unit has appeared at a predetermined frequency within a first time. An audio detection device characterized by the above.
2. The audio detection device according to claim 1, wherein the band division unit divides the frequency band using a low-pass filter and a high-pass filter. An audio detection device characterized by the above.
3. The audio detection device according to claim 1 or 2, wherein the power determination unit includes determination means that outputs "voice power has appeared" when the power value of the lowest-frequency band among the frequency bands divided by the band division unit is the largest and exceeds a predetermined threshold; the audio determination unit stores the output of the power determination unit and outputs an audio detection output when exceeding a predetermined frequency within the first time. An audio detection device characterized by the above.
4. An audio detection method for determining speech included in an input acoustic signal, comprising: a band division step of dividing the input acoustic signal into a predetermined frequency band; a power determination step of determining that the power value of the lowest-frequency band among the divided frequency bands is the largest and exceeds a predetermined threshold; an audio determination step of determining whether it is audio based on the output of the power determination step; and having the audio determination step has a count determination step of determining that the output of the power determination step has appeared at a predetermined frequency within a first time. and having An audio detection method characterized by the above.
5. The audio detection method according to claim 4, wherein the band division step divides the frequency band using a low-pass filter and a high-pass filter. An audio detection method characterized by the above.
6. The audio detection method according to claim 4 or 5, wherein the power determination step includes determination means that outputs "voice power has appeared" when the power value of the lowest-frequency band among the frequency bands divided in the band division step is the largest and exceeds a predetermined threshold. The voice determination step stores the output of the power determination step and outputs voice detection when the frequency exceeds a predetermined frequency within the first time period. A voice detection method characterized by the above.
Citation Information
Patent Citations
Sound processing device and program
JP2009175474A
Semiconductor device, system, electronic apparatus, and voice recognition method
JP2017068153A
Speech detection apparatus and speech detection program
JP2018180482A
Notification sound detection device and notification sound detection method
JP2021002013A
Voice detection device, voice detection program, and voice detection method
JP2022032721A