Howling Detection Method, Howling Suppression Device

The method addresses howling detection challenges by using multi-element determination and enhanced frequency analysis to accurately identify howling in voice signals, ensuring stable detection and minimal sound quality impact.

JP7713230B2Active Publication Date: 2025-07-25INITIATEC CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2021188273
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-11-19
Publication Date
2025-07-25
Estimated Expiration
2041-11-19

AI Technical Summary

Technical Problem

Existing howling detection methods face issues with excessive detection or overlooking howling, particularly in voice signals like singing, and struggle with accurate frequency analysis and multi-frequency howling, leading to sound quality degradation and potential hearing impairment.

Method used

A howling detection method using multi-element determination based on peak level, level ratios, temporal variations, and frequency analysis with enhanced frequency resolution and error correction, employing FFT and band aggregation to reduce calculation and improve accuracy.

Benefits of technology

The method provides stable and accurate howling detection with reduced over-detection, low power consumption, and minimal sound quality degradation, enabling effective howling suppression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007713230000009
    Figure 0007713230000009
  • Figure 0007713230000010
    Figure 0007713230000010
  • Figure 0007713230000011
    Figure 0007713230000011
Patent Text Reader

Abstract

To provide a howling suppression device for stably and accurately detecting howling without excessively detecting a voice signal such as singing and without overlooking actual howling in detecting howling that occurs in a public address system.SOLUTION: In a howling suppression device, an FFT 211 performs frequency analysis of a voice signal from a microphone, a bandwidth aggregation part 212 aggregates bandwidth information acquired by the analysis, a peak search part 213 sequentially searches peaks, and a multi-element determination part 220 determines howling from determination elements acquired from a peak level part 215, a peak-to-adjacent level comparison part 216, a peak-to-neighboring and whole level comparison part 217, a peak level time fluctuation part 218 and a peak frequency time fluctuation part 219. The peak frequency approximation and error correction part 214 increases a Q value of a howling removing dip filter, and increases the accuracy of a peak frequency from information of peak and their adjacent bands in order to reduce influence to be exerted on sound quality.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a detection method for howling caused by acoustic coupling between a speaker and a microphone, a howling suppression device, and a program in a voice amplification device where the speaker and the microphone are arranged in the same acoustic space.

Background Art

[0002] In the howling detection technology using digital signal processing technology, there has been conventionally known a method of detecting howling by performing frequency analysis on an audio signal from a microphone and comparing the signal level (power spectrum value) of the peak obtained from the analysis result with a predetermined threshold value. (For example, Patent Document 1)

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Patent Document 2

Patent Document 3

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the case of a method of determining howling based on whether the signal level of the peak obtained by frequency analysis is equal to or higher than a threshold value as in Patent Document 1, there is a possibility of excessive detection that howling is occurring even though howling is not occurring, such as during singing. On the contrary, if the threshold value is increased to reduce such excessive detection, the risk of overlooking howling increases even when howling actually occurs.

[0005] Howling not only gives discomfort to listeners, but also risks causing hearing impairment, so it is necessary to avoid overlooking howling. However, considering that the number of howling removal dip filters implemented in the howling suppression device is finite, excessive detection should be appropriately restricted. Also, unnecessarily activating the howling removal dip filter is not preferable from the perspective of sound quality degradation.

[0006] In order to suppress such false detections, instead of making a determination based on a single element such as the peak signal level, a multi-element determination combining a plurality of related elements including the howling frequency and its temporal variation should be made. In voice signals such as singing, the challenge is to detect howling more stably and accurately without excessive detection and without overlooking actual howling.

[0007] Also, as a howling suppression device, it is necessary to increase the Q value of the dip filter that performs howling removal in order to minimize the impact on sound quality. For that purpose, it is necessary to accurately grasp the howling frequency.

[0008] In the frequency analysis means of voice signals, several methods have been devised in the past, and they are mainly classified into a method using an adaptive filter and a method using a fast Fourier transform (FFT). The method using the LMS filter (Least Mean Squares Filter), which is a typical method of adaptive filters, generally has good accuracy in detecting frequencies, but has the problem that it may take time to converge. (For example, see Patent Document 2) On the other hand, in the method using the fast Fourier transform (FFT), although information in the frequency domain can be obtained in a certain period of time, it is necessary to increase the number of voice samples in order to improve the frequency resolution. In that case, it takes time to accumulate voice samples, which may increase the reaction time until howling detection and lead to an increase in the amount of calculation. (For example, see Patent Document 3) In the frequency analysis means, issues include the stability of analysis, shortening of response time, and reduction of calculation time.

[0009] Furthermore, considering that howling occurs under the conditions of acoustic coupling between the speaker and the microphone, multiple frequencies of howling may occur simultaneously, and dealing with such simultaneous multi-frequency howling is also an issue.

Means for Solving the Problem

[0010] To solve the above problems, the howling detection unit in the present invention includes a frequency analysis means for the audio signal from the microphone, a means for sequentially searching for peaks in the frequency domain from the frequency spectrum information obtained by the frequency analysis means, and means for determining howling based on the following elements obtained based on the searched peaks · Peak level, · Level ratio between the peak and the adjacent band, · Level ratio between the peak and the nearby band or the entire band, · Peak level time variation, · Peak frequency, and · Peak frequency time variation 、 It is characterized by comprising means for determining howling with these elements. (Claim 1)

[0011] said Among the frequency spectrum information of the frequency analysis means, information in a band with a frequency resolution higher than that required for howling detection is aggregated by this , and it is characterized by comprising band aggregation means for suppressing the calculation amount of howling detection. ( Claim 1 )

[0012] said The frequency analysis means uses the fast Fourier transform (FFT) yes the peak and the peak with the adjacent low frequency and high frequencyEnhance the accuracy of the peak frequency through polynomial approximation and error correction based on the information of three points comprising peak frequency approximation and error correction means characterized by doing so.( Claim 1 )

Advantages of the Invention

[0013] The present invention can provide a howling detection method with a low risk of over-detection for voice signals such as singing, while suppressing the amount of calculation, and can realize a howling suppression device with low power consumption, small size, and low cost.

Brief Description of the Drawings

[0014]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Embodiments for Carrying Out the Invention

[0015] Hereinafter, embodiments of the present invention will be described with reference to the drawings.

Example

[0016] FIG. 1 is a diagram showing an outline when one embodiment of a howling suppression device using a howling detection method according to the present invention is applied to a loudspeaker. The sound input to the microphone 101 arranged in the acoustic space 100 is converted into an electrical signal and input to the howling suppression device 200. In the howling suppression device 200, a detected howling and a suppressed voice signal are output, power-amplified by the power amplifier 104, and amplified by the speaker 103 arranged in the acoustic space 100. The sound amplified by the speaker 103 is input to the microphone 101 arranged in the same acoustic space 100 via the acoustic coupling 102. In this way, when the voice signal satisfies the persistence condition when circulating in the same system, it is a howling phenomenon.

[0017] In the howling suppression device 200, the voice signal input to the input terminal 201 is appropriately amplified by the pre-stage amplifier 202 and converted into a digital signal by an analog-to-digital converter (ADC) 203, for example, at a sampling frequency of 32 kHz. The digital signal is supplied to the howling removal dip filter 204 and the howling detection unit 210.

[0018] The howling removal dip filter 204 activates the filter with the information calculating the howling frequency when the howling detection unit 210 detects howling, and acts to remove howling. Note that the howling removal dip filter 204 can be configured by a plurality (for example, 16) of band removal filters as needed.

[0019] The digital signal that has passed through the howling removal dip filter 204 is converted into an analog audio signal by a digital-to-analog converter (DAC) 205, amplified by a subsequent-stage amplifier 206, and sent to a power amplifier 104 via an output terminal 207.

[0020] On the other hand, in the howling detection unit 210, the digital signal from the analog-to-digital converter (ADC) 203 is converted from the time domain to frequency domain data by a fast Fourier transform (FFT) 211 using, for example, 1024 audio samples.

[0021] The bandwidth of the frequency domain data from the fast Fourier transform (FFT) 211 is, when the sampling frequency is f s , and the number of samples in the time domain is N, Bandwidth = Sampling frequency (f s ) ÷ Number of samples (N) as shown. Also, the number of data in the frequency domain at this time, that is, the number of bands, is 1 / 2 of the number of samples (N). For example, when the sampling frequency is 32 kHz and the number of samples is 1024, the bandwidth is 31.25 Hz, and the number of data in the frequency domain, that is, the number of bands, is 512.

[0022] Since this bandwidth is the same for all low and high frequencies, when used for howling detection, it has excessive frequency resolution in the high frequency range. Therefore, the band aggregation unit 212 aggregates it to an appropriate resolution for howling detection. What is shown by the solid line (a) in the graph of FIG. 2 is the frequency resolution of each frequency band of the fast Fourier transform (FFT) 211. For example, when the sampling frequency is 32 kHz and the number of samples is 1024, it is 53.3 cents in the 1 kHz band and about 5.4 cents in the 10 kHz band in the higher frequency range. At high frequencies like this, since the frequency resolution is excessive compared to the mid - frequencies, as the frequency increases to high frequencies, by aggregating several bands, the frequency resolution is made to fall within a certain range, which is the function of the band aggregation unit 212 in FIG. 1. The example of the frequency resolution after aggregation is shown by the broken line (b) in the graph of FIG. 2. When the sampling frequency is 32 kHz and the number of data in the frequency domain after fast Fourier transform is 512, for 1 kHz and above, 2 bands are aggregated, for 2 kHz and above, 4 bands are aggregated, for 4 kHz and above, 8 bands are aggregated, and for 8 kHz and above, 16 bands are aggregated. By doing so, the frequency resolution is kept within the range of about 50 cents to 100 cents, and the number of data in the frequency domain is reduced to 96, which is 1 / 5 or less.( Claim 1 )

[0023] The data in the frequency domain aggregated by the band aggregation unit 212 in FIG. 1 is supplied to the peak search unit 213. In this peak search unit 213, the data in the frequency domain is sequentially compared to search for in which frequency band the peak is located. For the data in the frequency domain, for example, the difference is calculated sequentially from the low - frequency to the high - frequency direction, and the part where the sign changes from plus to minus is extracted as the peak. The search for the peak may also be performed from the high - frequency to the low - frequency direction.

[0024] The peak data extracted by the peak search unit 213 is supplied to the peak level unit 215, the peak - to - adjacent - level comparison unit 216, the peak - to - neighboring·whole - band - level comparison unit 217, the peak - level time - variation unit 218, and further to the peak - frequency approximation·error - correction unit 214 in FIG. 1. In the peak level unit 215, the level of the peak is extracted and supplied to the multi - element determination unit 220 in FIG. 1. In the peak - to - adjacent - level comparison unit 216, the ratios of the levels of the low - frequency side and the high - frequency side adjacent to the peak to the peak level are calculated and also supplied to the multi - element determination unit 220. The purpose of this peak - to - adjacent - level comparison unit 216 is to evaluate that the peak clearly protrudes with respect to the adjacent bands. In the peak-to-nearby / full-band level comparison unit 217, similar to the previous peak-to-adjacent level comparison unit 216, the ratio of the level in the nearby band or the full band to the peak level is calculated and is also supplied to the multi-element determination unit 220. The purpose of this peak-to-nearby / full-band level comparison unit 217 is to separate howling and singing by evaluating the level ratio between the peak and the levels of other voice signals. In the peak level time variation unit 218, the instantaneous value of the level of the previous peak level unit 215 is memorized and calculated as the peak level time variation value, and is also supplied to the multi-element determination unit 220. In the peak frequency time variation unit 219, the peak frequency calculated by the peak frequency approximation / error correction unit 214 described later is memorized and calculated as the peak frequency time variation value, and is also supplied to the multi-element determination unit 220.

[0025] The multi-element determination unit 220 determines howling based on the information of the supplied multi-elements. (Claim 1) The conditions for determining howling are as follows. (Condition 1) The peak level is equal to or higher than a preset threshold value. (Condition 2) The level ratio between the peak and the adjacent band is equal to or higher than a preset threshold value. (Condition 3) The level ratio between the peak and the nearby band or the full band is equal to or higher than a preset threshold value. (Condition 4) The peak level time variation is within a preset threshold value, or has an increasing tendency over time. (Condition 5) The peak frequency is within the voice band that can be handled. (Condition 6) The peak frequency time variation is within a preset threshold value.

[0026] Figure 3 is a reference diagram showing the levels of peak point A, peak point B, and peak point C and the situation of the peak-to-adjacent level ratio. The peak B point indicates a peak that satisfies the following two conditions: (Condition 1) the peak level is equal to or higher than a preset threshold, and (Condition 2) the level ratio between the peak and the adjacent band is equal to or higher than a preset threshold. The peak A point indicates a peak that satisfies (Condition 2), i.e., the level ratio between the peak and the adjacent band is equal to or higher than a preset threshold, but does not satisfy (Condition 1), i.e., the peak level is not equal to or higher than a preset threshold. The peak C point indicates a peak that satisfies (Condition 1), i.e., the peak level is equal to or higher than a preset threshold, but does not satisfy (Condition 2), i.e., the level ratio between the peak and the adjacent band is not equal to or higher than a preset threshold. By observing actual howling, it has been confirmed that the howling part shows a sharp peak. Therefore, the following two conditions, (Condition 1) the peak level is equal to or higher than a preset threshold, and (Condition 2) the level ratio between the peak and the adjacent band is equal to or higher than a preset threshold, are also candidates for the howling determination conditions. In the case of the reference example in Figure 3, the peak A point and the peak C point, which are ultimately searched as peaks, do not meet the howling determination. Note that the peak level threshold may vary for each frequency as shown in Figure 3. Furthermore, a hysteresis may be provided by setting a difference between the threshold when changing from a state where the condition is not satisfied to a state where the condition is satisfied and the threshold when changing from a state where the condition is satisfied to a state where the condition is not satisfied.

[0027] Figure 4 shows the occurrence of howling during speech, and Figure 5 shows an example of singing that may be over-detected as howling. Figure 4 and Figure 5(a) are graphs showing examples of (Condition 4) the peak level time variation and (Condition 6) the peak frequency time variation. Also, Figure 4 and Figure 5(b) are spectrograms synchronized with the graphs in (a).

[0028] Figure 4(b) is a spectrogram when howling occurs during speech, and Figure 5(b) is an example of a spectrogram seen in singing that may be over-detected as howling. In Figure 5(b), compared to Figure 4(b), there are many white or gray areas indicating the presence of signals in regions other than the peaks. Utilizing this feature, an attempt is made to separate howling from signals that are not howling. (Condition 3) states that the level ratio of the peak and the neighboring band or the entire band being equal to or greater than a preset threshold is used as the determination condition for howling.

[0029] In both Figure 4(a) and Figure 5(a), (Condition 6) the peak frequency time variation is shown by the solid line (i), and (Condition 4) the peak level time variation is shown by the dashed line (ii). The (i) in Figure 4(a) shows the peak frequency time variation when howling occurs during speech, that is, the frequency time variation of howling, which is generally within 100 ppm. In contrast, the (i) in Figure 5(a) shows the peak frequency time variation in singing, and in this example of singing, a frequency variation of 600 ppm or more is observed. Regarding the peak level time variation, in this example, no extreme difference is seen between the (ii) of the howling during speech in Figure 4(a) and the (ii) of the singing in Figure 4(b). Rather, the variation around the howling convergence in Figure 4 is prominent, and this is because the peak level variation is significantly caused by howling convergence at this level of time variation.

[0030] (Condition 5) Whether the peak frequency is within the frequency range that the howling suppression device can handle is the determination condition.

[0031] When the multi-element determination unit 220 determines howling using the supplied multi-element information and detects the occurrence of howling, it supplies information including the peak frequency calculated by the peak frequency approximation and error correction unit 214, which will be described later, to the howling removal dip filter 204 to suppress howling.

[0032] Now, when using the fast Fourier transform (FFT) in the frequency analysis means, as pointed out in Patent Document 3, there are problems with the accuracy of the howling frequency and the amount of calculation and calculation time. In order to improve the frequency accuracy with the fast Fourier transform (FFT), a method of increasing the number of samples is common. However, in that case, the amount of calculation of the fast Fourier transform (FFT) increases, and it is necessary to accumulate more samples of the audio signal before starting the frequency analysis by the fast Fourier transform (FFT), resulting in a long delay time. The problem of calculation time due to the amount of calculation is expected to be solved eventually because the calculation device is getting faster year by year with the development of semiconductor technology. On the other hand, the problem of delay time due to the accumulation of audio samples is determined by the sampling theorem, so the delay time increases proportionally as the number of samples increases. This is independent of the development of semiconductor technology and is constant. For example, when the sampling frequency is 32 kHz and the number of samples is 1024, the frequency resolution (bandwidth) is 31.25 Hz and the delay time is 32 ms. However, to increase the frequency resolution (bandwidth) by 16 times, it is necessary to increase the number of samples to 16,384, which is 16 times. At that time, the frequency resolution (bandwidth) is about 1.95 Hz and the delay time is 512 ms. In this way, increasing the frequency resolution only with the fast Fourier transform (FFT) increases the delay time, and as a result, the reaction time of howling detection increases, which is not a favorable state for a howling suppression device.

[0033] The peak frequency approximation and error correction unit 214 in FIG. 1 performs peak approximation from the frequency spectrum information using the fast Fourier transform (FFT) by passing through three points of the peak and its adjacent low frequency range, and further performs error correction to calculate a peak frequency with high accuracy. high frequency When the frequency is f and the level is L, the quadratic function can be expressed by Equation (1).

[0034] When the frequency is f and the level is L, the quadratic function can be represented by Equation (1).

Equation

Equation

Equation

Equation

[0035] Figure 6 is a graph showing an example of quadratic approximation when analyzing a 400 Hz audio signal with a sampling frequency of 32 kHz and 1024 samples of the fast Fourier transform (FFT). In this example, the peak frequency 406.25 Hz (f2) and level value 0.0306 (L2) are obtained through the frequency spectrum analysis of the fast Fourier transform (FFT). The frequency 375 Hz (f1) and level value 0.0186 (L1) adjacent to the lower frequency of the peak, and the frequency 437.5 Hz (f3) and level value 0.0088 (L3) adjacent to the higher frequency of the peak are obtained. From Equation 2, Equation 3, and Equation 4, the quadratic approximation frequency f a of the peak frequency is 401.73 Hz. This means an approximation error of 0.43%, which can be said to be a relatively good result. However, if the Q value of the howling removal dip filter 204 is increased to minimize the impact on sound quality, this approximation error of 0.43% is not sufficient, and it is desirable to further reduce the error.

[0036] The results of numerical simulation using a computer for the error of the quadratic approximation frequency in the previous section are shown by the solid line in FIG. 7. As a result, the error of the quadratic approximation frequency oscillates periodically according to the quadratic approximation frequency fa obtained in Equation 4, the sampling frequency fs, and the value calculated by the number of samples N of the fast Fourier transform (FFT). Its peak value is as shown by the broken line in FIG. 7, and it can be seen that it is inversely proportional to the quadratic approximation frequency. fa It can be seen that it is inversely proportional. When attempting error correction from this numerical simulation using a computer, if the error correction frequency is fb, it is as shown in Equation 5.

Equation

[0037] As described above, it is possible to improve the accuracy of the peak frequency by performing quadratic approximation and error correction. However, if there is no need to improve the accuracy, only quadratic approximation may be performed without error correction.

[0038] As described above, the embodiments of the present invention have been described in detail with reference to the drawings. However, the specific configuration is not limited to this embodiment, and design changes and the like within the scope not departing from the gist of the present invention are also included.

[0039] In addition, various processing functions in the above embodiments may be realized by a computer including a DSP (Digital Signal Processor). In that case, the processing contents of each function are described by a program, and by executing this program on a computer, each processing function is realized on the computer. Furthermore, the processing contents of each function may be described in HDL (Hardware Description Language) and realized by a semiconductor typified by an FPGA (Field Programmable Gate Array).

Industrial Applicability

[0040] It can be applied to the detection of howling caused by acoustic coupling between a speaker and a microphone, and to the use of howling suppression.

Explanation of Reference Numerals

[0041] 100: Acoustic space 101: Microphone 102: Acoustic coupling 103: Speaker 200: Howling suppression device 201: Input terminal 202: Pre-stage amplifier 203: Analog-to-digital converter (ADC) 204: Howling removal dip filter 205: Digital-to-analog converter (DAC) 206: Post-stage amplifier 207: Output terminal 210: Howling detection unit 211: Fast Fourier transform (FFT) 212: Band aggregation unit 213: Peak search section 214: Peak frequency approximation and error correction section 215: Peak level section 216: Peak to adjacent level comparison section 217: Peak to near - by and full - band level comparison section 218: Peak level time variation section 219: Peak frequency time variation section 220: Multi - element determination section

Claims

1. Frequency analysis means for an audio signal from a microphone, means for sequentially searching for peaks in the frequency domain from the frequency spectrum information obtained by the frequency analysis means, obtained based on the searched peaks - peak level, - level ratio between the peak and the adjacent band, - level ratio between the peak and the neighboring band or the entire band, - peak level time variation, - peak frequency, and - peak frequency time variation, comprising means for determining howling using these elements, characterized by comprising band aggregation means for suppressing the calculation amount of howling detection by aggregating information in a band having a frequency resolution equal to or higher than that required for howling detection among the frequency spectrum information of the frequency analysis means, using fast Fourier transform (FFT) for the frequency analysis means, characterized by comprising peak frequency approximation and error correction means for improving the accuracy of the peak frequency by polynomial approximation and error correction from the information of three points of the peak, the low frequency and the high frequency adjacent to the peak, a howling suppression device.

2. A program for functioning as the howling suppression device according to Claim 1.

Citation Information

Patent Citations

  • Peak frequency tracking device

    JP1992265998A

  • Howling suppressing device

    JP1994164278A

  • Howling detecting and preventing circuit and sound reinforcing device using the same

    JP1998145888A

  • Howling suppression device utilizing adaptive notch filter

    JP2001285986A

  • Voice quality converting device and program

    JP2006330343A