Howling suppression method and device, audio processing system, computer equipment and medium
By fusion-based time-frequency feature detection and dynamic adjustment of filter parameters, the high missed detection rate and high false positive rate problems of howling detection and suppression solutions in complex environments in existing technologies are solved, and efficient howling suppression is achieved on embedded devices, improving audio quality and user experience.
Patent Information
- Application Number
- CN202511027910.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-24
- Publication Date
- 2025-10-10
AI Technical Summary
Existing howling detection and suppression solutions have high missed detection rates and high false positive rates in complex environments with multiple howling frequencies, cannot effectively handle howling group bands, have high computational costs, poor real-time performance, and are difficult to deploy on embedded devices.
By obtaining the time domain level value and frequency morphology information of multiple consecutive frames of audio signals, combining filters to suppress howling frequencies, using time-frequency joint feature fusion to detect howling frequencies, and dynamically adjusting the filter parameters to suppress howling sounds.
Quickly locate and suppress howling frequencies in complex environments, reduce the impact on sound quality, improve user experience, reduce false positives, increase detection accuracy, and maximize sound quality.
Smart Images

Figure CN120766698A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of audio processing technology, and in particular to a howling suppression method, apparatus, audio processing system, computer equipment, and computer-readable storage medium. Background Art
[0002] In scenarios such as video and conference systems, multiple speakers, microphones, and mixers are usually used. The audio signal is transmitted in a loop between these devices, which can easily cause acoustic feedback problems and lead to howling. Howling not only seriously affects the system's audio quality and user experience, but may also interfere with the normal progress of meetings or entertainment, and even damage the speaker equipment. Therefore, these howling needs to be detected and eliminated.
[0003] Existing howling detection and suppression schemes, such as those that use a linear combination of multiple howling features to determine howling, can detect a single howling frequency to a certain extent, but their miss detection and false positive rates are high. The miss detection rate rises sharply when multiple howling frequencies are present simultaneously. Some technical solutions determine the waveform symmetry of each speech signal segment based on the correlation between all peaks and all valleys and the changing trend of all peaks in the time domain of each speech signal segment. They also determine the peak significance value of each speech signal segment by the distribution of all peaks in the spectrum and the difference between the maximum peak and all other peaks. This solution can identify howling frequencies to a certain extent, but it is difficult to find the specific howling frequency when there is loud background sound and many howling frequencies. Furthermore, the howling sound needs to reach a certain loudness to reach the peak difference value, and the detection time is also long. Another technical solution decomposes multiple signals into multiple channels, obtains the maximum energy channel, and performs an FFT on it. If this channel is a howling channel, the gain of the channel is attenuated to achieve the purpose of howling suppression. While this solution has low performance consumption, it has a wide channel width and significant sound loss.
[0004] Through analysis of the various existing technical solutions mentioned above, it can be seen that the detection capability of multiple howling frequency points is insufficient, the feature combination has a high miss rate for multiple howling frequency points, and is unable to process howling group bands; when the background noise is greater than -20dB, the misjudgment rate will rise rapidly, the low-frequency signal after frequency shifting processing is easily confused with human voices speaking or singing, and has weak resistance to interference in complex environments; the use of fixed filter parameters leads to the loss of high-frequency components of speech and significant damage to sound quality; the model-based method requires high computing power support, has high real-time performance and computing costs, and is difficult to deploy on embedded devices. Summary of the Invention
[0005] Based on this, it is necessary to provide a howling suppression method, device, audio processing system, computer equipment and computer-readable storage medium to address one of the above-mentioned defects, which can improve the detection and suppression effect of howling frequencies.
[0006] The present application provides a howling suppression method, comprising:
[0007] Acquire multiple frames of continuous audio signals, calculate the time domain level value of each frame of audio signals and determine the time domain energy change state;
[0008] Performing short-time Fourier transform on the audio signal of each frame in turn to obtain frequency morphology information of the audio signal;
[0009] Analyze the sound category of the audio signal according to the time domain energy change state and frequency morphology information to determine the howling frequency point;
[0010] A filter is used to suppress and filter the howling frequency points of the audio signal to eliminate the howling sound.
[0011] In one embodiment, analyzing the sound category of the audio signal according to the time domain energy change state and the frequency morphology information to determine the howling frequency point includes:
[0012] Calculate the inter-frame continuous energy difference of the audio signal;
[0013] Obtain a spectrum according to the frequency morphology information, analyze the maximum peak of the spectrum, and find and eliminate a pseudo fundamental frequency;
[0014] Searching for a howling peak value at the maximum peak point;
[0015] If the continuous energy difference between frames reaches a first threshold, or there are L consecutive frames of audio signals with the same frequency, it is determined for the first time that a howling frequency point exists.
[0016] In one embodiment, the howling suppression method further includes:
[0017] The fundamental wave adjacent peak ratio and the fundamental wave to harmonic power ratio of each howling peak are calculated. If the fundamental wave adjacent peak ratio and the fundamental wave to harmonic power ratio respectively reach the set second threshold, it is determined as a howling signal for the second time.
[0018] In one embodiment, the howling suppression method further includes:
[0019] The audio signal is subjected to pulse-like signal recognition and correlation recognition to determine whether the audio signal meets the sound characteristics of microphone collision. If so, the audio signal is determined to be a microphone collision signal. Otherwise, it is determined to be a howling signal for the third time.
[0020] In one embodiment, the howling suppression method further includes:
[0021] Detect the harmonic distribution of the audio signal, calculate the energy difference of each harmonic and the high-frequency component, and if the high-frequency component contains rich harmonics, it is determined to be a human voice, otherwise it is determined to be a howling signal.
[0022] In one embodiment, the howling suppression method further includes:
[0023] If the adjacent peak ratio of the fundamental wave and the fundamental wave to harmonic power ratio do not reach the set second threshold, the audio signal does not meet the sound characteristics of microphone collision, and the high-frequency component contains rich harmonics. Determine whether there are multiple adjacent frequency group bands in the audio signal. If so, calculate whether the audio signal in the group band meets the set fundamental wave to harmonic power ratio. If so, determine that the center of the group band is the howling frequency point.
[0024] In one embodiment, the howling suppression method further includes:
[0025] The number of energy-increasing frames is counted according to the time-domain energy change state. If the number of energy-increasing frames is greater than a set frame number threshold and meets a continuous increase condition or a rapid increase condition, it is determined that the audio signal may contain a howling signal, and the step of analyzing the sound category of the audio signal according to the time-domain energy change state and frequency morphology information to determine the howling frequency is entered.
[0026] In one embodiment, the howling suppression method further includes:
[0027] If the number of energy increasing frames does not meet the continuous rising condition or the rapid rising condition, and the current peak value and high-frequency component are less than the third threshold, it is determined that the probability of the howling signal is greater than the human voice, and the step of analyzing the sound category of the audio signal according to the time domain energy change state and frequency morphology information to determine the howling frequency is entered.
[0028] In one embodiment, the howling suppression method further includes:
[0029] The number of frequency points with the same peak value in each frame of the audio signal is determined. If the number of frequency points is greater than the set frequency band number threshold, it is determined that the audio signal is in a stable state, and the step of analyzing the sound category of the audio signal based on the time domain energy change state and frequency morphology information to determine the howling frequency point is entered.
[0030] In one embodiment, the howling suppression method further includes:
[0031] If the standard deviation of each frame of audio signal is less than the set standard deviation threshold, and the average power of the time domain energy is greater than the set power threshold, the step of analyzing the sound category of the audio signal according to the time domain energy change state and frequency morphology information to determine the howling frequency point is entered.
[0032] In one embodiment, before calculating the time domain level value of each frame of audio signal to determine the time domain energy change state, the method further includes:
[0033] grading noise in audio signals;
[0034] A dynamic gain strategy is used based on the noise level to eliminate the noise effect in the audio signal.
[0035] In one embodiment, before searching for a howling peak at the maximum peak point of the spectrum, the method further includes:
[0036] Using the fast Fourier transform values and the interpolation weights, the spectrum is inversely fast Fourier transformed to obtain a refined spectrum.
[0037] In one embodiment, using a filter to suppress and filter the howling frequency points of the audio signal to eliminate the howling sound includes:
[0038] When the howling frequency point is reached, the howling frequency points within the period are sorted and merged with the adjacent frequencies to generate a frequency point set to be processed;
[0039] If it is a single howling frequency point detected for the first time, the filter parameters are initialized and a narrowband notch filter is used to suppress the frequency points concentrated in the frequency points to be processed;
[0040] If the howling frequency point is detected multiple times at the same frequency point, the step value of the filter parameter is adjusted according to the preset strategy, and the maximum attenuation level is expanded to suppress the frequency bandwidth of the frequency points to be processed. When the upper limit of the number of filters is reached, the dynamic attenuation of the microphone gain is triggered.
[0041] In one embodiment, the howling suppression method further includes:
[0042] When the duration of no howling reaches the release period, the filters are released in order of energy from low to high;
[0043] If the microphone gain adjustment is turned on, the microphone gain will be restored first and then the filters will be released one by one;
[0044] If howling occurs again after subsequent release, the cycle interval time is extended.
[0045] In one embodiment, the howling suppression method further includes:
[0046] If the release fails multiple times in a row, it is determined to be a strong feedback scenario, and the global gain adjustment and broadband suppression combined strategy is enabled for filtering suppression. When the recurrence of the historical release frequency is detected, the corresponding filter parameters are reset first.
[0047] The present application provides a howling suppression device, comprising:
[0048] A time domain calculation unit is used to obtain a continuous multi-frame audio signal, calculate the time domain level value of each frame of audio signal and determine the time domain energy change state;
[0049] A frequency domain calculation unit, configured to sequentially perform short-time Fourier transform on the audio signal of each frame to obtain frequency morphology information of the audio signal;
[0050] a howling frequency detection unit, configured to analyze the sound category of the audio signal according to the time domain energy change state and the frequency morphology information to determine the howling frequency;
[0051] The filtering and suppressing unit is used to suppress and filter the howling frequency points of the audio signal using a filter to eliminate the howling sound.
[0052] The present application provides an audio processing system, comprising: a mixing module, a howling suppression module, an EQ module and an output module; wherein the howling suppression module is used to perform the steps of the howling suppression method;
[0053] The mixing module is used to mix the audio signals of multiple microphones;
[0054] The howling suppression module is used to detect the howling frequency and suppress the howling sound;
[0055] The EQ module is used to control the EQ parameters of the audio signal;
[0056] The output module is used to output the sound signal to the speaker.
[0057] In one embodiment, the audio processing system further includes:
[0058] The echo cancellation module is used to filter the audio signal.
[0059] In one embodiment, in the audio processing system, each microphone is connected to an EQ module and a howling suppression module; wherein each howling suppression module is connected to an EQ module; each EQ module receives an audio signal from a microphone respectively, and outputs the sound signal to the speaker after passing through the mixing module and the output module.
[0060] In one embodiment, in the audio processing system, each microphone is connected to a mixing module, and the sound signal is output to a speaker after passing through an EQ module, a howling suppression module and an output module.
[0061] The present application provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and is characterized in that when the processor executes the computer program, the steps of the howling suppression method are implemented.
[0062] The present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the howling suppression method are implemented.
[0063] The technical solution of the above embodiment calculates the time domain level value of each frame of the audio signal to determine the time domain energy change state and performs short-time Fourier transform on the audio signal of each frame in turn to obtain the frequency shape information of the audio signal, and then analyzes the sound category of the audio signal based on the time domain energy change state and frequency shape information to determine the howling frequency point, and finally uses a filter to suppress and filter the howling frequency point of the audio signal to eliminate the howling sound; this technical solution adopts a joint time-frequency feature fusion to detect the howling frequency point and filtering strategy, which can quickly locate the howling frequency point and suppress the energy of the frequency point in a musical background or other complex environment, so as to reduce the impact on the sound quality and improve the user experience.
[0064] Furthermore, by analyzing the characteristics of human voice speaking and singing, microphone slapping or collision characteristics, musical sound characteristics, spectrum characteristics when frequency shifting is turned on, and howling signal characteristics, the howling frequency points can be accurately distinguished, thereby achieving the purpose of improving the detection accuracy.
[0065] Furthermore, high-resolution frequency positioning and dynamic filter intelligent control are adopted to maximize the preservation of sound quality while also improving the suppression efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 is a flow chart of a howling suppression method according to an embodiment;
[0067] Figure 2 This is an example dynamic gain strategy flow chart;
[0068] Figure 3 This is an example flow chart for detecting howling frequencies;
[0069] Figure 4 This is an example sound category recognition flow chart;
[0070] Figure 5 This is an example dynamic filtering flow chart;
[0071] Figure 6 is a schematic structural diagram of a howling suppression device according to an embodiment;
[0072] Figure 7 is a framework diagram of an audio processing system according to an embodiment;
[0073] Figure 8 This is a block diagram of an audio processing system structure.
[0074] Figure 9 is another example of an audio processing system structure block diagram;
[0075] Figure 10 The diagram is a schematic diagram of an example computer device structure. DETAILED DESCRIPTION
[0076] In order to make the purpose, technical solutions and advantages of this application more clearly understood, the present application is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0077] In the embodiments of the present application, "at least one" means one or more, and "a plurality of" means two or more. For example, a plurality of objects refers to two or more objects. Words such as "include" or "comprise" and the like mean that the information appearing before "include" or "comprises" covers the information listed after "include" or "comprises" and its equivalents, and does not exclude other information. The "and / or" mentioned in the embodiments of the present application indicates that three relationships may exist. The character " / " generally indicates that the objects before and after are in an "or" relationship.
[0078] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a," "the," and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0079] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0080] refer to Figure 1 As shown, Figure 1 The figure is a flowchart of a howling suppression method according to an embodiment, including:
[0081] S10 , acquiring a plurality of continuous frames of audio signals, calculating the time domain level value of each frame of audio signals and determining a time domain energy change state.
[0082] In this step, the audio signal from the microphone can be collected, and L frames of audio signals can be stored in real time for calculating time domain features and frequency domain features. A single-channel audio signal (number of sampling points N) is input and cached in an internal storage unit (preset length M×L, supporting multi-frame continuous analysis). The current signal level RMS (Root Mean Square) value and the time domain level energy change state are calculated; among them, RMS is a common indicator for measuring signal amplitude, which represents the average power of a signal.
[0083] In one embodiment, in order to effectively avoid the influence of noise on howling judgment, the noise of the audio signal is graded, and the audio signal is processed using a dynamic gain strategy to eliminate the influence of noise in the audio signal.
[0084] refer to Figure 2 As shown, Figure 2 This is an example of a dynamic gain strategy flow chart, including the following:
[0085] s201, calculate the peak level within a frame;
[0086] s202, calculate the dynamic gain value;
[0087] s203, calculate the static gain coefficient;
[0088] s204, outputs the signal after gain.
[0089] For example, RMS can be used to calculate the dynamic level, and the calculation formula is as follows:
[0090]
[0091] Where RMS is the dynamic level value, x represents the time domain signal sampled by the audio signal, and i represents the sampling point.
[0092] For noise classification processing, it can be as follows:
[0093] Signal RMS < Threshold_agc_down: If Threshold_agc_down is in the range of [-70, -60], it is considered as noise and is not involved in subsequent processing.
[0094] Threshold_agc_down≤RMS<Threshold_agc_up: Gain is applied to improve the small signal signal-to-noise ratio.
[0095] RMS ≥ Threshold_agc_up: The Threshold_agc_up range is [-3, -1], and limiting processing is performed to prevent harmonic interference caused by clipping.
[0096] As in the above embodiment, the effective signal-to-noise ratio can be improved and the influence of environmental noise on sound feature detection can be reduced.
[0097] In one embodiment, considering that there are other audio signals input in the audio signal that do not generate howling, such as playing music, the input signal characteristics may be similar to howling and lead to misjudgment. Based on this, echo cancellation processing can be introduced to filter other audio signals before performing howling suppression processing to reduce misjudgment.
[0098] S20 , performing short-time Fourier transform on the audio signal of each frame in sequence to obtain frequency morphology information of the audio signal.
[0099] In this step, the frequency morphology information of the audio signal is obtained by performing short-time Fourier transform on the audio signal of each frame, thereby obtaining the spectrum of the audio signal.
[0100] In order to achieve higher frequency accuracy and reduce the loss of sound caused by the activation of the filter, in one embodiment, a spectrum refinement algorithm is used to process the spectrum. Specifically, a frequency domain interpolation method is used. After using the FFT (Fast Fourier Transform) value and the interpolation weight, an IFFT (Inverse Fast Fourier Transform) is performed on it to obtain a refined spectrum. The frequency range of the refined spectrum is calculated. The calculation formula can be as follows:
[0101]
[0102] The complex weights for frequencies are calculated as follows:
[0103]
[0104] The weight ww and coefficient a corresponding to each frequency value are as follows:
[0105]
[0106]
[0107] Finally, the complex fft value of each frequency point is multiplied by each coefficient a to obtain the complex spectrum complexX, and fft is performed to obtain fft_shift, and the value is multiplied by the complex weight, and finally IFFT is performed to convert it into the refined spectrum.
[0108] S30: Analyze the sound category of the audio signal according to the time domain energy change state and the frequency shape information to determine a howling frequency point.
[0109] In this step, the energy, duration, and morphological changes in the time and frequency domains are used to analyze the differences between various other sounds and the characteristics of the howling signal, and the howling frequency point is accurately located through sound category analysis. For example, other sounds can be human singing and talking, microphone slapping or collision, frequency shifting, and other audio processing functions.
[0110] In one embodiment, considering that howling signals are acoustic feedback phenomena, they have characteristics such as the same frequency point, continuous energy increase, rapid time, and small harmonic components. Based on this, in addition to using the traditional fundamental wave neighboring peak ratio (PNPR), fundamental wave to harmonic power ratio, and inter-frame continuous energy difference (IMSD), before analyzing the sound category of the audio signal, further feature analysis is performed based on the calculated time domain features and frequency domain features, respectively comparing the continuous level change pattern in the time domain space and the spectrum morphology pattern in the frequency domain. These include: howling features: PNPR, PHPR, and IMSD; human voice features: rich harmonics, low sound stability; microphone collision recognition: pulse-like signal recognition, and correlation recognition.
[0111] For example, by analyzing the energy state in the time domain and the spectrum in the frequency domain, combined with the joint judgment and combined analysis in the time domain and spectrum, the howling frequency points can be accurately distinguished. Figure 3 As shown, Figure 3 This is an example flow chart for detecting howling frequencies, which may include the following:
[0112] (a) Counting the number of frames with increasing energy based on the time-domain energy change state; if the number of frames with increasing energy is greater than a set frame number threshold and satisfies a continuous increase condition or a rapid increase condition, determining that the audio signal may contain a howling signal, and entering a step of analyzing the sound category of the audio signal to determine the howling frequency.
[0113] Specifically, if the energy increasing frame number k>frame number threshold Th_k, and the continuous increase condition or the rapid increase condition is met, then a howling signal may exist, and the sound category analysis process is entered.
[0114] (b) If the number of energy increasing frames does not meet the continuous increase condition or the rapid increase condition, and the current peak value and high-frequency component are less than the third threshold, it is determined that the probability of the howling signal is greater than the human voice, and the step of analyzing the sound category of the audio signal to determine the howling frequency is entered.
[0115] Specifically, if the continuous rising condition is not met, but the current peak value and high-frequency component hfM < high-frequency component threshold Th_hfm are met, it is considered that howling is likely to occur. Because human speech contains rich harmonic information, the sound category analysis process is entered.
[0116] (c) Determine the number of frequency points with the same peak value in each frame of the audio signal. If the number of frequency points is greater than a set frequency band number threshold, the audio signal is determined to be in a stable state, and the process proceeds to step S30 to analyze the sound type of the audio signal and determine the howling frequency point.
[0117] Specifically, if the current audio signal is howling but is not recognized in time, resulting in the howling being in a stable state, the number of energy increment frames k is less than the frame number threshold Th_k, and the frequency stability of the frequency point in the frequency domain is used as the standard for suspected howling. When the number of the same peak frequency points C in L frames is greater than the minimum threshold The_c, the sound is considered to be in a stable state and enters the sound category analysis.
[0118] (d) If the standard deviation of each frame of the audio signal is less than the set standard deviation threshold, and the average power of the time domain energy is greater than the set power threshold, the process proceeds to step S30 to analyze the sound category of the audio signal and determine the howling frequency.
[0119] Specifically, since the audio signal may experience oscillations in frequency when a sound processor such as frequency shifting is enabled, if the standard deviation std of L consecutive frames is less than the minimum threshold Th_std and the sound energy is greater than the minimum power threshold Th_rms, and if the above requirements are met, the sound category analysis process is entered.
[0120] For example, if none of the above (a), (b), (c), and (d) can determine that it is necessary to enter the sound category analysis process, if the peak standard deviation std < standard deviation threshold Th_std is satisfied, and the judgment conditions of low-frequency howling are met, it is considered that a howling signal may exist and the sound category analysis process is entered.
[0121] In one embodiment, reference Figure 4 As shown, Figure 4 This is an example flow chart of sound category recognition. Step S30 of analyzing the sound category of the audio signal to determine the howling frequency may include:
[0122] (1) Calculate the continuous energy difference between frames of the audio signal, obtain a spectrum based on the frequency morphology information, analyze the maximum peak of the spectrum, and find and eliminate the pseudo fundamental frequency; search for the howling peak at the maximum peak point. If the continuous energy difference between frames meets a first threshold, or there are L consecutive frames of audio signals with the same frequency, it is first determined that a howling frequency point may exist.
[0123] Specifically, first calculate IMSD and refine the spectrum of the maximum peak point and search for possible howling peak f_how. Due to the complexity of the environment and the fact that the human voice is in the middle and low frequency position when speaking, the harmonics will be higher than the fundamental peak. The pseudo fundamental frequency will cause PHPR to misjudge. Therefore, first analyze the maximum peak to determine whether there is a pseudo fundamental frequency. For example, you can look for 1 / 2, 1 / 3, and 1 / 4 times the current peak f_max. If there is a peak at this multiple, and the peak energy difference is small, and the sound energy is large, that is,
[0124]
[0125] The peak is determined to be a pseudo fundamental frequency. After eliminating the influence of the pseudo fundamental frequency, if the IMSD threshold Th_imsd is met, or there are L consecutive frames with the same frequency, then the first determination is that a howling point may exist.
[0126] (2) Calculate the fundamental-to-harmonic power ratio PNPR and fundamental-to-harmonic power ratio PHPR of each howling peak f_how. If the PNPR and PHPR reach the set second thresholds, respectively, where the second thresholds include a PHPR threshold and a PHPR threshold, then the second judgment is that it may be a howling signal.
[0127] Specifically, the PNPR value and the PHPR value of each howling peak value f_how are calculated. If the PNPR value and the PHPR value reach the set threshold value, it is determined that the second determination is that the signal may be a howling signal.
[0128] (3) Perform pulse-like signal recognition and correlation recognition on the audio signal to determine whether the audio signal meets the sound characteristics of microphone collision. If so, the audio signal is determined to be a microphone collision signal. Otherwise, the third judgment may be a howling signal.
[0129] Specifically, to determine whether the sound is a microphone collision, in order to prevent misjudgment caused by microphone tapping, a pulse judgment method can be added. Since the microphone tapping signal has the characteristics of sudden change and no periodicity, the inter-frame energy difference or correlation judgment is used. Both can distinguish whether the sound is caused by microphone misoperation. The judgment formula can be as follows:
[0130] C = real(ifft(fft.*conj(fft))
[0131] If the microphone tapping requirements are not met, the third judgment is that it may be a howling signal.
[0132] (4) Detect the harmonic distribution of the audio signal, calculate the energy difference of each harmonic and the high-frequency component, and if the high-frequency component contains rich harmonics, it is determined to be a human voice. Otherwise, it is determined to be a howling signal for the fourth time, and finally confirmed as the howling frequency point.
[0133] Specifically, check the harmonic distribution, including calculating the energy difference and high-frequency components of each harmonic, as follows:
[0134]
[0135] in, is the harmonic energy difference, hfm is the high frequency component; if it contains rich harmonics, it means it is a human voice speaking or singing, otherwise, the fourth time it is judged as a howling signal and is output as the final howling signal frequency.
[0136] For example, if (1) is satisfied but (2) is not, but there is only one peak and the PHPR reaches the second threshold, the second judgment is that it may be a howling signal, and the judgment continues to (3). If the howling frequency point cannot be determined after (2), (3) and (4), that is, if the fundamental wave's neighboring peak ratio PNPR and the fundamental wave to harmonic power ratio PHPR do not reach the set second threshold, the audio signal does not meet the sound characteristics of microphone collision, and the high-frequency component contains rich harmonics, it is necessary to further determine whether there is a howling group.
[0137] In one embodiment, when determining whether a howling group exists, it is possible to determine whether the audio signal contains multiple adjacent frequency group bands. If so, it is calculated whether the audio signal in the group band meets the requirement of the set fundamental wave to harmonic power ratio PHPR. If so, the center of the group band is determined to be the howling frequency point.
[0138] Specifically, the howling group is to determine whether there are multiple adjacent frequencies due to long-term howling, which does not meet the adjacent peak ratio. First, check whether there is a howling group band, and consider the adjacent peaks less than the set peak threshold as members of the howling group band. If so, determine whether the harmonic ratio PHPR is met in the group band. If so, determine that the center of the group band is the howling frequency point.
[0139] S40: Utilize a filter to suppress and filter the howling frequency points of the audio signal to eliminate the howling sound.
[0140] In this step, after the howling frequency is detected by time-frequency joint feature fusion, the howling frequency of the audio signal is suppressed and filtered using a filter to eliminate the howling sound. In a musical background or other complex environment, the howling frequency can be quickly located and the energy of the frequency can be suppressed to reduce the impact on sound quality and improve the user experience.
[0141] For example, this embodiment may adopt a dynamic filter design, such as using a notch filter to suppress howling, or may combine other high-pass and low-pass filters to effectively resolve the impact of the howling group.
[0142] To ensure filtering quality, in step S40 of this embodiment, a dynamic filter strategy can be adopted to suppress and filter the howling frequency points of the audio signal using a filter to eliminate the howling sound. By setting different filter parameters, a single or multiple filter combinations can be dynamically adjusted based on the principle of minimizing sound loss to achieve the purpose of quickly suppressing howling. When howling no longer occurs, the filter is dynamically released to preserve the original sound quality.
[0143] Accordingly, in one embodiment, reference Figure 5 As shown, Figure 5 Here is an example dynamic filtering flow chart, which may include the following:
[0144] 1) When the howling frequency point is reached, the howling frequency points within the period are sorted and merged with adjacent frequencies to generate a frequency point set to be processed.
[0145] Specifically, when a howling point is detected and the corresponding time T is reached, the howling points in the period (including frequency, number of occurrences, and energy intensity) are sorted and merged with adjacent frequencies to generate a frequency point set to be processed.
[0146] 2) Implement different response strategies based on different howling frequencies.
[0147] Primary response: If a single howling frequency point is detected for the first time, the filter parameters are initialized and a narrowband notch filter is used to suppress the frequencies in the frequency concentration to be processed. Specifically, when a single howling point is detected for the first time, the filter parameters (center frequency f, attenuation gain g, quality factor Q) are initialized and a narrowband notch filter is used for suppression (the Q value is adaptively matched to the frequency bandwidth).
[0148] Multi-level enhancement: If howling is detected multiple times at the same frequency point, the filter parameter step value is adjusted according to the preset strategy, and the maximum attenuation level expands and suppresses the frequency bandwidth of the frequency points to be processed. When the upper limit of the number of filters is reached, the microphone gain is dynamically attenuated. Specifically, when howling is detected multiple times at the same frequency point, the filter parameters are adjusted according to the preset strategy (such as gradually reducing the Q value step value ΔQ, expanding the suppression bandwidth when the maximum attenuation level Q is less than the minimum threshold Th_Q, or adding a parallel filter to handle group-band howling). When the upper limit of the number of filters is reached, the microphone gain is dynamically attenuated, with an attenuation step value Δg and a maximum attenuation level g less than the minimum threshold Th_g.
[0149] 3) A release mechanism can also be set during the howling suppression process.
[0150] Release mechanism: When the duration without howling reaches the release period T < the minimum threshold Th_release, the filters are released in order from low to high energy; if the microphone gain adjustment is turned on, the microphone gain is restored first and then the filters are released one by one; if howling occurs again after subsequent releases, the period interval time is extended to avoid oscillation.
[0151] 4) You can also set an exception handling strategy during the howling suppression process.
[0152] Exception handling: If multiple consecutive release failures occur, it is determined to be a strong feedback scenario. A combined global gain adjustment and broadband suppression strategy is enabled for filtering and suppression. When a recurrence of a historical release frequency is detected, the corresponding filter parameters are reset first. Specifically, if multiple consecutive release failures occur (exceeding the preset number N < the minimum threshold Th_fil_n), it is determined to be a strong feedback scenario. The combined global gain adjustment and broadband suppression strategy is directly enabled. When a recurrence of a historical release frequency is detected, the corresponding filter parameters are reset first to avoid delays caused by repeated initialization.
[0153] In summary, the technical solutions of the above embodiments are used to accurately locate the howling frequency points through time-frequency joint feature fusion detection and adaptive dynamic filter strategy. The dynamic change effects of audio processing functions such as human singing and speaking, microphone slapping or collision, and frequency shift are analyzed through energy, duration, and morphological changes in the time and frequency domains, and the difference in howling signal characteristics is distinguished. By using the dynamic filter strategy, by setting different filter parameters, a single or multiple filter combinations are dynamically adjusted based on the principle of minimum sound loss to achieve the purpose of quickly suppressing howling. When howling no longer occurs, the filter is dynamically released to preserve the original sound quality. By analyzing the characteristics of different sound types, whether it is a howling sound is accurately distinguished to achieve the purpose of improving accuracy. High-resolution frequency positioning and dynamic filter intelligent control are used to maximize the preservation of sound quality while also improving suppression efficiency. This solves the problem of quickly locating the howling frequency points and suppressing the energy of the frequency points in a musical background or other complex environment, thereby reducing the impact on sound quality and improving user experience.
[0154] An embodiment of the howling suppression device is described below.
[0155] refer to Figure 6 As shown, Figure 6 The figure is a schematic structural diagram of a howling suppression device according to an embodiment, comprising:
[0156] The time domain calculation unit 10 is used to obtain a continuous multi-frame audio signal, calculate the time domain level value of each frame of the audio signal to determine the time domain energy change state;
[0157] The frequency domain calculation unit 20 is configured to perform short-time Fourier transform on the audio signal of each frame to obtain frequency pattern information of the audio signal;
[0158] The howling frequency point detection unit 30 is configured to analyze and determine the sound category of the audio signal according to the time domain energy change state and the frequency pattern information to obtain the howling frequency point.
[0159] The filter suppression unit 40 is configured to perform filter suppression on the howling frequency point of the audio signal by using a filter to eliminate the howling sound.
[0160] The howling suppression device of the embodiment can perform the howling suppression method provided by the embodiment, and the implementation principle is similar. The actions performed by each module in the howling suppression device in each embodiment of the present application correspond to the steps in the howling suppression method in each embodiment of the present application. For detailed functions of each module of the howling suppression device, refer to the description of the corresponding howling suppression method in the foregoing, which will not be repeated here.
[0161] The following describes an embodiment of an audio processing system.
[0162] The audio processing system of the embodiment, as shown in Figure 7 Figure 7 is a framework diagram of the audio processing system of an embodiment, which includes a mixing module, a howling suppression module, an EQ module, and an output module. The howling suppression module is configured to perform the steps of the howling suppression method of any of the above embodiments.
[0163] Specifically, the mixing module is configured to mix the audio signals of multiple microphones (1-n, n≥2). The howling suppression module is configured to detect the howling frequency point and suppress the howling sound. The EQ module is configured to control the EQ parameters of the audio signal. The output module is configured to output the sound signal to a loudspeaker.
[0164] Further, the audio processing system can further include an echo cancellation module configured to filter the audio signal. Specifically, when other audio signals that do not generate howling are input, such as playing music, the characteristics of the input signal can be similar to the howling, leading to false judgment. Therefore, the echo cancellation module can be introduced to filter other audio signals and then perform the howling suppression processing to reduce false judgment.
[0165] The solution of the above-mentioned embodiment is applicable to scenarios with multiple other audio signal inputs. Since the howling suppression module adopts the time-frequency joint feature fusion detection and filtering strategy, the howling frequency can be quickly located and the energy of the frequency can be suppressed in a musical background or other complex environment, so as to reduce the impact on the sound quality and improve the user experience. By analyzing the characteristics of human voice speaking and singing, the characteristics of microphone slapping or collision, the characteristics of music sound, the spectrum characteristics when frequency shifting is turned on, and the characteristics of howling signals, the howling frequency can be accurately distinguished, so as to achieve the purpose of improving the detection accuracy. High-resolution frequency positioning and dynamic filter intelligent control are used to maximize the preservation of sound quality while also improving the suppression efficiency.
[0166] In one embodiment, the audio processing system of the present application, such as Figure 8 As shown, Figure 8 This is an example of an audio processing system structure block diagram, with n microphones, each of which is connected to an EQ module and a howling suppression module; wherein each howling suppression module is connected to an EQ module; each EQ module receives an audio signal from one microphone, and outputs the sound signal to the speaker after passing through the mixing module and the output module.
[0167] Specifically, each audio input signal has its own independent EQ module and howling suppression module. The howling suppression module detects howling and controls the EQ module parameters. Each audio input signal performs howling suppression independently without affecting each other. If microphone 1 generates howling, howling suppression will not affect the signals of other microphones, thus providing a better experience.
[0168] In one embodiment, the audio processing system of the present application, such as Figure 9 As shown, Figure 9 This is another example of an audio processing system structure block diagram. Each microphone is connected to a mixing module, and the sound signal is output to the speaker after passing through an EQ module, a howling suppression module, and an output module.
[0169] Specifically, each audio input signal is mixed into one audio signal through the mixing module, and then output to the EQ module, howling suppression module, etc. The howling suppression module detects howling and controls the EQ module parameters, thereby reducing the use of computing power and making it suitable for use in more types of low-end devices.
[0170] The following describes embodiments of the computer device and computer-readable storage medium of the present application.
[0171] refer to Figure 10 As shown, Figure 101 is a schematic diagram of an exemplary computer device structure, which may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc. The computer device 100 may include one or more of the following components: a processing component 102, a memory 104, a power component 106, a multimedia component 108, an audio component 109, an input / output (I / O) interface 112, a sensor component 114, and a communication component 116.
[0172] Processing component 102 generally controls the overall operation of computer device 100, such as operations associated with display, telephone calls, data communications, camera operations, and recording operations.
[0173] The memory 104 is configured to store various types of data to support operations in the computer device 100, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0174] The power supply assembly 106 provides power to the various components of the computer device 100 .
[0175] The multimedia component 109 includes a screen that provides an output interface between the computer device 100 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). In some embodiments, the multimedia component 108 includes a front camera and / or a rear camera.
[0176] The audio component 109 is configured to output and / or input audio signals.
[0177] I / O interface 112 provides an interface between processing component 102 and peripheral interface modules, such as a keyboard, click wheel, buttons, etc. These buttons may include, but are not limited to, a home button, volume buttons, a start button, and a lock button.
[0178] The sensor assembly 114 includes one or more sensors for providing various aspects of status assessment for the computer device 100. The sensor assembly 114 may include a proximity sensor configured to detect the presence of a nearby object without any physical contact.
[0179] The communication component 116 is configured to facilitate wired or wireless communication between the computer device 100 and other devices. The computer device 100 can access a wireless network based on a communication standard, such as WiFi, a carrier network (such as 2G, 3G, 4G or 5G), or a combination thereof.
[0180] Those skilled in the art can understand that the computer device structure provided by the above-mentioned embodiments is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. A specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0181] The present application also provides a computer readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the steps in the methods of the above-mentioned embodiments. Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by a computer program instructing related hardware, and the computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In the embodiments of the present application, any reference to a memory, a database or other medium can include at least one of a non-volatile and a volatile memory. The non-volatile memory can include a read-only memory (ROM), a magnetic tape, a floppy disk, a flash memory, an optical memory, a high-density embedded non-volatile memory, a resistive memory (ReRAM), a magnetoresistive random access memory (MRAM), a ferroelectric memory (FRAM), a phase change memory (PCM), a graphene memory, etc. The volatile memory can include a random access memory (RAM) or an external cache memory, etc. As an illustration but not as a limitation, the RAM can be in various forms, such as a static random access memory (SRAM) or a dynamic random access memory (DRAM), etc. The database involved in the embodiments of the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments of the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without being limited thereto.
[0182] It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are information and data authorized by the user or authorized by all parties.
[0183] The technical features of the above embodiments can be combined in any manner. In order to make the description simple, all possible combinations of the technical features in the above embodiments are not described, but as long as the combinations of the technical features do not exist, they should be considered as the scope of the present disclosure.
[0184] The above embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of the patent. It should be noted that for ordinary skilled persons in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are within the scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. A howling suppression method, characterized in that: include: Acquire multiple frames of continuous audio signals, calculate the time domain level value of each frame of audio signals and determine the time domain energy change state; Performing short-time Fourier transform on the audio signal of each frame in turn to obtain frequency morphology information of the audio signal; Analyze the sound category of the audio signal according to the time domain energy change state and frequency morphology information to determine the howling frequency point; A filter is used to suppress and filter the howling frequency points of the audio signal to eliminate the howling sound.
2. The howling suppression method according to claim 1, wherein: Analyzing the sound category of the audio signal according to the time domain energy change state and the frequency morphology information to determine the howling frequency point includes: Calculate the inter-frame continuous energy difference of the audio signal; Obtain a spectrum according to the frequency morphology information, analyze the maximum peak of the spectrum, and find and eliminate a pseudo fundamental frequency; Searching for a howling peak value at the maximum peak point; If the continuous energy difference between frames reaches a first threshold, or there are L consecutive frames of audio signals with the same frequency, it is determined for the first time that a howling frequency point exists.
3. The howling suppression method according to claim 1, wherein: Also includes: The fundamental wave adjacent peak ratio and the fundamental wave to harmonic power ratio of each howling peak are calculated. If the fundamental wave adjacent peak ratio and the fundamental wave to harmonic power ratio respectively reach the set second threshold, it is determined as a howling signal for the second time.
4. The howling suppression method according to claim 3, wherein: Also includes: The audio signal is subjected to pulse-like signal recognition and correlation recognition to determine whether the audio signal meets the sound characteristics of microphone collision. If so, the audio signal is determined to be a microphone collision signal. Otherwise, it is determined to be a howling signal for the third time.
5. The howling suppression method according to claim 4, characterized in that: Also includes: Detect the harmonic distribution of the audio signal, calculate the energy difference of each harmonic and the high-frequency component, and if the high-frequency component contains rich harmonics, it is determined to be a human voice, otherwise it is determined to be a howling signal.
6. The howling suppression method according to claim 5, characterized in that: Also includes: If the adjacent peak ratio of the fundamental wave and the fundamental wave to harmonic power ratio do not reach the set second threshold, the audio signal does not meet the sound characteristics of microphone collision, and the high-frequency component contains rich harmonics. Determine whether there are multiple adjacent frequency group bands in the audio signal. If so, calculate whether the audio signal in the group band meets the set fundamental wave to harmonic power ratio. If so, determine that the center of the group band is the howling frequency point.
7. The howling suppression method according to any one of claims 2 to 6, characterized in that: Also includes: The number of energy-increasing frames is counted according to the time-domain energy change state. If the number of energy-increasing frames is greater than a set frame number threshold and meets a continuous increase condition or a rapid increase condition, it is determined that the audio signal may contain a howling signal, and the step of analyzing the sound category of the audio signal according to the time-domain energy change state and frequency morphology information to determine the howling frequency is entered.
8. The howling suppression method according to claim 7, characterized in that: Also includes: If the number of energy increasing frames does not meet the continuous rising condition or the rapid rising condition, and the current peak value and high-frequency component are less than the third threshold, it is determined that the probability of the howling signal is greater than the human voice, and the step of analyzing the sound category of the audio signal according to the time domain energy change state and frequency morphology information to determine the howling frequency is entered.
9. The howling suppression method according to claim 8, characterized in that: Also includes: The number of frequency points with the same peak value in each frame of the audio signal is determined. If the number of frequency points is greater than the set frequency band number threshold, it is determined that the audio signal is in a stable state, and the step of analyzing the sound category of the audio signal based on the time domain energy change state and frequency morphology information to determine the howling frequency point is entered.
10. The howling suppression method according to claim 9, characterized in that: Also includes: If the standard deviation of each frame of audio signal is less than the set standard deviation threshold, and the average power of the time domain energy is greater than the set power threshold, the step of analyzing the sound category of the audio signal according to the time domain energy change state and frequency morphology information to determine the howling frequency point is entered.
11. The howling suppression method according to claim 1, wherein: Before calculating the time domain level value of each frame of audio signal and determining the time domain energy change state, the method further includes: grading noise in audio signals; A dynamic gain strategy is used based on the noise level to eliminate the noise effect in the audio signal.
12. The howling suppression method according to claim 1, characterized in that: Before searching for a howling peak at the maximum peak point of the spectrum, the method further includes: Using the fast Fourier transform values and the interpolation weights, the spectrum is inversely fast Fourier transformed to obtain a refined spectrum.
13. The howling suppression method according to claim 1, wherein: The method comprises: using a filter to suppress and filter the howling frequency of the audio signal to eliminate the howling sound, comprising: When the howling frequency point is reached, the howling frequency points within the period are sorted and merged with the adjacent frequencies to generate a frequency point set to be processed; If it is a single howling frequency point detected for the first time, the filter parameters are initialized and a narrowband notch filter is used to suppress the frequency points concentrated in the frequency points to be processed; If the howling frequency point is detected multiple times at the same frequency point, the step value of the filter parameter is adjusted according to the preset strategy, and the maximum attenuation level is expanded to suppress the frequency bandwidth of the frequency points to be processed. When the upper limit of the number of filters is reached, the dynamic attenuation of the microphone gain is triggered.
14. The howling suppression method according to claim 13, characterized in that: Also includes: When the duration of no howling reaches the release period, the filters are released in order of energy from low to high; If the microphone gain adjustment is turned on, the microphone gain will be restored first and then the filters will be released one by one; If howling occurs again after subsequent release, the cycle interval time is extended.
15. The howling suppression method according to claim 14, characterized in that: Also includes: If the release fails multiple times in a row, it is determined to be a strong feedback scenario, and the global gain adjustment and broadband suppression combined strategy is enabled for filtering suppression. When the recurrence of the historical release frequency point is detected, the corresponding filter parameters are reset first.
16. A howling suppression device, characterized in that: include: A time domain calculation unit is used to obtain a continuous multi-frame audio signal, calculate the time domain level value of each frame of audio signal and determine the time domain energy change state; A frequency domain calculation unit, configured to sequentially perform short-time Fourier transform on the audio signal of each frame to obtain frequency morphology information of the audio signal; a howling frequency detection unit, configured to analyze the sound category of the audio signal according to the time domain energy change state and the frequency morphology information to determine the howling frequency; The filtering and suppressing unit is used to suppress and filter the howling frequency points of the audio signal using a filter to eliminate the howling sound.
17. An audio processing system, characterized in that: include: A mixing module, a howling suppression module, an EQ module and an output module; wherein the howling suppression module is used to perform the steps of the howling suppression method according to any one of claims 1 to 15; The mixing module is used to mix the audio signals of multiple microphones; The howling suppression module is used to detect the howling frequency and suppress the howling sound; The EQ module is used to control the EQ parameters of the audio signal; The output module is used to output the sound signal to the speaker.
18. The audio processing system according to claim 17, characterized in that Each microphone is connected to an EQ module and a howling suppression module; each howling suppression module is connected to an EQ module; each EQ module receives the audio signal from one microphone, and outputs the sound signal to the speaker after passing through the mixing module and the output module; or Each microphone is connected to the mixing module, and the sound signal is output to the speaker after passing through an EQ module, a howling suppression module and an output module.
19. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the howling suppression method according to any one of claims 1 to 15 are implemented.
20. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the howling suppression method according to any one of claims 1 to 15 are implemented.
Citation Information
Cited By
Audio processing method and device, electronic equipment and storage medium
CN121122326A