Howling suppression method, device, audio system and sound amplification system
Through the method of updating the filter coefficients of the frequency domain adaptive filter and remote audio signals, the problems of adaptive filter estimation deviation and environmental changes of howling suppressed adaptive filters in the amplified sound system are solved, and the stability and adaptability are improved.
Patent Information
- Application Number
- CN202210772248.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-30
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2042-06-30
AI Technical Summary
The howling suppression method in existing amplification systems has problems of adaptive filter estimation deviation and performance degradation during environmental changes, especially in acoustic loop feedback. The existing methods increase operational overhead or do not work well when environmental changes.
Frequency domain adaptive filter is used to process the audio signal frame by frame, frequency point by frequency, and filter coefficients are updated with the remote audio signal. Combined with the initial state estimation feedback path and adaptive filter real-time tracking, the correlation between the filter input signal and the reference signal is reduced, and howling suppression is achieved.
Effectively suppress feedback signals, improve the stability of the audio processing system and the ability to adapt to changes in the sound field environment, reduce the correlation between the filter input signal and the reference signal, and improve the howling suppression effect.
Smart Images

Figure CN115175063B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of digital signal processing, and in particular to a howling suppression method, device, speaker, and sound amplification system. Background Art
[0002] In a sound reinforcement system, the signal collected by the microphone is transmitted to the speaker for amplification and broadcast. The audio signal played by the speaker is then picked up again by the microphone. The transmission and feedback of the audio signal between the speaker and the microphone form an acoustic loop. During this transmission process, when the volume is high, the sound feedback loop forms positive feedback, that is, the acoustic loop gain is greater than 1. This continuous feedback process amplifies the sound step by step, producing a harsh howling sound, which seriously affects the user's listening experience.
[0003] Currently, methods for suppressing howling in sound reinforcement systems include frequency and phase shifting, notch suppression, and adaptive feedback suppression. Frequency and phase shifting, during the sound processing process, disrupts the phase characteristics required for positive feedback by changing the frequency or phase of the sound in real time. Notch suppression, targeting the frequency point where howling occurs, forcibly lowers the acoustic loop gain at that frequency point using a notch filter. However, both methods alter the frequency response of the sound signal or system, causing some distortion to the sound.
[0004] Adaptive feedback suppression uses an adaptive filter to track the feedback path and offset its effects, effectively preventing howling. However, due to the high correlation between the input signal and the reference signal, the adaptive filter estimation may be biased. There are two common approaches to addressing this bias: decorrelation techniques and fixed-coefficient filters. However, the former requires increased computational overhead, while the latter significantly degrades performance with subtle changes in the sound field environment. Summary of the Invention
[0005] The embodiments of the present application provide a howling suppression method, device, speaker, and sound amplification system, aiming to solve at least one technical problem in the prior art.
[0006] According to a first aspect of an embodiment of the present application, a howling suppression method is provided, the method comprising:
[0007] Preprocessing the audio signal in the sound reinforcement system and converting the audio signal into the frequency domain;
[0008] Based on the filter coefficients of the frequency-domain adaptive filter, the converted audio signal is processed frame by frame and frequency by frequency point to obtain the corresponding output signal. At the same time, the filter coefficients are updated using the far-end audio signal in the current frame signal as a reference signal for use in the next frame signal processing;
[0009] All the obtained output signals are converted into the time domain to obtain the target audio.
[0010] In one possible implementation, before processing the initial frame signal at any frequency point, the method further includes:
[0011] determining initial filter coefficients corresponding to a transfer function of a loudspeaker-to-microphone path in the sound reinforcement system;
[0012] The process of processing the initial frame signal at any frequency point includes:
[0013] Based on the initial filter coefficients, the initial frame signal is processed to obtain a corresponding output signal, and the initial filter coefficients are updated at the same time;
[0014] The initial frame is determined based on the number of the frequency domain adaptive filters.
[0015] In another possible implementation, the process of processing a non-initial frame signal at any frequency point includes:
[0016] Based on the updated filter coefficients obtained after processing the previous frame signal, the current frame signal is processed to obtain a corresponding output signal, and the filter coefficients are updated at the same time.
[0017] In another possible implementation, the audio signal in the sound reinforcement system includes a remote audio signal and an audio signal collected by a microphone. The process of processing each frame signal at any frequency point to obtain a corresponding output signal includes:
[0018] Determine a residual signal of the audio signal collected by the microphone in the current frame signal based on the current frame signal and the corresponding filter coefficients, as well as the output signal of the previous frame;
[0019] Performing local amplification processing on the residual signal to obtain a corresponding output signal.
[0020] In another possible implementation, the process of updating the filter coefficients using the far-end audio signal in the current frame signal as a reference signal includes:
[0021] Based on a preset update step size, the current frame signal and the corresponding filter coefficients, a residual signal, and a far-end audio signal in the current frame signal, new filter coefficients are determined and the filter coefficients corresponding to the current frame signal are updated.
[0022] In another possible implementation, the process of determining the residual signal includes:
[0023] If the echo suppression amount of the current frame signal is greater than or equal to a preset threshold, copy the filter coefficient corresponding to the current frame signal, and determine the residual signal based on the current frame signal, the corresponding filter coefficient, and the previous frame output signal;
[0024] If the echo suppression amount is less than the preset threshold, determining the residual signal based on the last copied filter coefficient, the current frame signal and the corresponding filter coefficient, and the last frame output signal;
[0025] The echo suppression amount is determined based on a power ratio of the output signal to the audio signal collected by the microphone.
[0026] In another possible implementation, if the preprocessing is short-time Fourier transform, converting all the obtained output signals into the time domain to obtain the target audio includes:
[0027] All the obtained output signals are subjected to inverse short-time Fourier transform back to the time domain to obtain the target audio.
[0028] According to a second aspect of an embodiment of the present application, a howling suppression device is provided, comprising: a local sound reinforcement system, an adder, a frequency domain adaptive filter, and a short-time Fourier transform module and an inverse transform module, wherein an input of the frequency domain adaptive filter is connected to a loudspeaker, an output of the frequency domain adaptive filter is connected to an input of the adder, an output of the adder is connected to an input of the local sound reinforcement system, and an output of the local sound reinforcement system is connected to the loudspeaker;
[0029] The short-time Fourier transform module is used to pre-process the audio signal collected by the microphone and convert the audio signal into the frequency domain;
[0030] The frequency domain adaptive filter is used to perform echo suppression and howling suppression processing on the converted audio signal and then output it to the adder;
[0031] The adder is used to subtract the converted audio signal from the signal output by the frequency domain adaptive filter and output the resultant signal to the local sound reinforcement system;
[0032] The local sound amplification system performs local sound amplification processing on the received signal and then transmits it to the inverse short-time Fourier transform module;
[0033] The inverse short-time Fourier transform module is used to convert the received signal into the time domain, obtain the target audio and transmit it to the speaker for playback.
[0034] According to a third aspect of the embodiments of the present application, there is provided a sound system comprising: a speaker and the howling suppression device according to the embodiment of the second aspect, wherein:
[0035] The loudspeaker is connected to the input end of the frequency domain adaptive filter in the howling suppression device, the output end of the frequency domain adaptive filter is connected to the input end of the adder in the howling suppression device, the output end of the adder is connected to the input end of the local sound amplification system in the howling suppression device, and the output end of the local sound amplification system is connected to the loudspeaker.
[0036] According to a fourth aspect of the embodiments of the present application, a sound amplification system is provided, comprising: a microphone, a loudspeaker, and the howling suppression device described in the embodiment of the second aspect above, wherein the howling suppression device is arranged between the microphone and the loudspeaker, and the howling suppression device is used to receive the audio signal collected by the microphone and output the generated target audio to the loudspeaker for playback.
[0037] The beneficial effects of the technical solution provided by the embodiments of the present application are:
[0038] Based on the filter coefficients of the frequency-domain adaptive filter, the frequency-domain audio signal is processed frame by frame and frequency by frequency point to obtain the corresponding output signal. At the same time, the filter coefficients are updated using the far-end audio signal in the current frame signal as the reference signal for use in processing the next frame signal. Since howling suppression reuses the filter structure used in echo cancellation, and the reference signal when updating the filter coefficients is the far-end signal, it not only reduces the correlation between the filter input signal and the reference signal, but also effectively suppresses the feedback signal. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments of the present application.
[0040] Figure 1 Schematic diagram of the signal transmission process in the related art;
[0041] Figure 2 A schematic diagram of a signal transmission process corresponding to a howling suppression method provided in an embodiment of the present application;
[0042] Figure 3 A flowchart of a howling suppression method provided in an embodiment of the present application.
[0043] Icon: 10-speaker; 20-microphone; 30-local sound reinforcement system; 40-adder; 50-frequency domain adaptive filter. DETAILED DESCRIPTION
[0044] The following describes the embodiments of the present application in conjunction with the accompanying drawings. It should be understood that the embodiments described below in conjunction with the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions of the embodiments of the present application.
[0045] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0046] When a sound reinforcement system uses a microphone to pick up sound, the sound signal collected by the microphone is transmitted to the speaker for amplification and playback, and the sound signal played by the speaker is transmitted through space and collected by the microphone again. Since it is impossible to completely isolate the sound pickup area of the microphone from the playback area of the speaker, the sound played by the speaker can easily be transmitted to the microphone through space and cause feedback howling. Figure 1 The figure shows the signal transmission process in the sound reinforcement system. x is the near-end speech signal, that is, the actual speaking sound, u is the audio signal finally played by the loudspeaker, k is the feedback signal after the transfer function H, that is, the audio signal played by the loudspeaker is collected by the microphone again after being transmitted through the space, y is the sound signal collected by the microphone, and G is the local sound reinforcement system. It can be seen that an acoustic feedback loop is formed between the microphone and the loudspeaker. When the loop enters positive feedback, the signal is gradually amplified in the continuous feedback, and eventually howling is generated.
[0047] Adaptive feedback suppression uses an adaptive filter to track the feedback path and offset its effects, effectively preventing howling. However, due to the high correlation between the input signal and the reference signal, the adaptive filter estimation is subject to bias. A common approach is to use decorrelation technology to reduce the correlation between the filter input and the reference signal. Decorrelation methods include injecting noise, increasing delay, adding nonlinear processing, and pre-filtering. However, pre-filtering requires the addition of coefficient estimation and inverse filtering circuits, increasing computational overhead, while the other methods have a less pronounced effect on suppression gain.
[0048] A simpler method for reducing adaptive filter estimation bias is to use a section of white noise to estimate the acoustic feedback path in the initial state. After the filter is set, the filter coefficients are fixed and remain unchanged during real-time use. However, the performance of the fixed filter method degrades significantly when the sound field environment undergoes subtle changes, such as when a window opens or people move around.
[0049] In order to solve the above-mentioned technical problems existing in the prior art, the embodiments of the present application provide a howling suppression method, device, speaker and sound amplification system.
[0050] The following describes several exemplary embodiments to illustrate the technical solutions of the embodiments of the present application and the technical effects produced by the technical solutions of the present application. It should be noted that the following embodiments can refer to, draw on, or combine with each other, and the same terms, similar features, and similar implementation steps in different embodiments will not be repeated.
[0051] Figure 2 The schematic diagram of the signal transmission process of the first howling suppression method provided by the embodiment of the present application is shown. n, k represent time and frequency respectively. Among them, X(n, k) represents the near-end speech signal, Y(n, k) represents the audio signal collected by the microphone, U(n, k) represents the local amplification signal, F(n, k) represents the far-end audio signal, and E(n, k) represents the residual signal used for adaptive filter update calculation. The filter coefficient corresponding to the transmission function of the feedback path from the speaker to the microphone is represented by H(n, k), and the filter coefficient of the frequency domain adaptive filter is represented by H Est (n, k) represents, G(n, k) represents the local sound reinforcement system processing, including automatic gain control, signal amplification, power amplification, etc.
[0052] Specifically, in this embodiment, the following steps are included:
[0053] Step 1: Based on the debugging audio, the initial filter coefficients corresponding to the transfer function of the loudspeaker-to-microphone path in the sound reinforcement system can be estimated.
[0054] Step 2: Initialize the frequency domain adaptive filter according to the initial filter coefficients. The filter coefficients of the frequency domain adaptive filter are updated starting from 0.
[0055] Step 3: The audio signal collected by the microphone is converted to the frequency domain through short-time Fourier transform, and then processed frequency-by-frequency and frame-by-frame in the frequency domain.
[0056] Step 4: During the processing, the signal is vectorized. Specifically, assuming that the number of frequency points is K and the number of filters is M, the vector corresponding to the nth frame of the frequency point k in the audio signal collected by the microphone is:
[0057]
[0058] Similarly, we can obtain X(n, k), F(n, k), and U(n, k) in vector form.
[0059] The vector form of the filter coefficients is:
[0060]
[0061] Similarly, we can obtain H(n, k) in vector form.
[0062] The input signal of the microphone can be expressed as:
[0063] Y(n, k)=X(n, k)+H H (n, k)[F(n, k)+U(n-1, k)]
[0064] Among them, H H(n, k) represents the conjugate transpose of H(n, k), Y(n, k) and F(n, k) are known quantities, and the others are unknown quantities.
[0065] The residual signal can be expressed as:
[0066]
[0067] In other words, after performing echo and howling suppression on a frame of audio signals collected by a microphone, a corresponding residual signal can be obtained. That is, the residual signal of a frame of audio signals collected by a microphone is equal to the difference between the audio signal of that frame and the echo and howling cancellation amounts.
[0068] Based on the Minimum Mean Square Error (MMSE) method, |E(n, k)| 2 The expectation of H is minimized, and H Est (n, k) update formula.
[0069] However, if the MMSE criterion is directly applied to the above formula, the reference signal F(n, k) + U(n, k) has a large correlation with the filter output E(n, k), which will lead to estimation bias in the filter. Therefore, this application changes the above formula to:
[0070]
[0071] According to the MMSE criterion, the update formula of the adaptive filter is:
[0072]
[0073] The superscript * indicates conjugation, is the power spectrum of the far-end signal. γ is the update step size, which is generally updated in a variable step size manner. For details, please refer to the relevant technology and will not be described here.
[0074] The actual application scenarios include the following three situations:
[0075] (1) The far-end audio signal is not 0 and the local amplification is not turned on. In this case, G(n, k) = 0, U(n, k) = 0, and the above formula is used to calculate H Est (n, k) is updated to perform echo cancellation. The residual signal is:
[0076]
[0077] (2) When the far-end audio signal is 0 and the local amplification is on, use the above formula to update H Est (n, k) performs howling suppression, and sets H Copy(n, k) = 0 to prevent the fixed coefficient from being unable to track the changes of H(n, k) when the environment changes. The residual signal is:
[0078]
[0079] (3) The remote audio signal is not 0, and the local amplification is turned on. Est (n, k) is continuously updated using the above formula. Copy (n, k) is H Est (n, k) copy, when H Est When (n, k) is updated and meets certain conditions (for example, the echo suppression amount is greater than a certain threshold), the copy is performed.
[0080] Using the above update formula, the filter estimation deviation is theoretically 0. The specific proof process is as follows:
[0081]
[0082] Where, E{|E(n, k)| 2} means to find |E(n, k)| 2 expectations.
[0083] make:
[0084] X′(n,k)=[X(n,k)+H H (n, k)U(n-1, k)]
[0085] but:
[0086]
[0087] in, is the inverse matrix of the autocorrelation matrix, r X,F is the cross-correlation vector. It is generally assumed that the correlation between X(n, k), U(n, k) and the far-end signal F(n, k) is 0, so the filter coefficient H Est There is no deviation between (n, k) and the filter coefficient H(n, k) corresponding to the actual feedback path.
[0088] Finally, the output signal of the sound reinforcement system is:
[0089] U(n, k) = G(n, k) E(n, k)
[0090] Step 5: Convert the output signal U(n, k) obtained in step 4 back to the time domain to obtain the target audio.
[0091] The solution proposed in this embodiment of the application improves the stability of the audio processing system and adapts to changes in the sound field environment through an initial state estimation feedback path combined with real-time tracking of an adaptive filter. The filter structure used for howling suppression and echo cancellation uses the far-end signal F(n, k) as the reference signal when updating the filter coefficients, and the far-end signal F(n, k) and the near-end speech echo U(n-1, k) as the reference signal during filtering. This significantly reduces the correlation between the filter input and the reference signal, while effectively suppressing the feedback signal.
[0092] Figure 3 This is a flow chart of a howling suppression method provided in an embodiment of the present application. Figure 3 The methods shown include:
[0093] S101 : Pre-process the audio signal in the sound reinforcement system and convert the audio signal into the frequency domain.
[0094] S102. Based on the filter coefficients of the frequency-domain adaptive filter, the converted audio signal is processed frame by frame and frequency by frequency point to obtain a corresponding output signal. At the same time, the filter coefficients are updated using the far-end audio signal in the current frame signal as a reference signal for use in processing the next frame signal.
[0095] S103: Convert all the obtained output signals into the time domain to obtain the target audio.
[0096] In this embodiment, the preprocessing in S101 is a short-time Fourier transform (STFT). The specific implementation process of S101 is as follows: the audio signal is framed, typically 10 to 30 ms per frame, with a typical overlap ratio of 50%. A time-domain window function (such as a Hanning window) is selected, the window function is moved, and the time-domain audio signal is windowed. Then, a fast Fourier transform is performed to convert the time-domain signal into the frequency domain.
[0097] Accordingly, in S103, all the obtained output signals can be subjected to an inverse short-time Fourier transform (ISTFT) back to the time domain to obtain the target audio. Specifically, after the output signal is subjected to an inverse fast Fourier transform, each frame of the signal is multiplied by a window function, and then overlap-added to obtain the target audio.
[0098] Obviously, by adopting the above-mentioned method of the embodiment of the present application, the audio signal in the frequency domain is processed frame by frame and frequency by frequency point based on the filter coefficient of the frequency domain adaptive filter, and the corresponding output signal is obtained while updating the filter coefficient for use in the next frame signal processing. Since the howling suppression reuses the filter structure used in the echo cancellation, and the reference signal when the filter coefficient is updated is the far-end signal, it not only reduces the correlation between the filter input signal and the reference signal, but also can effectively suppress the feedback signal.
[0099] In an embodiment of the present application, a possible implementation is provided. Before processing the initial frame signal of any frequency point in S102, the following steps may be further included:
[0100] S100 (not shown in the drawings): determining initial filter coefficients corresponding to a transfer function from a loudspeaker to a microphone in a sound reinforcement system.
[0101] The process of processing the initial frame signal of any frequency point in S102 may specifically include: processing the initial frame signal based on the initial filter coefficients to obtain a corresponding output signal, and updating the initial filter coefficients at the same time.
[0102] Specifically, in this embodiment, the initial filter coefficients corresponding to the transfer function of the speaker-to-microphone path in the sound reinforcement system can be estimated based on the debugging audio. The estimation method can adopt an offline filter coefficient calculation method. For the sake of brevity, the specific calculation process will not be repeated here.
[0103] In an embodiment of the present application, the adaptive filter used is a frequency domain blocking filter, which includes a group of filters. The initial frame can be determined based on the number of filters in the group. For example, if the group of filters includes 10, the initial frame of any frequency point is the 10th frame. When processing the 10th frame signal, the 1st to 10th frame signals are required. While processing the 10th frame signal to obtain the corresponding output signal, it is necessary to update the coefficients of the initial filter to obtain the first filter coefficients for use in processing the 11th frame signal. Similarly, when processing the 11th frame signal, it is necessary to use the 2nd to 11th frame signals. While processing the 11th frame signal to obtain the corresponding output signal, it is necessary to update the coefficients of the first filter for use in processing the 12th frame signal, and thus the audio signal is processed frame by frame and frequency point by frequency point.
[0104] It should be noted that in this embodiment, each frame of signal is processed with reference to the M frames of signal preceding the frame of signal, which can make the echo and howling elimination more clean, that is, the howling suppression effect is better. Wherein, M is an integer greater than 1 and is the number of filters in the filter structure.
[0105] In the above embodiment, the initial state estimation feedback path is combined with the adaptive filter real-time tracking, which not only improves the stability of the audio processing system but also enables it to adapt to changes in the sound field environment.
[0106] A possible implementation method is provided in an embodiment of the present application. The process of processing the non-initial frame signal of any frequency point in S102 includes: processing the current frame signal based on the filter coefficient updated after processing the previous frame signal to obtain the corresponding output signal, and updating the filter coefficient at the same time.
[0107] In this embodiment, if the current frame signal is the 20th frame signal of a certain frequency point, the filter coefficient obtained after processing the 19th frame signal of the frequency point is updated, and the 20th frame signal is processed to obtain the corresponding output signal. At the same time, the filter coefficient is updated for use when processing the 21st frame signal.
[0108] In the above embodiment, the audio signal in the sound reinforcement system includes a remote audio signal and an audio signal collected by a microphone. The process of processing each frame signal of any frequency point to obtain a corresponding output signal in S102 includes:
[0109] The audio signal collected by a microphone is subjected to echo suppression, howling suppression, and local sound amplification processing to obtain a corresponding output signal.
[0110] Specifically, in this embodiment, echo suppression and howling suppression are performed on a frame of audio signals collected by a microphone to obtain a corresponding residual signal, and local amplification is performed on the residual signal to obtain a corresponding output signal.
[0111] Specifically, the residual signal of the audio signal collected by the microphone in the current frame signal can be determined based on the current frame signal and the corresponding filter coefficients, as well as the output signal of the previous frame. For example, for the nth frame signal at frequency k, the corresponding residual signal can be determined according to the following formula:
[0112]
[0113] U(n,k)=G(n,k)E(n,k) where Y(n,k) represents the nth frame of the frequency k in the audio signal collected by the microphone, F(n,k) represents the nth frame of the remote audio signal of the frequency k, and H Est (n, k) is the filter coefficient corresponding to the n-th frame signal of frequency point k, and U(n-1, k) represents the output signal after processing the n-1-th frame signal of frequency point k.
[0114] It should be noted that, in this embodiment, during the calculation of the residual signal, if n=1, then U=0, that is, when processing the first frame of input signal, the speaker has no output signal yet.
[0115] Based on the determination of the residual signal, the corresponding output signal can be determined according to the following formula:
[0116] U(n, k) = G(n, k) E(n, k)
[0117] Wherein, G(n, k) represents the local sound reinforcement system processing of the residual signal E(n, k) of the nth frame at frequency point k, including automatic gain control, signal amplification, power amplification, etc., and U(n, k) represents the output signal after the processing of the signal of the nth frame at frequency point k.
[0118] An embodiment of the present application provides a possible implementation method. The process of updating the filter coefficient using the far-end audio signal in the current frame signal as a reference signal in S102 may specifically include:
[0119] Based on a preset update step size, the current frame signal and the corresponding filter coefficients, the residual signal, and the far-end audio signal in the current frame signal, new filter coefficients are determined and the filter coefficients corresponding to the current frame signal are updated.
[0120] Specifically, in this embodiment, the filter coefficients may be updated according to the following formula:
[0121]
[0122] Among them, E * (n, k) represents the conjugate vector of the residual signal E(n, k) of the n-th frame signal at frequency k. is the power spectrum of the far-end audio signal, γ is the update step size, F(n, k) represents the nth frame of the far-end audio signal at frequency k, H Est (n, k) is the filter coefficient corresponding to the nth frame signal of frequency k, H Est (n+1, k) is the filter coefficient corresponding to the (n+1)th frame signal of frequency point k.
[0123] It should be noted that, in this embodiment, the minimum mean square error (MMSE) method can be used to make |E(n, k)| determined by the following formula for determining the residual signal 2 The expectation of is minimized, and the formula for updating the filter coefficients is obtained.
[0124] In some feasible embodiments of the present application, the process of determining the residual signal includes:
[0125] If the echo suppression amount of the current frame signal is greater than or equal to the preset threshold, the filter coefficient corresponding to the current frame signal is copied, and a residual signal is determined based on the current frame signal, the corresponding filter coefficient and the previous frame output signal.
[0126] If the echo suppression amount is less than a preset threshold, a residual signal is determined based on the last copied filter coefficient, the current frame signal and the corresponding filter coefficient, and the last frame output signal.
[0127] The amount of echo suppression is determined based on the power ratio of the output signal to the audio signal collected by the microphone.
[0128] Specifically, in this embodiment, the residual signal can be determined according to the following formula:
[0129]
[0130] in, Indicates the echo suppression amount of the nth frame signal at frequency k. If the echo suppression amount is greater than or equal to the preset threshold, Copied Among them, U(n-1, k) represents the output signal after the n-1th frame signal of frequency point k is processed, and Y(n, k) represents the nth frame of frequency point k in the audio signal collected by the microphone.
[0131] If the echo suppression amount of the nth frame signal at frequency k is less than the preset threshold, then The last copied value is used. It should be noted that the last copied value refers to the filter coefficient copied when the echo suppression amount of the ni-th frame signal is greater than or equal to the preset threshold, where i=1, 2, 3, ..., n-1.
[0132] It should be noted that, in this embodiment, the amount of echo suppression of the nth frame signal of frequency point k can be determined based on the power ratio of the output signal corresponding to the multiple frame signals before the nth frame signal of frequency point k and the audio signal collected by the multiple frame microphones. For example, the amount of echo suppression of the 10th frame signal can be based on the power ratio of the output signal corresponding to the 1st to 9th frame signals and the 1st to 9th frame signals collected by the microphone. Assuming that the number of filters in the frequency domain adaptive filter structure used in this embodiment is 10, the 1st to 9th frame signals are not processed using the method in the embodiment of the present application, and can be processed using related technologies to obtain corresponding output signals.
[0133] For another example, the amount of echo suppression for the 15th frame signal can be based on the power ratio of the output signal corresponding to the 6th to 14th frame signals to the 6th to 14th frame signals collected by the microphone. Assuming that the number of filters in the frequency-domain adaptive filter structure used in this embodiment is 10, the 6th to 9th frame signals are not processed using the method in the embodiment of the present application, but can be processed using related technologies to obtain corresponding output signals. The 10th to 14th frame signals can be processed using the method in the embodiment of the present application to obtain corresponding output signals.
[0134] In summary, the howling suppression method provided in the embodiments of the present application, through an initial state estimation feedback path combined with real-time tracking of an adaptive filter, not only improves the stability of the audio processing system but also adapts to changes in the sound field environment. Howling suppression reuses the filter structure used in echo cancellation. When updating the filter coefficients, the reference signal is the far-end signal, and during filtering, the reference signal is the echo of the far-end signal and the near-end speech. This significantly reduces the correlation between the filter input and the reference signal, while effectively suppressing the feedback signal.
[0135] The present application also provides a howling suppression device, comprising: a local sound reinforcement system, an adder, a frequency-domain adaptive filter, and a short-time Fourier transform module and an inverse transform module. The input of the frequency-domain adaptive filter is connected to a speaker, the output of the frequency-domain adaptive filter is connected to the input of the adder, the output of the adder is connected to the input of the local sound reinforcement system, and the output of the local sound reinforcement system is connected to the speaker.
[0136] The short-time Fourier transform (STFT) module preprocesses the audio signal collected by the microphone, converting it into the frequency domain. The frequency-domain adaptive filter performs echo and howling suppression on the converted audio signal before outputting it to the adder. The adder subtracts the converted audio signal from the signal output by the frequency-domain adaptive filter and outputs it to the local sound reinforcement system. The local sound reinforcement system performs local sound reinforcement processing on the received signal before transmitting it to the inverse short-time Fourier transform (ISFT) module. The inverse short-time Fourier transform (ISFT) module converts the received signal into the time domain, generating the target audio signal to drive the speakers for playback.
[0137] An embodiment of the present application provides a sound system comprising: a loudspeaker and the howling suppression device provided in the above embodiment. The loudspeaker is connected to the input of a frequency-domain adaptive filter in the howling suppression device, the output of the frequency-domain adaptive filter is connected to the input of an adder in the howling suppression device, the output of the adder is connected to the input of a local sound reinforcement system in the howling suppression device, and the output of the local sound reinforcement system is connected to the loudspeaker.
[0138] The audio signal collected by the microphone can be transmitted to the speaker in wireless or wired ways. For example, the speaker in this embodiment can be a Bluetooth speaker, which transmits the audio signal data to the microphone via Bluetooth, or it can be connected to the microphone via WiFi or other local area network access methods. When the microphone collects the audio signal, it is transmitted to the howling suppression device in the speaker. After the howling suppression device performs howling analysis on the audio signal, the audio signal and the generated reference signal are sent to the speaker for playback.
[0139] An embodiment of the present application provides a sound amplification system, comprising: a microphone, a speaker, and the howling suppression device provided in the above embodiment. The howling suppression device is arranged between the microphone and the speaker, and is used to receive the audio signal collected by the microphone and output the generated target audio to the speaker for playback.
[0140] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similarities between the various embodiments can be referred to in conjunction with each other. For device embodiments, since they are generally similar to method embodiments, their description is relatively simple, and for relevant details, reference can be made to the description of the method embodiments.
[0141] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of the devices, methods, and computer program products according to the multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment, or a portion of code, and the module, program segment, or a portion of code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the boxes can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or can be implemented using a combination of dedicated hardware and computer instructions.
[0142] In addition, the functional modules in each embodiment of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0143] If the function is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0144] The above are merely examples of the present application and are not intended to limit the scope of protection of the present application. For those skilled in the art, the present application may have various modifications and variations. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application. It should be noted that similar numbers and letters represent similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined or explained in subsequent figures.
[0145] The above are only optional implementation methods for some implementation scenarios of this application. It should be pointed out that for ordinary technicians in this technical field, without departing from the technical concept of the solution of this application, other similar implementation methods based on the technical ideas of this application also fall within the protection scope of the embodiments of this application.
Claims
1. A howling suppression method, characterized in that: include: Preprocessing the audio signal in the sound reinforcement system and converting the audio signal into the frequency domain; Based on the filter coefficients of the frequency-domain adaptive filter, the converted audio signal is processed frame by frame and frequency by frequency point to obtain the corresponding output signal. At the same time, the filter coefficients are updated using the far-end audio signal in the current frame signal as a reference signal for use in the next frame signal processing; Convert all the obtained output signals into the time domain to obtain the target audio; The audio signal in the sound reinforcement system includes a remote audio signal and an audio signal collected by a microphone. The process of processing each frame signal at any frequency point to obtain a corresponding output signal includes: Determine a residual signal of the audio signal collected by the microphone in the current frame signal based on the current frame signal and the corresponding filter coefficients, as well as the output signal of the previous frame; Performing local amplification processing on the residual signal to obtain a corresponding output signal; The process of determining the residual signal includes: If the echo suppression amount of the current frame signal is greater than or equal to a preset threshold, copy the filter coefficient corresponding to the current frame signal, and determine the residual signal based on the current frame signal, the corresponding filter coefficient, and the previous frame output signal; If the echo suppression amount is less than the preset threshold, determining the residual signal based on the last copied filter coefficient, the current frame signal and the corresponding filter coefficient, and the last frame output signal; The echo suppression amount is determined based on a power ratio of the output signal to the audio signal collected by the microphone.
2. The method according to claim 1, characterized in that Before processing the initial frame signal of any frequency point, the method further includes: determining initial filter coefficients corresponding to a transfer function of a loudspeaker-to-microphone path in the sound reinforcement system; The process of processing the initial frame signal at any frequency point includes: Based on the initial filter coefficients, the initial frame signal is processed to obtain a corresponding output signal, and the initial filter coefficients are updated at the same time; The initial frame is determined based on the number of the frequency domain adaptive filters.
3. The method according to claim 2, characterized in that The process of processing the non-initial frame signal at any frequency point includes: Based on the updated filter coefficients obtained after processing the previous frame signal, the current frame signal is processed to obtain a corresponding output signal, and the filter coefficients are updated at the same time.
4. The method according to claim 1, wherein The process of updating the filter coefficients using the far-end audio signal in the current frame signal as a reference signal includes: Based on a preset update step size, the current frame signal and the corresponding filter coefficients, a residual signal, and a far-end audio signal in the current frame signal, new filter coefficients are determined and the filter coefficients corresponding to the current frame signal are updated.
5. The method according to any one of claims 1 to 4, characterized in that If the preprocessing is short-time Fourier transform, converting all the obtained output signals into the time domain to obtain the target audio includes: All the obtained output signals are subjected to inverse short-time Fourier transform back to the time domain to obtain the target audio.
6. A howling suppression device, characterized in that: include: A local sound reinforcement system, an adder, a frequency domain adaptive filter, and a short-time Fourier transform module and an inverse transform module, wherein the input of the frequency domain adaptive filter is connected to the speaker, the output of the frequency domain adaptive filter is connected to the input of the adder, the output of the adder is connected to the input of the local sound reinforcement system, and the output of the local sound reinforcement system is connected to the speaker; The short-time Fourier transform module is used to pre-process the audio signal collected by the microphone and convert the audio signal into the frequency domain; The frequency domain adaptive filter is used to perform echo suppression and howling suppression processing on the converted audio signal and then output it to the adder; The adder is used to subtract the converted audio signal from the signal output by the frequency domain adaptive filter and output the resultant signal to the local sound reinforcement system; The local sound amplification system performs local sound amplification processing on the received signal and then transmits it to the inverse short-time Fourier transform module; The inverse short-time Fourier transform module is used to convert the received signal into the time domain, obtain the target audio and transmit it to the speaker for playback; The audio signal in the sound reinforcement system includes a remote audio signal and an audio signal collected by a microphone. The adder is used to process each frame signal of any frequency point to obtain a corresponding output signal. The process includes: determining a residual signal of the audio signal collected by the microphone in the current frame signal based on the current frame signal and the corresponding filter coefficients, as well as the previous frame output signal; The local sound amplification system performs local sound amplification processing on the residual signal to obtain a corresponding output signal; The process of determining the residual signal includes: If the echo suppression amount of the current frame signal is greater than or equal to a preset threshold, copy the filter coefficient corresponding to the current frame signal, and determine the residual signal based on the current frame signal, the corresponding filter coefficient, and the previous frame output signal; If the echo suppression amount is less than the preset threshold, determining the residual signal based on the last copied filter coefficient, the current frame signal and the corresponding filter coefficient, and the last frame output signal; The echo suppression amount is determined based on a power ratio of the output signal to the audio signal collected by the microphone.
7. A sound system, characterized in that: include: A speaker, and a howling suppression device as claimed in claim 6, wherein: The loudspeaker is connected to the input end of the frequency domain adaptive filter in the howling suppression device, the output end of the frequency domain adaptive filter is connected to the input end of the adder in the howling suppression device, the output end of the adder is connected to the input end of the local sound amplification system in the howling suppression device, and the output end of the local sound amplification system is connected to the loudspeaker.
8. A sound amplification system, characterized in that: include: A microphone, a speaker, and the howling suppression device as claimed in claim 6, wherein the howling suppression device is arranged between the microphone and the speaker, and the howling suppression device is used to receive the audio signal collected by the microphone and output the target audio obtained by signal processing to the speaker for playback.
Citation Information
Patent Citations
Echo cancellation method, device and equipment and storage medium
CN111199748A