A Method for Reverberation and Feedback Suppression in Indoor Sound Reinforcement Systems Based on Speech Cepstral Analysis
By employing a feedback suppression method based on speech cepstral analysis, combined with speech activity detection and differential cepstral spike probability statistics, the parameters of the feedback suppression filter are updated in real time. This solves the problems of poor reverberation and feedback suppression in existing technologies, and improves the stability and speech intelligibility of indoor sound reinforcement systems.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HANGZHOU CANGHAI GUANZHI TECHNOLOGY CO LTD
- Filing Date
- 2026-02-09
- Publication Date
- 2026-05-26
AI Technical Summary
In existing indoor sound reinforcement systems, reverberation and feedback suppression methods such as notch filtering, frequency shifting, and adaptive feedback suppression (AFC) cannot effectively improve the system's stable gain. Furthermore, the adaptive feedback suppression method (AFC-CE) suffers from inaccurate estimation of the feedback path impulse response parameters due to the influence of the correlation of the speech input signal, resulting in poor feedback suppression performance.
A speech cepstral analysis-based approach is adopted. By combining feedback suppression filtering with speech activity detection, differential cepstral spike probability statistics, and estimation of the room impulse response function of the acoustic feedback path, the parameters of the feedback suppression filter are updated in real time. The optimal feedback suppression filter parameters are selected by using the gain-corrected audio input-output variance ratio.
It significantly improves reverberation and feedback suppression, enhances system stability gain, reduces the negative impact on speech intelligibility, and effectively suppresses ringing and howling.
Smart Images

Figure CN121665169B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of indoor sound reinforcement technology, specifically relating to a method for reverberation and feedback suppression in indoor sound reinforcement systems based on speech cepstral analysis. Background Technology
[0002] In indoor sound reinforcement systems, there is a direct acoustic feedback path and multiple reflection paths between the microphone and the speaker. The superposition of sound due to this multipath feedback is called reverberation. Reverberation introduces color distortion and distortion into the sound source signal, and prolonged reverberation negatively impacts speech intelligibility. Regenerative reverberation in indoor sound reinforcement systems exacerbates this negative impact, making it difficult for listeners to hear the speaker clearly. More severe acoustic feedback produces ringing and howling, rendering the sound reinforcement system unusable.
[0003] Current technical solutions include notch filtering, frequency shifting, and adaptive feedback suppression (AFC). Notch filtering controls howling by reducing the gain at the frequency where howling occurs. Frequency shifting suppresses howling by changing the frequency relationship between the microphone signal and the speaker output signal. Neither of these methods improves the reduction of source signal distortion, nor does it significantly improve the system's stability gain. Adaptive feedback suppression (AFC) originates from the adaptive echo cancellation (AEC) algorithm in communication systems. Its main steps involve iteratively obtaining an estimate of the acoustic feedback path impulse response through the least mean square (LMS) algorithm and eliminating this acoustic feedback signal from the microphone signal. The main problem with this method is the correlation between the feedback signal and the speech input signal, which leads to a bias in the estimation of the impulse response parameters of the feedback path. Even after eliminating this correlation using prediction error method (PEM), the accuracy of the feedback path impulse response estimation generated by adaptive filtering is still lower than that using cepstral analysis.
[0004] Cepstral analysis was initially applied to echo detection of seismic waves. After continuous improvement and development, this method has been widely used to process convolutional signals. Acoustic feedback is the superposition of the convolution of the source signal and the impulse response function of the feedback path; therefore, estimating the impulse response function of the acoustic feedback channel using cepstral analysis is theoretically feasible. The papers *Acoustic Feedback and Echo Cancellation in Speech Communication Systems* and *On the use of cepstral analysis in acoustic feedback cancellation* propose a method for using cepstral analysis to adaptive feedback suppression of differential signals in public address systems. Through multiple asymptotic iterations, a relatively accurate estimate of the feedback path impulse response function is obtained, thereby eliminating the influence of acoustic feedback from the microphone signal and improving the system's stability gain. This method is called Differential Cepstral Adaptive Feedback Suppression (AFC-CE).
[0005] The Adaptive Feedback Suppression (AFC-CE) algorithm directly ignores the effects of the input speech cepstral spectrum and the high-order convolution of the difference function. Although the progressive approach can eliminate some of the interference of the speech cepstral spectrum on the estimation of the impulse response function of the acoustic feedback path, this elimination is limited by the forgetting factor, and therefore its feedback suppression effect still needs to be improved. Summary of the Invention
[0006] In view of the above, the purpose of this invention is to provide a method for reverberation and feedback suppression in indoor sound reinforcement systems based on speech cepstral analysis, comprising the following steps:
[0007] Collect and store microphone audio data from indoor sound reinforcement systems;
[0008] The audio data stored by feedback suppression filtering is processed to obtain differential data frames. The differential data frames are amplified by gain control, stored and output as audio. The parameters of the feedback suppression filter used in the feedback suppression filtering process are updated in real time after being estimated by speech activity detection, differential cepstral peak probability statistics and room impulse response function of acoustic feedback path.
[0009] The feedback suppression effect is evaluated in real time using the gain-corrected audio input-output variance ratio, and the optimal feedback suppression filter parameters are selected based on the feedback suppression effect.
[0010] Preferably, a computer ring buffer is used to store audio data, wherein the length of the ring buffer should accommodate at least one input / output audio data frame length plus half the length of the differential cepstral calculation data frame;
[0011] Among them, half the length of the differential cepstral calculation data frame should be greater than the total acoustic feedback delay length plus the length of the impulse response function, and should be an integer multiple of the input / output data frame;
[0012] The total acoustic feedback delay includes the propagation delay caused by the direct propagation path of sound from the speaker to the microphone, as well as the input and output delay of the audio processor.
[0013] The length of the impulse response function specifically refers to the finite length of the room impulse response starting from the direct acoustic feedback response impulse.
[0014] Preferably, the speech activity detection is used to determine whether the differential data frame after feedback suppression filtering contains a speech signal. Specifically, the energy threshold method can be used to determine this, and the differential data frame containing the speech signal is stored in the cepstral analysis buffer for differential cepstral analysis.
[0015] Preferably, differential cepstral spike probability statistics are used to estimate the primary update of the feedback suppression filter parameters, including the following process:
[0016] Transformation processing steps: Randomly select differential data frames containing speech signals from the cepstral analysis buffer, and sequentially perform windowing, Fourier transform, logarithmic transform, and inverse Fourier transform to obtain the absolute amplitude and angle values of the differential cepstral. Divide the absolute amplitude values into positive and negative groups according to the angle values.
[0017] After repeating the above transformation process multiple times, the mean absolute value of the amplitude and the probability of local spikes at each point in the two sets of data are calculated. Then, some data in the probability of local spikes are set to zero and multiplied pairwise with the mean absolute value of the amplitude and subtracted to obtain the initial feedback filter parameters.
[0018] By combining the impulse response position of the direct acoustic feedback and the total delay of the acoustic feedback, some data in the initial feedback filter parameters are set to zero, and the amplification ratio coefficients in the remaining initial feedback filter parameter values are removed to obtain the primary feedback suppression filter parameters.
[0019] Preferably, setting a portion of the data in the probability of local spikes to zero includes:
[0020] Set the peak probability values that are less than a preset threshold to zero, and at the same time set the peak probability values whose left side length of the frequency inversion axis is not less than the input-output delay of the audio processor to zero.
[0021] Preferably, by combining the impulse response position of the direct acoustic feedback and the total delay of the acoustic feedback, some data in the initial feedback filter parameters are set to zero, including:
[0022] The position of the maximum value in the initial feedback filter parameters is determined as the impulse response position of the direct acoustic feedback. Starting from this position, a window with a length not greater than the total delay length of the acoustic feedback is set, and the initial feedback filter parameter values outside the window are set to zero.
[0023] Preferably, the initial feedback filter parameters are obtained by multiplying and subtracting the pairwise product of the peak probabilities after setting both positive and negative sets to zero and the mean absolute value of the amplitude of the differential cepstrum, including:
[0024] h pre (n) = P pk+ (n)·M c+ (n) - P pk- (n)·M c- (n)
[0025] Among them, P pk+ (n) and P pk- (n) represents the positive and negative peak probabilities, respectively, M c+ (n) and M c- (n) represents the mean of the absolute values of the amplitudes at each point in the positive and negative cepstrum, where n is the index of each point.
[0026] Preferably, the estimation of the room impulse response function of the acoustic feedback path is used for iterative updates of the feedback suppression filter parameters in rounds other than the initial update, including the following process:
[0027]
[0028] in, This represents the parameters of the secondary feedback suppression filter updated in round j+1. This represents the feedback suppression filter parameters updated in round j. When j=1, The parameter h of the primary feedback suppression filter is represented. pri (n), where α represents the amplitude parameter used to adjust the approximation in each round. This represents the residual estimate of the impulse response function at round j+1;
[0029] in, The estimation process is as follows: using the primary feedback suppression filter parameter h pri The estimation method of (n) applies to the parameters of the feedback suppression filter. The processed differential data frames are then estimated.
[0030] Preferably, when amplifying differential data frames through gain control, the gain value is adjusted by gain step increment;
[0031] The real-time update of the feedback suppression filter parameters is repeated multiple times with each increase in gain step increment.
[0032] Preferably, the feedback suppression effect is evaluated in real time using the gain-corrected audio input-output variance ratio, and the optimal feedback suppression filter parameters are selected based on the feedback suppression effect, including:
[0033]
[0034] h(n)* = argmin h (ESR g )
[0035] Among them, Y 2 E represents the sum of squares of the audio input data frames. 2 This represents the sum of squares of the differential data frames output after the Y data frame has undergone feedback suppression filtering. Indicates gain. The parameter h(n)* represents the feedback suppression effect, and argmin represents the optimal feedback suppression filter parameters. h (ESR g ) indicates that the feedback suppression effect ESR is achieved. g The minimum parameters for the feedback suppression filter.
[0036] Compared with the prior art, the beneficial effects of the present invention include at least the following:
[0037] The feedback suppression filter parameters used in this invention are updated in real time after speech activity detection, differential cepstral peak probability statistics, and estimation of the impulse response function of the acoustic feedback path. In other words, the mean of the differential cepstral is weighted by the peak probability of the differential cepstral, and the acoustic feedback suppression effect is monitored and optimized in real time during the iteration process. This significantly improves the estimation accuracy of the room impulse response function, thereby improving the reverberation and feedback suppression effect. Attached Figure Description
[0038] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0039] Figure 1 This is a system principle block diagram of the AFC-CE method provided in the embodiment;
[0040] Figure 2 This is a flowchart of the reverberation and feedback suppression method for indoor sound reinforcement systems based on speech cepstral analysis provided in the embodiment.
[0041] Figure 3This is the cepstral mean of the reverberant speech provided in the embodiment;
[0042] Figure 4 This is a simulation graph of the forgetting factor and h(n) asymptotic regression provided in the example;
[0043] Figure 5 The example provides a comparison chart of 100 asymptotic iterations and 100 mean cepstrum values for speech when λ is 0.7, 0.8, and 0.9. The yellow line represents the asymptotic method, and the blue line represents the mean method.
[0044] Figure 6 This is a comparison of the speech cepstral after 151,227,458 asymptotic iterations with the mean of the speech cepstral after 100 iterations when λ is 0.97, 0.98, and 0.99, as provided in the example. The yellow line represents the asymptotic method, and the blue line represents the mean method.
[0045] Figure 7 This is a comparison chart of the asymptotic cepstral values of speech at different iterations when λ=0.99 and the average cepstral value when n=100, provided in the example. In the chart, the yellow line represents the asymptotic method and the blue line represents the mean method.
[0046] Figure 8 The examples show the AFC-CE method provided in the embodiments before and after ESRg optimization, as well as the ARC-CE method of the present invention. ppwm The result after 1000 asymptotic iterations. Detailed Implementation
[0047] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and do not limit the scope of protection of this invention.
[0048] In the existing AFC-CE method, the principle block diagram of the loudspeaker system is as follows: Figure 1 As shown in the diagram. Here, u is the speaker's voice signal picked up by the microphone, and the microphone output signal y, in addition to the u signal, also includes the acoustic feedback signal propagating from the speaker output through the acoustic feedback path to the microphone. F is the acoustic feedback path. H is an adaptive feedback suppression filter. e is the differential signal of the microphone output signal y after acoustic feedback is eliminated by H, G is the total gain of the amplification system, and D is the total delay of the feedback path, including the input and output delays of the amplification system and the audio processor. i / o Delay D of direct acoustic feedback path s / m .
[0049] The differential digital input signal e(n) = y(n) - h(n) * x(n) is simply called the differential signal. y(n) is the digital signal output from the microphone, h(n) is the parameter sequence of the adaptive feedback suppression filter, x(n) is the digital signal output to the speaker, and * represents the convolution operator. The real-valued differential cepstral of e(n) is... e (n) is the sum of two parts, the first part being the real cepstral of the input speech signal. u (n), the second part is the k-fold convolution sum of the 1 / k weighted system gain g(n), system delay d(n), and [f(n)-h(n)], and the difference cepstral expression is:
[0050] c e (n) = c u (n)+
[0051] Here, [f(n)-h(n)] is simply called the difference function. The system gain g(n) is adjustable. If the loudspeaker system has undergone good equalization, g(n) can also be represented by a constant G. The system delay d(n) can be expressed as a sequence of unit pulses of a single pulse. For example, a delay of D sampling periods can be expressed as the sequence d(n) = [0 0 …0 1], where the number of 0s is D. e (n) is abbreviated as differential cepstral, where the speech cepstral is... u (n) is a random function, while the difference function is relatively stable. The speech cepstral plays an interfering role in the difference cepstral analysis, which is equivalent to random noise.
[0052] When the indoor space is large and the reverberation is long, the room impulse response function f(n) of the acoustic feedback path is expressed using a longer impulse sequence length. f(n) can be divided into three components: direct feedback, early reverberation, and late reverberation. Direct feedback is the impulse response generated by the straight feedback path between the speaker and the microphone, and its amplitude is the largest. Early reverberation and late reverberation are feedback paths formed after one or more reflections, and their feedback amplitude decreases exponentially with the increase of the number of reflections and the path length.
[0053] The existing AFC-CE method mentioned above has speech cepstral limitations. u (n) interferes with the estimated value of f(n). Therefore, this embodiment of the invention provides a method for reverberation and feedback suppression of indoor sound reinforcement system based on speech cepstral analysis. The purpose of the cepstral analysis is to use the estimated value of impulse response function f(n) as h(n) to gradually minimize [f(n)-h(n)], and to combine h(n) and delay D as parameters of feedback suppression filter to remove acoustic feedback components in the input signal. The impulse response function f(n) specifically refers to direct feedback and the impulse response sequence of a certain length thereafter. Specific improvements include: (1) controlling the delay D of the audio processing system.i / o Make the difference cepstral expression c e The difference function [f(n)-h(n)] in (n) is separated from the left head region with greater cepstral energy and the fundamental frequency cepstral signal on the cepstral axis to reduce interference. (2) The probability and statistics method is used to improve the estimation accuracy of f(n). (3) Real-time evaluation feedback suppression performance is added to optimize the estimation process.
[0054] This invention provides a method for reverberation and feedback suppression in indoor sound reinforcement systems based on speech cepstral analysis. The indoor sound reinforcement system includes an audio processor and connected microphones and speakers. The audio processor includes a computer system, sound card hardware for audio processing, and audio processing software. The speakers include frequency band equalization and power amplification functions. The microphones and speakers are fixedly installed indoors, therefore their feedback paths are also fixed relative to the speech cepstral analysis. u In this scenario, the impulse response function f(n) of the feedback path is stable or changes slowly. Therefore, the update rate of the feedback suppression filter parameters can be much lower than the feedback suppression filtering processing speed of the audio input data frames.
[0055] like Figure 2 As shown in the figure, an embodiment of the present invention provides a method for reverberation and feedback suppression in an indoor sound reinforcement system based on speech cepstral analysis, comprising the following steps:
[0056] S1: Initialize and collect audio data from the indoor sound reinforcement system and store it.
[0057] In this embodiment, the initialization process includes setting the size of the audio processor's input / output data buffer. One purpose of setting and adjusting the buffer size is to control the audio processor's input / output latency D. i / o The larger the buffer, the longer the data frames read and written in the audio stream, and the greater the audio processing latency. Besides the data buffer, factors affecting audio processor latency include the complexity of the audio processing algorithm and the computer's performance. Actual latency can be measured using system identification methods.
[0058] Besides the input / output delay D caused by audio processing i / o In addition, the total acoustic feedback delay D also includes the delay D caused by the direct propagation distance from the speaker to the microphone. s / m That is, D = D i / o + D s / m Among them, delay D s / m It can be estimated based on the propagation distance and the speed of sound propagation.
[0059] In practice, D should be controlled within the range of 20 to 50 milliseconds, i.e., 20ms ≤ D ≤ 50ms. This is because a delay greater than 50 milliseconds results in a noticeable auditory difference, which may affect the listener's experience. When the delay is less than 20 milliseconds, f(n) is located in the cepstral region of the speech spectrum. u (n) is a region with high energy and may overlap with the cepstral position of the speech fundamental frequency, which will cause significant interference to the estimation of f(n).
[0060] After initialization, the audio acquisition module obtains audio data frames through the sound card's audio input buffer connected to the microphone. These audio data frames are stored in a circular buffer on the computer, with a length L. rbuf It should accommodate at least one input / output audio data frame of length L. i / o Adding differential cepstral calculations to half the data frame length L c / 2 . Referred to as: L rbuf ≥ L i / o + L c / 2 .
[0061] Among them, half the length L of the differential cepstral calculation data frame c / 2 It should be greater than the total acoustic feedback delay length D plus the sequence length L of the impulse response function f(n). f And for input / output data frames L i / o Integer multiples of L are denoted as: L c / 2 ≥ D + L f L c / 2 = N·L i / o N is a positive integer; the length of the impulse response function specifically refers to the finite length of the room impulse response starting from the direct acoustic feedback response impulse.
[0062] The data frames of the audio input buffer are stored sequentially in the circular buffer. The parameters of the circular buffer and the feedback suppression filter are set to zero during initialization.
[0063] S2, the feedback suppression filter processes the stored audio data to obtain differential data frames, and the differential data frames are amplified, stored and output as audio through gain control.
[0064] In this embodiment, feedback suppression filtering is performed after each frame of audio data is stored. The specific steps are as follows:
[0065] (1) Reverse the order of the parameters of the feedback suppression filter. The expression for the feedback suppression filter H(q,n) is:
[0066] H(q,n) = h1(n)q -1 + h2(n)q -2 + … + h Dn+Lf (n)q-(Dn+Lf)
[0067] Where q is the delay operator, for example: q -1 ·y(n) = y(n-1), where Dn represents the number of sampling cycles for the total acoustic feedback delay, and Lf is the same as L f Let h represent the sequence length of f(n), and the data sequence after reversing the order of the feedback suppression filter parameters be denoted as: rev (n) = [h Dn+Lf [(n),…,h2(n),h1(n)];
[0068] (2) Read an audio data frame Y(n) from the input / output buffer and store it in the circular buffer in order, denoted as Y(n) = [y(n) y(n+1) … y(n+L)]. i / o -1)];
[0069] (3) The Dn+Lf data points before the first data y(n) of the input buffer data frame stored in the circular buffer are compared with the feedback suppression filter parameter h. rev (n) Multiply each pairwise and sum them, subtract the result from the first data y(n), and update the difference at the position of the first data y(n). This is denoted as: Differential data frame e(n) = y(n) - [h1·e(n-1) + h2·e(n-2) + ... + h Lc / 2 ·e(n-(Dn+Lf))], repeat this step until all data of Y(n) is updated; the differential data frame E(n) after feedback suppression filtering is = [e(n) e(n+1) … e(n+L)] i / o -1)].
[0070] (4) After gain control of E(n), X(n) is generated and sent to the audio output buffer for playback;
[0071] (5) Read the next input buffer data frame and repeat steps (2) to (4) above.
[0072] In this embodiment, the gain control process performs a proportional amplification of the differential data frame E(n), i.e., X(n) = G·E(n). As can be seen from the above differential cepstral expression, a larger gain value G enhances the effect of the differential function [f(n)-h(n)] on the speech cepstral. u The salience of (n) is more conducive to the accurate estimation of the difference function. Therefore, the value of G can be gradually increased as the feedback suppression filter parameter h(n) asymptotically approximates the feedback path impulse response function f(n). This is denoted as:
[0073] G m+1 = G m + ΔG
[0074] Where m is the gain step order, m = 0, 1, ... m max G0 is the initial gain, which is the gain that maintains system stability when h(n) is zero. ΔG is the gain step increment. The gain step can be linear, meaning ΔG remains constant, or it can be nonlinearly compressed, meaning when G... m When it reaches a certain value, ΔG decreases, or ΔG decreases with G. m The value decreases as G increases. Nonlinear compression helps maintain system stability during parameter estimation of the feedback suppression filter. m As the value gradually approaches a certain system stable gain, ΔG can decrease to 0.
[0075] In the embodiment, the feedback suppression filter parameter h(n) is iteratively updated after being estimated by speech activity detection, differential cepstral spike probability statistics, and the impulse response function of the acoustic feedback path.
[0076] In this embodiment, voice activity detection is used to determine whether the differential data frame E(n) after feedback suppression filtering contains a voice signal. This is because signals without voice are easily blocked by the noise reduction function of the loudspeaker system and cannot be subjected to differential cepstral analysis. Specifically, an energy threshold method can be used to determine this, and differential data frames containing voice signals are stored in a cepstral analysis buffer for differential cepstral analysis. The length of the cepstral analysis buffer can hold several seconds (e.g., 3 to 5 seconds) of data frames E(n) containing voice.
[0077] In this embodiment, the purpose of differential cepstral spike probability statistics is to obtain information about the difference function [f(n)-h(n)] from the differential cepstral and to minimize interference from the speech cepstral. Since the speech cepstral is a random function, while [f(n)-h(n)] is a relatively stable function, the amplitude fluctuations of the speech cepstral in a single differential cepstral calculation are significant, especially since some data frames contain the fundamental frequency cepstral, which are the main sources of interference when resolving the difference function. Using statistical observation methods can greatly reduce this interference. Therefore, the method of this invention uses relatively long speech segments to randomly select multiple data frames for cepstral statistics.
[0078] from Figure 3The comparison shows the difference between a single speech cepstral and the average of multiple speech cepstral values. n=1 represents a randomly selected speech segment from a continuous speech stream, and n=100 represents 100 randomly selected speech segments from that stream, with the average cepstral values calculated. It can be seen that the average cepstral values at points to the right of the 20ms cepstral axis are significantly lower than individual cepstral values. 20ms is also the period length of the fundamental frequency (50Hz) of adult male speech, which has a relatively long fundamental frequency period. Speech cepstral energy is mainly located to the left of the fundamental frequency cepstral. The initialization phase ensured that the difference function was located to the right of the 20ms cepstral axis by setting the input / output buffer size. Since the difference function is a relatively stable value, its mean is equal to the difference function itself; therefore, calculating the average of the difference cepstral values can greatly reduce speech cepstral interference. However, apart from a portion of the impulse response values in the direct feedback and early reverberation that are significantly more prominent than the differential cepstral mean at that frequency inversion position, most of the differential function values are comparable to the differential cepstral mean at that point. Therefore, using an amplitude threshold to eliminate speech cepstral interference is still not accurate enough.
[0079] To further improve the analytical accuracy of the impulse response value, this invention introduces the differential cepstral peak probability as a weighted average of the differential cepstral mean. The principle of the differential cepstral peak probability is as follows: since the speech cepstral is a random function, except for the region with higher energy on the left side of 20ms which has a higher peak probability, the peak probabilities at the right side are generally lower. When a relatively stable differential function is superimposed on the speech cepstral at a certain point, the peak probability at that point is increased, especially at the cepstral position of direct acoustic feedback where the peak probability is significantly higher than in other nearby regions. Therefore, compared with the amplitude threshold, using the peak probability as the threshold and weighting the cepstral mean allows h(n) to more accurately approximate f(n). Even if a higher peak probability at a certain point may mask the peak probabilities of neighboring points, since the differential cepstral is calculated based on the difference between h(n) and f(n), the impulse response at that point is suppressed by the h(n) filter in subsequent iterations, and its peak probability will decrease, while the masked peak probabilities of neighboring points will increase. In this way, even if the impulse response amplitude and the cepstral mean of the speech are close, the impulse response at that point can still be identified by using peak probability weighting, and a higher peak probability corresponds to a larger estimated impulse response value.
[0080] In fact, the cepstral spike probability method preferentially suppresses high-energy acoustic feedback in the early stages of iteration, while low-energy acoustic feedback is suppressed later. This method allows the system to achieve a significant increase in system stability gain in the early stages of feedback suppression, and the accuracy increases with the number of iterations.
[0081] In this embodiment, differential cepstral spike probability statistics are used to estimate the initial update of the feedback suppression filter parameters, including the following process:
[0082] 1. Transformation Processing Steps: Randomly select differential data frames containing speech signals from the cepstral analysis buffer, and sequentially perform windowing, Fourier transform, logarithmic transform, and inverse Fourier transform to obtain the absolute amplitude and angle values of the differential cepstral spectrum. Divide the absolute amplitude values into positive and negative groups according to the angle values. The specific process is as follows:
[0083] 1-1. Fill the entire cepstral analysis buffer with the data frame E(n) containing the speech difference function. The length of this buffer is the length L of E(n). c Several times, usually more than 10 times or with a time length of 3 to 5 seconds, to obtain more accurate statistical results, a random sample of length L is drawn from the cepstral analysis buffer. c The differential data frame E(n), where L c = 2·L c / 2 ;
[0084] 1-2, Apply a window (e.g., a Hanning window) to E(n). After windowing, the data E... w (n) = E(n)·W(n), where W(n) is a window function; for E w Perform a Fourier transform (fft) on (n) to obtain the absolute value of the amplitude spectrum A(n) = |fft(E w (n))|; Perform a logarithmic transformation on the absolute value of the amplitude spectrum A(n) to obtain A log (n) = log10(A(n)); for the logarithm A log (n) Perform an inverse Fourier transform (ifft) to obtain the absolute value of the cepstral amplitude |C e (n)| and the angle value ∠C e (n), i.e., |C e (n)| = |ifft(A log (n))|,∠C e (n) = ∠ifft(A log (n));
[0085] 1-3, divide the absolute value of the amplitude into positive and negative groups according to the angle value. Those with an angle value of 0 are divided into the positive group C. e+ (n), those with an angle value of π are divided into negative group C. e- (n), in practice, 1 can be used as the dividing line between the two sets of data. Because |C e (n)| is even-symmetric, and only the first half of each data set is taken, with a length of L. c / 2 The value is denoted as:
[0086] When ∠C e (n)=0 or ∠C e When (n) < 1, C e+ (n) = |C e (n)[:L c / 2]|;
[0087] When ∠C e (n) = π or ∠C e When (n)>1, C e+ (n) = 0;
[0088] When ∠C e (n) = π or ∠C e When (n)>1, C e- (n) = |C e (n)[:L c / 2 ]|;
[0089] When ∠C e (n)=0 or ∠C e When (n) < 1, C e- (n) = 0.
[0090] Among them, [:L c / 2 ] indicates taking C e The first half of the length L in (n) c / 2 The value.
[0091] 2. After repeating the transformation process multiple times, calculate the mean absolute amplitude and the probability of local spikes for each point in both sets of data. Then, set some data points in the local spike probability to zero, multiply each pairwise by the mean absolute amplitude, and subtract the results to obtain the initial feedback filter parameters. The specific process is as follows:
[0092] 2-1. Repeat the above transformation process P times, and calculate the mean absolute value of the amplitude and the probability of local spikes for each point in both sets of data. Let this be denoted as:
[0093] M c+ (n) = / P;
[0094] M c- (n) = [ ] / P;
[0095] P pk+ (n) = [ / P;
[0096] P pk- (n) = [ ] / P;
[0097] Among them, M c+ (n), M c- (n) represents the mean of the absolute values of the cepstral differences between the positive and negative groups. (n) (n) represents the absolute amplitude of the difference cepstrum of the positive and negative groups at point n, where i is the random sampling number and n is the cepstrum index, n = 0, 1, ... L c / 2 -1. P pk+ (n), P pk- (n) represents the probability of a local peak in the absolute value of the cepstral difference between the positive and negative groups, when the absolute value of the cepstral difference between the positive and negative groups is a local peak at point n. (n) (n) is counted as 1, otherwise it is counted as 0.
[0098] 2-2, set the peak probability values that are less than a preset threshold to zero, denoted as: when P pk+ (n) ≤ P thr+ At that time, P pk+ (n) = 0; when P pk- (n) ≤ P thr- At that time, P pk- (n) = 0; Preset threshold P thr+ P thr- The specific value can be obtained by statistically analyzing the cepstral spikes of pure speech data. Generally, it is taken as the average of the spike probabilities of all points on the right side of the cepstral spectrum 20ms plus twice the standard deviation.
[0099] 2-3, Set the input / output delay D of the frequency inverted axis audio processor. i / o The peak probability value on the left is set to zero, denoted as: when n≤N d At that time, P pk+ (n) = 0, P pk- (n) = 0. Where N d The number of sampling cycles for the input-output delay of the audio processor, when D i / o When N is less than 20 milliseconds d It can be appropriately increased to the 20-millisecond position of the frequency reversal axis.
[0100] The input and output delay of the audio processor can be adjusted by setting the input and output buffer size. The purpose is to ensure that the starting position of the impulse response function on the cepstral axis is located to the right of the high-energy region of the speech cepstral and the speech fundamental frequency pulse, so as to avoid overlap with the high-energy signal of the speech cepstral and the fundamental frequency pulse signal.
[0101] 2-4. Multiply and subtract the peak probabilities and mean absolute values of the amplitudes after completing steps 2-2 and 2-3 respectively to obtain the initial feedback filter parameters, denoted as: h pre (n) = P pk+ (n)·M c+ (n) - P pk- (n)·M c- (n).
[0102] 3. Combining the impulse response position of the direct acoustic feedback and the total delay of the acoustic feedback, after setting some data in the initial feedback filter parameters to zero, the remaining initial feedback filter parameter values, after removing the amplification ratio coefficients, are used as the primary feedback suppression filter parameters. The specific process is as follows:
[0103] 3-1, The position of the impulse response argmax for direct acoustic feedback is determined by searching for the maximum value position in the initial feedback filter parameters. n (h pre And set the initial feedback filter parameter value to zero to the left of the pulse response position, denoted as: when n <argmax n (h pre When h pre (n) = 0.
[0104] 3-2. At this point, another issue arises: the higher-order convolutions of the difference function in the differential cepstral expression, i.e., the portion where k > 1, appear at 2D (twice the total delay) on the cepstral axis and at integer multiples of D on the right (3D, 4D, ...). These higher-order convolutions are actually parasitic within the lower-order or initial-order convolution signals and are spurious signals in cepstral analysis. Once the initial convolution is eliminated by the feedback suppression filter, the higher-order convolutions also disappear. Therefore, the initial feedback filter parameter value h, located to the right of the total acoustic feedback delay 2D, can be... pre (n) is set to zero to eliminate spurious signal interference, denoted as: when n ≥ 2Dn, h pre (n) = 0. Dn represents the number of sampling periods for the total acoustic feedback delay.
[0105] In practice, a window with a length not greater than the total delay of the acoustic feedback is set starting from the position of the maximum impulse response, and the initial feedback filter parameter values outside the window are set to zero.
[0106] 3-3, Remove the remaining initial feedback filter parameter values h pre The amplification ratio coefficient G in (n) is used as the parameter h of the primary feedback suppression filter. pri (n), denoted as: h pri (n) = h pre (n) / G.
[0107] Using h pri (n) Perform feedback suppression filtering on the input audio data. Denoted as: E pri (n) = y(n) - h pri (n)*X(n) / G, where E pri (n) represents the number of times h is passed. pri (n) The difference function after feedback filtering. For details of the feedback filtering process, please refer to the steps above.
[0108] In this embodiment, the estimation of the room impulse response function of the acoustic feedback path is used to perform iterative updates of the feedback suppression filter parameters in rounds other than the initial update. This includes the following process:
[0109] Although h pri The probability weighting in the (n) algorithm only calculates a portion of f(n), but this method is effective for eliminating speech cepstral interference. The residual [f(n)-h(n)] can be obtained by repeatedly applying E... pri (n) is eliminated by feedback suppression filtering, that is: ,in This represents the parameters of the secondary feedback suppression filter updated in round j+1. This represents the feedback suppression filter parameters updated in round j. When j=1, The parameter h of the primary feedback suppression filter is represented. pri (n), where α represents the adjustment parameter, 0≤α≤1, used to adjust the amplitude of each approximation, similar to λ. When α=0, the filter parameters are no longer updated.
[0110] This represents the residual estimate of the impulse response function in round j+1, and its estimation process is similar to h. pri (n) are the same, the only difference is that the differential data frame e(n) being processed is processed by the previous filter parameter h. pri (n) The result generated and stored in the cepstral analysis buffer after the action and playback of the filter (n). That is, the result using the feedback suppression filter parameter h. pri The estimation method of (n) applies to the parameters of the feedback suppression filter. The processed differential data frames are then estimated.
[0111] exist In the estimation process, the location of the maximum value in step 3-1 may not necessarily still be the location of the direct acoustic feedback impulse response, because this impulse response has already been suppressed in the previous filtering. Therefore, h sec (n) The calculation needs to be repeated multiple times. As the main impulse response is eliminated by the feedback suppression filter, the system gain G can be gradually increased so that the smaller impulse response at the tail of f(n) gradually gains an increased probability of cepstral spikes and then increases in h. sec The expression is obtained in (n). In the specific implementation, each increase of ΔG can be used to repeat the real-time update of the feedback suppression filter parameters multiple times, that is, repeat h several times. sec (n) is calculated, and the number of repetitions can increase as G increases.
[0112] S3 uses the gain-corrected audio input-output variance ratio to evaluate the feedback suppression effect in real time, and selects the optimal feedback suppression filter parameters based on the feedback suppression effect.
[0113] In the embodiment, due to the randomness of the speech cepstral, in h sec (n) Fluctuations can occur during the approximation process, and the increase in gain G can also cause system instability. Therefore, it is necessary to evaluate the feedback suppression effect in real time. This evaluation helps to improve the convergence speed and accuracy of the filter.
[0114] The real-time evaluation method is based on the minimum mean square error (LMS) principle of adaptive filtering. Although LMS has bias problems in speech processing, practice shows that real-time evaluation parameters designed according to the LMS principle can effectively improve h(n) accuracy and maximum stable gain (MSG).
[0115] The evaluation parameter ESR is the ratio of the squares of the audio output and input data frames, denoted as: ESR = E 2 / Y 2 Among them, Y 2 E is the sum of squares of the data at each point in the audio input data frame. 2 The sum of squares of each data point in the difference function data frame output after the Y data frame has undergone feedback suppression filtering:
[0116]
[0117]
[0118] Where k represents the index of the point.
[0119] When the initial state of the system is stable and the feedback suppression filter is in effect, the ESR gradually decreases. The smaller |f(n)-h(n)| is, the smaller the ESR is, and ESR < 1. When ESR >> 1, it can be determined that the system is losing stability or has experienced howling. When the gain G remains constant, h is chosen when the ESR is smaller. sec (n) is used as the feedback suppression filter parameter under this gain, and this is used as the next gain G. m The parameters during stepping. To bias the system towards selecting feedback suppression filter parameters with a larger G, this method uses the variance ratio of the gain correction as a measure of the feedback suppression effect:
[0120]
[0121] The optimal result h(n)* of the feedback suppression filter parameters is:
[0122] h(n)* = argmin h (ESR g )
[0123] Where, argmin h (ESR g ) indicates that the feedback suppression effect ESR is achieved.g The minimum parameters for the feedback suppression filter.
[0124] The method proposed in this invention uses the peak probability of the differential cepstral as a weighted average to calculate the parameters of the feedback suppression filter, and uses the gain-corrected speech input-output variance ratio as a method for evaluating the real-time feedback suppression effect. These are key steps that differ from the existing AFC-CE method.
[0125] The existing AFC-CE method does not involve evaluating the effectiveness of real-time feedback suppression or selecting optimal feedback suppression filter parameters. Instead of directly using the estimated value f^(n) of f(n) as a parameter of h(n), the AFC-CE method asymptotically obtains h(n) by introducing a forgetting factor λ:
[0126] h(n) = λ·h(n-1) + (1-λ)·f^(n)
[0127] In fact, λ plays a role in suppressing speech cepstral c. u The interference effect of h(n) on the estimate of f(n) can be observed through simulations to see how h(n) approximates f(n) at different λ values, and how it suppresses the cepstral spectrum of random speech. Figure 4 As shown in (1), when f(n) is simplified to 1 and λ (i.e., Iamda) is 0.7, 0.8, and 0.9, the asymptotic regression of h(n) requires 12, 20, and 43 iterations to approach 0.99f(n) respectively. Figure 4 As shown in Figure (2), when λ is 0.97, 0.98, and 0.99, the asymptotic regression of h(n) requires 151, 227, and 458 iterations to approximate 0.99f(n), respectively. From the above observations, it can be inferred that when λ is small, h(n) approximates f(n) faster and requires fewer iterations. Conversely, when λ is large, h(n) approximates f(n) slower and requires more iterations.
[0128] Figure 5 The figure shows the comparison between the speech cepstral values after 100 asymptotic iterations and the mean of the 100 cepstral iterations when λ is 0.7, 0.8, and 0.9. In each iteration, a 312ms (16kHz sampling rate, 5000 sample length) speech data frame was randomly selected from a 5-second broadcast speech for cepstral calculation. It can be seen that with the same number of asymptotic iterations, the speech cepstral amplitude fluctuates more after iteration when λ is smaller. However, even with larger λ values, the fluctuation amplitude is still larger compared to the mean speech cepstral value. Therefore, it can be inferred that the asymptotic iteration of the speech cepstral significantly interferes with the estimation of the difference function.
[0129] Figure 6The figure shows the comparison between the speech cepstral values after 151,227, and 458 asymptotic iterations and the mean of the speech cepstral values after 100 iterations when λ is 0.97, 0.98, and 0.99, respectively. Each iteration uses the same speech data and sampling method as described above. Figure 5 The results are the same. Comparing the average cepstral values of speech after 100 random samplings, it can be seen that the fluctuation amplitude of the speech cepstral values after the required number of asymptotic iterations is greater than the average fluctuation amplitude of the cepstral values after 100 iterations. Therefore, it is inferred that the asymptotic iteration method using the forgetting factor will have certain limitations in improving the accuracy of difference function estimation.
[0130] Figure 7 The figure shows a comparison between the asymptotic cepstral values of speech at different iteration numbers when λ=0.99 and the average cepstral value when n=100. It can be seen that the fluctuation amplitude of the asymptotic cepstral values after iteration increases with the number of iterations, thus increasing the random interference caused by the estimation of the difference function. This is detrimental to long-term tracking of changes in the difference function.
[0131] In summary, when the forgetting factor λ is small, the regression speed of h(n) is faster, but the random interference caused by the speech cepstral is greater. When the forgetting factor λ is close to 1, the regression speed is slower, requiring more regression iterations, but the increase in the number of iterations also amplifies the random interference of the differential cepstral. In comparison, the method of using the differential cepstral mean in this invention has more advantages than the asymptotic iteration method with forgetting factor λ in the prior art.
[0132] Compared to the AFC-CE method, the method of this invention can significantly improve the estimation accuracy of the impulse response function and can stably track the changes in the impulse response function over a long period of time. The accuracy evaluation of the impulse response function uses a normalized unaligned function. Its definition is: .
[0133] Figure 8 The image shows the result after ESR. g Existing AFC-CE method and the method of this invention before and after optimization. ppwm A comparison of the squared MIS results after 1000 iterations. It can be seen that without ESR... g The optimized AFC-CE method suffers from large and unstable estimation errors for f(n). After optimization using ESRg, and especially with the AFC-CE method of this invention... ppwm The improved method significantly reduces estimation error and stability. Data shows that the improved AFC-CE method reduces the squared error of MIS by at least half and increases the maximum stable gain (MSG) by at least 10 dB.
[0134] The specific embodiments described above illustrate the technical solution and beneficial effects of the present invention in detail. It should be understood that the above description is only the most preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, additions, and equivalent substitutions made within the scope of the principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for reverberation and feedback suppression in an indoor sound reinforcement system based on speech cepstral analysis, characterized in that, Includes the following steps: Collect and store microphone audio data from indoor sound reinforcement systems; The audio data stored by feedback suppression filtering is processed to obtain differential data frames. The differential data frames are amplified by gain control, stored and output as audio. The parameters of the feedback suppression filter used in the feedback suppression filtering process are updated in real time after being estimated by speech activity detection, differential cepstral peak probability statistics and room impulse response function of acoustic feedback path. The differential cepstral spike probability statistics are used to estimate the initial update of the feedback suppression filter parameters, and include the following process: Transformation processing steps: Randomly select differential data frames containing speech signals from the cepstral analysis buffer, and sequentially perform windowing, Fourier transform, logarithmic transform, and inverse Fourier transform to obtain the absolute amplitude and angle values of the differential cepstral. Divide the absolute amplitude values into positive and negative groups according to the angle values. After repeating the above transformation process multiple times, the mean absolute value of the amplitude and the probability of local spikes at each point in the two sets of data are calculated. Then, some data in the probability of local spikes are set to zero and multiplied pairwise with the mean absolute value of the amplitude and subtracted to obtain the initial feedback filter parameters. Combining the impulse response position of the direct acoustic feedback and the total delay of the acoustic feedback, after setting some data in the initial feedback filter parameters to zero, the amplification ratio coefficient in the remaining initial feedback filter parameter values is removed and used as the primary feedback suppression filter parameters. The feedback suppression effect is evaluated in real time using the gain-corrected audio input-output variance ratio, and the optimal feedback suppression filter parameters are selected based on the feedback suppression effect.
2. The method for reverberation and feedback suppression of indoor sound reinforcement systems based on speech cepstral analysis according to claim 1, characterized in that, Audio data is stored in a computer ring buffer, the length of which should accommodate at least half the length of the input / output audio data frame plus the length of the differential cepstral calculation data frame. Among them, half the length of the differential cepstral calculation data frame should be greater than the total acoustic feedback delay length plus the length of the impulse response function, and should be an integer multiple of the input / output data frame; The total acoustic feedback delay includes the propagation delay caused by the direct propagation path of sound from the speaker to the microphone, as well as the input and output delay of the audio processor. The length of the impulse response function specifically refers to the finite length of the room impulse response starting from the direct acoustic feedback response impulse.
3. The method for reverberation and feedback suppression of indoor sound reinforcement systems based on speech cepstral analysis according to claim 1, characterized in that, Speech activity detection is used to determine whether differential data frames processed by feedback suppression filtering contain speech signals. Specifically, the energy threshold method can be used to determine this. Differential data frames containing speech signals are stored in the cepstral analysis buffer for differential cepstral analysis.
4. The method for reverberation and feedback suppression of indoor sound reinforcement systems based on speech cepstral analysis according to claim 1, characterized in that, Set some data in the probability of local spikes to zero, including: Set the peak probability values that are less than a preset threshold to zero, and at the same time set the peak probability values whose left side length of the frequency inversion axis is not less than the input-output delay of the audio processor to zero.
5. The method for reverberation and feedback suppression of indoor sound reinforcement systems based on speech cepstral analysis according to claim 1, characterized in that, Combining the impulse response position of the direct acoustic feedback and the total delay of the acoustic feedback, some data in the initial feedback filter parameters are set to zero, including: The position of the maximum value in the initial feedback filter parameters is determined as the impulse response position of the direct acoustic feedback. Starting from this position, a window with a length not greater than the total delay length of the acoustic feedback is set, and the initial feedback filter parameter values outside the window are set to zero.
6. The method for reverberation and feedback suppression of indoor sound reinforcement systems based on speech cepstral analysis according to claim 1, characterized in that, The initial feedback filter parameters are obtained by multiplying and subtracting the pairwise product and mean absolute values of the amplitudes of the positive and negative sets of peak probabilities after being set to zero, and then multiplying each pairwise: h pre (n) = P pk+ (n)·M c+ (n) - P pk- (n)·M c- (n) Among them, P pk+ (n) and P pk- (n) represents the positive and negative peak probabilities, respectively, M c+ (n) and M c- (n) represents the mean of the absolute values of the amplitudes at each point in the positive and negative cepstrum, where n is the index of each point.
7. The method for reverberation and feedback suppression of indoor sound reinforcement systems based on speech cepstral analysis according to claim 1, characterized in that, The estimation of the room impulse response function of the acoustic feedback path is used for iterative updates of the feedback suppression filter parameters in rounds other than the initial update, including the following process: in, This represents the parameters of the secondary feedback suppression filter updated in round j+1. This represents the feedback suppression filter parameters updated in round j. When j=1, This represents the parameters of the primary feedback suppression filter, where α represents the amplitude parameter used to adjust the approximation in each round. This represents the residual estimate of the impulse response function at round j+1; in, The estimation process is as follows: the estimation method of the primary feedback suppression filter parameters is used to estimate the parameters of the feedback suppression filter. The processed differential data frames are then estimated.
8. The method for reverberation and feedback suppression of indoor sound reinforcement systems based on speech cepstral analysis according to claim 1, characterized in that, When amplifying differential data frames through gain control, the gain value is adjusted by gain step increment; The real-time update of the feedback suppression filter parameters is repeated multiple times with each increase in gain step increment.
9. The method for reverberation and feedback suppression of indoor sound reinforcement systems based on speech cepstral analysis according to claim 1, characterized in that, The feedback suppression effect is evaluated in real time using the gain-corrected audio input-output variance ratio, and the optimal feedback suppression filter parameters are selected based on the feedback suppression effect, including: h(n)* = argmin h (ESR g ) Among them, Y 2 E represents the sum of squares of the audio input data frames. 2 This represents the sum of squares of the differential data frames output after the Y data frame has undergone feedback suppression filtering. Indicates gain. The parameter h(n)* represents the feedback suppression effect, and argmin represents the optimal feedback suppression filter parameters. h (ESR g ) indicates that the feedback suppression effect ESR is achieved. g The minimum parameters for the feedback suppression filter.