Local sound amplification method based on acoustic feedback suppression technology
By combining bone conduction sensors and air conduction microphones with neural network prediction models and adaptive filtering technology, parameters are dynamically adjusted to generate potential howling frequencies and perform adaptive notch filtering and micro-delay perturbation. This solves the problems of inaccurate howling prediction and fixed parameters in traditional acoustic feedback suppression methods, and achieves high-gain, lossless audio quality feedback suppression.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN POROS TECH CO LTD
- Filing Date
- 2026-04-02
- Publication Date
- 2026-05-05
AI Technical Summary
Traditional acoustic feedback suppression methods in public address systems suffer from problems such as inaccurate howling prediction, fixed parameters that cannot adapt to changes in the acoustic environment, and impure reference signals leading to poor feedback suppression effects.
The system uses a bone conduction sensor and an air conduction microphone to simultaneously acquire pure speech reference signals and mixed audio signals. Combined with a neural network prediction model and adaptive filtering technology, parameters are dynamically adjusted to generate potential howling frequency points, which are then suppressed by adaptive notch filtering and random micro-delay perturbations.
It achieves high-gain, lossless audio quality and strong environmental adaptability feedback suppression, significantly improving the system's stability and feedback suppression capability, and reducing speech impairment and misjudgment rate.
Smart Images

Figure CN121985259A_ABST
Abstract
Description
Technical Field
[0001] This invention discloses a local amplification method based on acoustic feedback suppression technology, belonging to the field of acoustic feedback suppression in amplification systems. Background Technology
[0002] Local amplification systems are widely used in education, conferences, stage performances, and other scenarios. Their core function is to pick up the speaker's voice through a microphone, amplify it, and then play it through speakers to compensate for sound attenuation during spatial propagation. However, when the sound reproduced by the speakers is picked up again by the microphone, forming a positive feedback loop, the system will produce feedback (howling). Feedback not only severely damages the listening experience and interferes with normal amplification, but can also overload and damage the amplifier or speaker units.
[0003] Traditional acoustic feedback suppression methods each have their limitations. Two-stage notch filters require pre-detection of the feedback frequency and setting of the notch before system use, but this method is time-consuming and cannot quickly adapt to changes in the amplification environment, easily generating unexpected feedback when the sound field changes. Adaptive notch filters can detect and suppress feedback in real time, but usually only start processing after feedback has already occurred, and the auditory interference and impact on the system at the moment of feedback cannot be avoided. While frequency shifting and phase shifting methods can improve system stability to some extent, their gain enhancement capability is limited and cannot meet the requirements of high-fidelity amplification.
[0004] Existing feedback suppression techniques based on adaptive filtering typically use the speaker drive signal as a reference, attempting to subtract the reentry sound component from the microphone-picked signal. However, the fundamental drawback of this method is that the reference signal and the desired signal are highly correlated, leading to difficulties in filter convergence, large offset, and difficulty in accurately separating the speaker's original speech from the reentry sound. Furthermore, the processing parameters of traditional methods are mostly fixed settings, unable to be dynamically adjusted according to the actual acoustic characteristics of the sound reinforcement space, resulting in unstable performance under different acoustic environments, and residual feedback still affecting speech intelligibility. Summary of the Invention
[0005] The purpose of this invention is to provide a local amplification method based on acoustic feedback suppression technology. This method utilizes a bone conduction sensor and an air conduction microphone to simultaneously acquire a clean speech reference signal and a mixed audio signal. Acoustic measurements are performed on the amplified space based on the transmitted detection signal, and a neural network prediction model is adjusted according to the measurement results to obtain an environment-adaptive critical frequency predictor. The mixed audio signal is adaptively filtered using the clean speech reference signal to obtain the speaker's speech component. The speaker's speech component is subjected to feedback risk analysis based on the environment-adaptive critical frequency predictor to generate potential feedback frequencies. Adaptive notch filtering is applied to the speaker's speech component using these potential feedback frequencies to form an initially suppressed amplified signal. Random micro-delay perturbations are applied to the initially suppressed amplified signal based on the potential feedback frequencies to generate a final suppressed feedback amplified signal. This invention achieves high-gain, lossless audio quality, and strong environmental adaptability through multi-dimensional fusion of bone conduction reference signals, environment-adaptive prediction, notch filtering, and micro-delay perturbations, resulting in local amplified feedback suppression.
[0006] The objective of this invention can be achieved through the following technical solutions: A local sound amplification method based on acoustic feedback suppression technology includes the following steps: S1. Simultaneous acquisition of the speaker's voice and the surrounding sound field is performed using a bone conduction sensor and an air conduction microphone to obtain a clean speech reference signal and a mixed audio signal. S2. Based on the transmitted detection signal, perform acoustic measurements on the amplified space and adjust the neural network prediction model according to the measurement results to obtain an environment-adaptive critical frequency predictor; S3. Adaptive filtering is performed on the mixed audio signal using the pure speech reference signal to obtain the speaker's speech component; S4. Based on the environmental adaptive critical frequency predictor, perform feedback risk analysis on the speaker's voice components to generate potential feedback frequency points. S5. Using the potential howling frequency point, perform adaptive notch filtering on the speaker's voice component to form an amplified signal with initial suppression. S6. Apply a random micro-delay perturbation to the initially suppressed amplified signal based on the potential howling frequency point to generate the final suppressed feedback amplified signal.
[0007] Preferably, step S1 includes: S1.1. A pure speech reference signal containing only the speaker's voice is picked up by a bone conduction sensor attached to the speaker's skull or mandible. S1.2. A mixed audio signal containing the speaker's voice, speaker re-entry sound, and ambient noise is picked up by an air conduction microphone placed at the speaker's collar or headgear. S1.3 Perform time delay estimation and compensation on the pure speech reference signal and the mixed audio signal to align the speaker's voice components in the two signals on the time axis.
[0008] Preferably, step S2 includes: S2.1. A broadband detection signal that covers the target frequency band of the loudspeaker array of the loudspeaker system and is not easily detected by the human ear is emitted into the loudspeaker space, and the reflected sound of the detection signal is received by the air conduction microphone. The reverberation time and frequency response characteristics of the loudspeaker space are calculated based on the impulse response between the emitted signal and the received signal. S2.2 Using the calculated reverberation time and frequency response characteristics as input parameters, dynamically adjust the internal weights and thresholds of the pre-trained neural network prediction model to adapt the neural network prediction model to the acoustic environment of the current sound reinforcement space. The neural network prediction model is used to output the probability of howling at each frequency point based on the spectrum of the input audio signal.
[0009] Preferably, step S2 further includes dynamically adjusting the internal weights and thresholds of the neural network prediction model: The calculated reverberation time and frequency response characteristics are introduced as constraints into the backpropagation algorithm of the neural network, so that the model is guided by the goal of matching the current acoustic environment during the weight update process. Guided by the constraints, the connection weights and bias thresholds of each layer of the neural network are iteratively corrected through the backpropagation algorithm, so that the acoustic features extracted by the feature extraction layer of the model can be adaptively matched with the acoustic modes of the current sound reinforcement space.
[0010] Preferably, step S3 includes: S3.1. Use the clean speech reference signal after time delay compensation and spectrum matching as the desired input of the adaptive filter, and use the mixed audio signal as the original input of the adaptive filter to initialize the filter coefficients. S3.2. The error signal between the filter output signal and the desired input is calculated iteratively through an adaptive filtering algorithm, and the filter coefficients are updated in real time according to the error signal, so that the filter output gradually approaches the speaker's voice component in the mixed audio signal. S3.3. The estimated signal output by the converged adaptive filter is used as the extracted speaker speech component, and the error signal is used as the residual reentry sound component and input to the feedback suppression effect evaluation module. The feedback suppression effect evaluation module judges the current feedback suppression state according to the energy change trend of the residual reentry sound component. When the energy of the residual reentry sound component continues to rise, the notch depth enhancement of the adaptive notch processing in the subsequent steps is triggered.
[0011] Preferably, step S3.2 includes: The error signal between the filter output signal and the desired input is calculated iteratively using an adaptive filtering algorithm, and the convergence step size and leakage factor of the adaptive filter are dynamically adjusted according to the amplitude and spectral characteristics of the error signal. When a sudden pulse interference is detected in the error signal, the update of the filter coefficients is temporarily frozen to avoid coefficient divergence, and the update is resumed after the pulse interference ends.
[0012] Preferably, step S3.3 includes: The estimated signal output by the converged adaptive filter is used as the extracted speaker speech component, and the error signal is used as the residual re-entry sound component and input to the feedback suppression effect evaluation module. The feedback suppression effect evaluation module determines the current feedback suppression state based on the energy change trend of the residual reentry sound component. When the energy of the residual reentry sound component continues to rise, it triggers the notch depth enhancement of the adaptive notch processing in subsequent steps.
[0013] Preferably, step S4 includes: S4.1 Input the spectral data of the speaker's voice components into the environment adaptive critical frequency predictor, and calculate and output the howling risk probability value corresponding to each frequency point through the forward propagation of the neural network prediction model. S4.2. Compare the howling risk probability value of each frequency point with multiple preset risk thresholds. Based on the comparison results, mark the frequency points that exceed the first threshold as potential howling frequency points, and sort the potential howling frequency points according to their risk probability. S4.3 Perform time series analysis on the marked potential howling frequencies to detect the trend of risk probability of each frequency over time, mark the frequencies with continuously rising risk probability as emergency howling frequencies and increase their processing priority.
[0014] Preferably, step S5 includes: S5.1 According to the potential howling frequency points and their corresponding risk levels, configure the corresponding notch depth, notch width and notch quality factor for each potential howling frequency point. The higher the risk level of the frequency point, the larger the notch depth and the narrower the notch width are allocated. S5.2 Connect the configured adaptive notch filters in series to the signal processing link according to the frequency priority order, and perform notch processing on the speaker's voice components in sequence. Each notch filter only attenuates the frequency components near its corresponding potential howling frequency. S5.3 During the notch filtering process, monitor the actual energy changes of each potential howling frequency point in real time, and make a comprehensive judgment based on the energy change trend of the residual reentry sound component. When the energy of a certain frequency point drops below the safe threshold, gradually reduce the notch depth of the corresponding notch filter until it is completely released, so as to avoid unnecessary damage to the sound quality.
[0015] Preferably, step S6 includes: S6.1 Based on the distribution range and risk level of the potential howling frequency points, generate a set of pseudo-random time delay sequences independently for each channel in the multi-channel loudspeaker array. The maximum time delay amplitude of the pseudo-random time delay sequence is negatively correlated with the wavelength of the potential howling frequency points covered by the channel, so that the high-frequency sound waves can obtain more refined time perturbation, so as to effectively destroy their spatial coherence. S6.2. The initially suppressed amplified signal is copied and distributed to each channel of the multi-channel speaker array, and a pseudo-random time delay perturbation is applied to the audio signal of each channel, so that the phase relationship of the sound waves emitted by each channel speaker in space presents a randomized distribution. S6.3 Real-time monitoring of energy changes at potential howling frequencies in the amplified signal after perturbation. When an upward trend in energy at a certain frequency is detected, the perturbation amplitude and rate of change of the pseudo-random time delay sequence of the corresponding channel are dynamically adjusted.
[0016] The beneficial effects of this invention are: This invention acquires a pure speech reference signal using a bone conduction sensor and uses it as the desired input for adaptive filtering, fundamentally solving the technical problems of impure reference signals and high correlation with the desired signal in traditional methods. Because it can accurately extract the speaker's speech components and remove re-entry voice components, this invention significantly reduces speech damage during feedback suppression, preserving the naturalness and clarity of the original speech, and achieving a balance between high-gain amplification and high-fidelity sound quality.
[0017] This invention measures the reverberation time and frequency response characteristics of a loudspeaker space by actively transmitting detection signals, and dynamically adjusts the internal parameters of a neural network prediction model based on the measurement results, enabling the model to adaptively match the current acoustic environment. Compared with traditional feedback suppression methods with fixed parameters, this invention can accurately predict the critical howling frequency in loudspeaker scenarios with different reverberation conditions and different space sizes, effectively solving the problems of misjudgment and missed judgment caused by environmental changes, and significantly improving the system's environmental adaptability and stability.
[0018] This invention integrates four techniques: adaptive filtering purification, neural network risk prediction, adaptive notch filtering, and multi-channel micro-delay perturbation, forming a complete suppression chain from source separation, pre-prediction, active notch filtering to physical layer disruption. The steps work synergistically: a clean reference ensures extraction accuracy, environmental prediction enables early warning, notch filtering provides precise attenuation, and delay perturbation disrupts phase superposition, thus blocking the formation conditions of positive feedback loops at multiple levels and significantly improving the system's overall feedback suppression capability and anti-interference robustness. Attached Figure Description
[0019] Figure 1 This is a schematic diagram of the structure of a local amplification method based on acoustic feedback suppression technology according to the present invention. Detailed Implementation
[0020] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0021] Example: Figure 1 As shown, a local amplification method based on acoustic feedback suppression technology includes the following steps: S1. Simultaneous acquisition of the speaker's voice and the surrounding sound field is performed using a bone conduction sensor and an air conduction microphone to obtain a clean speech reference signal and a mixed audio signal. S2. Based on the transmitted detection signal, perform acoustic measurements on the amplified space and adjust the neural network prediction model according to the measurement results to obtain an environment-adaptive critical frequency predictor; S3. Adaptive filtering is performed on the mixed audio signal using the pure speech reference signal to obtain the speaker's speech component; S4. Based on the environmental adaptive critical frequency predictor, perform feedback risk analysis on the speaker's voice components to generate potential feedback frequency points. S5. Using the potential howling frequency point, perform adaptive notch filtering on the speaker's voice component to form an amplified signal with initial suppression. S6. Apply a random micro-delay perturbation to the initially suppressed amplified signal based on the potential howling frequency point to generate the final suppressed feedback amplified signal.
[0022] Preferably, a bone conduction sensor and an air conduction microphone are used to simultaneously acquire the speaker's voice and the surrounding sound field, obtaining a clean speech reference signal and a mixed audio signal. The specific implementation method is as follows: A bone conduction sensor and an air conduction microphone are integrated into the microphone device worn by the speaker. The bone conduction sensor is attached to the speaker's skull or mandible. Utilizing the sound wave conduction characteristics of human tissue, it is only sensitive to the speaker's own vocal vibrations, and is largely unresponsive to airborne speaker re-entry sounds and environmental noise. This allows it to pick up a pure speech reference signal containing only the speaker's voice, denoted as [missing information]. , where n is the discrete-time sampling point index. The air-conduction microphone is positioned at the speaker's collar or headband, picking up a mixed audio signal containing the speaker's voice, speaker re-entry sound, and ambient noise via air conduction, denoted as . Because bone conduction signals travel faster through human tissue than airborne sound, and their propagation paths are different, there is a certain time delay between the speaker's vocal components in the two signals. Therefore, it is necessary to estimate and compensate for the time delay of the two signals to ensure the accuracy of the reference signal in subsequent adaptive filtering processing.
[0023] To achieve precise alignment of the two signals, the relative time delay between the speaker's vocal components in the bone conduction and air conduction signals must first be estimated. This invention employs a time delay estimation method based on the cross-correlation function to calculate the cross-correlation sequence of the two signals within the analysis frame; the peak position of this sequence is the estimated time delay value. The specific calculation formula is as follows: ,in, Let be the cross-correlation function of the two signals, and N be the number of sampling points in the analysis frame. The candidate time delay is represented by the number of sampling points, and its value range is set according to the maximum possible time delay, such as the maximum time difference between human propagation and airborne propagation. The search is used to... To reach the maximum The value is used to obtain a coarse time delay estimate. Since the actual time delay may not be an integer multiple of the sampling period, to obtain a more accurate time delay estimate, parabolic interpolation can be performed near the cross-correlation peak based on the coarse estimate to obtain a time delay value with sub-sampling accuracy. To obtain accurate time delay estimates Afterwards, the bone conduction signal Perform time delay compensation to synchronize it with the air conduction signal. The speaker's vocal components are aligned on the timeline. Because... It may contain a decimal part, requiring a fractional delay filter to compensate for the delay of non-integer sampling points. This invention uses Lagrange interpolation to design a fractional delay filter, whose time-domain response is: Where M is the filter order, ranging from 4 to 8, and D is the desired fractional delay, expressed in units of sampling points. Filter coefficients It is obtained by polynomial approximation of the ideal sinc function. (The bone conduction signal is then used.) The pure speech reference signal after time delay compensation is obtained by passing through this fractional delay filter. .at this time, and The speaker's vocal volume is synchronized with the timing.
[0024] Due to the differences in frequency response characteristics between bone conduction sensors and air conduction microphones, bone conduction signals have stronger energy in the low-frequency range but may attenuate in the high-frequency range, while air conduction signals have a wider natural frequency band. To further improve the accuracy of the reference signal, the time-delay-compensated bone conduction signal needs to undergo spectral matching processing to ensure that its spectral envelope matches the spectral envelope of the speaker's vocal component in the air conduction signal. This invention uses an adaptive equalization filter to achieve spectral matching, updating the filter coefficients by minimizing the mean square error between the equalized bone conduction signal and the speech component in the air conduction signal. The equalization filter is typically an FIR structure, and its coefficient update formula is based on the normalized least mean square algorithm: ,in, Let L be the coefficient vector of the equalization filter at time n, and L be the filter order. The input signal vector; Let be the error signal, where For the desired signal, the gas conduction signal is taken here. Clean speech segments extracted after speech activity detection; This is the step size factor, which controls the convergence speed; The value should be a small positive integer to prevent the denominator from being zero. After iterative convergence, the equalization filter outputs... This refers to the clean speech reference signal whose spectral characteristics match the speaker's voice in the air conduction signal, denoted as... .
[0025] Preferably, acoustic measurements are performed on the amplified space based on the transmitted detection signal, and the neural network prediction model is adjusted according to the measurement results to obtain an environment-adaptive critical frequency predictor. The specific implementation method is as follows: This invention transmits a broadband detection signal, covering the target frequency band and imperceptible to the human ear, into the amplified space via a loudspeaker array of a sound reinforcement system. Preferably, a logarithmic sweep signal is used as the detection signal, whose frequency changes exponentially with time, exhibiting good anti-interference capability and signal-to-noise ratio. An air-conduction microphone simultaneously receives the sound waves reflected from the detection signal through the amplified space, obtaining the room impulse response. Let the transmitted detection signal be... The received signal is The relationship between the two can be expressed as: ,in For the room impulse response of the sound reinforcement space, This represents a linear convolution operation. The room impulse response can be obtained by deconvolving the received and transmitted signals. Specifically, this invention uses the frequency domain least squares method to solve the problem, and its calculation expression is as follows: ,in, This represents the frequency domain of the room impulse response, where k is the frequency index. and These are the Fourier transforms of the received signal and the transmitted signal, respectively. Represents the complex conjugate of the transmitted signal. A small regularization constant is used to prevent the denominator from being zero. The frequency domain room impulse response is obtained. Then, the time-domain room impulse response can be obtained through inverse Fourier transform. .based on This allows for further calculation of the reverberation time and frequency response characteristics of the amplified space. Reverberation time Defined as the time required for the sound energy density to decrease by 60 dB, through the... The Schroeder curve is obtained by inverse integration of the squared envelope, and then the attenuation rate is obtained by linear fitting of the curve, thereby calculating the reverberation time: ,in, This represents the attenuation rate of the Schroder curve in the range of -5dB to -35dB, expressed in dB / s. The frequency response characteristic is directly taken as... , which represents the gain or attenuation characteristics of the amplified space for each frequency component.
[0026] The calculated reverberation time and frequency response characteristics As input parameters, the internal weights and thresholds of the pre-trained neural network prediction model are dynamically adjusted to adapt the model to the acoustic environment of the current sound reinforcement space. This neural network prediction model employs a multilayer perceptron structure; the input layer receives the spectral characteristics of the audio signal, and the output layer represents the probability of howling at each frequency point. To achieve environmental adaptation, this invention incorporates reverberation time and frequency response characteristics as constraints into the backpropagation algorithm of the neural network. Specifically, let the objective function of the neural network model be... Where L is the loss function for howling risk prediction, using binary cross-entropy, and R is the environmental constraint regularization term. Here, represents the regularization coefficient. The environmental constraint regularization term is defined as the deviation between the model output and the expected output of the current acoustic environment, and its expression is: Where M is the total number of frequency points, For the model to the first The predicted risk probability output at each frequency point This serves as the benchmark for the expected environmental risk probability calculated based on the current reverberation time and frequency response characteristics. During backpropagation, the objective function J is related to the neural network's... Layer weights The gradient is: ,in, Calculated using the standard error backpropagation algorithm Then, according to the chain rule, it can be expanded as follows: , The derivative of the model output with respect to the weights can also be obtained through backpropagation. By introducing an environmental constraint regularization term, the neural network model aims to match the current acoustic environment during weight updates, enabling adaptive matching between the acoustic features extracted by the model's feature extraction layer and the acoustic modes of the current amplified space. The iterative update formula uses stochastic gradient descent with momentum: Where t is the number of iterations. For learning rate, Momentum factor This represents the weight update amount from the previous iteration. After multiple iterations, the neural network model converges to a state that matches the current acoustic environment of the loudspeaker space. At this point, the probability of howling risk at each frequency point output by the model can accurately reflect the possibility of howling at each frequency point under the current reverberation and frequency response conditions, thus obtaining an environmentally adaptive critical frequency predictor.
[0027] Preferably, the pure speech reference signal is used to perform adaptive filtering on the mixed audio signal to obtain the speaker's speech component. The specific implementation method is as follows: A clean speech reference signal with time delay compensation and spectrum matching is obtained. and the original mixed audio signal Then, the speaker's voice component is accurately extracted from the mixed audio signal through adaptive filtering. This invention employs a normalized least mean square adaptive filter to extract the pure speech reference signal. As a desired response, the mixed audio signal will be As input to the filter, the filter coefficients are iteratively updated to make the filter output approximate the part of the mixed audio signal that is correlated with the reference signal, i.e., the speaker's speech component. The coefficient vector of the adaptive filter at time n is defined as... ,in Let the filter order be . Let represent the p-th filter coefficient. The filter input vector consists of current and historical samples of the mixed audio signal, and is... Filter output This refers to the speaker's speech component estimated from the mixed audio signal, and its calculation formula is as follows: ,in, Let n be the output signal of the filter at time n, representing the currently estimated speaker speech components; This is the transpose of the coefficient vector; The input signal vector; For the p-th filter coefficient; To mix audio signals at time The sample value; P is the filter order, which determines the filter's memory length and modeling capability, and is usually set to an integer between 128 and 512 based on the room impulse response length.
[0028] Filter output Expected response Error signals between This reflects the accuracy of the filter estimation and also represents the portion of the mixed audio signal that cannot be explained by the reference signal, namely the residual reentry component and ambient noise. The formula for calculating the error signal is: ,in, The error signal represents the clean speech reference signal. With filter output The difference between them; This is the preprocessed, clean speech reference signal. When the filter converges... should approach Zhongyu The relevant part, and It then tends to be the reentry sound and ambient noise components that are unrelated to the reference signal.
[0029] To achieve adaptive updating of filter coefficients, this invention employs a normalized least mean square algorithm. This algorithm drives coefficient adjustment through instantaneous error signals and incorporates input signal power for normalization to improve convergence stability. The filter coefficient update formula is as follows: ,in, This is the updated filter coefficient vector; The leakage factor, ranging from 0.999 to 1, is used to control the filter's memory attenuation characteristics. When a sudden pulse interference is detected in the error signal, it can be temporarily... Set to 0 to update with the freeze coefficient, and restore after the interference ends; The convergence step size is dynamically adjustable, ranging from 0.1 to 0.5. It can be adaptively adjusted according to the spectral characteristics of the error signal. For example, when there are many high-frequency components in the error signal, the step size can be appropriately reduced to avoid coefficient divergence. This represents the error signal at the current moment; The input signal vector; The square of the Euclidean norm of the input signal vector is used to normalize and update the step size. Let be a small positive integer, with a range of values of 1. to This is used to prevent the algorithm from diverging when the denominator is zero.
[0030] During the filter's iterative convergence process, the system monitors the error signal in real time. The energy change trend. Define the short-time energy of the error signal at time n as... , where Q is the frame length of the short-time energy analysis. When detected When the residual reentry sound component shows a continuous upward trend, it indicates that the system is about to enter a state of high risk of howling. At this point, the feedback suppression effect evaluation module triggers the notch depth enhancement mechanism of the adaptive notch processing in subsequent step S5 to intervene in advance to suppress potential howling. After a sufficient number of iterations, the adaptive filter reaches a convergent state, at which point the filter output... This refers to the extracted speaker's voice components, denoted as... And error signal The residual reentry sound component is then input to the subsequent processing module to dynamically adjust the notch filtering intensity. Through the above adaptive filtering process, this invention achieves high-precision speaker speech extraction guided by a clean reference signal, laying a high-quality signal foundation for subsequent feedback suppression.
[0031] Preferably, the speaker's voice components are analyzed for howling risk based on the environmental adaptive critical frequency predictor to generate potential howling frequency points. The specific implementation method is as follows: Obtaining the speaker's voice component Next, step S4 first performs spectral analysis on the speech signal, extracting the energy features of each frequency point as input to the neural network prediction model. The speaker's speech components are then... Processed frame by frame, each frame is [length missing]. Each frame has 1000 sampling points with a 50% inter-frame overlap. A Hanning window is applied to each frame signal, followed by a Fast Fourier Transform to obtain the frequency domain representation. Where k is the frequency index, and its value ranges from 0 to... , t represents the total number of frequency points; t is the frame index, indicating the time sequence. Further calculation of the power spectral density at each frequency point is then performed. And use it as the input feature vector of the environment adaptive critical frequency predictor. Input feature vector The neural network prediction model, dynamically adjusted in step S2, is fed into the model. The forward propagation of the model calculates the probability value of howling risk corresponding to each frequency point. The neural network model uses a multilayer perceptron structure, including an input layer, several hidden layers, and an output layer. The number of nodes in the output layer is the same as the total number of frequency points K, and a sigmoid activation function is used to map the output values to a probability range between 0 and 1. The forward propagation calculation process can be represented as follows: , ,in, For the first The output vector of the hidden layer, ; The input layer vector, i.e., the input features. ; and The first The weight matrix and bias vector of the layer have been dynamically adjusted in step S2 according to the environmental acoustic characteristics; For the first The activation function for a layer is typically the ReLU function. ; The risk probability vector output by the model, where This represents the probability of a howling sound occurring at the k-th frequency point at time t; The sigmoid activation function is defined as follows: L represents the total number of layers in the neural network, including the output layer. Through the forward propagation calculation described above, the real-time howling risk probability for each frequency point in the current speech frame can be obtained.
[0032] Obtain the risk probability vector for each frequency point Then, step S4.2 compares the risk probability value of each frequency point with multiple preset risk thresholds to filter out potential howling frequency points and determine their priority. This invention sets three risk thresholds: a low-risk threshold... Medium risk threshold and high risk threshold ,satisfy For each frequency point k, based on its risk probability... The relationship between the risk level and the threshold values determines the risk level. and priority score The calculation formula is as follows: ,in, The priority score for frequency point k at time t is given; the higher the score, the more priority the frequency point needs to be processed. Let be the weighting coefficient, satisfying This is used to amplify the priority differences between different risk levels, and can be adopted. , , ; The preset risk threshold is set according to the system's requirements for howling sensitivity; a value of [value] can be selected. , , Priority scores will be assigned. The frequency points are marked as potential howling frequencies and sorted from high to low according to priority scores to generate a list of potential howling frequencies. This provides a basis for subsequent adaptive notch filtering.
[0033] To further improve the accuracy of howling prediction and avoid misjudgments caused by transient noise or speech harmonics, step S4.3 performs time series analysis on the marked potential howling frequencies to detect the trend of risk probability changes over time for each frequency. For each potential howling frequency k, its historical risk probability sequence over consecutive T frames is recorded. The trend of risk probability is determined by calculating the linear regression slope or the cumulative sum of first-order differences of the sequence. This invention uses the cumulative sum of first-order differences within a sliding window as a trend measure, calculated using the following formula: ,in, The risk probability trend value of frequency point k at time t is a positive value indicating that the overall risk probability is increasing, and a negative value indicates that it is decreasing; T is the window length of the time series analysis, that is, the number of historical frames considered, which is usually 5 to 10 frames. Let k be the risk probability of frequency point k in the m-th frame before the current time. These are weighting coefficients, used to assign higher weight to recent changes, and can be expressed in an exponential decay form. ,in As the attenuation factor, if taken .when Greater than the preset trend threshold When this occurs, it indicates that the risk probability of that frequency point is continuously rising and is about to enter a high-risk state. At this time, the frequency point is marked as an emergency howling frequency point, and its processing priority is further increased. Specifically, the priority score of the emergency howling frequency point is updated to... ,in This is the trend enhancement factor. Updated list of potential howling frequency points. This will guide the adaptive notch filtering and micro-delay perturbation in subsequent steps S5 and S6. Through the above time series analysis, the present invention can effectively distinguish between sudden speech peaks and real howling precursors, significantly reduce the probability of false triggering, and improve the intelligence and reliability of the feedback suppression system.
[0034] Preferably, adaptive notch filtering is performed on the speaker's speech component using the potential howling frequency point to form an initially suppressed amplified signal. The specific implementation method is as follows: After obtaining the list of potential howling frequencies and their corresponding risk levels, step S5.1 first configures the corresponding notch filter parameters for each potential howling frequency. The core parameters of the notch filter include notch depth, notch width, and quality factor (Q). The quality factor Q is defined as the ratio of the notch filter's center frequency to its bandwidth, reflecting the notch filter's selectivity. For frequencies with higher risk levels, a larger notch depth needs to be allocated to ensure effective suppression of impending howling, while a narrower notch width is allocated to minimize the impact on the sound quality of adjacent frequencies. Let the center frequency of the i-th potential howling frequency be... The corresponding risk level is The value is an integer from 1 to 3, corresponding to low, medium, and high risk, respectively. The notch depth is... Notch width and quality factor The configuration relationship can be represented as: ,in, Let be the notch depth of the i-th notch filter, representing the amount of attenuation of the signal at the center frequency; The reference notch depth is typically a constant between 6dB and 12dB. This is a depth adjustment factor, ranging from 0.2 to 0.5, used to control the gain factor of risk level on notch depth; The risk level of frequency point i is quantified by the priority score output in step S4; The notch width of the i-th notch filter, in Hz, is defined as the bandwidth at which the amplitude-frequency response of the notch filter drops by 3dB. The reference notch width is typically a constant between 10Hz and 20Hz. This is the width adjustment factor, ranging from 0.3 to 0.6, used to control the compression factor of the notch width based on the risk level; Let be the quality factor of the i-th notch filter, dimensionless, and derived from the center frequency. With notch width The ratio is calculated, and a higher quality factor indicates better selectivity of the notch filter and less impact on adjacent frequency points.
[0035] After completing the parameter configuration, step S5.2 connects the configured multiple adaptive notch filters in series to the signal processing link according to frequency priority, and processes the speaker's voice components. Notch filtering is performed sequentially. Each notch filter adopts a second-order IIR notch filter structure, and its transfer function can be expressed as: ,in, Let z be the transfer function of the i-th notch filter, and z be a complex variable; For normalized digital angular frequency, The center frequency of the notch filter. The system sampling rate; The distance parameter from the pole to the unit circle, and the notch width. and quality factor Relevant, satisfy ; The cosine value corresponding to the center frequency; and These are the unit delay and two-unit delay operators, respectively. The amplitude-frequency response of this second-order IIR notch filter at the center frequency... A notch is formed at the notch, and the depth of the notch is determined by the relationship between the coefficients of the numerator and denominator. The actual notch depth is... With parameters Relationship satisfaction Multiple notch filters are connected in series according to priority, meaning the input of the (i+1)th notch filter is the output of the ith notch filter. The final output signal is... Where K is the total number of currently active notch filters. This is the initial suppression of the amplified signal after processing by all notch filters. By using a series connection, each notch filter works independently and without interference, enabling precise attenuation of multiple potential howling frequencies simultaneously.
[0036] During notch filtering, step S5.3 monitors the actual energy changes of each potential howling frequency in real time, and combines this with the residual reentry sound component output from step S3.3. The energy change trend is comprehensively judged, and the release timing of the notch filter is dynamically adjusted to avoid unnecessary damage to the sound quality. The short-time energy of the i-th potential howling frequency at time t is defined as... ,in The signal after notch filtering At frequency The nearby spectral components, where M is the frame length for short-time energy analysis. Simultaneously, the short-time energy of the residual reentry sound component is defined as... Comprehensive judgment indicators Defined as a weighted combination of the two: ,in, This is a comprehensive risk index for the i-th frequency point at time t, used to determine whether to release the notch filter at that frequency point; Let i be the short-time energy at frequency i at the current moment; The reference energy level of frequency point i under normal amplification conditions can be obtained through historical data statistics; This represents the short-time energy of the residual re-entry component at the current moment; This is the reference energy level for the residual reentry sound component; This is a weighting coefficient, ranging from 0.3 to 0.7, used to balance the importance of the frequency's own energy and residual reentry sound energy in the judgment. When continuous Frames below a preset safety threshold When this time is reached, it indicates that the whistling risk at that frequency point has been eliminated. At this point, the notch depth of the corresponding notch filter is gradually reduced until it is completely released. The notch depth release adopts a linear attenuation method, with an attenuation amount per frame. ,in The release time constant is typically set to 50 to 100 frames to avoid sudden changes in sound quality caused by the abrupt release of the notch filter. Through the aforementioned real-time monitoring and dynamic release mechanism, this invention effectively suppresses feedback while preserving the original sound quality to the maximum extent, achieving a balance between feedback suppression and sound fidelity.
[0037] Preferably, a random micro-delay perturbation is applied to the initially suppressed amplified signal based on the potential howling frequency point to generate the final suppressed amplified signal. The specific implementation method is as follows: The initial suppressed amplification signal after adaptive notch processing is obtained. Subsequently, step S6 further disrupts the conditions for constructing the feedback loop at the physical sound field level through multi-channel random micro-delay perturbation technology. When the sound waves emitted by the loudspeaker array superimpose in space, if the signals of each channel are in phase, an in-phase enhancement region will be formed at a specific location, increasing the risk of feedback; conversely, if the phases of the signals of each channel are randomly distributed, stable in-phase superposition can be effectively avoided, thereby cutting off the physical basis of the positive feedback loop. Based on the distribution range and risk level of potential howling frequencies, this invention independently generates a set of pseudo-random time delay sequences for each channel in the multi-channel loudspeaker array, so that the phase relationship of the sound waves of each channel in space presents a random distribution.
[0038] The core of step S6.1 lies in determining the amplitude range of the time delay perturbation based on the wavelength characteristics of the potential howling frequencies. High-frequency sound waves have short wavelengths and strong spatial coherence, requiring more precise time perturbations to effectively disrupt their phase consistency; low-frequency sound waves have long wavelengths and are relatively less sensitive to time delay perturbations. Therefore, this invention makes the maximum time delay amplitude of the pseudo-random time delay sequence negatively correlated with the wavelength of the potential howling frequencies covered by the channel. Let the set of potential howling frequencies corresponding to the j-th speaker channel be denoted as . ,in This represents the number of frequencies that require close monitoring for this channel. The pseudo-random delay sequence for this channel. The value of time t is generated by the following formula: ,in, Let be the time delay perturbation value applied to j channels at time t, in seconds; The maximum delay of the channel is determined by the highest frequency component among the potential howling frequencies covered by the channel; Let be the normalized pseudorandom number generator function for the j-th channel, with a value range of . Furthermore, they are independent of each other in different channels; The highest frequency among the potential howling frequencies covered by channel j is the frequency with the shortest wavelength and the one that requires the most precise perturbation; k is a scaling factor, ranging from 0.1 to 0.5, used to control the overall magnitude of the maximum delay amplitude, ensuring that the perturbation is imperceptible to the human ear within the range of 0.1 milliseconds to 1 millisecond.
[0039] After obtaining the pseudo-random time delay sequence for each channel, step S6.2 will initially suppress the amplified signal. The audio signal is copied to each channel of the multi-channel speaker array, and a pseudo-random time delay perturbation corresponding to that channel is applied to the audio signal of each channel. Due to the time delay value... Typically, the time delay is not an integer multiple of the sampling period, requiring a fractional delay filter to achieve sub-sampling level precision time delay compensation. This invention employs a fractional delay filter based on Lagrange interpolation, whose time-domain response is similar to the fractional delay filter structure used in step S1, but here it is applied to multi-channel signal allocation rather than signal alignment. For channel j, the output signal after applying the time delay perturbation... It can be represented as: ,in, Let be the output signal of the j-th speaker channel at time n; L is the order of the fractional delay filter, ranging from 4 to 8; Based on the delay value at the current moment The k-th filter coefficient of the j-th channel is calculated in real time using the Lagrange interpolation formula. This is a historical sample of the initial suppression of the amplified signal. Through the above processing, the sound waves emitted by each channel loudspeaker exhibit small, independent, and randomly changing shifts in time, thus making their phase relationship in space randomized and effectively avoiding the formation of stable in-phase superposition regions.
[0040] During the disturbance application process, step S6.3 monitors the energy changes of each potential howling frequency point in the amplified signal after the disturbance in real time. When an upward trend in the energy of a certain frequency point is detected, the disturbance amplitude and rate of change of the pseudo-random time delay sequence of the corresponding channel are dynamically adjusted to form closed-loop feedback control. The j-th channel is defined as the channel at the m-th potential howling frequency point. The short-time energy at the location is This can be obtained through bandpass filtering and short-time energy calculation. A comprehensive assessment of the overall howling risk trend of this channel is then conducted. The calculation formula is as follows: ,in, Let be the comprehensive risk index of the j-th channel at time t. It is dimensionless and used to determine whether the disturbance intensity of the channel needs to be increased. The number of potential howling frequency points corresponding to channel j; The weighting coefficient for frequency point m is proportional to the risk level of that frequency point and can be determined by the priority score output in step S4. The short-time energy of frequency point m at the current moment; This represents the reference energy level of frequency point m under normal amplification conditions. This represents the energy gradient at frequency point m, with a positive value indicating an increase in energy. The gradient enhancement coefficient, taking a positive value between 2 and 5, is used to amplify the impact of the upward energy trend on risk indicators; exponential function This is used to ensure that when energy rises rapidly, the risk indicator grows non-linearly, thereby triggering a more aggressive disturbance adjustment. Exceeding the preset disturbance trigger threshold At that time, the system dynamically adjusts the pseudo-random time delay sequence parameters of the channel: increasing the maximum time delay amplitude. To enhance the perturbation strength and simultaneously increase the rate of change of the time delay sequence, until... The signal returns to a safe range. Through the aforementioned adaptive adjustment mechanism, this invention achieves dynamic control of the spatial sound field coherence, effectively blocking the formation of feedback loops at the physical level. Working in conjunction with front-end signal processing methods, it ultimately produces a high-quality, feedback-suppressed amplification signal. .
[0041] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0042] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A local amplification method based on acoustic feedback suppression technology, characterized in that, Includes the following steps: S1. Simultaneous acquisition of the speaker's voice and the surrounding sound field is performed using a bone conduction sensor and an air conduction microphone to obtain a clean speech reference signal and a mixed audio signal. S2. Based on the transmitted detection signal, perform acoustic measurements on the amplified space and adjust the neural network prediction model according to the measurement results to obtain an environment-adaptive critical frequency predictor; S3. Adaptive filtering is performed on the mixed audio signal using the pure speech reference signal to obtain the speaker's speech component; S4. Based on the environmental adaptive critical frequency predictor, perform feedback risk analysis on the speaker's voice components to generate potential feedback frequency points. S5. Using the potential howling frequency point, perform adaptive notch filtering on the speaker's voice component to form an amplified signal with initial suppression. S6. Apply a random micro-delay perturbation to the initially suppressed amplified signal based on the potential howling frequency point to generate the final suppressed feedback amplified signal.
2. The local amplification method based on acoustic feedback suppression technology according to claim 1, characterized in that, Step S1 includes: S1.
1. A pure speech reference signal containing only the speaker's voice is picked up by a bone conduction sensor attached to the speaker's skull or mandible. S1.
2. A mixed audio signal containing the speaker's voice, speaker re-entry sound, and ambient noise is picked up by an air conduction microphone placed at the speaker's collar or headgear. S1.3 Perform time delay estimation and compensation on the pure speech reference signal and the mixed audio signal to align the speaker's voice components in the two signals on the time axis.
3. The local amplification method based on acoustic feedback suppression technology according to claim 1, characterized in that, Step S2 includes: S2.
1. A broadband detection signal that covers the target frequency band of the loudspeaker array of the loudspeaker system and is not easily detected by the human ear is emitted into the loudspeaker space, and the reflected sound of the detection signal is received by the air conduction microphone. The reverberation time and frequency response characteristics of the loudspeaker space are calculated based on the impulse response between the emitted signal and the received signal. S2.2 Using the calculated reverberation time and frequency response characteristics as input parameters, dynamically adjust the internal weights and thresholds of the pre-trained neural network prediction model to adapt the neural network prediction model to the acoustic environment of the current sound reinforcement space. The neural network prediction model is used to output the probability of howling at each frequency point based on the spectrum of the input audio signal.
4. The local amplification method based on acoustic feedback suppression technology according to claim 3, characterized in that, Step S2 further includes dynamically adjusting the internal weights and thresholds of the neural network prediction model: The calculated reverberation time and frequency response characteristics are introduced as constraints into the backpropagation algorithm of the neural network, so that the model is guided by the goal of matching the current acoustic environment during the weight update process. Guided by the constraints, the connection weights and bias thresholds of each layer of the neural network are iteratively corrected through the backpropagation algorithm, so that the acoustic features extracted by the feature extraction layer of the model can be adaptively matched with the acoustic modes of the current sound reinforcement space.
5. A local amplification method based on acoustic feedback suppression technology according to claim 1, characterized in that, Step S3 includes: S3.
1. Use the clean speech reference signal after time delay compensation and spectrum matching as the desired input of the adaptive filter, and use the mixed audio signal as the original input of the adaptive filter to initialize the filter coefficients. S3.
2. The error signal between the filter output signal and the desired input is calculated iteratively through an adaptive filtering algorithm, and the filter coefficients are updated in real time according to the error signal, so that the filter output gradually approaches the speaker's voice component in the mixed audio signal. S3.
3. The estimated signal output by the converged adaptive filter is used as the extracted speaker speech component, and the error signal is used as the residual reentry sound component and input to the feedback suppression effect evaluation module. The feedback suppression effect evaluation module judges the current feedback suppression state according to the energy change trend of the residual reentry sound component. When the energy of the residual reentry sound component continues to rise, the notch depth enhancement of the adaptive notch processing in the subsequent steps is triggered.
6. A local amplification method based on acoustic feedback suppression technology according to claim 5, characterized in that, Step S3.2 includes: The error signal between the filter output signal and the desired input is calculated iteratively using an adaptive filtering algorithm, and the convergence step size and leakage factor of the adaptive filter are dynamically adjusted according to the amplitude and spectral characteristics of the error signal. When a sudden pulse interference is detected in the error signal, the update of the filter coefficients is temporarily frozen to avoid coefficient divergence, and the update is resumed after the pulse interference ends.
7. A local amplification method based on acoustic feedback suppression technology according to claim 5, characterized in that, Step S3.3 includes: The estimated signal output by the converged adaptive filter is used as the extracted speaker speech component, and the error signal is used as the residual re-entry sound component and input to the feedback suppression effect evaluation module. The feedback suppression effect evaluation module determines the current feedback suppression state based on the energy change trend of the residual reentry sound component. When the energy of the residual reentry sound component continues to rise, it triggers the notch depth enhancement of the adaptive notch processing in subsequent steps.
8. A local amplification method based on acoustic feedback suppression technology according to claim 1, characterized in that, The S4 step includes: S4.1 Input the spectral data of the speaker's voice components into the environment adaptive critical frequency predictor, and calculate and output the howling risk probability value corresponding to each frequency point through the forward propagation of the neural network prediction model. S4.
2. Compare the howling risk probability value of each frequency point with multiple preset risk thresholds. Based on the comparison results, mark the frequency points that exceed the first threshold as potential howling frequency points, and sort the potential howling frequency points according to their risk probability. S4.3 Perform time series analysis on the marked potential howling frequencies to detect the trend of risk probability of each frequency over time, mark the frequencies with continuously rising risk probability as emergency howling frequencies and increase their processing priority.
9. A local amplification method based on acoustic feedback suppression technology according to claim 1, characterized in that, Step S5 includes: S5.1 According to the potential howling frequency points and their corresponding risk levels, configure the corresponding notch depth, notch width and notch quality factor for each potential howling frequency point. The higher the risk level of the frequency point, the larger the notch depth and the narrower the notch width are allocated. S5.2 Connect the configured adaptive notch filters in series to the signal processing link according to the frequency priority order, and perform notch processing on the speaker's voice components in sequence. Each notch filter only attenuates the frequency components near its corresponding potential howling frequency. S5.3 During the notch filtering process, monitor the actual energy changes of each potential howling frequency point in real time, and make a comprehensive judgment based on the energy change trend of the residual reentry sound component. When the energy of a certain frequency point drops below the safe threshold, gradually reduce the notch depth of the corresponding notch filter until it is completely released, so as to avoid unnecessary damage to the sound quality.
10. A local amplification method based on acoustic feedback suppression technology according to claim 1, characterized in that, Step S6 includes: S6.1 Based on the distribution range and risk level of the potential howling frequency points, generate a set of pseudo-random time delay sequences independently for each channel in the multi-channel loudspeaker array. The maximum time delay amplitude of the pseudo-random time delay sequence is negatively correlated with the wavelength of the potential howling frequency points covered by the channel, so that the high-frequency sound waves can obtain more refined time perturbation, so as to effectively destroy their spatial coherence. S6.
2. The initially suppressed amplified signal is copied and distributed to each channel of the multi-channel speaker array, and a pseudo-random time delay perturbation is applied to the audio signal of each channel, so that the phase relationship of the sound waves emitted by each channel speaker in space presents a randomized distribution. S6.3 Real-time monitoring of energy changes at potential howling frequencies in the amplified signal after perturbation. When an upward trend in energy at a certain frequency is detected, the perturbation amplitude and rate of change of the pseudo-random time delay sequence of the corresponding channel are dynamically adjusted.