Howling / feedback suppression method and device in sound reinforcement scene, medium and product

Through the iterative step length of frequency domain adaptive filtering and recursive neural network model control, the problem of howling/feedback suppression in sound-stretching scenarios is solved, significantly improving signal quality and maximum gain.

CN120034813AActive Publication Date: 2025-05-23TRUE SPACE (ZHUHAI) TECH CO LTD

Patent Information

Application Number
CN202510479617.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-05-23
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

In sound reinforcement scenarios, the prior art is difficult to effectively suppress howling/feedback, especially in complex acoustic environments, resulting in limited signal distortion and maximum gain improvement.

Method used

Through frequency domain adaptive filtering processing, combined with the pre-trained recursive neural network model to control iteration step length, the acoustic transfer function is estimated and iterated, thereby achieving effective suppression of howling/feedback.

Benefits of technology

It significantly improves the distortion and maximum gain in sound-reinforced scenes, and can cover the accuracy of acoustic feedback estimation in more actual scenes, ensuring high-quality sound restoration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120034813A_ABST
    Figure CN120034813A_ABST
Patent Text Reader

Abstract

The invention provides a howling / feedback suppression method and device in a sound reinforcement scene, a medium and a product. The method comprises the following steps: acquiring a microphone signal of a current frame; frequency domain self-adaptive filtering processing is carried out on the microphone to be processed, and a predicted target signal is output to the loudspeaker, and the frequency domain self-adaptive filtering processing comprises the following steps: multiplying the estimated acoustic transfer function of the previous frame by the loudspeaker signal of the previous frame in the frequency domain to obtain a feedback estimation signal, and then subtracting the feedback estimation signal from the microphone signal to obtain a predicted target signal; a residual signal is obtained; the residual signal is used as a prediction target signal to be output to a loudspeaker after being subjected to inverse transformation, meanwhile, iteration is carried out on an estimated acoustic transfer function of a previous frame, an estimated acoustic transfer function of a current frame is obtained, and the step length of iteration is controlled through a recurrent neural network model. According to the method, the accuracy of acoustic feedback estimation in various actual scenes can be improved, the distortion degree in a sound reinforcement scene is improved, and the maximum gain is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of audio processing technology, and in particular to a howling / feedback suppression method, device, medium and product in a sound reinforcement scenario. Background Art

[0002] Howling is a common problem in audio systems, and is common in conference systems, classroom sound reinforcement, lectures, KTV, and hearing aids. The root cause of howling is that the sound output by the speaker is picked up and amplified by the microphone again, forming a positive feedback loop. In this loop, when certain specific frequency signals are rapidly accumulated and amplified in the feedback loop, a harsh howling sound is generated. The current howling suppression technologies mainly include frequency shifting, notching, deep speech enhancement, and adaptive filtering.

[0003] Frequency shifting technology appeared earlier and is relatively simple to implement. Frequency shifting can shift the peak of the feedback signal, so that the signal is repeatedly superimposed in the feedback loop at different peaks, thereby slowing down the speed of feedback amplification. Frequency shifting can slightly increase the gain when howling occurs, usually by 1dB-2dB. Frequency shifting will cause significant distortion, and the gain that can be increased is limited.

[0004] Notch technology includes static notch and adaptive notch. Static notch requires measurement of the specific acoustic environment and design of a series of notch filters to filter the frequency band where howling is most likely to occur. Adaptive notch automatically estimates the howling frequency through howling detection, and automatically designs filters to suppress the howling frequency. The detection methods of howling frequencies can be divided into two categories: fixed rules based on spectral features and deep learning methods. However, the use of notch filters will cause signal distortion. For complex acoustic environments, such as multiple speakers, multiple microphones, and scenes with moving microphones, there may be too many frequency points that produce howling. The use of a large number of notch filters will cause serious distortion and make the voice unintelligible. When howling detection is relatively reliable, the gain when howling occurs can usually be increased by 6dB-8dB through notching.

[0005] Deep speech enhancement technology is a method of using deep learning to enhance speech to achieve howling suppression. These methods do not rely on known speaker signals or predictions of acoustic transfer functions. They directly remove howling sounds from microphone signals by eliminating noise through neural network models. Because they do not rely on the estimation of acoustic transfer functions, they are very robust to changes in the acoustic environment. However, since known speaker signals are not used as references, it is difficult for the model to accurately estimate the feedback signal and stably restore the original sound source. In particular, when the feedback is strong, the microphone signal-to-noise ratio is low, and the signal after speech enhancement is often severely distorted. In addition, sounds with characteristics similar to howling, such as some musical instruments, are easily eliminated by mistake.

[0006] Adaptive filtering technology is a signal processing technology that can continuously iterate its parameters through the real-time input and output of the system to minimize the predefined error. It estimates the feedback signal and eliminates it according to the estimated feedback signal to achieve the purpose of restoring the target voice. Since the speaker signal is known, when the acoustic transfer function is completely known, not only can the occurrence of howling be completely avoided, but the feedback can also be perfectly eliminated to truly restore the target sound source. Therefore, there is no upper limit to the maximum gain improvement of perfect adaptive filtering. The difficulty of adaptive filtering lies in the estimation of the acoustic transfer function. Adaptive filtering continuously iterates and adjusts the estimate of the acoustic transfer function based on the result of feedback elimination to slowly approach the real acoustic transfer function. Adaptive filtering is divided into time domain adaptive filtering and frequency domain adaptive filtering. According to the domain of the estimated acoustic transfer function, time domain adaptive filtering is implemented through time domain convolution, which usually has a small delay, but is not able to fit longer acoustic transfer functions. It is widely used in the hearing aid field. Frequency domain adaptive filtering processes the signal in the frequency domain, replacing the convolution operation in the time domain adaptive filtering with multiplication. The time complexity increases slowly with the filter length. Therefore, the filter that can be used is longer in time and has a stronger fitting ability for long acoustic transfer functions, which is suitable for sound reinforcement scenarios.

[0007] In traditional frequency domain adaptive filtering, the control of acoustic transfer function iteration can be classified into Wiener filtering, Kalman filtering, etc. based on different theoretical assumptions. Assumptions include that the acoustic transfer function will not change suddenly, and the feedback signal is not related to the current target signal. In practical applications, these assumptions are often not met, such as when the direction and position of the microphone may change suddenly, and the feedback signal is highly correlated with the target signal when the sound is prolonged in singing. At this time, the estimation of the acoustic transfer function and even the feedback will produce huge errors, and even the sound reinforcement will be seriously distorted or there will be no howling suppression effect. An existing solution proposes to predict the sudden change of the feedback path through a neural network to overcome the problem of locking in the Kalman filter, which solves a problem of statistical adaptive filtering in a targeted manner, but cannot systematically solve the defects of statistical adaptive filtering. Another existing solution uses a neural network to assist adaptive filtering, but it still uses a statistical method for adaptive filtering, but only uses a neural network to judge a step lock situation and resets it in a targeted manner, which cannot meet the requirements of multi-sound scenes. Summary of the invention

[0008] The first object of the present invention is to provide a howling / feedback suppression method in a sound reinforcement scenario, which can improve the accuracy of acoustic feedback estimation in a variety of actual scenarios, improve the distortion in the sound reinforcement scenario and enhance the maximum gain.

[0009] A second object of the present invention is to provide a computer device for implementing the howling / feedback suppression method in the above-mentioned sound reinforcement scenario.

[0010] A third object of the present invention is to provide a computer-readable storage medium for implementing the howling / feedback suppression method in the above-mentioned sound reinforcement scenario.

[0011] A fourth object of the present invention is to provide a computer program product for implementing the howling / feedback suppression method in the above-mentioned sound reinforcement scenario.

[0012] In order to achieve the above-mentioned first purpose, the present invention provides a howling / feedback suppression method in a sound reinforcement scenario, which includes the following steps: obtaining a microphone signal of a current frame; performing frequency domain adaptive filtering processing on the signal of the current frame, and outputting a predicted target signal to a loudspeaker: wherein the frequency domain adaptive filtering processing includes the following steps: in the frequency domain, multiplying the estimated acoustic transfer function of the previous frame by the loudspeaker signal of the previous frame to obtain a feedback estimation signal; in the frequency domain, subtracting the feedback estimation signal from the microphone signal of the current frame to obtain a residual signal; the residual signal is inversely transformed as the predicted target signal, and is output to the loudspeaker through a feedforward path, and at the same time, the estimated acoustic transfer function of the previous frame is iterated in the direction in which the residual signal of the current frame and the loudspeaker signal of the previous frame tend to have a smaller cross-correlation, to obtain the estimated acoustic transfer function of the current frame, and the iteration step size is controlled by a pre-trained recursive neural network model.

[0013] It can be seen from the above scheme that the present invention suppresses howling through frequency domain adaptive filtering, wherein the step size is controlled by a pre-trained recursive neural network model to iteratively estimate the acoustic transfer function, replacing the traditional adaptive filtering method based on statistical assumptions to control the adaptive process. It can cover the accuracy of acoustic feedback estimation in more actual scenarios and significantly improve the distortion and maximum gain in the sound reinforcement scenario.

[0014] A further solution is that the recursive neural network model includes the following steps during the training process: simulating the changing impulse responses on different moving trajectories in rooms of different sizes; simulating the feedback path of the collected voice and music data through the generated impulse response, and processing it using adaptive filtering, wherein the iterative step size of the adaptive filtering is controlled by the recursive neural network model, and the loss function is the mean square error between the simulated feedback signal and the amplitude spectrum of the feedback signal estimated in the adaptive filtering in the overall simulation process.

[0015] It can be seen that the collected speech and music data as training data covers diversified scenarios such as sudden changes in acoustic functions, prolonged sounds, strong background noise, etc. In the adaptive process, the iterative step size of estimating the acoustic transfer function depends on the minimization regression of the feedback signal estimation error in the training data, ensuring the lowest feedback estimation and the most accurate sound restoration in various real scenarios. This solves the problem that traditional adaptive filtering does not model real audio data and is not optimized enough in many real scenarios.

[0016] A further solution is that, when the recursive neural network model is trained, the collected voice and music data are used to simulate the feedback path through the generated impulse response, and the feedback gain corresponding to the feedback path is random within a set range, wherein at the beginning of training, the feedback gain needs to be set below the critical gain.

[0017] It can be seen that avoiding premature divergence of the acoustic transfer function estimate at the beginning of the training process allows the model to learn better.

[0018] A further solution is that the input of the recursive neural network model is a vector formed by combining the amplitude spectrum of the microphone signal of each frame with the amplitude spectrum of the residual signal, and the output is a vector of the same length as the input as a hidden layer. The hidden layer passes through a linear layer and is activated using a logistic function, and outputs a single scalar as the shared step size for iteration of all subbands of the current frame.

[0019] A further solution is that the input of the recursive neural network model is a vector formed by combining the amplitude spectrum of the microphone signal of each frame with the amplitude spectrum of the residual signal, and the output is a vector of the same length as the input as a hidden layer. The hidden layer passes through a linear layer and is activated using a logistic function to output the step size adopted in each subband iteration of the current frame.

[0020] It can be seen from this that the recursive neural network model can also output the step size adopted when iterating each subband of the current frame, thereby improving the iterative effect of estimating the acoustic transfer function.

[0021] A further solution is that the model structure of the recurrent neural network model is LSTM or GRUs.

[0022] A further solution is to iterate the estimated acoustic transfer function of the previous frame in the direction where the residual signal of the current frame and the speaker signal of the previous frame tend to have a smaller cross-correlation, which is expressed as: , is the gradient of the mth subband during the iteration of the estimated acoustic transfer function of the k-1th frame, represents the mth subband of the residual signal of the kth frame, is the mth subband of the loudspeaker signal of the kth frame, for The complex conjugate of is the recursive smoothing of the loudspeaker spectral energy of the m-th subband of the k-th frame, ,in, is the smoothing coefficient; get the estimated acoustic transfer function of the current frame: ,in, is the mth subband of the estimated acoustic transfer function of the kth frame, is the mth subband of the estimated acoustic transfer function of the k-1th frame, is the iteration step size.

[0023] In order to achieve the above-mentioned second purpose, the present invention provides a computer device, including a processor and a memory, wherein: a computer program is stored in the memory, and when the computer program is executed by the processor, the above-mentioned howling / feedback suppression method based on neural network controlled adaptive filtering is implemented.

[0024] In order to achieve the third objective mentioned above, the present invention provides a computer-readable storage medium, on which a computer program is stored, wherein: when the computer program is executed by a processor, the howling / feedback suppression method in the sound reinforcement scenario mentioned above is implemented.

[0025] In order to achieve the fourth objective mentioned above, the present invention provides a computer program product, including computer instructions, wherein: when the computer instructions are executed by a processor, the howling / feedback suppression method in the sound reinforcement scenario mentioned above is implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 It is a flow chart of an embodiment of a howling / feedback suppression method in a sound reinforcement scenario of the present invention.

[0027] Figure 2 It is a principle block diagram of an embodiment of a howling / feedback suppression method in a sound reinforcement scenario of the present invention.

[0028] The present invention is further described below in conjunction with the accompanying drawings and embodiments. DETAILED DESCRIPTION

[0029] The method for suppressing howling / feedback in a sound reinforcement scenario of the present invention improves the prediction accuracy of the estimated acoustic transfer function in a variety of actual scenarios by iterating the control step length of a recursive neural network trained on a large amount of real data, thereby improving the maximum gain and distortion that can be achieved in the sound reinforcement scenario. The present invention also provides a computer device, a computer-readable storage medium, and a computer program product for implementing the above-mentioned method for suppressing howling.

[0030] Embodiment of the method for suppressing howling / feedback in a sound reinforcement scenario: This embodiment is described in an indoor sound reinforcement scenario. A microphone and a speaker are set up indoors, and the signal collected by the microphone is processed by the adaptive filtering system and then output to the speaker. In this scenario, the target sound source is collected by the microphone and then output to the speaker. The sound played by the speaker is collected by the microphone after being reflected by the surrounding environment. The feedback signal collected by the microphone is analyzed and eliminated in real time by the adaptive filtering system, so that the howling phenomenon will not occur during the sound reinforcement process of the target sound source, achieving the purpose of howling suppression.

[0031] Assume that the speaker signal is played, the sound wave is picked up by the microphone after being reflected in the room, and is played again by the speaker to form a closed-loop system. The acoustic transfer function is usually simplified to a linear time-invariant system (Linear Time-Invariant, LTI), based on the assumption that the actual physical environment is stable and the positions of the sound source and receiver are fixed. The Acoustic Transfer Function (ATF) is a mathematical representation that describes the changes in the audio signal during the process of sound propagating from the sound source to the receiving point. It basically describes how the acoustic system "transfers" or "transforms" the input signal.

[0032] Therefore, in this closed-loop system with feedback, the microphone signal to be processed is expressed as ,in, represents the parameters of a linear system of length L at time point n, represents the last L samples of the loudspeaker signal at time point n, The signal representing the target sound source is what the adaptive filtering system needs to obtain from The target of the restore.

[0033] The adaptive filtering system implements howling suppression by executing the howling / feedback suppression method in the sound reinforcement scenario of this embodiment. Specifically, the adaptive filtering system is implemented by a computer program, see Figure 1 , when the computer program is executed, it includes implementing the following steps: S11: Obtain the microphone signal of the current frame.

[0034] S12: Perform frequency domain adaptive filtering on the microphone signal of the current frame.

[0035] S13: Output the predicted target signal to the speaker, and iterate the estimated acoustic transfer function of the previous frame to obtain the estimated acoustic transfer function of the current frame.

[0036] See also Figure 2 In the above step S11, the microphone signal of the current frame in the frequency domain collected by the microphone 20 Includes the target sound source signal , and the speaker signal of the previous frame After entering the speaker 30, the speaker 30 outputs the actual feedback signal obtained by the actual acoustic transfer function H .

[0037] The duration of each frame of microphone signal is the same as that of each frame of loudspeaker signal and they are divided into the same number of sub-bands in the frequency domain after short-time Fourier transform. The m-th sub-band of the microphone signal of the current frame (the k-th frame) is represented as , the mth subband of the speaker signal of the previous frame (k-1th frame) is expressed as The mth subband of the estimated acoustic transfer function of the previous frame (k-1th frame) is expressed as In the above step S12, the feedback estimation signal is obtained by multiplying the estimated acoustic transfer function of the previous frame with the speaker signal of the previous frame in the frequency domain, which specifically includes: in the frequency domain, the mth subband of the estimated acoustic transfer function of the previous frame is obtained. The mth subband of the loudspeaker signal of the previous frame Multiply them together to get the mth subband of the feedback estimation signal of the kth frame Then, in the frequency domain, the microphone signal of the current frame is subtracted from the feedback estimation signal to obtain the residual signal, which specifically includes: subtracting the corresponding subband of the feedback estimation signal from each subband of the microphone signal of the current frame, that is, obtaining the residual signal corresponding to the subband. The mth subband of the residual signal of the kth frame is expressed as , .

[0038] In the above step S13, the residual signal is subjected to inverse short-time Fourier transform to obtain a predicted target signal in the time domain, and the predicted target signal is output to the speaker 30 through the feedforward path G. The feedforward path G refers to the channel through which the signal is transmitted from the system input to the output, and usually includes post-processing such as a limiter. At the same time, the estimated acoustic transfer function of the previous frame is iterated in the direction in which the residual signal of the current frame and the speaker signal of the previous frame tend to have a smaller cross-correlation, and the estimated acoustic transfer function of the current frame is obtained, and the step size of the iteration is controlled by the pre-trained recursive neural network model 10.

[0039] In the process of iterating the estimated acoustic transfer function of the previous frame, each subband is independently processed. Specifically, while determining the speaker output, the estimated acoustic transfer function of the previous frame The residual signal towards the current frame The speaker signal of the previous frame Iterate in the direction that approaches smaller cross-correlation. The process is expressed as: , is the gradient of the mth subband during the iteration of the estimated acoustic transfer function of the kth frame, represents the residual signal of the mth subband of the kth frame, is the mth subband of the loudspeaker signal of the kth frame, for The complex conjugate of is the recursive smoothing of the loudspeaker spectral energy of the m-th subband of the k-th frame, ,in, is the smoothing coefficient, which is empirically set between 0.8 and 0.9, so that the mth subband of the estimated acoustic transfer function of the kth frame can be obtained. ,in, is the mth subband of the estimated acoustic transfer function of the kth frame, is the mth subband of the estimated acoustic transfer function of the k-1th frame, is the iteration step size, that is, in the iterative process of the estimated acoustic transfer function of the k-1th frame, the step size of each subband is used for iteration The step size of each frame iteration is controlled by a pre-trained recurrent neural network model to find the most suitable step size relative to the actual feedback signal.

[0040] The input of the recursive neural network model of this embodiment is a vector formed by combining the amplitude spectrum of the microphone signal of each frame with the amplitude spectrum of the residual signal, and the output is a vector of the same length as the input as a hidden layer. The hidden layer passes through a linear layer and is activated using a logistic function, and a single scalar is output as the shared step size for iteration of all subbands of the current frame, as shown below: , where RNN is the above recursive neural network model, It is the hidden state of the recurrent neural network, which is used to remember the previous information in the time series and pass it backward.

[0041] In various embodiments, the output of the pre-trained RNN model is a vector , the vector Indicates the step size used for each subband iteration in the current frame, that is, , so that in the process of iterating the estimated acoustic transfer function of the previous frame to obtain the estimated acoustic transfer function of the current frame, the iteration is performed with the step size corresponding to each subband in the current frame. For example, the mth subband of the estimated acoustic transfer function of the kth frame , the m-1th subband of the estimated acoustic transfer function of the kth frame , the m-2th subband of the estimated acoustic transfer function of the kth frame .

[0042] Optionally, the model structure of the recurrent neural network model is Long Short-Term Memory (LSTM) or Gated recurrent units (GRUs).

[0043] The input and output of the recursive neural network model during the training process are generated in real time through the simulation feedback environment. The training process includes the following steps: Simulate the changing impulse responses on different moving trajectories in rooms of different sizes, which can be specifically achieved through the existing IMAGE method. The collected voice and music data are simulated through the generated impulse response feedback path and processed using adaptive filtering, where the iterative step size of the adaptive filtering is controlled by a recursive neural network model, and the loss function is the mean square error between the simulated feedback signal and the amplitude spectrum of the feedback signal estimated in the adaptive filtering during the entire simulation process. The loss function comes from the entire simulation process to avoid overfitting the model's feedback prediction for a certain frame. The feedback path refers to the path of the signal output by the system from the output end to the input end. The feedback path in the sound reinforcement scenario is the actual acoustic transfer function mentioned above.

[0044] When the recursive neural network model is trained, the collected voice and music data are simulated through the generated impulse response in the feedback path, and the feedback gain corresponding to the feedback path is random within the set range. Among them, in the initial stage of training, the feedback gain needs to be set below the critical gain to avoid the estimated acoustic function from diverging too early, so that the model cannot learn. As the accuracy of the model effect improves, howling can be avoided at higher gains, and the feedback gain range in the simulation can be continuously improved to allow the model to gradually have the ability to cope with higher feedback gains. In the context of howling suppression, the critical gain refers to the maximum gain value that the acoustic system can withstand before the howling begins to occur. It is usually expressed in decibels (dB). The critical gain is also called the maximum stable gain (MSG), which is the upper limit of the gain at which the system can operate safely in practical applications.

[0045] In summary, the present invention suppresses howling by means of frequency domain adaptive filtering, wherein the step size is controlled by a pre-trained recursive neural network model to iterate the estimated acoustic transfer function, replacing the traditional adaptive filtering method based on statistical assumptions to control the adaptive process. The advantages are that there is no theoretical upper limit for the maximum stable gain (ASG) improvement, and no processing method that damages the sound quality is used, ensuring that the sound after processing has a high degree of restoration. The howling suppressor or feedback suppressor implemented based on the present invention can cover the accuracy of acoustic feedback estimation in more actual scenarios, and significantly improve the distortion and maximum gain in the sound reinforcement scenario.

[0046] Computer device embodiment: The computer device of this embodiment includes a processor and a memory. The memory stores a computer program. When the processor executes the computer program, the howling / feedback suppression method embodiment in the sound reinforcement scenario is implemented.

[0047] The computer device may include but is not limited to a processor and a memory. Those skilled in the art will appreciate that the computer device may include more or fewer components, or a combination of certain components, or different components, for example, the computer device may also include input and output devices, network access devices, buses, etc.

[0048] For example, the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microcontroller or any conventional processor, etc. The processor is the control center of a computer device, and uses various interfaces and lines to connect various parts of the entire computer device.

[0049] The memory can be used to store computer programs and / or modules. The controller realizes various functions of the computer device by running or executing the computer programs and / or modules stored in the memory, and calling the data stored in the memory. For example, the memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, an application required for at least one function (such as a sound receiving function, a sound conversion to text function, etc.), etc.; the data storage area can store data created according to the use of the mobile phone (such as audio data, text data, etc.), etc. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (FlashCard), at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0050] Computer readable storage medium embodiment: If the module integrated in the computer device of the above embodiment is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the implementation of all or part of the process of the howling / feedback suppression method embodiment in the sound reinforcement scene can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the controller, the steps of the howling / feedback suppression method embodiment in the above sound reinforcement scene can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The storage medium may include: any entity or device capable of carrying computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media does not include electrical carrier signals and telecommunications signals.

[0051] Computer program product embodiment: The computer program product of this embodiment includes computer instructions, which are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes each step of the embodiment of the howling / feedback suppression method in the above-mentioned sound reinforcement scenario.

[0052] Finally, it should be emphasized that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various changes and modifications. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for suppressing howling / feedback in a sound reinforcement scenario, characterized in that: The following steps are involved: Get the microphone signal of the current frame; Perform frequency domain adaptive filtering on the microphone signal of the current frame and output the predicted target signal to the speaker: The frequency domain adaptive filtering process comprises the following steps: In the frequency domain, the feedback estimation signal is obtained by multiplying the estimated acoustic transfer function of the previous frame with the speaker signal of the previous frame; In the frequency domain, subtracting the feedback estimation signal from the microphone signal of the current frame to obtain a residual signal; The residual signal is used as the predicted target signal after inverse transformation and is output to the speaker through a feedforward path. At the same time, the estimated acoustic transfer function of the previous frame is iterated in a direction in which the residual signal of the current frame and the speaker signal of the previous frame tend to have a smaller cross-correlation, so as to obtain the estimated acoustic transfer function of the current frame. The step size of the iteration is controlled by a pre-trained recursive neural network model. The recursive neural network model includes the following steps during training: Simulate impulse responses of different moving trajectories in rooms of different sizes; The collected speech and music data are simulated through the generated impulse response feedback path and processed using adaptive filtering, wherein the iteration step size of the adaptive filtering is controlled by the recursive neural network model, and the loss function is the mean square error between the simulated feedback signal and the amplitude spectrum of the feedback signal estimated in the adaptive filtering in the overall simulation process.

2. The method for suppressing howling / feedback in a sound reinforcement scenario according to claim 1, characterized in that: When the recursive neural network model is trained, the collected voice and music data are used to simulate the feedback path through the generated impulse response, and the feedback gain corresponding to the feedback path is random within a set range, wherein at the beginning of training, the feedback gain is set below the critical gain.

3. The method for suppressing howling / feedback in a sound reinforcement scenario according to claim 1, characterized in that: The model structure of the recursive neural network model is LSTM or GRUs.

4. The method for suppressing howling / feedback in a sound reinforcement scenario according to claim 1, characterized in that: The input of the recursive neural network model is a vector formed by combining the amplitude spectrum of the microphone signal of each frame with the amplitude spectrum of the residual signal, and the output is a vector with the same length as the input as a hidden layer. The hidden layer passes through a linear layer and is activated using a logistic function to output a single scalar, which is used as the shared step size for iterating all subbands of the current frame.

5. The method for suppressing howling / feedback in a sound reinforcement scenario according to claim 1, characterized in that: The input of the recursive neural network model is a vector formed by combining the amplitude spectrum of the microphone signal of each frame with the amplitude spectrum of the residual signal, and the output is a vector with the same length as the input as a hidden layer. The hidden layer passes through a linear layer and is activated using a logistic function to output the step size adopted in each subband iteration of the current frame.

6. The method for suppressing howling / feedback in a sound reinforcement scenario according to any one of claims 1 to 4, characterized in that: The estimated acoustic transfer function of the previous frame is iterated in a direction in which the residual signal of the current frame and the speaker signal of the previous frame tend to have a smaller cross-correlation, and the expression is: , is the gradient of the mth subband during the iteration of the estimated acoustic transfer function of the k-1th frame, represents the mth subband of the residual signal of the kth frame, is the mth subband of the loudspeaker signal of the kth frame, for The complex conjugate of is the recursive smoothing of the loudspeaker spectral energy of the m-th subband of the k-th frame, ,in, is the smoothing coefficient; Get the estimated acoustic transfer function of the current frame: ,in, is the mth subband of the estimated acoustic transfer function of the kth frame, is the mth subband of the estimated acoustic transfer function of the k-1th frame, is the iteration step size.

7. A computer device comprising a processor and a memory, characterized in that: The memory stores a computer program, and when the computer program is executed by the processor, the howling / feedback suppression method in the sound reinforcement scene described in any one of claims 1 to 6 is implemented.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the howling / feedback suppression method in a sound reinforcement scenario described in any one of claims 1 to 6 is implemented.

9. A computer program product comprising computer instructions, characterized in that: When the computer instructions are executed by the processor, the howling / feedback suppression method in the sound reinforcement scenario described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Howling suppression method based on feedback signal spectrum estimation

    CN102740214A

  • Method and apparatus for speech enhancement using deep learning model in inverse frequency domain

    CN117789742A

  • Howling canceler apparatus and sound amplification system

    EP1703767A2

  • Acoustic feedback suppression for audio amplification systems

    US20070104335A1

Cited By

  • Sound feedback control method, device and system for low-time-delay sound reinforcement system

    CN120676290A

  • Acoustic feedback path prediction system, prediction method, hearing aid and medium

    CN120812502A