Sound feedback control method, device and system for low-time-delay sound reinforcement system
By combining the full-subband grouped long short-term memory network with the frequency shift method, a sound reinforcement system model was constructed and trained and fine-tuned, which solved the acoustic feedback problem in the sound reinforcement system and achieved efficient voice quality improvement and enhanced system stability under low latency.
Patent Information
- Application Number
- CN202510777360.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-06-11
AI Technical Summary
Existing sound reinforcement systems are prone to acoustic feedback when the gain exceeds the maximum stable gain, resulting in damage to the output voice quality. In addition, traditional acoustic feedback control methods have poor robustness and are difficult to effectively suppress the impact of environmental noise.
A full-subband grouped long short-term memory network combined with the frequency shift method is adopted. By constructing a sound reinforcement system model, an open-loop training dataset is generated, and a complex spectrum mapping target is designed. The network is trained using a mixed amplitude spectrum and complex spectrum loss function, and fine-tuned in a simulated closed-loop system to optimize the signal reconstruction strategy to reduce latency.
It effectively suppresses acoustic feedback at low latency, improves voice quality and intelligibility, significantly increases the maximum stable gain of the sound reinforcement system, and ensures system stability and security.
Smart Images

Figure CN120676290A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of sound reinforcement technology, and in particular to a method, device and system for controlling sound feedback of a low-delay sound reinforcement system. Background Art
[0002] Sound reinforcement systems are used to amplify sound, and their typical applications include multimedia classrooms and local conferencing systems. However, when the system gain exceeds the maximum stable gain, serious acoustic feedback can occur, severely impairing the quality of the output voice and even causing the system to crash. Therefore, acoustic feedback control has become an essential component of sound reinforcement systems, improving both their sound reinforcement performance and ensuring their stability and security. Some known acoustic feedback control methods, such as adaptive filtering, require modeling the acoustic feedback path, providing very limited additional stable gain, making it difficult to effectively suppress the effects of ambient noise, and resulting in relatively poor robustness. Summary of the Invention
[0003] In order to solve the problems existing in the prior art, the embodiments of the present application provide a method, apparatus, system, computing device, computer storage medium and product containing a computer program for acoustic feedback control of a low-latency sound reinforcement system, which can effectively suppress acoustic feedback at a low latency, improve the quality and intelligibility of speech, and significantly increase the maximum stable gain of the sound reinforcement system.
[0004] In a first aspect, an embodiment of the present application provides a low-latency sound reinforcement system acoustic feedback control method, comprising: constructing a sound reinforcement system model based on an acquired input voice signal of the sound reinforcement system, an acoustic signal picked up by a microphone, and an output signal of a loudspeaker, establishing a closed-loop transfer function, and determining a system stability condition; constructing an acoustic feedback path and generating an open-loop training dataset to simulate an acoustic feedback signal in a critical unstable state; extracting the time-frequency features of the input signal and designing a complex spectrum mapping target for a neural network; constructing a full-subband grouped long short-term memory network, comprising an encoder, full-band and sub-band processing modules, and a decoder, wherein the full-band module models global frequency domain dependencies through grouped LSTMs, and the sub-band module processes local time-frequency features through frequency downsampling and independent LSTMs; using linear layers and overlap-addition instead of deconvolution operations to reduce computational complexity; training the network using a mixed amplitude spectrum and complex spectrum loss function to optimize the signal reconstruction strategy to reduce latency; fine-tuning the network in a simulated closed-loop system to adapt to the actual acoustic environment and nonlinear distortion, and limiting the output signal amplitude through hard clipping; and deploying the fine-tuned model to a real sound reinforcement system to collaboratively suppress acoustic feedback with a frequency shift method.
[0005] In some possible implementations, in the sound reinforcement system modeling step, the maximum stable gain is determined by using the Nyquist stability criterion, and the room impulse response is simulated based on the virtual source method to construct the feedback path.
[0006] In some possible implementations, generating an open-loop training data set includes: setting the amplifier gain within a range of 0.5 to 0.999 times the maximum stable gain; applying frequency shift processing to the microphone signal containing feedback to generate input data; and using a pure amplifier output signal without feedback as a training target.
[0007] In some possible implementations, the optimized signal reconstruction strategy to reduce latency includes: using a short-time Fourier transform with a frame shift of 2ms; designing a synthetic window function that relies only on the current frame and the next frame data for signal reconstruction; and controlling the overall algorithm latency to within 4ms.
[0008] In some possible implementations, the network is trained using a mixed amplitude spectrum and complex spectrum loss function, including: the mixed loss function combines the amplitude spectrum mean square error and the complex spectrum real / imaginary part error, and eliminates the influence of window function spectrum leakage by recalculating the time spectrum.
[0009] In some possible implementations, the mixed magnitude spectrum and complex spectrum loss function is expressed as:
[0010]
[0011] Where, Represents the total loss function, Characterizes the mean square error of the amplitude spectrum, Characterizes the complex spectrum mean square error, Characterize the network to predict the time-frequency domain signal, S represents the target time-frequency domain signal, r and· i Represent the real and imaginary parts of the time-frequency domain features, respectively, ‖·‖ F Characterize the norm of the matrix F.
[0012] In some possible implementations, constructing the sound reinforcement system model includes: determining a maximum stable gain using a Nyquist stability criterion, and simulating a room impulse response based on a virtual source method to construct a feedback path.
[0013] In some possible implementations, when the method is deployed, the output of the full-subband grouped long short-term memory network is combined with an adaptive filtering method to form a multi-stage acoustic feedback suppression system.
[0014] In the second aspect, the embodiment of the present application provides a low-latency sound reinforcement system sound feedback control device, including: a construction module for constructing a sound reinforcement system model based on the acquired input voice signal of the sound reinforcement system, the sound signal picked up by the microphone, and the output signal of the speaker, establishing a closed-loop transfer function, and determining the system stability condition; the construction module is also used to construct an acoustic feedback path and generate an open-loop training data set to simulate the acoustic feedback signal under a critical unstable state; a processing module is used to extract the time-frequency characteristics of the input signal and design the complex spectrum mapping target of the neural network; the construction module is also used to construct a full-subband grouped long short-term memory network, which includes an encoder, a full-band and sub-band The processing module and decoder, the full-band module models the global frequency domain dependency through grouped LSTM, and the sub-band module processes the local time-frequency features through frequency downsampling and independent LSTM; linear layers and overlap-addition are used instead of deconvolution operations to reduce computational complexity; the processing module is also used to train the network using a mixed amplitude spectrum and complex spectrum loss function, and optimize the signal reconstruction strategy to reduce latency; the processing module is also used to fine-tune the network in a simulated closed-loop system to adapt to the actual acoustic environment and nonlinear distortion, and limit the output signal amplitude through hard clipping; the processing module is also used to deploy the fine-tuned model to a real sound reinforcement system, and cooperate with the frequency shift method to suppress acoustic feedback.
[0015] In a third aspect, an embodiment of the present application provides a sound reinforcement system, characterized in that it is equipped with the acoustic feedback control device described in the second aspect.
[0016] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium comprising computer-readable instructions. When a computer reads and executes the computer-readable instructions, the computer executes the method as described in any one of the first aspects.
[0017] In a fifth aspect, an embodiment of the present application provides a computing device comprising a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the method described in any one of the first aspects is executed.
[0018] In a sixth aspect, an embodiment of the present application provides a product comprising a computer program, which, when the computer program product runs on a processor, enables the processor to execute the method as described in any one of the first aspects. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0020] Figure 1 This is a schematic diagram of a sound reinforcement system architecture provided by an embodiment of the present application;
[0021] Figure 2 1 is a flow chart of a method for controlling acoustic feedback in a low-latency sound reinforcement system according to an embodiment of the present application;
[0022] Figure 3 This is a schematic diagram of an abstract representation of a sound reinforcement system model provided in an embodiment of the present application;
[0023] Figure 4 This is a schematic diagram of an acoustic feedback signal simulation provided by an embodiment of the present application;
[0024] Figure 5 This is a schematic diagram of a full-subband grouped long short-term memory network provided by an embodiment of the present application;
[0025] Figure 6 This is a schematic diagram of a closed-loop training of a full-subband grouped long short-term memory memory network provided by an embodiment of the present application;
[0026] Figure 7 This is a schematic diagram of an embodiment of the present application providing a method of adding a trained model to a closed-loop system to suppress acoustic feedback;
[0027] Figure 8 It is a structural diagram of an acoustic feedback control device for a low-latency sound reinforcement system provided in an embodiment of the present application. DETAILED DESCRIPTION
[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0029] The term "and / or" as used herein describes an association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. The symbol " / " as used herein indicates that the related objects are in an "or" relationship, for example, A / B means either A or B.
[0030] The terms "first" and "second" in this specification and claims are used to distinguish different objects rather than to describe a specific order of objects. For example, "first response message" and "second response message" are used to distinguish different response messages rather than to describe a specific order of response messages.
[0031] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0032] In the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more, for example, multiple processing units means two or more processing units, etc.; multiple elements means two or more elements, etc.
[0033] To facilitate understanding of the embodiments of the present application, further explanation will be given below with reference to specific embodiments in conjunction with the accompanying drawings. The embodiments do not constitute a limitation on the embodiments of the present invention.
[0034] Figure 1 FIG. 1 shows a schematic diagram of a sound reinforcement system architecture provided by an embodiment of the present application. Figure 1 As shown, a sound reinforcement system usually includes a microphone, an amplifier, and a loudspeaker. Among them, the microphone is used to receive sound signals, perform sound-to-electric conversion and perform the required signal processing. The sound signal is then amplified by the amplifier and output through the loudspeaker after electro-acoustic conversion. During use, since the microphone and the loudspeaker are in the same sound field, the sound signal played by the loudspeaker is inevitably picked up again, amplified, and output by the microphone, which forms a closed-loop system between the microphone and the loudspeaker, resulting in an acoustic feedback problem. Although the existing deep learning-based acoustic feedback control method can provide a greater stable gain and effectively suppress the influence of environmental noise, combining it with the frequency shift method will produce severe voice distortion.
[0035] In view of this, an embodiment of the present application provides a two-stage low-latency acoustic feedback control method for a sound reinforcement system, using a deep neural network as a post-processing network for the traditional frequency shift method to jointly perform acoustic feedback control on the sound reinforcement system. For indoor sound reinforcement systems, first, a large number of acoustic feedback paths are constructed using simulated room impulse responses, and the maximum stable gain of the simulated system is calculated based on the Nyquist stability criterion. Next, a large amount of data with critical acoustic feedback characteristics processed by the frequency shift method is generated in a closed-loop system for offline open-loop training of the deep neural network. Then, a low-latency strategy is designed to ensure that the network latency is controlled within 4ms. In addition, a new loss function is introduced to further improve the performance of the algorithm. Finally, the trained network is fine-tuned within the closed-loop system to further improve the network's acoustic feedback suppression capability.
[0036] For example, Figure 2FIG. 1 is a flow chart showing a method for controlling acoustic feedback of a low-delay sound reinforcement system according to an embodiment of the present application. Figure 2 As shown, the acoustic feedback control method may include the following steps:
[0037] S21: Based on the acquired input voice signal of the sound reinforcement system, the sound signal picked up by the microphone, and the output signal of the speaker, a sound reinforcement system model is constructed, a closed-loop transfer function is established, and a system stability condition is determined.
[0038] In this embodiment, a mathematical model of a closed-loop sound reinforcement system is established. The model needs to include several key parts: the input signal, the sound output by the loudspeaker, the sound received by the microphone, and the path for the sound to return from the loudspeaker to the microphone (i.e., the feedback path). The time domain and frequency domain representations are considered separately during modeling. In the time domain, the sound is represented by the signal value at a discrete time point, and the main focus is on how the signal changes over time. In the frequency domain, the signal is decomposed into different frequency components using a short-time Fourier transform, so that it is possible to more clearly see which frequencies are prone to howling. By establishing a mathematical model, the conditions for system stability, that is, the conditions for no howling, can be derived. Specifically, when the system open-loop gain (the amplification factor of the sound after one cycle) satisfies both the amplitude and phase conditions at certain frequencies, howling will occur.
[0039] For more information on the abstract representation of the sound reinforcement system model, please refer to Figure 3 , v(n) is the input speech signal, y(n) is the sound signal picked up by the microphone, u(n) is the loudspeaker output signal, f(n) is the feedback path transfer function, and g(n) is the amplifier transfer function of the sound reinforcement system, then it can be expressed as follows:
[0040] y(n)=v(n)+u(n)*f(n)
[0041] u(n)=y(n)*g(n)
[0042] Where * represents the convolution operation, n represents the number of discrete time points, and n is a natural number, that is, n = [0, 1, ..., n]. The time-frequency domain expression of the above equation is:
[0043] Y(ω,l)=V(ω,l)+U(ω,l)F(ω,l)
[0044] U(ω,l)=Y(ω,l)G(ω,l)
[0045] Where Y(ω,l) represents the time-frequency domain characteristics of the acoustic signal picked up by the microphone, V(ω,l) represents the time-frequency domain characteristics of the input speech signal, U(ω,l) represents the time-frequency domain characteristics of the loudspeaker output signal, F(ω,l) represents the time-frequency domain characteristics of the feedback path transfer function, G(ω,l) represents the time-frequency domain characteristics of the amplifier transfer function of the sound reinforcement system, ω represents the number of frequency points and is a natural number, l represents the number of frames and is a natural number.
[0046] Combining the above equations, we can get the closed-loop transfer function from microphone to speaker, which is expressed as:
[0047]
[0048] Where G(ω,l)F(ω,l) represents the open-loop transfer function.
[0049] Generally, assuming the transfer function of the feedback path is fixed or slowly varying, then F(ω,l)≈F(ω). A sound reinforcement system is stable only when all poles of the closed-loop system lie within the unit circle. The presence of acoustic feedback is a sign of instability in a closed-loop sound reinforcement system. According to the Nyquist stability criterion, a closed-loop system is unstable when the open-loop transfer function satisfies the following conditions:
[0050]
[0051] Where, represents the set of integers, |·| represents the modulus value, and ∠· represents the phase.
[0052] When the phase of the system's open-loop transfer function is an integer multiple of 2π and the transfer function modulus is greater than 1, the closed-loop sound reinforcement system becomes unstable at the corresponding frequency, resulting in howling. If we assume that the amplifier only has amplification function, that is, G(ω,l)=G, where G is the amplification factor, then the maximum stable gain G that keeps the closed-loop sound reinforcement system stable can be calculated. max for:
[0053]
[0054] Where Ω(l) = {ω|∠F(ω,l) = 2kπ}, representing the set of frequencies that satisfy the corresponding phase conditions. In practice, the gain of the sound reinforcement system must be less than the maximum closed-loop stable gain. To avoid ambiguity, the following description of the scheme will omit the frame index l and frequency index ω.
[0055] S22: Construct an acoustic feedback path and generate an open-loop training dataset to simulate the acoustic feedback signal in a critical unstable state.
[0056] In this embodiment, the first step is to simulate the path of sound from the loudspeaker back to the microphone, known as the feedback path. This simulation takes into account the room's acoustic characteristics, such as wall reflections. The virtual source method is typically used to calculate the room's impulse response. With this feedback path model, the system's behavior at different gain levels can be simulated, particularly near critical stability. Next, a large amount of training data must be generated. This process involves generating a clean speech signal as input. This signal is then passed through a simulated closed-loop sound reinforcement system, where the amplifier gain is deliberately set near instability, resulting in a signal with a tendency toward howling. Frequency shifting of these signals, while suppressing howling, introduces distortion. The processed signal serves as the input data for the neural network, while the corresponding target data is the ideal output signal without the feedback loop. This creates the data pairs required for open-loop training: the input is the signal processed by frequency shifting but still retaining residual artifacts, and the output is the desired ideal signal. This data generation process can be carried out on a large scale, providing sufficient samples for subsequent deep learning training.
[0057] Specifically, it can be seen from step S21 that when the system does not meet the Nyquist stability criterion, the signal at the frequency ω will be continuously amplified to form howling. In an indoor sound reinforcement system, the transfer function of the feedback path is the sum of all room impulse responses from the microphone to the speaker. In this embodiment, the virtual source method is used to simulate the room impulse response, and then the maximum stable gain G can be obtained. max . Figure 4 A schematic diagram of an acoustic feedback signal simulation provided by an embodiment of the present application is shown. Figure 3 The difference is that a frequency shifter is connected after the transfer function of the amplifier in the sound reinforcement system. In order to generate a speech signal with acoustic feedback, the following Figure 4 The acoustic feedback signal simulation system shown generates an analog acoustic feedback signal, where the amplifier gain G∈[0.5G max ,0.999G max ], so that the system is in a state where howling is likely to occur. The frequency shift method is a traditional acoustic feedback control method, which keeps the closed-loop sound reinforcement system stable by controlling the phase condition of the Nyquist stability criterion. The specific method is to control the phase of the microphone receiving signal so that after the sound signal passes through a cycle of the closed-loop sound reinforcement system, the frequency component of the acoustic feedback signal has a different phase each time it reaches the microphone. Therefore, the maximum stable gain of the closed-loop sound reinforcement system can be significantly improved. However, the frequency shift method does not cut off the acoustic feedback path, and this method will seriously damage the quality of the output speech. Therefore, this embodiment uses a deep neural network as a post-processing network of the traditional frequency shift method to jointly perform acoustic feedback control of the sound reinforcement system. The frequency shift method can be implemented by constructing an analytical signal of the microphone receiving signal y(n). By designing a Hilbert filter, the analytical signal of y(n) can be obtained as:
[0058]
[0059] Where, represents the Hilbert transform of the signal received by the microphone, and j represents the imaginary unit.
[0060] Since the Hilbert filter is non-causal, it will introduce a time delay of half the filter length. In this embodiment, the Hilbert filter length is 33. Then, the output signal of the frequency shifter can be obtained as:
[0061]
[0062] Where, Characterize the real part, H(ω,n) represents the frequency response of the Hilbert filter, ω m According to the frequency shift, f m Characterize the frequency shift, and f m =ω m (f s / 2π), f s Characterizes the sampling frequency. In this embodiment, the frequency shift can be 5 Hz.
[0063] After frequency shift processing, a large amount of data with critical acoustic feedback characteristics can be obtained for offline open-loop training of deep neural networks (DNN). Through deep neural networks, the feedback-free signal s(n) can be directly estimated.
[0064] s(n)=Gv(t-τf s )
[0065] Where τ represents the delay of the sound reinforcement system, and t represents the time.
[0066] A large number of (d(n), s(n)) data pairs are open-loop data sets, where d(n) is the feedback signal processed by the frequency shift method, which serves as the input of the DNN, and s(n) is the non-feedback signal, that is, the signal that should be output when there is no feedback effect, which is the target that the DNN needs to predict.
[0067] S23: Extract the time-frequency features of the input signal and design the complex spectrum mapping target of the neural network.
[0068] In this embodiment, a time-frequency analysis is performed on the sound signal processed by the frequency shift method. The short-time Fourier transform is used to decompose the sound into representations of different time segments and frequency components, resulting in a complex spectrum containing real and imaginary parts. This spectrum data is reorganized into a format suitable for neural network input, with the real and imaginary parts separated as two channels and then arranged into a three-dimensional tensor according to the time and frequency dimensions. Next, the learning objectives of the neural network are designed. Here, the network is chosen to directly predict the time-frequency representation of the sound signal that the speaker should output under ideal conditions. To enable the network to learn better, the spectrum amplitude is compressed, which balances the importance of different frequency components. The network needs to simultaneously predict the amplitude and phase information of the output signal, as both aspects are important for sound quality. This design requires the neural network to not only eliminate the residual howling components of the frequency shift method, but also repair the sound quality damage caused by the frequency shift method, ultimately outputting a clear and natural sound signal.
[0069] Specifically, during feature extraction, the frequency shift method output signal d(n) is subjected to a short-time Fourier transform (STFT) to obtain a complex time-frequency spectrum. The real and imaginary parts of the complex spectrum are concatenated along the channel dimension to form a tensor of dimension 2xTxF. T is the total number of frames of the input speech, and F is the total number of frequency points. This tensor serves as the input of the neural network. The goal of DNN is to directly predict the time-frequency spectrum of the target signal. The complex spectrum mapping method is used, that is,
[0070]
[0071]
[0072] in,· r and· i Represent the real and imaginary parts of the time-frequency domain features respectively. for The time-frequency domain characteristics of is the network prediction value, Characterize the complex spectrum mapping, D c Characterizes the time-frequency domain characteristics of the frequency shift output signal after amplitude spectrum compression, Φ represents the parameters of the DNN model, D is the input feature, is the predicted value, β c is the amplitude spectrum compression coefficient, usually 0.5.
[0073] In some more specific embodiments, the sampling frequency of the speech signal is 16 kHz, the length of each frame is 20 ms, the frame shift is 2 ms, the number of Fourier transform points is 320, and the total number of frequency points is 161.
[0074] S24: Construct a full-subband grouped long short-term memory network, including an encoder, full-band and sub-band processing modules, and a decoder
[0075] In this embodiment, a deep neural network structure is designed specifically for dealing with acoustic feedback problems. The neural network adopts an architecture that combines global and local information processing to capture full-band and sub-band information. It mainly consists of three parts: an encoder that converts input features into high-dimensional representations, a core processing module for learning sound features, and a decoder that converts the processed features back to the time-frequency domain. The encoder uses a two-dimensional convolutional layer to extract preliminary features. The core processing module combines two different long and short-term memory networks, one responsible for capturing the global information of the entire frequency band, and the other focusing on processing the details of each sub-band. In order to reduce computational complexity, a group processing strategy is adopted to divide the features into several groups and then process them separately before merging them. The decoding part uses a deconvolution layer to reconstruct the output signal. Low latency also needs to be considered when designing the network. Operations in the time dimension do not need to cache information from multiple frames in the past, but only rely on information from the current frame and the previous frame.
[0076] Specifically, Figure 5 This is a schematic diagram of a full-subband grouped long short-term memory network provided by the embodiment of the present application. Please refer to Figure 5This embodiment uses a full-subband grouped long short-term memory (FSB-GLSTM) network for acoustic feedback suppression. It consists of three parts: an encoder, a full-subband grouped long short-term memory (FSB-GLSTM) module, and a decoder. The encoder includes a two-dimensional convolutional layer (Conv2D) with a convolution kernel size of 1×3 along the time and frequency dimensions, and a normalization layer to extract a D-dimensional embedding for each time-frequency point. The decoder is a corresponding 1×3 two-dimensional deconvolution layer (Deconv2D) to obtain the real and imaginary parts of the output speech time-frequency spectrum. The full-subband grouped long short-term memory module can be divided into a full-band module and a sub-band module. In the full-band module, for a given D×T×F dimensional input embedding, the T×F embedding corresponding to each frame is first compressed into a frame-level embedding based on the Conv2D layer. The frame-level embedding is then modeled using a grouped LSTM. Finally, the D-dimensional embedding of each time-frequency point is restored based on the corresponding Deconv2D layer. In this way, the grouped LSTM layer can simultaneously model all frequency components at the current moment, thereby capturing full-band information. In this embodiment, a grouped LSTM layer is used to reduce the computational complexity of the model. Specifically, the grouped LSTM layer divides the input embedding along the feature dimension, and then uses an independent LSTM to model each part, and finally splices them together. This can significantly reduce the computational complexity of the model with a small performance loss. At the same time, considering that in the actual process, deconvolution will first be based on step-size zero padding, and then a conventional convolution operation will be performed. Therefore, when the step size is greater than 1, due to the information redundancy in the zero-padding operation, this will significantly increase the computational complexity. Therefore, this embodiment uses a linear layer and an overlap-add method to replace the Deconv2D layer. At this time, the computational complexity can be reduced to approximately 1 / J of the original Deconv2D layer, where J is the deconvolution step size. Considering that acoustic feedback generally only appears at some frequency components, this embodiment also uses a sub-band module to model each frequency component separately to capture the time-frequency information therein. In addition, in order to further reduce the computational complexity, the Conv2D layer is used to downsample the frequency to reduce the sub-band LSTM input and output feature dimensions. Note that in FSB-GLSTM, the convolution kernel size along the time dimension is 1. Therefore, in this embodiment, no additional buffer is required to store information of past frames. Only the cell state and hidden state of the previous frame of LSTM need to be cached to model time information. Therefore, the network only requires a very small memory footprint at runtime and is more suitable for streaming inference.
[0077] In some more specific embodiments, the specific network hyperparameter descriptions and settings are shown in Table 1. In this embodiment, B = 6, D = 8, E = 4, I = 8, J = 4, H = 128, E' = 8, I' = 5, J' = 5, and H' = 16 are selected. The final parameter count is 723.67K, and the computational complexity (MACs) is 716.57M / s. In addition, the network is trained with 32 batches and 150 iterations. The Adam optimizer is used for training with a learning rate of 0.001. If the validation set loss does not decrease after 3 consecutive iterations, the learning rate is halved. If it does not decrease after 6 consecutive iterations, training is terminated.
[0078] Table 1
[0079]
[0080] S25: Use mixed amplitude spectrum and complex spectrum loss functions to train the network and optimize the signal reconstruction strategy to reduce latency.
[0081] In this embodiment, the neural network needs to learn how to repair the sound signal while ensuring that the entire processing process meets low latency requirements. First, a suitable loss function must be selected to guide the neural network's learning direction. This embodiment uses a hybrid loss function that combines the magnitude spectrum and the complex spectrum. This requires that the network output sound energy distribution is consistent with the target, while also ensuring that the real and imaginary parts are as close to the true values as possible. This ensures both clarity and naturalness of the speech. To prevent spectral leakage caused by the analysis window from affecting the training effect, the network's predicted spectrum is first converted back to the time domain signal using a synthesis window. The spectrum is then recalculated using another analysis window, and the loss is finally evaluated based on this recalculated spectrum. In terms of low latency strategy, by optimizing the synthesis window function and overlapping and adding, the system only needs information from the current and next frames to complete reconstruction, keeping the overall latency very low. In specific implementation, a small frame shift parameter is selected, combined with a specific window function design, to compress the algorithm latency to 4 milliseconds while maintaining sufficient frequency resolution.
[0082] Specifically, according to step S24, the time-frequency spectrum of the network output predicted speech is subjected to an inverse short-time Fourier transform to obtain the output signal of the speaker. Since the overlap-add method has a time delay, reducing the window length can further reduce the algorithm delay, but it will reduce the frequency resolution of the time-frequency feature, thereby negatively affecting the acoustic feedback suppression performance. Therefore, this patent uses a smaller synthesis window to reduce the algorithm delay to 4ms while maintaining the frequency resolution. Specifically, the analysis window selected in this patent is:
[0083]
[0084] Where N is the frame length, R is the frame shift, and t is the number of discrete time points within a frame. When reconstructing the signal, the current frame and the next frame are overlapped and added, and the output of the current frame is:
[0085]
[0086] Where t = [1,…,R], is the time domain signal of each frame, n is the number of frames, k is the number of frequency points, and is the synthetic window function, which is defined as:
[0087]
[0088] At this time, if the frame shift is selected as 2ms, that is, R=32, the final network delay is 4ms. In this embodiment, the loss function of network training is selected as the mean square error (MSE) loss function between the amplitude spectrum and complex spectrum of the predicted speech and the target speech, that is:
[0089]
[0090] Where, ‖·‖ F Represents the norm of matrix F.
[0091] Considering w m (t) The spectrum leakage of the window function is more serious, so the time domain signal is first synthesized when calculating the loss function. Then convert it into the time-frequency domain, that is Then solve the MSE loss, where can be any type of window function. In this patent, we select w h (t) is a Hanning window with a frame length of 20ms and a frame shift of 10ms.
[0092] S26: Fine-tune the network in a simulated closed-loop system to adapt to the actual acoustic environment and nonlinear distortion.
[0093] In this embodiment, the adaptability problem of neural networks in actual application scenarios is solved. Since previous training was carried out in an open-loop environment, and in real scenarios the network output will be played through a speaker and picked up again by a microphone to form a closed-loop feedback, this difference will lead to a decrease in model performance. In order to eliminate this effect, it is necessary to establish a simulated closed-loop system for fine-tuning training. This simulation system contains a complete signal loop, that is, the input voice is first mixed with the output signal after network processing through the simulated feedback path to form a closed-loop microphone input, and then pre-processed by the frequency shift method and sent to the neural network. At the same time, the simulation system also takes into account the nonlinear characteristics of real speakers, such as hard limiting of signal amplitude. The previous mixed loss function continues to be used during fine-tuning, but the training data becomes the signal pairs generated by this closed-loop simulation system.
[0094] Specifically, considering that there is a model mismatch problem when the open-loop trained network is used in an actual closed-loop sound reinforcement system, this patent fine-tunes the open-loop trained model in a closed-loop simulation system. Figure 6 The diagram of closed-loop training of full-subband grouped long short-term memory network is shown. Figure 4 The difference is that after the frequency shifter, FSB-GLSTM is connected. Figure 6 In the simulated closed-loop system, it is assumed that f(n) is the simulated room impulse response, g(n) only plays an amplification role, and the delay is τ g , then g(n)=Gδ(n-τ g f s ), where δ(n) is the unit impulse response. Considering that actual speakers have nonlinear distortion, it is assumed that the simulated speaker has a saturation region, that is, hard clipping is performed on the input audio signal:
[0095]
[0096] in, is the maximum value that the speaker can output and is set to 1 during the simulation. Then, the input signal of the microphone at this time can be obtained as:
[0097]
[0098] By predicting the loudspeaker output signal in a simulated closed-loop system, the model can be fine-tuned based on the loss function of step S25 to further enhance the acoustic feedback suppression capability of the present invention.
[0099] S27: Deploy the fine-tuned model to a real sound reinforcement system and use it in conjunction with the frequency shift method to suppress acoustic feedback.
[0100] In this embodiment, the fully trained neural network model is deployed in a real sound reinforcement system for practical application. Figure 7 At this point, the entire system begins operating in a closed-loop manner. The sound signal collected by the microphone is first pre-processed using frequency shifting to eliminate most of the howling. It is then fed into a trained neural network for detailed processing. The network conducts in-depth analysis of the signal, identifying and suppressing any residual feedback components that the frequency shifting method cannot completely eliminate, while also correcting any speech distortion caused by the frequency shifting process. The processed signal is then passed through a power amplifier to drive the speakers for playback. The sound is then fed back to the microphone through the room's acoustic environment, forming a complete closed-loop. In actual operation, the network must process the continuous audio stream in real time, maintaining extremely low processing latency to ensure system synchronization.
[0101] The above is an introduction to the low-latency acoustic feedback control method for a sound reinforcement system provided by an embodiment of the present application. To address the problem of acoustic feedback in a sound reinforcement system, a model is trained using open-loop data as a training set, then fine-tuned in a closed-loop system. Finally, the trained model is placed in the closed-loop sound reinforcement system to suppress the acoustic feedback signal. In practical applications, the sound reinforcement system receives a signal with acoustic feedback, first preprocessing it using the frequency shift method. The trained network is then used to remove the residual acoustic feedback component, resulting in an accurate estimate of the loudspeaker output signal. Furthermore, since hearing aid systems and assisted listening systems also fall within the scope of sound reinforcement systems, the method proposed in an embodiment of the present application is also applicable to hearing aids and other devices. The advantage of this method is that the neural network does not work alone, but rather works in conjunction with the traditional frequency shift method. The frequency shift method provides basic stability assurance, while the neural network further improves sound quality and suppression effects. If necessary, it can also be combined with other acoustic feedback suppression techniques, such as using the output of the neural network as a reference signal for an adaptive filter to form a more powerful hybrid suppression system. The ultimate effect achieved is to achieve more natural and clear speech output quality than traditional methods while maintaining system stability.
[0102] It is understandable that the size of the sequence number of each step in the above-mentioned embodiments does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. In addition, in some possible implementations, the steps in the above-mentioned embodiments can be selectively executed according to actual conditions, and can be partially executed or fully executed, which is not limited here. All or part of any features of any embodiment of the present application can be freely and arbitrarily combined without contradiction. The combined technical solution is also within the scope of the present application.
[0103] Based on the method in the above embodiment, the embodiment of the present application also provides a low-delay sound reinforcement system sound feedback control device. For example, Figure 8 FIG. 1 shows a schematic diagram of the structure of a low-delay sound reinforcement system sound feedback control device provided by the present application. Figure 8As shown, the device 800 includes 801 and a processing module 802.
[0104] The construction module 801 is used to construct a sound reinforcement system model based on the acquired input voice signal of the sound reinforcement system, the sound signal picked up by the microphone, and the output signal of the speaker, establish a closed-loop transfer function, and determine the system stability condition;
[0105] The construction module 801 is further used to construct an acoustic feedback path and generate an open-loop training data set to simulate an acoustic feedback signal in a critical unstable state;
[0106] Processing module 802, used to extract the time-frequency characteristics of the input signal and design the complex spectrum mapping target of the neural network;
[0107] The construction module 801 is further used to construct a full-subband grouped long short-term memory network, which includes an encoder, full-band and sub-band processing modules, and a decoder. The full-band module uses grouped LSTM to model global frequency domain dependencies, and the sub-band module uses frequency downsampling and independent LSTM to process local time-frequency features. Linear layers and overlap-add are used instead of deconvolution operations to reduce computational complexity.
[0108] The processing module 802 is further configured to train the network using a mixed amplitude spectrum and complex spectrum loss function to optimize the signal reconstruction strategy to reduce latency;
[0109] The processing module 802 is further used to fine-tune the network in the simulated closed-loop system to adapt to the actual acoustic environment and nonlinear distortion, and to limit the output signal amplitude through hard clipping;
[0110] The processing module 802 is further configured to deploy the fine-tuned model to a real sound reinforcement system, and cooperate with the frequency shift method to suppress acoustic feedback.
[0111] Based on the methods in the above embodiments, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a processor, the processor executes the methods in the above embodiments.
[0112] Based on the methods in the above embodiments, an embodiment of the present application provides a computer program product. When the computer program product runs on a processor, the processor executes the methods in the above embodiments.
[0113] It is understood that the processor in the embodiments of the present application may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.
[0114] The method steps in the embodiments of the present application can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, mobile hard disks, CD-ROMs or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC.
[0115] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted via the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state drive (SSD)).
[0116] It will be understood that the various numerical numbers involved in the embodiments of the present application are merely distinctions for the convenience of description and are not intended to limit the scope of the embodiments of the present application.
Claims
1. A low-latency sound reinforcement system acoustic feedback control method, characterized in that: The method comprises: Based on the acquired sound reinforcement system input voice signal, the sound signal picked up by the microphone, and the output signal of the speaker, a sound reinforcement system model is constructed, a closed-loop transfer function is established, and the system stability conditions are determined; Construct an acoustic feedback path and generate an open-loop training dataset to simulate the acoustic feedback signal in a critical unstable state; Extract the time-frequency features of the input signal and design the complex spectrum mapping target of the neural network; A full-subband grouped long short-term memory network is constructed, consisting of an encoder, full-band and sub-band processing modules, and a decoder. The full-band module uses grouped LSTMs to model global frequency-domain dependencies, while the sub-band modules process local time-frequency features through frequency downsampling and independent LSTMs. Linear layers and overlap-add are used instead of deconvolution operations to reduce computational complexity. A hybrid amplitude spectrum and complex spectrum loss function is used to train the network and optimize the signal reconstruction strategy to reduce latency. Fine-tune the network in a simulated closed-loop system to adapt to the actual acoustic environment and nonlinear distortion, and limit the output signal amplitude through hard clipping; The fine-tuned model is deployed in a real sound reinforcement system and used in conjunction with the frequency shift method to suppress acoustic feedback.
2. The method according to claim 1, characterized in that In the sound reinforcement system modeling step, the maximum stable gain is determined by the Nyquist stability criterion, and the room impulse response is simulated based on the virtual source method to construct a feedback path.
3. The method according to claim 1, characterized in that Generating an open-loop training data set includes: Set the amplifier gain within the range of 0.5 to 0.999 times the maximum stable gain; Applying frequency shift processing to the microphone signal containing feedback to generate input data; The pure amplifier output signal without feedback is used as the training target.
4. The method according to claim 1, wherein The optimized signal reconstruction strategy to reduce latency includes: Use short-time Fourier transform with a frame shift of 2ms; Design a synthetic window function that only relies on the current frame and the next frame data for signal reconstruction; The overall algorithm delay is controlled within 4ms.
5. The method according to claim 1, wherein The network training using the mixed amplitude spectrum and complex spectrum loss function includes: The hybrid loss function combines the mean square error of the amplitude spectrum with the real / imaginary part errors of the complex spectrum, and eliminates the influence of the window function spectrum leakage by recalculating the time spectrum.
6. The method according to claim 1, characterized in that The mixed amplitude spectrum and complex spectrum loss function is expressed as: Where, Represents the total loss function, Characterizes the mean square error of the amplitude spectrum, Characterizes the complex spectrum mean square error, Characterize the network to predict the time-frequency domain signal, S represents the target time-frequency domain signal, r and· i Represent the real and imaginary parts of the time-frequency domain features, respectively, ‖·‖ F Characterize the norm of the matrix F.
7. The method according to claim 1, characterized in that The constructing of the sound reinforcement system model includes: The maximum stable gain is determined by the Nyquist stability criterion, and the room impulse response is simulated based on the virtual source method to construct the feedback path.
8. The method according to claim 1, characterized in that When the method is deployed, the output of the full-subband grouped long short-term memory network is combined with an adaptive filtering method to form a multi-stage acoustic feedback suppression system.
9. A low-delay sound reinforcement system acoustic feedback control device, characterized in that: The device comprises: A construction module is used to construct a sound reinforcement system model based on the acquired input voice signal of the sound reinforcement system, the sound signal picked up by the microphone, and the output signal of the speaker, establish a closed-loop transfer function, and determine the system stability condition; The building module is further used to construct an acoustic feedback path and generate an open-loop training data set to simulate an acoustic feedback signal in a critical unstable state; The processing module is used to extract the time-frequency characteristics of the input signal and design the complex spectrum mapping target of the neural network; The building module is also used to construct a full-subband grouped long short-term memory network, which includes an encoder, full-band and sub-band processing modules, and a decoder. The full-band module uses grouped LSTM to model global frequency domain dependencies, and the sub-band module uses frequency downsampling and independent LSTM to process local time-frequency features. Linear layers and overlap-add are used instead of deconvolution operations to reduce computational complexity. The processing module is further used to train the network using a mixed amplitude spectrum and complex spectrum loss function to optimize the signal reconstruction strategy to reduce latency; The processing module is also used to fine-tune the network in the simulated closed-loop system to adapt to the actual acoustic environment and nonlinear distortion, and to limit the output signal amplitude through hard clipping; The processing module is further used to deploy the fine-tuned model to a real sound reinforcement system, and cooperate with the frequency shift method to suppress acoustic feedback.
10. A sound reinforcement system, characterized in that: The acoustic feedback control device according to claim 9 is deployed.
Citation Information
Patent Citations
channel equalization method based on an NARX neural network and block feedback
CN109905337A
Snakelike robot climbing control method, system and device and storage medium
CN112731810A
Code index spread spectrum underwater acoustic communication method based on recurrent neural network
CN113541726A
Closed-loop system acoustic feedback suppression method based on deep learning
CN115243162A
Acoustic feedback cancellation method based on deep learning
CN115881148A