Planned sampling non-autoregressive learning method for acoustic feedback suppression
By employing a non-autoregressive learning method based on planned sampling and a finite-order Newman series approximation operator, the computational complexity and stability issues of deep learning acoustic feedback suppression methods under high system gain conditions are addressed. This improves the robustness and speech intelligibility of the acoustic feedback suppression model, making it suitable for hearing aids, wireless headphones, and in-vehicle voice systems.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NANJING INST OF TECH
- Filing Date
- 2026-02-05
- Publication Date
- 2026-05-08
AI Technical Summary
Existing deep learning-based acoustic feedback suppression methods suffer from high computational complexity and unstable training processes. Furthermore, there are significant differences between autoregressive training and closed-loop inference, leading to performance degradation under high system gain conditions.
A non-autoregressive planned sampling learning method is adopted. By constructing a non-autoregressive open-loop training framework, a two-stage finite-boundary planned sampling mechanism, and a finite-order Newman series approximation operator, it gradually adapts to the real closed-loop operating conditions. A controlled planned sampling mechanism is introduced to approximate the infinite impulse response of acoustic feedback based on the finite-order Newman series.
Under conditions of high system gain and complex acoustic coupling, the robustness and stability of the acoustic feedback suppression model are improved, speech intelligibility is enhanced, and it is suitable for audio amplification devices such as hearing aids, wireless headphones and in-vehicle voice systems.
Smart Images

Figure CN121999792A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of audio signal processing and deep learning technology, specifically relating to a planned sampling non-autoregressive learning method for acoustic feedback suppression. Background Technology
[0002] In audio amplification systems, including hearing aids, wireless headphones, and in-vehicle voice systems, the speaker output signal is coupled back to the microphone via an acoustic path, forming a closed-loop feedback structure. When the system gain is high, this can easily lead to problems such as spectral coloration, howling, and speech distortion, severely impacting the user experience. As modern audio devices increasingly adopt high amplification gain, open-fit structures, and transparent listening modes, the acoustic coupling effect is further amplified, significantly reducing the system stability margin and making traditional feedback suppression methods inadequate for practical needs.
[0003] Existing acoustic feedback suppression techniques mainly include phase modulation feedback control, notch filter-based howling suppression, and adaptive feedback cancellation methods. Phase modulation methods introduce frequency-dependent phase perturbations to disrupt the positive feedback path; notch filtering targets the narrowband resonant frequencies that trigger howling; and adaptive feedback cancellation methods estimate the feedback path and eliminate feedback components to achieve suppression. However, these traditional methods often experience a sharp performance degradation under high system gain conditions, easily introducing spectral artifacts or unintended deletion of speech components, thereby reducing speech intelligibility.
[0004] In recent years, with the widespread application of deep learning in speech-related tasks, acoustic feedback suppression methods based on deep neural networks have gradually attracted attention. These methods leverage the modeling ability of neural networks to understand complex nonlinear acoustic mappings, thereby improving feedback suppression performance to some extent. However, existing deep learning-based feedback suppression methods still have significant limitations in their training strategies, mainly in the following two aspects:
[0005] 1) Some methods employ a recursive learning strategy, using the model output recursively as subsequent input during training to simulate closed-loop feedback behavior in real systems, thereby constructing the feedback signal required for training. Although this approach can realistically reflect the cumulative characteristics of the feedback signal, it suffers from high computational complexity, unstable training process, and requires unfolding long time series, making it difficult to scale to large-scale datasets and long speech scenarios, which severely restricts its practical application.
[0006] 2) To improve training efficiency, another approach is to use a teacher-forced training strategy, which assumes that feedback has been completely suppressed in each frame, thereby avoiding recursive execution and achieving non-autoregressive training similar to speech enhancement models. Although this strategy significantly reduces training costs, due to the significant difference between the training phase and the real closed-loop inference phase, the model needs to rely entirely on its own historical prediction results when inferring, which is prone to exposure bias. This causes the prediction error to accumulate and propagate continuously in the closed-loop feedback, resulting in poor performance in real systems.
[0007] Therefore, designing a training method that can maintain the efficiency of non-autoregressive training while effectively narrowing the gap between it and autoregressive closed-loop inference has become the key to improving the actual performance of acoustic feedback suppression models. Summary of the Invention
[0008] To address the technical problems existing in the prior art, this invention proposes a planned sampling non-autoregressive learning method for acoustic feedback suppression. This method maintains the efficiency of teacher-mandated training while introducing a controlled planned sampling mechanism and an approximation of the infinite impulse response of acoustic feedback based on a finite-order Newman series, enabling the model to gradually adapt to real closed-loop operating conditions during training.
[0009] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:
[0010] A non-autoregressive learning method for planned sampling aimed at acoustic feedback suppression includes the following steps:
[0011] Step A: Construct a non-autoregressive open-loop training framework: Under the set boundary stable gain condition, generate reference input features that do not contain recursive dependencies based on the teacher-forced policy, and construct initial training sample pairs.
[0012] Step B: Construct a two-stage finite boundary plan sampling mechanism: The training process is divided into two stages; the first stage uses teacher-forced input; the second stage dynamically selects the input source for the current frame between teacher-forced input and model prediction results based on a preset probability scheduling strategy.
[0013] Step C: Construct a finite-order Newman series approximation operator to approximate the recursive structure of the infinite impulse response of acoustic feedback, and generate feedback estimation features consistent with the closed-loop operation behavior.
[0014] Step D, phased model training, includes:
[0015] Step D1, First Stage: Use the training samples constructed in Step A to perform non-autoregressive preliminary training on the acoustic feedback suppression model structure to obtain the basic model;
[0016] Step D2, Second Stage: Based on the basic model, the input signal selected according to the preset probability scheduling strategy in step B, and combined with the feedback estimation features generated by the finite-order Newman series approximation operator in step C, are used to construct new training samples; and the constructed new training samples are used to continue training the basic model, so that the model gradually adapts to the closed-loop inference conditions and obtains the final trained acoustic feedback suppression model.
[0017] Furthermore, the non-autoregressive open-loop training framework constructed in step A specifically includes:
[0018] A1. Determine the boundary stability gain, which should be 2-4 dB lower than the maximum stability gain;
[0019] A2. Based on the boundary stabilization gain and the clean target speech, construct a hybrid sample containing the acoustic feedback and the clean target speech by simulating acoustic feedback through convolution, and construct sample pairs of the hybrid sample and the clean target speech;
[0020] A3. Determine the use of multi-scale Fourier transform loss function during training, and use clean target speech as labels for supervised learning.
[0021] Furthermore, the two-stage finite boundary plan sampling mechanism constructed in step B specifically includes:
[0022] B1. In the first stage, the structural parameters of the acoustic feedback suppression model are frozen, and the neural network is driven by teacher-forced input to generate non-autoregressive acoustic feedback signals under non-recursive conditions, which serve as candidate feedback signals for the training process.
[0023] B2. In the second stage, a dynamic selection is made between the teacher-mandated input and the model prediction result using a preset sampling probability; the sampling probability varies with the training rounds, specifically expressed as follows:
[0024] ,
[0025] in, For the first Sampling probability of each round and These represent the initial probability and lower bound probability of using teacher-forced input, respectively, with the lower bound probability being non-zero; parameters and Used to control the offset position and decay rate of the sampling probability curve.
[0026] Furthermore, the construction of the finite-order Newman series approximation operator in step C is achieved by expanding the output signal of the closed-loop feedback system into an infinite-order Newman series and truncating it at higher orders. The truncation order is determined by the acoustic feedback path characteristics and the system gain. The input speech sequence is convolved with the finite-order Newman series approximation operator and superimposed to obtain the acoustic feedback signal of the approximate closed-loop feedback system.
[0027] Furthermore, the acoustic feedback suppression model structure is an asymmetric encoding and decoding model structure, including: an asymmetric dual encoder, both of which extract and downsample the input signal through two layers of convolutional modules; an attention fusion module, which fuses the output features of the two encoders and outputs the fused features; a parallel time-frequency-LSTM module, in which multiple parallel-connected time-frequency-LSTM modules simultaneously use the fused features as input to perform multi-angle time-frequency modeling, and concatenate the output features of each time-frequency-LSTM module to form rich joint time-frequency features; and a decoder, which gradually upsamples to recover the time-frequency dependency through two cascaded deconvolutional modules, estimates the spectral mask and applies it to the input microphone mixed signal, and outputs enhanced speech with acoustic feedback removed.
[0028] Furthermore, the phased training in step D is as follows:
[0029] Step D1, First Stage: Input the mixed samples constructed in Step A into the first encoder and the sample pairs into the second encoder. Use clean target speech as the training target to supervise the training of the acoustic feedback suppression model structure and obtain the basic model.
[0030] Step D2, Second Stage: Based on the basic model, select the input signal source according to the preset probability, and use the feedback estimation features generated by the finite-order Newman series to replace the candidate feedback estimation features to construct new training samples. Use the training samples and the sample pairs composed of the training samples and the clean target speech as the input of the two encoders, and use the clean target speech as the training target to continue training the obtained basic model, so that the model gradually adapts to the closed-loop inference conditions and obtains the final trained acoustic feedback suppression model.
[0031] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0032] The method proposed in this invention is a training method applicable to continuous acoustic feedback suppression systems. It introduces prediction-related feedback information into model training without explicitly performing autoregressive recursive inference, effectively reducing the mismatch between the training and closed-loop inference stages. Through a restricted two-stage planned sampling strategy, the dependency between model output and input is gradually restored during training. Furthermore, by setting a non-zero sampling lower bound, the model is prevented from completely relying on its own prediction results, thereby improving the stability of the training process and preventing performance degradation or model collapse. Finally, a finite-order Newman series is used to approximate the recursive structure of the infinite impulse response of acoustic feedback, transforming the originally difficult problem of infinite-level feedback accumulation into a computable finite impulse response form, significantly reducing computational complexity while maintaining approximate accuracy.
[0033] This invention maintains good robustness and stability under high system gain and complex acoustic coupling conditions, improves acoustic feedback suppression performance and speech intelligibility, and is applicable to various audio amplification devices such as hearing aids, wireless headphones, and in-vehicle voice systems. It has good engineering practicality and application prospects. Attached Figure Description
[0034] Figure 1 A flowchart of the method provided by the present invention;
[0035] Figure 2 A schematic diagram of the training process for the acoustic feedback suppression model structure provided by this invention;
[0036] Figure 3 This is a structural diagram of the acoustic feedback suppression model used in this invention. Detailed Implementation
[0037] To make the technical solution of the present invention clearer, the technical solution of the present invention will be described in further detail below with reference to the accompanying drawings and specific embodiments.
[0038] like Figures 1 to 2 As shown, the present invention provides a non-autoregressive learning method for planned sampling aimed at acoustic feedback suppression, comprising the following steps:
[0039] Step A: Construct a non-autoregressive open-loop training framework: Under the set boundary stable gain condition, generate reference input features that do not contain recursive dependencies based on the teacher-forced policy, and construct initial training sample pairs.
[0040] Step B: Construct a two-stage finite boundary plan sampling mechanism: The training process is divided into two stages; the first stage uses teacher-forced input; the second stage dynamically selects the input source for the current frame between teacher-forced input and model prediction results based on a preset probability scheduling strategy.
[0041] Step C: Construct a finite-order Newman series approximation operator to approximate the recursive structure of the infinite impulse response of acoustic feedback, and generate feedback estimation features consistent with the closed-loop operation behavior.
[0042] Step D, phased model training, includes:
[0043] Step D1, First Training Stage: Use the training samples constructed in Step A to perform non-autoregressive preliminary training on the acoustic feedback suppression model structure to obtain the basic model;
[0044] Step D2, Second Training Stage: Based on the basic model, new training samples are constructed by selecting the input signal according to the preset probability scheduling strategy in step B and combining it with the feedback estimation features generated by the finite-order Newman series approximation operator in step C; and the constructed new training samples are used to continue training the basic model so that the model gradually adapts to the closed-loop inference conditions and obtains the final trained acoustic feedback suppression model.
[0045] Specifically, the non-autoregressive open-loop training framework constructed in step A includes:
[0046] A1. Determine the boundary stability gain :
[0047] For a given room impulse response, the maximum critical steady-state gain Represented as:
[0048] ,
[0049] ,
[0050] in, Angular frequency, For a given room impulse response, the frequency domain representation is given. Indicates system latency. Indicates the sampling frequency. This is the set of all dangerous frequencies that would produce acoustic feedback howling due to phase alignment. It is an integer;
[0051] Gain margin Represented as:
[0052] ,
[0053] To avoid howling and excessive reverberation, the boundary stabilization gain should be set 2-4 dB lower than the maximum stabilization gain.
[0054] A2. Based on the aforementioned boundary-stabilized gain and clean target speech, a hybrid sample containing acoustic feedback and clean target speech is constructed by simulating acoustic feedback through convolution. This leads to the construction of a system containing mixed samples. and pure target voice sample pairs Specifically, it is expressed as:
[0055] , ,
[0056] in, Indicates the first Frame mixed sample signal, Indicates the first A frame of clean target speech signal, For room shock response, Indicates system latency. Indicates the first The acoustic feedback signal of the frame; because the teacher forces the assumption that the model can completely filter the acoustic feedback in each frame, the acoustic feedback of the next frame is entirely composed of the input of the previous frame.
[0057] A3. During training, a multi-scale Fourier transform loss function is used, with clean target speech as the label for supervised learning, and the enhanced speech signal is output; the multi-scale Fourier transform loss function is expressed as:
[0058] ,
[0059] in, This represents the short-time Fourier transform loss. The clean target speech signal output by the model. The enhanced output signal for model estimation Indicates the spectral convergence loss. Indicates logarithmic magnitude loss, and and The specific representation is as follows:
[0060] , ,in This indicates the number of time frames output by the short-time Fourier transform.
[0061] Specifically, the two-stage finite boundary plan sampling mechanism constructed in step B includes:
[0062] B1. In the first stage, the structural parameters of the acoustic feedback suppression model are frozen, and the neural network is driven by teacher-forced input to generate a non-autoregressive acoustic feedback signal under non-recursive conditions. This signal serves as the candidate feedback estimation signal for the training process, and is expressed as follows:
[0063] ,
[0064] in, Indicates the first The teacher forces the input of the frame. This represents the currently frozen model structure parameters. By performing forward inference on the previous frame's input under the frozen parameter conditions, a model with similar closed-loop behavior can be constructed without explicit recursive execution. Frame candidate feedback estimation signal ;
[0065] B2. In the second stage, a sampling probability curve is constructed to ensure that the sampling probability does not decay to zero during training, thus preventing the model from relying entirely on its own prediction results during the convergence phase. The sampling probability changes with the training rounds, specifically as follows:
[0066] ,
[0067] in, For the first Sampling probability of each round and These represent the initial probability and lower bound probability of using teacher-forced input, respectively, with the lower bound probability being non-zero; parameters and Used to control the offset position and decay rate of the sampling probability curve; by setting a non-zero lower limit for the sampling probability, the model always retains a certain proportion of teacher-forced input throughout the training process, thereby preventing the training process from degenerating into a fully self-conditional state and improving the stability and robustness of the model under closed-loop operating conditions.
[0068] By introducing a Bernoulli random variable to select between teacher-forced input and model prediction results using a preset sampling probability, the source of the input signal is determined, specifically as follows:
[0069] ,
[0070] in, Indicates the determined first The source of the frame input signal, Let be a Bernoulli random variable; through the aforementioned mechanism, the model gradually incorporates prediction-related input conditions during the training process.
[0071] Specifically, in step C, the finite-order Newman series approximation operator is constructed, and the specific process is as follows:
[0072] C1. Infinite Impulse Response Newman Series Approximation: Consider a linear acoustic feedback system operating under boundary stable gain conditions, where the input signal in the closed-loop system is... With output signal The relationship between them is represented as follows:
[0073] ,
[0074] in, , This represents the scalar gain corresponding to the boundary stability gain. Indicates the impulse response via the acoustic feedback path The corresponding convolution operator allows us to obtain the analytical form of the closed-loop system as follows:
[0075] ,
[0076] in, Denotes the inverse operator; under the boundary stable gain condition, the operator norm satisfies Therefore, the inverse operator can be expanded in the form of a Newman series, i.e. ,
[0077] Therefore, the output signal of the closed-loop system can be represented as an infinite series formed by the superposition of multiple feedback components:
[0078] ,
[0079] In the discrete-time domain, the above expression can be equivalently represented as a multiple convolution form, where the output signal is the superposition of the convolution results of the input signal and the impulse responses of different order acoustic feedback paths, where higher-order feedback components correspond to multiple self-convolutions of the acoustic feedback paths; thus, the infinite impulse response Newman series approximate convolution operator is expressed as:
[0080] ;
[0081] C2. Newman series approximation of finite impulse response: due to... Under certain conditions, the influence of higher-order feedback operators gradually diminishes with increasing order; therefore, infinite-order Newman series can be applied to finite-order series. Truncation is performed at the specified point, and the truncation order is determined by this. The choice of the acoustic feedback path and system gain is determined by both factors, allowing for control over computational complexity while maintaining approximate accuracy. Therefore, a finite-order Newman series is used to approximate the closed-loop feedback system, and its output signal is expressed as:
[0082] ;
[0083] Under the finite-order Newman series approximation, the th The feedback components of a frame can be obtained from the input speech sequence: With the following kernel function The approximate convolution operator, constructed using a finite-order Newman series, is used to calculate the following:
[0084] ,
[0085] in, This represents the direct path component, corresponding to the current input signal that does not propagate through feedback. Represents the acoustic feedback path impulse response Second self-convolution, and .
[0086] Specifically, such as Figure 3 As shown, the acoustic feedback suppression model structure is an asymmetric encoding and decoding model structure, including: an asymmetric dual encoder, both of which extract and downsample the input signal through two layers of convolutional modules. One encoder takes only the microphone mixed signal as input, while the other encoder takes a combination of the microphone mixed signal and the clean target speech signal as input; an attention fusion module, which fuses the output features of the two encoders and outputs the fused features; a parallel time-frequency-LSTM module, in which multiple parallel-connected time-frequency-LSTM modules simultaneously take the fused features as input to perform multi-angle time-frequency modeling, and concatenate the output features of each time-frequency-LSTM module to form rich joint time-frequency features; and a decoder, which gradually upsamples to recover the time-frequency dependency through two cascaded deconvolutional modules, estimates the spectral mask and applies it to the input microphone mixed signal, and outputs enhanced speech with acoustic feedback removed.
[0087] Specifically, for such Figure 3 The acoustic feedback suppression model structure shown is trained in stages, specifically as follows:
[0088] Step D1, First Stage: Using the mixed samples constructed in Step A, and sample pairs consisting of mixed samples and clean target speech as inputs to the two encoders, with clean target speech as the training target, for example... Figure 3 The asymmetric encoding / decoding model structure shown is used for training to obtain the basic model;
[0089] Step D2, Second Stage: Based on the basic model, select the input signal source according to the preset probability, and use the feedback estimation features generated by the finite-order Newman series to replace the candidate feedback estimation features to construct new training samples. Use the training samples and the sample pairs composed of the training samples and the clean target speech as the input of the two encoders, and use the clean target speech as the training target to continue training the obtained basic model, so that the model gradually adapts to the closed-loop inference conditions and obtains the final trained acoustic feedback suppression model.
[0090] In one specific embodiment, comparative experiments were conducted on the publicly available DNS Challenge and VCTK datasets to verify the denoising effect of the method provided by the present invention.
[0091] First, all speech data in the datasets were processed using a 16kHz sampling rate. The DNS Challenge dataset contains clean speech signals from 2150 speakers, and the VCTK dataset contains speech recordings from 110 English speakers with different accents. Each speaker reads approximately 400 sentences from news texts. The VCTK dataset is divided according to speaker identity, with the speech from speakers numbered 225-234 used as the test set and the speech from the remaining speakers used as the training set, thus ensuring that there is no overlap between the training and test sets.
[0092] Room impulse responses were generated using an image-based method: 200 rooms were simulated, with room length and width uniformly sampled within the range of 10-15 meters and height within the range of 3-5 meters. Seven sets of room impulse responses were generated for each room, with five sets used for training and evaluation under known acoustic conditions, and the remaining two sets used for generalization performance evaluation under unknown acoustic conditions. In each simulated room, microphones were placed near the center of the room, with a horizontal offset of 0.3 meters and a height sampled within the range of 1.6-1.8 meters. Speakers were placed 0.2-0.5 meters away from the microphones. The sound absorption coefficients of the room walls were selected from a uniform distribution of 0.8-1.0, and the reflection order was set to 10 to simulate a real indoor acoustic environment.
[0093] The following existing models were used for experimental verification: 1) DTLN-AEC, a dual-path acoustic echo cancellation model, which models the speech waveform in the time domain and enhances it in the frequency domain using spectral masking, achieving joint modeling of time and frequency domain features. This method can simultaneously utilize temporal structure information and spectral feature information to achieve effective echo suppression while ensuring low system latency. It has been widely used in real-time acoustic echo and acoustic feedback suppression tasks and serves as a representative baseline method; 2) MTFAA-Net, a multi-scale time-frequency convolutional neural network structure that introduces an axial attention mechanism to enhance the time-frequency feature modeling capability. Originally used for acoustic echo cancellation tasks, it can be used as a time-frequency feature model under open-loop conditions. The frequency domain baseline method replaces the deep filtering synthesis module in this embodiment to meet the experimental constraints, and adopts a mask-based signal reconstruction method to complete the speech enhancement processing; 3) DeepMFC is an acoustic feedback suppression method based on deep neural networks. It addresses the feedback problem under boundary stability conditions in hearing aid systems. It directly maps the speech signal containing feedback interference to a clean speech signal under open-loop conditions through supervised learning to improve the stability of the system under high-gain conditions; 4) MSA-DPCRN is a frequency domain convolutional recurrent neural network structure. It adopts an asymmetric dual-path encoder design to extract features from the microphone signal and the reference signal respectively, thereby obtaining the spectral representation and cross-domain correlation features.
[0094] Tables 1 and 2 present the performance comparison results of training on the DNS Challenge dataset and the VCTK dataset, respectively, using the planned sampling method and the method of the present invention on the existing model architecture. Planned sampling refers to the sampling probability curve using the traditional exponential decay form, where the sampling probability forced by the teacher decays to zero, and no finite-order Newman series approximation operator is introduced to approximate the acoustic feedback of the closed-loop system.
[0095] Table 1
[0096]
[0097] Table 2
[0098]
[0099] As can be seen from the table, the model trained using the method of this invention significantly outperforms the control model trained using only traditional planned sampling in all four performance metrics: Perceptual Speech Quality Assessment (PESQ), Scale Invariant Signal-to-Noise Ratio (SI-SNR), Extended Short-Time Objective Intelligibility (ESTOI), and Transcription Error Rate (TER).
[0100] In summary, this invention, by introducing a controlled planned sampling training mechanism and finite-order Newman series approximation feedback, effectively alleviates the mismatch between the training phase and the closed-loop inference phase while maintaining training efficiency. This improves the stability and robustness of the acoustic feedback suppression model under high-gain closed-loop conditions, stably and effectively enhancing the model's performance. The method is ingenious and novel, possessing good engineering practicality and application prospects.
[0101] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Any modifications, equivalent substitutions, or improvements made by those skilled in the art within the scope of the technology disclosed in this invention, based on the technical solution and inventive concept of the present invention, should be covered within the protection scope of this invention. Therefore, the protection scope of this invention should be determined by the scope of the claims.
Claims
1. A non-autoregressive learning method for planned sampling aimed at acoustic feedback suppression, characterized in that, Includes the following steps: Step A: Construct a non-autoregressive open-loop training framework: Under the set boundary stable gain condition, generate reference input features that do not contain recursive dependencies based on the teacher-forced policy, and construct initial training sample pairs. Step B: Construct a two-stage finite boundary plan sampling mechanism: The training process is divided into two stages; the first stage uses teacher-forced input; the second stage dynamically selects the input source for the current frame between teacher-forced input and model prediction results based on a preset probability scheduling strategy. Step C: Construct a finite-order Newman series approximation operator to approximate the recursive structure of the infinite impulse response of acoustic feedback, and generate feedback estimation features consistent with the closed-loop operation behavior. Step D, phased model training, includes: Step D1, First Stage: Use the training samples constructed in Step A to perform non-autoregressive preliminary training on the acoustic feedback suppression model structure to obtain the basic model; Step D2, Second Stage: Based on the basic model, the input signal selected according to the preset probability scheduling strategy in step B, and combined with the feedback estimation features generated by the finite-order Newman series approximation operator in step C, are used to construct new training samples; and the constructed new training samples are used to continue training the basic model, so that the model gradually adapts to the closed-loop inference conditions and obtains the final trained acoustic feedback suppression model.
2. The non-autoregressive learning method for planned sampling aimed at acoustic feedback suppression according to claim 1, characterized in that, The non-autoregressive open-loop training framework constructed in step A specifically includes: A1. Determine the boundary stability gain, which should be 2-4 dB lower than the maximum stability gain; A2. Based on the boundary stabilization gain and the clean target speech, construct a hybrid sample containing the acoustic feedback and the clean target speech by simulating acoustic feedback through convolution, and construct sample pairs of the hybrid sample and the clean target speech; A3. Determine the use of multi-scale Fourier transform loss function during training, and use clean target speech as labels for supervised learning.
3. The planned sampling non-autoregressive learning method for acoustic feedback suppression according to claim 2, characterized in that, Step B constructs a two-stage finite boundary plan sampling mechanism, which specifically includes: B1. In the first stage, the structural parameters of the acoustic feedback suppression model are frozen, and the neural network is driven by teacher-forced input to generate non-autoregressive acoustic feedback signals under non-recursive conditions, which serve as candidate feedback signals for the training process. B2. In the second stage, a dynamic selection is made between the teacher-forced input and the model prediction result using a preset sampling probability; the sampling probability changes with the training rounds, specifically as follows: , in, For the first Sampling probability of each round and These represent the initial probability and lower bound probability of using teacher-forced input, respectively, with the lower bound probability being non-zero; parameters and Used to control the offset position and decay rate of the sampling probability curve.
4. The non-autoregressive learning method for planned sampling aimed at acoustic feedback suppression according to claim 3, characterized in that, In step C, the finite-order Newman series approximation operator is constructed by expanding the output signal of the closed-loop feedback system into an infinite-order Newman series and truncating it at higher orders. The truncation order is determined by the acoustic feedback path characteristics and the system gain. The input speech sequence is convolved with the finite-order Newman series approximation operator and superimposed to obtain the acoustic feedback signal of the approximate closed-loop feedback system.
5. The planned sampling non-autoregressive learning method for acoustic feedback suppression according to claim 4, characterized in that, The acoustic feedback suppression model structure is an asymmetric encoding and decoding model structure, including: an asymmetric dual encoder, each of which performs feature extraction and downsampling on the input microphone mixed signal through two layers of convolutional modules; an attention fusion module, which fuses the output features of the two encoders and outputs the fused features; a parallel time-frequency-LSTM module, in which multiple parallel-connected time-frequency-LSTM modules simultaneously use the fused features as input to perform multi-angle time-frequency modeling, and concatenate the output features of each time-frequency-LSTM module to form rich joint time-frequency features; and a decoder, which gradually upsamples to recover the time-frequency dependency through two cascaded deconvolutional modules, estimates the spectral mask and applies it to the input microphone mixed signal, and outputs enhanced speech with acoustic feedback removed.
6. The planned sampling non-autoregressive learning method for acoustic feedback suppression according to claim 5, characterized in that, Step D, the phased training, is as follows: Step D1, First Stage: Input the mixed samples constructed in Step A into the first encoder and the sample pairs into the second encoder. Use clean target speech as the training target to train the acoustic feedback suppression model structure and obtain the basic model. Step D2, Second Stage: Based on the basic model, select the input signal source according to the preset sampling probability, and use the feedback estimation features generated by the finite-order Newman series to replace the candidate feedback estimation features to construct new training samples. Use the training samples and the sample pairs composed of the training samples and the clean target speech as the input of the two encoders, and use the clean target speech as the training target to continue training the obtained basic model, so that the model gradually adapts to the closed-loop inference conditions and obtains the final trained acoustic feedback suppression model.