Active noise reduction method based on improved Wave-U-Net model
By improving the Wave-U-Net model, combining the mirror source method and SEF function, the divergence problem of traditional FxLMS algorithm under high variability and speaker saturation is solved, and efficient and robust active noise reduction effect is achieved.
Patent Information
- Application Number
- CN202510688662.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-08-29
AI Technical Summary
Traditional FxLMS algorithms have divergence risks when secondary paths are highly variable or output saturated, have slow adaptive characteristics, and require high sampling frequency operation to solve the causal relationship problem, limiting the noise suppression effect.
Using the improved Wave-U-Net model, the main path and secondary path are generated by constructing a low-frequency hybrid sinusoidal noise dataset using the mirror source method, and the speaker saturation characteristics are simulated by the SEF function, combined with the improved gated convolutional layer and the LSTM module, model training and evaluation are performed, and reference microphone and error microphone are used for supervised learning.
It realizes efficient noise control under different nonlinearity and noise types, significantly improving noise reduction performance and robustness, especially in nonlinear scenarios, improving noise reduction effect compared with traditional methods.
Smart Images

Figure CN120564684A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of active noise reduction, and in particular to an active noise reduction method based on an improved Wave-U-Net model. Background Art
[0002] Active Noise Control (ANC) technology is widely used in scenarios such as active noise-canceling headphones, car cabin noise reduction, and active sound barriers, and can effectively achieve noise control.
[0003] Traditional active noise reduction methods primarily rely on adaptive filters. With the development of high-performance, high-speed processors, digital ANC processors are becoming increasingly popular due to their robustness, ease of design, and ability to adapt to environmental changes. Furthermore, these digital processors have paved the way for the implementation of adaptive control algorithms. Among these adaptive algorithms, the FxLMS algorithm is widely used in various applications in the field of active noise reduction. The FxLMS algorithm compensates for secondary path delays and achieves optimal noise control around the error sensor. However, the adaptive nature of FxLMS risks divergence when there is high variability in the secondary path or when the output is saturated. Furthermore, the adaptive algorithm's slow convergence significantly limits its noise suppression capabilities. Furthermore, the traditional FxLMS algorithm requires operation at a higher sampling frequency to minimize causality issues in ANC systems. To address this, we propose an active noise reduction method based on a modified Wave-U-Net model. Summary of the Invention
[0004] The object of the present invention is to provide an active noise reduction method based on an improved Wave-U-Net model to solve the problems raised in the above background technology.
[0005] To achieve the above objectives, the present invention provides the following technical solution: an active noise reduction method based on an improved Wave-U-Net model, comprising the following steps:
[0006] S1. Dataset setup: Construct a low-frequency mixed sinusoidal noise dataset based on the characteristics of transformer noise;
[0007] S2, simulation path construction: both the primary path and the secondary path are generated by the mirror source method;
[0008] S3, nonlinear compensation: the output is subjected to the SEF function to simulate the speaker saturation characteristics;
[0009] S4. Model training and evaluation: The GWUNet model is trained using the constructed low-frequency mixed sinusoidal noise dataset, and the model is evaluated using untrained nonlinearity and untrained noise.
[0010] Optionally, S1 includes: constructing a low-frequency mixed sinusoidal noise data set based on the characteristics of transformer noise, and in order to more realistically simulate the noise conditions in the industrial environment, Gaussian white noise with different signal-to-noise ratios can be randomly added to the pure sinusoidal signal data set;
[0011] The signal-to-noise ratio range is set to [5, 10, 15, 20] dB, which can simulate noise interference of different intensities. The signal-to-noise ratio is an important indicator to measure the signal strength relative to the background noise. The unit is decibel (dB), and its mathematical expression is:
[0012]
[0013] Where, P signal Represents the power of the signal, P noise It represents the power of the noise. As can be seen from the formula, the smaller the SNR value, the greater the proportion of noise in the signal.
[0014] Optionally, the S2 includes: the path and the secondary path are both generated by the mirror source method; the mirror source method is used to set a room with a length, width and height of 3 meters, 4 meters and 2 meters respectively, and the reverberation time (T60) is 0.2s; during the training process, the reference microphone (I), secondary speaker (J) and error microphone (K) are set to a 1×1×1 pattern, the coordinates of the reference microphone are (1.5, 1, 1) meters, the coordinates of the secondary speaker are (1.5, 2.5, 1) meters, and the coordinates of the error microphone are (1.5, 3, 1), and the main path and secondary path are generated based on this configuration.
[0015] Optionally, the S3 includes:
[0016] In the study of loudspeakers in nonlinear ANC (NANC) systems, nonlinearity is usually represented by a single saturation model, where the limiting effect can be represented by a nonlinear function, namely the scaled error function (SEF). The formula of SEF is:
[0017]
[0018] Where y represents the cancellation signal of the input speaker, η 2 Indicates the degree of nonlinearity, the smaller the value, the higher the degree of nonlinearity;
[0019] When η 2 As η approaches infinity, SEF becomes linear; 2 When it approaches zero, SEF behaves as a hard limit. In order to study the impact of nonlinearity on ANC performance, this paper selects four different speaker nonlinearities for model training, namely η 2 =0.1 (severe nonlinearity), η2 =1 (moderate nonlinearity), η 2 = 10 (weak nonlinearity) and η 2 =∞(linear). By comparing the normalized mean square error of the denoising results of the GWUNet method and the FxLMS method under different speaker nonlinearities, the difference in noise reduction performance between the two methods is evaluated.
[0020] Optionally, S4 includes: during the model training process, using the AMSGard optimizer to update the network weights; the initial learning rate of the training is 0.001, and then the learning rate is dynamically adjusted through an exponentially decreasing regulator, and the learning rate attenuation coefficient is 0.98; the batch size is set to 60, and the total number of iterations is 100.
[0021] Optionally, the gated convolution layer in S1 performs data processing, and the gated convolution layer includes: a convolution layer, a BN layer and a gated linear unit;
[0022] The improved gated convolutional layer uses one-dimensional convolution to process the transformer noise signal and extract its time and frequency domain features. The batch normalization layer is connected after the convolutional layer to ensure the consistency of the input data distribution of each layer, thereby accelerating training and improving model stability. The gated linear unit is a core component that dynamically adjusts the information flow and determines the probability of features being passed to the next layer. The improved gated convolutional layer uses a dual-branch structure to process input data.
[0023] Among them, branch A is used to calculate the threshold value, and branch B is used to calculate the characteristic value;
[0024] In branch A, data is processed sequentially through the convolutional layer and the BN layer, and a threshold value is generated using the Sigmoid activation function. This value reflects the importance of the feature, with more significant features corresponding to higher values. In branch B, data is also processed through the convolutional layer and the BN layer, and the eigenvalue is directly output. By multiplying the threshold value with the eigenvalue, the system filters out important information and suppresses irrelevant or redundant data. The calculation formula is as follows:
[0025]
[0026] Where BN(·) represents batch normalization; W and V are the weight matrices of the convolution kernels in the two branches, respectively; b and c are the bias terms in the two branches, respectively; * represents the convolution operation; σ represents the sigmoid activation function.
[0027] Optionally, the GWUNet model uses an improved gated convolution as the basic building block of the network model, which can more effectively extract features;
[0028] The input and output of the model are both single-channel waveform signals, and the single-channel waveform signal is output after calculation by the model;
[0029] In this model, skip connections are used to pass the input of the encoding module to the decoding module at the same level, ensuring that the features extracted by the encoder can be fully utilized. This method can effectively avoid the loss of input information during transmission, thereby optimizing the propagation of model information and gradients; the LSTM module embedded in the middle layer can effectively capture the long-term dependencies of acoustic signals and enhance the model's ability to model temporal features; except for the output layer and the middle layer, each module uses the LeakyReLU activation function to introduce nonlinear characteristics, and adds a batch normalization layer after convolution or deconvolution to accelerate training and improve stability; the output layer uses convolution combined with the Tanh activation function to generate the output signal.
[0030] Compared with the prior art, the present invention provides an active noise reduction method based on an improved Wave-U-Net model, which has the following beneficial effects:
[0031] This active noise reduction method based on the improved Wave-U-Net model improves the Wave-U-Net network, which is composed of convolutions, and proposes an active noise reduction method using improved gated convolutional layers as the basic unit. By adding functional units such as the primary path, secondary path, and loudspeaker to the model, this method effectively utilizes ANC technology to achieve noise reduction. The method is defined as a supervised learning mechanism, using the noise signal collected by the reference microphone as the training feature value and the noise collected by the error microphone as the training target value. The model can learn how to effectively suppress noise from the training data, thereby achieving efficient active noise control. In experiments, the nonlinear distortion caused by the loudspeaker is simulated by scaling the error function, and the model is trained using four nonlinearities. The performance of the proposed method is then evaluated on real recorded noise. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 This is a schematic diagram of the structure of the improved gated convolutional layer of the present invention;
[0033] Figure 2 Schematic diagram of the GWUNet model structure of the present invention;
[0034] Figure 3 For the present invention in η 2 =0.5, the time domain and power density spectrum changes of the 200Hz sinusoidal signal and transformer noise spliced noise;
[0035] Figure 4 For the present invention in η 2 Schematic diagram of the time domain and power density spectrum changes of the mixed noise of 200Hz sinusoidal signal and transformer noise under the condition of =0.5. DETAILED DESCRIPTION
[0036] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0037] like Figures 1-4 As shown, the present invention provides a technical solution: an active noise reduction method based on an improved Wave-U-Net model, comprising the following steps:
[0038] S1. Dataset setup: Construct a low-frequency mixed sinusoidal noise dataset based on the characteristics of transformer noise;
[0039] A low-frequency mixed sinusoidal noise dataset was constructed based on the characteristics of transformer noise. To more realistically simulate the noise conditions in an industrial environment, Gaussian white noise with different signal-to-noise ratios was randomly added to the pure sinusoidal signal dataset.
[0040] The signal-to-noise ratio range is set to [5, 10, 15, 20] dB, which can simulate noise interference of different intensities. The signal-to-noise ratio is an important indicator to measure the signal strength relative to the background noise. The unit is decibel (dB), and its mathematical expression is:
[0041]
[0042] Where, P signal Represents the power of the signal, P noise It represents the power of the noise. As can be seen from the formula, the smaller the SNR value, the greater the proportion of noise in the signal.
[0043] S2, simulation path construction: both the primary path and the secondary path are generated by the mirror source method;
[0044] Both the primary and secondary paths are generated by the mirror source method. The mirror source method is used to set up a room with a length, width, and height of 3 meters, 4 meters, and 2 meters, respectively, and a reverberation time (T60) of 0.2s. During the training process, the reference microphone (I), secondary speaker (J), and error microphone (K) are set in a 1×1×1 pattern. The coordinates of the reference microphone are (1.5, 1, 1) meters, the coordinates of the secondary speaker are (1.5, 2.5, 1) meters, and the coordinates of the error microphone are (1.5, 3, 1). The primary and secondary paths are generated based on this configuration.
[0045] S3, nonlinear compensation: the output is subjected to the SEF function to simulate the speaker saturation characteristics;
[0046] In the study of loudspeakers in nonlinear ANC (NANC) systems, nonlinearity is usually represented by a single saturation model, where the limiting effect can be represented by a nonlinear function, namely the scaled error function (SEF). The formula of SEF is:
[0047]
[0048] Where y represents the cancellation signal of the input speaker, η 2 Indicates the degree of nonlinearity, the smaller the value, the higher the degree of nonlinearity;
[0049] When η 2 As η approaches infinity, SEF becomes linear; 2 When it approaches zero, SEF behaves as a hard limit. In order to study the impact of nonlinearity on ANC performance, this paper selects four different speaker nonlinearities for model training, namely η 2 =0.1 (severe nonlinearity), η 2 =1 (moderate nonlinearity), η 2 = 10 (weak nonlinearity) and η 2 =∞(linear). By comparing the normalized mean square error of the denoising results of the GWUNet method and the FxLMS method under different speaker nonlinearities, the difference in noise reduction performance between the two methods is evaluated.
[0050] S4. Model training and evaluation: The GWUNet model is trained using the constructed low-frequency mixed sinusoidal noise dataset, and the model is evaluated using untrained speaker nonlinearities and untrained noise.
[0051] During model training, the AMSGard optimizer is used to update the network weights. The initial learning rate of training is 0.001, and then the learning rate is dynamically adjusted through an exponentially decreasing regulator with a learning rate decay coefficient of 0.98. The batch size is set to 60, and the total number of iterations is 100.
[0052] In S1, an improved gated convolutional layer is used for data processing. The improved gated convolutional layer consists of a convolutional layer, a batch normalization layer, and a gated linear unit. The convolutional layer uses one-dimensional convolution to process the transformer noise signal and extract its time and frequency domain features. The batch normalization layer is connected after the convolutional layer to ensure the consistency of the input data distribution across layers, thereby accelerating training and improving model stability. The gated linear unit is a core component that dynamically adjusts the information flow and determines the probability of features being passed to the next layer. The improved gated convolutional layer uses a dual-branch structure to process input data.
[0053] Among them, branch A is used to calculate the threshold value, and branch B is used to calculate the eigenvalue. In branch A, the data is processed by the convolution layer and the BN layer in sequence, and the threshold value is generated using the Sigmoid activation function. This value reflects the importance of the feature, and the more significant the feature, the higher the value. In branch B, the data is also processed by the convolution layer and the BN layer, and the eigenvalue is directly output. By multiplying the threshold value and the eigenvalue, the system filters out important information and suppresses irrelevant or redundant data. The calculation formula is as follows:
[0054]
[0055] Where BN(·) represents batch normalization; W and V are the weight matrices of the convolution kernels in the two branches, respectively; b and c are the bias terms in the two branches, respectively; * represents the convolution operation; σ represents the sigmoid activation function.
[0056] The GWUNet model uses improved gated convolution as the basic building block of the network model, which can extract features more effectively. The input and output of the model are both single-channel waveform signals, which are output after calculation by the model. In this model, jump connections are used to pass the input of the encoding module to the decoding module at the same level to ensure that the features extracted by the encoder can be fully utilized. This method can effectively avoid the loss of input information during transmission, thereby optimizing the propagation effect of model information and gradients. The LSTM module embedded in the middle layer can effectively capture the long-term dependencies of acoustic signals and enhance the model's ability to model temporal features. Except for the output layer and the middle layer, each module uses the Leaky ReLU activation function to introduce nonlinear characteristics, and adds a batch normalization layer after convolution or deconvolution to accelerate training and improve stability. The output layer uses convolution combined with the Tanh activation function to generate the output signal.
[0057] As an application of this embodiment:
[0058] The GWUNet model parameter configuration shows that the encoding module consists of five layers of gated convolutional units (DownGLUs). The number of input channels increases with each layer, and each layer uses convolution operations with a kernel size of 15 and a stride of 1 to maintain temporal resolution and extract features. The intermediate layers use LSTMs (of dimension 120) to model temporal features, effectively capturing long-term dependencies in acoustic signals. The decoding module uses five layers of deconvolutional gated units (UpGLUs) to gradually restore feature details. The output features of the corresponding layers in the encoding path are concatenated with the input of the decoding layer via skip connections. For example, UpGLU_5 receives the concatenation of 120 channels from the encoding layer and 120 channels from the skip connection (for a total of 240 channels). Each deconvolution layer uses a kernel size of 5 and a stride of 1. The final output layer uses a convolutional layer with a kernel size of 1 to compress the channels and uses the Tanh activation function to map the features into a single-channel waveform signal. In this design, the symmetrical structure of the encoding and decoding paths, the LSTM module, and the skip connection significantly improve the model's ability to model noise characteristics. The parameter configuration of the GWUNet model is shown in the following table:
[0059]
[0060] Evaluation Metrics: Normalized mean square error (NMSE) is used as an indicator to evaluate the performance of the ANC system, and the experimental results are analyzed and compared based on this indicator. As a commonly used indicator for ANC evaluation, NMSE is defined as follows:
[0061]
[0062] Where L represents the signal length, d(n) represents the primary noise signal, and e(n) represents the error signal (i.e., the residual signal after the primary noise signal and the anti-noise signal a(n) cancel each other out). NMSE is usually a negative number; lower values indicate better noise reduction.
[0063] Results: Three untrained noises were tested in a linear system and two nonlinear systems. For 200Hz sinusoidal noise, mixed sinusoidal noise 1, and transformer noise 1, the step sizes for updating FxLMS were 0.00004, 0.00006, and 0.0007, respectively. To verify the superiority of the proposed method, it was compared with current classic active noise reduction methods, including the FxLMS algorithm, Wave-U-net, and SEWUNet.
[0064] The NMSE (dB) data of the ANC system under different settings and untrained noise are shown in the following table:
[0065]
[0066] The table above shows the NMSE of the test noise. From the table above, the NMSE value is far less than 0. It can be seen that under different nonlinearities and untrained noise scenarios, the four methods can effectively attenuate noise, but their noise reduction performance varies. 2 =∞), the GWUNet method achieves NMSE values of -25.27dB and -17.89dB for 200Hz sinusoidal signal and mixed sinusoidal signal 1, respectively, under the MSE loss function, which are approximately 0.43dB and 0.32dB higher than the suboptimal model Wave-U-Net. After adopting the MAE loss function, the noise reduction performance of the three deep learning-based active noise reduction methods is improved, with the GWUNet method achieving the lowest NMSE value. In the severe nonlinear scenario (η 2 =0.1), GWUNet achieved an NMSE of -9.62dB for transformer noise 1, representing improvements of 0.72dB and 3.70dB compared to Wave-U-Net and SEWUNet, respectively, and a 3.19dB improvement over the traditional FxLMS algorithm, demonstrating its strong nonlinear processing capabilities. Overall, GWUNet maintained stable performance across various noise types and speaker nonlinearities, demonstrating a particularly strong noise reduction advantage when using the MAE loss function.
[0067] like Figure 3 As shown in Figure 3, the noise reduction effects of different methods on splicing noise are demonstrated when the speaker nonlinearity is 0.5. Figure 3 In the time domain variation diagram, the blue curve represents the signal before ANC control, and the red curve represents the signal after control. It can be seen from the figure that the GWUNet method can quickly adapt to the sudden change of noise type and achieve a noise reduction of 16.56dB, while the FxLMS algorithm requires a longer adaptation time and only reaches 10.67dB. Figure 3 The power density curve further shows that GWUNet has better noise suppression capability in a wider frequency band.
[0068] like Figure 4 As shown in Figure 2, the test results of mixed noise under the same nonlinear conditions are presented. Figure 4 The time domain analysis shows that GWUNet has a faster convergence speed; Figure 4 The power spectrum curve shows that this method achieves a noise reduction effect of 16.65dB across a wide frequency band, significantly outperforming the 9.59dB achieved by the FxLMS algorithm. Experimental results demonstrate that GWUNet can effectively cope with different types of noise variations, demonstrating superior noise reduction performance and robustness.
[0069] The above generally describes the present invention in detail. However, it is obvious to those skilled in the art that modifications or improvements may be made based on the present invention. Therefore, modifications or improvements that do not depart from the spirit of the present invention are within the scope of protection of the present invention.
Claims
1. An active noise reduction method based on an improved Wave-U-Net model, characterized by: The steps include: S1. Dataset setup: Construct a low-frequency mixed sinusoidal noise dataset based on the characteristics of transformer noise; S2, simulation path construction: both the primary path and the secondary path are generated by the mirror source method; S3, nonlinear compensation: the output is subjected to the SEF function to simulate the speaker saturation characteristics; S4. Model training and evaluation: The GWUNet model is trained using the constructed low-frequency mixed sinusoidal noise dataset, and the model is evaluated using untrained nonlinearity and untrained noise.
2. The active noise reduction method based on the improved Wave-U-Net model according to claim 1, characterized in that: The S1 includes: constructing a low-frequency mixed sinusoidal noise data set based on the characteristics of transformer noise. In order to more realistically simulate the noise conditions in the industrial environment, Gaussian white noise with different signal-to-noise ratios can be randomly added to the pure sinusoidal signal data set; The signal-to-noise ratio range is set to [5, 10, 15, 20] dB, which can simulate noise interference of different intensities. The signal-to-noise ratio is an important indicator to measure the signal strength relative to the background noise. The unit is decibel (dB), and its mathematical expression is: Where, P signal Represents the power of the signal, P noise It represents the power of the noise. As can be seen from the formula, the smaller the SNR value, the greater the proportion of noise in the signal.
3. The active noise reduction method based on the improved Wave-U-Net model according to claim 1, characterized in that: The S2 includes: the path and the secondary path are both generated by the mirror source method; the mirror source method is used to set a room with a length, width and height of 3 meters, 4 meters and 2 meters respectively, and the reverberation time (T60) is 0.2s; during the training process, the reference microphone (I), the secondary speaker (J) and the error microphone (K) are set to a 1×1×1 pattern, the coordinates of the reference microphone are (1.5, 1, 1) meters, the coordinates of the secondary speaker are (1.5, 2.5, 1) meters, and the coordinates of the error microphone are (1.5, 3, 1), and the main path and secondary path are generated based on this configuration.
4. The active noise reduction method based on the improved Wave-U-Net model according to claim 1, characterized in that: The S3 includes: In the study of loudspeakers in nonlinear ANC (NANC) systems, nonlinearity is usually represented by a single saturation model, where the limiting effect can be represented by a nonlinear function, namely the scaled error function (SEF). The formula of SEF is: Where y represents the cancellation signal of the input speaker, η 2 Indicates the degree of nonlinearity. The smaller the value, the higher the degree of nonlinearity.
5. The active noise reduction method based on the improved Wave-U-Net model according to claim 1, characterized in that: The S4 includes: during the model training process, using the AMSGard optimizer to update the network weights; the initial learning rate of the training is 0.001, and then the learning rate is dynamically adjusted through an exponentially decreasing regulator with a learning rate attenuation coefficient of 0.98; the batch size is set to 60, and the total number of iterations is 100.
6. The active noise reduction method based on the improved Wave-U-Net model according to claim 1, characterized in that: The gated convolution layer in S1 performs data processing, and the gated convolution layer includes: a convolution layer, a BN layer and a gated linear unit; The improved gated convolutional layer uses one-dimensional convolution to process the transformer noise signal and extract its time and frequency domain features. The batch normalization layer is connected after the convolutional layer to ensure the consistency of the input data distribution of each layer, thereby accelerating training and improving model stability. The gated linear unit is a core component that dynamically adjusts the information flow and determines the probability of features being passed to the next layer. The improved gated convolutional layer uses a dual-branch structure to process input data. Among them, branch A is used to calculate the threshold value, and branch B is used to calculate the characteristic value; In branch A, data is processed sequentially through the convolutional layer and the BN layer, and a threshold value is generated using the Sigmoid activation function. This value reflects the importance of the feature, with more significant features corresponding to higher values. In branch B, data is also processed through the convolutional layer and the BN layer, and the eigenvalue is directly output. By multiplying the threshold value with the eigenvalue, the system filters out important information and suppresses irrelevant or redundant data. The calculation formula is as follows: Where BN(·) represents batch normalization; W and V are the weight matrices of the convolution kernels in the two branches, respectively; b and c are the bias terms in the two branches, respectively; * represents the convolution operation; σ represents the sigmoid activation function.
7. The active noise reduction method based on the improved Wave-U-Net model according to claim 1, characterized in that: The GWUNet model uses improved gated convolution as the basic building block of the network model, which can extract features more effectively. The input and output of the model are both single-channel waveform signals, and the single-channel waveform signal is output after calculation by the model; In this model, skip connections are used to transmit the input of the encoding module to the decoding module at the same level, ensuring that the features extracted by the encoder are fully utilized. This approach effectively prevents input information from being lost during transmission, thereby optimizing the propagation of model information and gradients. The LSTM modules embedded in the middle layer can effectively capture the long-term dependencies of acoustic signals, enhancing the model's ability to model temporal features. Except for the output layer and the middle layer, all modules use the Leaky ReLU activation function to introduce nonlinear characteristics, and a batch normalization layer is added after convolution or deconvolution to accelerate training and improve stability. The output layer uses convolution combined with Tanh activation function to generate the output signal.