Audio noise reduction method and system based on local attention and feature fusion

Through the audio denoising method of local attention and feature fusion, using multi-level stacked denoising units and local attention modules, the shortcomings of existing audio denoising methods in noise adaptability and computational efficiency are solved, and efficient and real-time audio denoising effects are achieved.

CN120656471APending Publication Date: 2025-09-16FUJIAN NEWLAND SOFTWARE ENGINEERING CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510698622.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing audio noise reduction methods have deficiencies in processing the temporal correlation of noise and computational efficiency, resulting in poor adaptability in sudden noise scenarios, excessive consumption of computing resources, and inability to meet real-time processing requirements.

Method used

An audio denoising method based on local attention and feature fusion is adopted. An audio denoising model is constructed through an encoder, a denoising module and a decoder. Multi-level stacked denoising units and local attention modules are used for audio denoising. Combined with one-dimensional depthwise separable convolution and sinusoidal position encoding, progressive noise elimination and lightweight operation are achieved.

Benefits of technology

It improves the quality, generalization, and efficiency of audio noise reduction, meets real-time processing requirements, is suitable for deployment on embedded devices, adapts to complex noise scenarios, reduces computational complexity, and maintains the ability to capture noise features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120656471A_ABST
    Figure CN120656471A_ABST
Patent Text Reader

Abstract

The invention provides an audio noise reduction method and system based on local attention and feature fusion in the technical field of crossing of audio signal processing and artificial intelligence, and the method comprises the steps: S1, building an audio noise reduction model based on an encoder, a noise reduction module and a decoder, and setting a loss function of the audio noise reduction model; s2, acquiring a large amount of historical audio data, preprocessing each piece of historical audio data, and labeling each piece of preprocessed historical audio data at least including clean audio, noise type, noise intensity, audio scene and audio content to construct a data set; s3, dividing the data set into a training set, a verification set and a test set to train, verify and test the audio noise reduction model; and S4, deploying the audio noise reduction model passing the test, and executing audio noise reduction operation through the deployed audio noise reduction model. The method has the advantages that the quality, generalization and efficiency of audio noise reduction are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Claims

1. An audio noise reduction method based on local attention and feature fusion, characterized by: The steps include: Step S1: creating an audio noise reduction model based on the encoder, the denoising module, and the decoder, and setting a loss function of the audio noise reduction model; The encoder is used to perform high-dimensional encoding on the input audio data to obtain a high-dimensional encoded feature vector; the denoising module is used to perform multi-level noise reduction operations on the high-dimensional encoded feature vector to obtain a denoised feature vector; the decoder is used to restore the denoised feature vector to a noise-free frequency. Step S2: Acquire a large amount of historical audio data, preprocess each of the historical audio data, and annotate each of the preprocessed historical audio data with at least clean audio, noise type, noise intensity, audio scene, and audio content to construct a dataset; Step S3: dividing the data set into a training set, a validation set, and a test set, and respectively training, validating, and testing the audio noise reduction model using the training set, validation set, and test set; Step S4: deploy the audio noise reduction model that has passed the test, and perform an audio noise reduction operation through the deployed audio noise reduction model.

2. The audio noise reduction method based on local attention and feature fusion according to claim 1, characterized in that: In step S1, the encoder is constructed based on one-dimensional convolution and a first nonlinear module, and the formula is: F en =ReLU1(Conv1D(A noisy )); Among them, F en Represents a high-dimensional encoding feature vector; ReLU1() represents the ReLU activation function, i.e., the first nonlinear module; Conv1D() represents a one-dimensional convolution; A noisy Represents the input audio data; The denoising module is constructed based on a multi-level stacked denoising unit, each of which is constructed based on a first-layer normalization module, a sinusoidal position encoding module, a point-by-point convolution, a first one-dimensional depth-separable convolution, a second nonlinear module, a local attention module, and a third nonlinear module; The first-layer normalization module and the sinusoidal position encoding module are used to regularize the input high-dimensional encoding feature vector and add audio position information to obtain a position-enhanced regularized feature vector: F pos =SPE(LayeNorm1(F en ))+LayeNorm1(F en ); Among them, F pos Represents the regularized feature vector; SPE() represents the sinusoidal position encoding module; LayeNorm1() represents the first layer normalization module; The point-by-point convolution, the first one-dimensional depth-separable convolution, and the second nonlinear module are used to perform a convolution operation on the regularized feature vector to obtain a depth feature vector: F depth =ReLU2(Dw_Conv1(Pointwise_Conv(F pos ))+Pointwise_Conv(F pos )); Among them, F depth Represents the depth feature vector; ReLU2() represents the ReLU activation function, that is, the second nonlinear module; Dw_Conv1() represents the first one-dimensional depth-separable convolution; Pointwise_Conv() represents point-by-point convolution; The local attention module and the third nonlinear module are used to perform feature mining on the deep feature vector to obtain a denoised feature vector; The decoder is built based on one-dimensional deconvolution to restore the denoised feature vector to the noise-free frequency: A clean =Transposed_Conv1 D(F de,R ); Among them, A clean Indicates noise-free frequency; Transposed_Conv1D() indicates one-dimensional deconvolution; F de,R represents the denoised feature vector.

3. The audio noise reduction method based on local attention and feature fusion according to claim 2, characterized in that: The local attention module is constructed based on a second normalization module, a second one-dimensional depth-separable convolution, a fourth nonlinear module, a first linear layer, a second linear layer, and a third linear layer; The second layer normalization module, the second one-dimensional depth-separable convolution and the fourth nonlinear module are used to extract local features from the depth feature vector: F middle =ReLU4(Dw_Conv2(LayeNorm2(F depth ))+LayeNorm2(F depth )); Among them, F middle Represents local features; ReLU4() represents the ReLU activation function, that is, the fourth nonlinear module; Dw_Conv2() represents the second one-dimensional depth-separable convolution; LayeNorm2() represents the second layer normalization module; The local features are divided into H non-overlapping feature segments of size P. Each of the feature segments is converted into a query vector Q, a key vector K, and a value vector V through the first linear layer, the second linear layer, and the third linear layer. Local attention is calculated based on the vectors Q, K, and V, and a local attention set is constructed based on each of the local attentions: Among them, Linearlayer1() represents the first linear layer; Linearlayer2() represents the second linear layer; Linearlayer3() represents the third linear layer; F middle,h Indicates the hth feature segment; F locaI,h represents the local attention of the hth feature segment; softmax() represents the normalized exponential function; T represents transposition; F local,H Indicates the local attention of the Hth feature segment; F local Represents a local attention set; The third nonlinear module is used to calculate the local attention set to generate a denoising mask, multiply the denoising mask by the input audio data to obtain a denoising feature subvector, and perform a multi-level denoising operation on each of the denoising feature subvectors to obtain a denoising feature vector: Among them, F de,1 Represents the first denoising feature subvector; ReLU3() represents the ReLU activation function, that is, the third nonlinear module; represents the first-level local attention set; Indicates element-by-element multiplication; F de,R Represents the denoised feature vector after repeated R denoising; represents the local attention set of level R; F de,R-1 Represents the denoised feature vector after repeating denoising for R-1 times.

4. The audio noise reduction method based on local attention and feature fusion according to claim 1, characterized in that: The step S2 is specifically as follows: Acquire a large amount of historical audio data, perform preprocessing on each of the historical audio data, including at least format conversion, signal cropping, resampling, normalization, audio segmentation, and feature extraction, annotate each of the preprocessed historical audio data with at least clean audio, noise type, noise intensity, audio scene, and audio content, and construct a dataset based on the annotated historical audio data.

5. The audio noise reduction method based on local attention and feature fusion according to claim 1, characterized in that: The step S3 is specifically as follows: Based on the time series, the dataset is divided into a training set, a validation set, and a test set in a ratio of 8:1:1, and the audio noise reduction model is trained using the training set until the loss value of the loss function is less than a preset loss threshold; The hyperparameters of the trained audio noise reduction model are adjusted using the validation set, and the performance indicators of the audio noise reduction model are verified. If the verification fails, the training set is expanded to continue training. If the verification passes, then: The PSNR value is calculated using the test set to test the audio noise reduction model that has passed the verification. If the test fails, the training set is expanded to continue training; if the test passes, the training is terminated.

6. An audio noise reduction system based on local attention and feature fusion, characterized by: Includes the following modules: An audio noise reduction model creation module is used to create an audio noise reduction model based on the encoder, the denoising module and the decoder, and set a loss function of the audio noise reduction model; The encoder is used to perform high-dimensional encoding on the input audio data to obtain a high-dimensional encoded feature vector; the denoising module is used to perform multi-level noise reduction operations on the high-dimensional encoded feature vector to obtain a denoised feature vector; the decoder is used to restore the denoised feature vector to a noise-free frequency. A data set construction module is used to obtain a large amount of historical audio data, preprocess each of the historical audio data, and annotate each of the preprocessed historical audio data with at least clean audio, noise type, noise intensity, audio scene, and audio content to construct a data set; An audio noise reduction model training module is used to divide the data set into a training set, a validation set, and a test set, and to train, validate, and test the audio noise reduction model using the training set, validation set, and test set, respectively; The audio noise reduction module is used to deploy the audio noise reduction model that has passed the test and perform audio noise reduction operations through the deployed audio noise reduction model.

7. The audio noise reduction system based on local attention and feature fusion according to claim 6, characterized in that: In the audio noise reduction model creation module, the encoder is constructed based on one-dimensional convolution and the first nonlinear module, and the formula is: F en =ReLU1(Conv1D(A noisy )); Among them, F en Represents a high-dimensional encoding feature vector; ReLU1() represents the ReLU activation function, i.e., the first nonlinear module; Conv1D() represents a one-dimensional convolution; A noisy Represents the input audio data; The denoising module is constructed based on a multi-level stacked denoising unit, each of which is constructed based on a first-layer normalization module, a sinusoidal position encoding module, a point-by-point convolution, a first one-dimensional depth-separable convolution, a second nonlinear module, a local attention module, and a third nonlinear module; The first-layer normalization module and the sinusoidal position encoding module are used to regularize the input high-dimensional encoding feature vector and add audio position information to obtain a position-enhanced regularized feature vector: F pos =SPE(LayeNorm1(F en ))+LayeNorm1(F en ); Among them, F pos Represents the regularized feature vector; SPE() represents the sinusoidal position encoding module; LayeNorm1() represents the first layer normalization module; The point-by-point convolution, the first one-dimensional depth-separable convolution, and the second nonlinear module are used to perform a convolution operation on the regularized feature vector to obtain a depth feature vector: F depth =ReLU2(Dw_Conv1(Pointwise_Conv(F pos ))+Pointwise_Conv(F pos )); Among them, F depth Represents the depth feature vector; ReLU2() represents the ReLU activation function, that is, the second nonlinear module; Dw_Conv1() represents the first one-dimensional depth-separable convolution; Pointwise_Conv() represents point-by-point convolution; The local attention module and the third nonlinear module are used to perform feature mining on the deep feature vector to obtain a denoised feature vector; The decoder is built based on one-dimensional deconvolution to restore the denoised feature vector to the noise-free frequency: A clean =Transposed_Conv1 D(F de,R ); Among them, A clean Indicates noise-free frequency; Transposed_Conv1D() indicates one-dimensional deconvolution; F de,R represents the denoised feature vector.

8. The audio noise reduction system based on local attention and feature fusion according to claim 7, characterized in that: The local attention module is constructed based on a second normalization module, a second one-dimensional depth-separable convolution, a fourth nonlinear module, a first linear layer, a second linear layer, and a third linear layer; The second layer normalization module, the second one-dimensional depth-separable convolution and the fourth nonlinear module are used to extract local features from the depth feature vector: F middle =ReLU4(Dw_Conv2(LayeNorm2(F depth ))+LayeNorm2(F depth )); Among them, F middle Represents local features; ReLU4() represents the ReLU activation function, that is, the fourth nonlinear module; Dw_Conv2() represents the second one-dimensional depth-separable convolution; LayeNorm2() represents the second layer normalization module; The local features are divided into H non-overlapping feature segments of size P. Each of the feature segments is converted into a query vector Q, a key vector K, and a value vector V through the first linear layer, the second linear layer, and the third linear layer. Local attention is calculated based on the vectors Q, K, and V, and a local attention set is constructed based on each of the local attentions: Among them, Linearlayer1() represents the first linear layer; Linearlayer2() represents the second linear layer; Linearlayer3() represents the third linear layer; F middle,h Indicates the hth feature segment; F local,h represents the local attention of the hth feature segment; softmax() represents the normalized exponential function; T represents transposition; F local,H Indicates the local attention of the Hth feature segment; F local Represents a local attention set; The third nonlinear module is used to calculate the local attention set to generate a denoising mask, multiply the denoising mask by the input audio data to obtain a denoising feature subvector, and perform a multi-level denoising operation on each of the denoising feature subvectors to obtain a denoising feature vector: Among them, F de,1 Represents the first denoising feature subvector; ReLU3() represents the ReLU activation function, that is, the third nonlinear module; represents the first-level local attention set; Indicates element-by-element multiplication; F de,R Represents the denoised feature vector after repeated R denoising; represents the local attention set of level R; F de,R-1 Represents the denoised feature vector after repeating denoising for R-1 times.

9. The audio noise reduction system based on local attention and feature fusion according to claim 6, characterized in that: The dataset construction module is specifically used for: Acquire a large amount of historical audio data, perform preprocessing on each of the historical audio data, including at least format conversion, signal cropping, resampling, normalization, audio segmentation, and feature extraction, annotate each of the preprocessed historical audio data with at least clean audio, noise type, noise intensity, audio scene, and audio content, and construct a dataset based on the annotated historical audio data.

10. The audio noise reduction system based on local attention and feature fusion according to claim 6, characterized in that: The audio noise reduction model training module is specifically used to: Based on the time series, the dataset is divided into a training set, a validation set, and a test set in a ratio of 8:1:1, and the audio noise reduction model is trained using the training set until the loss value of the loss function is less than a preset loss threshold; The hyperparameters of the trained audio noise reduction model are adjusted using the validation set, and the performance indicators of the audio noise reduction model are verified. If the verification fails, the training set is expanded to continue training. If the verification passes, then: The PSNR value is calculated using the test set to test the audio noise reduction model that has passed the verification. If the test fails, the training set is expanded to continue training; if the test passes, the training is terminated.

Citation Information

Cited By

  • Lightweight robust unsupervised feature selection audio denoising method based on algorithm

    CN121331150A