Audio noise reduction method, device and system based on deep learning
Through the deep learning-based audio noise reduction method, the combination technology of multi-scale time-frequency decomposition and dynamic core generation networks is used, combined with dual-path processing and differentiable acoustic equation constraint adversarial training, the problems of incomplete noise feature capture and insufficient adaptive adjustment capabilities in the existing technology are solved, and more efficient and high-quality audio noise reduction effect is achieved.
Patent Information
- Application Number
- CN202510478994.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing audio noise reduction methods are difficult to fully capture the characteristic distribution of different noise types on multiple scales, and lack the ability to adaptively adjust the dynamic characteristics of noise, resulting in limited noise reduction effect.
The audio noise reduction method based on deep learning is adopted to obtain the mixed time-frequency characteristics and noise fingerprint maps through multi-scale time-frequency decomposition, and the preset dynamic core generation network is used for parameter parallel processing, combining the dual-path processing structure and dynamic time-frequency domain cross-fusion, and finally perform differentiable acoustic equation constraint adversarial training.
It significantly improves the overall efficiency and effect of audio signal processing, can more effectively deal with complex and changing noise environments, ensure the acoustic consistency of reconstructed audio, and reduce distortion or artifacts in noise reduction results.
Smart Images

Figure CN120148537A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of audio noise reduction, and particularly relates to an audio noise reduction method, device and system based on deep learning. Background Art
[0002] As an important research direction in the field of acoustic signal processing, audio noise reduction technology has broad application prospects in fields such as communication systems, speech recognition, medical auscultation, and multimedia applications. The current research hotspots focus on how to more accurately capture noise characteristics, improve the time-frequency domain representation ability of audio signals, and efficiently remove noise components while retaining useful information. There are several key problems in existing audio noise reduction methods. First, traditional noise reduction algorithms are usually based on single-time domain or frequency domain analysis, making it difficult to comprehensively capture the characteristic distributions of different noise types at multiple scales, resulting in limited noise reduction effects. Most existing methods use static convolutional kernel structures when processing noise characteristics, lacking the ability to adaptively adjust to the dynamic characteristics of noise and being unable to effectively handle complex and changing noise environments. Existing technologies often process amplitude information and phase information separately, ignoring the coupling relationship between the two, resulting in insufficient acoustic consistency in the reconstructed audio. Summary of the Invention
[0003] The main object of the present invention is to provide an audio noise reduction method, device and system based on deep learning, which can effectively improve the overall efficiency and effect of audio signal processing.
[0004] To achieve the above object, the present invention provides an audio noise reduction method based on deep learning, including: Obtaining an input noisy audio signal and performing multi-scale time-frequency decomposition to obtain mixed time-frequency features and a noise fingerprint map; Performing parameter parallel processing on the noise fingerprint map through a preset dynamic kernel generation network, and performing preliminary noise reduction processing on the mixed time-frequency features to obtain noise-reduced mixed data; Constructing a dual-path processing structure for the noise-reduced mixed data to obtain amplitude-optimized data and phase-optimized data; Performing dynamic time-frequency domain cross-fusion on the amplitude-optimized data and the phase-optimized data to obtain fused audio data; Performing differentiable acoustic equation-constrained adversarial training on the fused audio data and performing inverse time-frequency transform processing to obtain a target noise-reduced audio signal.
[0005] Further, the obtaining an input noisy audio signal and performing multi-scale time-frequency decomposition to obtain mixed time-frequency features and a noise fingerprint map includes: Performing multi-band decomposition on the noisy audio signal to obtain a set of time-domain sequences; Performing spectral entropy analysis on the time domain sequence set to obtain a spectrum activity index; Performing adaptive threshold filtering according to the spectrum activity index to obtain a noise feature enhancement matrix; Extracting Mel-frequency cepstral coefficients from the noise feature enhancement matrix to obtain a noise feature vector; Performing self-encoding processing on the noise feature vector to obtain the noise fingerprint spectrum; Performing discrete cosine transform on the noise fingerprint spectrum to obtain time-varying spectrum features; The time-varying spectrum feature and the noise feature vector are subjected to tensor concatenation and normalization processing to obtain the mixed time-frequency feature.
[0006] Furthermore, the noise fingerprint spectrum is processed in parallel by using a preset dynamic kernel generation network, and the mixed time-frequency features are subjected to preliminary noise reduction processing to obtain noise-reduced mixed data, including: Extracting the noise level of the noise fingerprint to obtain a time domain feature sequence and a frequency domain feature sequence; Performing nonlinear transformation on the time domain feature sequence to obtain the time domain dynamic convolution kernel; Performing frequency band filtering on the frequency domain feature sequence to obtain the frequency domain dynamic filter bank; Performing noise suppression mask analysis on the mixed time-frequency features according to the time-domain dynamic convolution kernel and the frequency-domain dynamic filter bank to obtain a refined mask matrix; The mixed time-frequency features are subjected to denoising according to the refined mask matrix to obtain the denoised mixed data.
[0007] Furthermore, the noise suppression mask analysis is performed on the mixed time-frequency features according to the time-domain dynamic convolution kernel and the frequency-domain dynamic filter bank to obtain a refined mask matrix, including: Performing a time-domain convolution operation on the mixed time-frequency feature according to the time-domain dynamic convolution kernel to obtain a time-domain suppression feature; Performing frequency domain filtering operation on the mixed time-frequency features according to the frequency domain dynamic filter group to obtain frequency domain suppression features; Performing feature fusion on the time domain suppression feature and the frequency domain suppression feature to obtain an initial mask feature map; Performing a cross attention operation on the initial mask feature map and the mixed time-frequency feature to obtain a context-aware feature; Performing nonlinear activation mapping on the context-aware features to obtain an activation feature matrix; Performing bidirectional gated recursive processing on the activation feature matrix to obtain time-frequency correlation features; Perform a sparse constraint transformation on the time-frequency correlation feature to obtain a mask coefficient vector; Perform a frequency band adaptive expansion on the mask coefficient vector to obtain an expanded mask matrix; Perform an edge smoothing process on the expanded mask matrix to obtain the refined mask matrix.
[0008] Furthermore, construct a dual-path processing structure for the noise-reduced mixed data to obtain amplitude-optimized data and phase-optimized data Separate the noise-reduced mixed data in the complex domain to obtain an amplitude component and a phase component; Perform a multi-layer residual mapping process on the amplitude component to obtain an amplitude feature map; Perform an adaptive threshold constraint process on the amplitude feature map to obtain the amplitude-optimized data; Extract the phase component in the complex domain to obtain a phase feature representation; Perform a cyclic consistency constraint process on the phase feature representation to obtain preliminary phase correction data; Extract complementary information from the preliminary phase correction data and the amplitude-optimized data to obtain complementary phase features; Perform self-calibration adjustment on the preliminary phase correction data according to the complementary phase features to obtain the phase-optimized data.
[0009] Furthermore, the dynamic time-frequency domain cross-fusion of the amplitude-optimized data and the phase-optimized data to obtain fused audio data includes: Calculate a weight matrix for the amplitude-optimized data to obtain an amplitude weight feature map; Perform a phase transformation on the phase-optimized data to obtain a phase enhancement feature; Perform an interaction construction according to the amplitude weight feature map and the phase enhancement feature to obtain an interaction weight matrix; Perform a multi-head attention decomposition on the interaction weight matrix to obtain time-frequency attention coefficients; Perform a re-weighting process on the amplitude-optimized data and the phase-optimized data according to the time-frequency attention coefficients to obtain time-frequency enhancement features; Perform a residual connection on the time-frequency enhancement features to obtain a context-dependent representation; Perform a complex domain reconstruction and amplitude-phase consistency optimization on the context-dependent representation to obtain time-frequency alignment fusion features; Perform a segmented overlapping process on the time-frequency alignment fusion features to obtain fused audio data.
[0010] Further, performing differentiable acoustic equation-constrained adversarial training on the fused audio data and performing inverse time-frequency transformation processing to obtain a target noise-reduced audio signal includes: Performing differentiable constraint construction on the fused audio data according to a preset acoustic propagation equation to obtain a differentiable constraint structure; Constructing a gradient penalty term for the differentiable constraint structure based on a preset clear audio reference sample to obtain a spectral distance loss function; Evaluating the authenticity of the fused audio data according to the spectral distance loss function to obtain a discrimination score; Performing waveform reconstruction on the fused audio data according to the discrimination score and the spectral distance loss function to obtain corrected audio data; Performing residual noise component reduction processing on the corrected audio data to obtain a pure time-frequency representation; Performing segmented inverse short-time Fourier transform on the pure time-frequency representation and performing time-domain signal synthesis to obtain an initial target audio; Performing artifact filtering on the initial target audio to obtain a target noise-reduced audio signal.
[0011] Further, performing differentiable constraint construction on the fused audio data according to a preset acoustic propagation equation to obtain a differentiable constraint structure includes: Extracting acoustic propagation characteristic parameters from the fused audio data to obtain an acoustic characteristic matrix; Constructing an expression for the acoustic characteristic matrix according to the acoustic propagation equation to obtain an acoustic field propagation expression; Performing discretization processing on the acoustic field propagation expression to obtain a grid acoustic field distribution map; Constructing a differentiable boundary condition according to the grid acoustic field distribution map to obtain an acoustic boundary constraint expression; Calculating the gradient of the acoustic boundary constraint expression to obtain an acoustic gradient field; Performing tensor fusion on the acoustic gradient field and the acoustic characteristic matrix to obtain an initial constraint structure; Performing operator constraint on the initial constraint structure to obtain a second-order partial differential constraint; Calculating constraint parameters for the second-order partial differential constraint to obtain dynamic constraint parameters; Modulating the initial constraint structure according to the dynamic constraint parameters to obtain an enhanced constraint structure; Performing Taylor series expansion on the enhanced constraint structure to obtain a linear approximation representation; Performing constraint construction on the fused audio data according to the linear approximation representation and the enhanced constraint structure to obtain a differentiable constraint structure.
[0012] The present invention also provides an audio noise reduction device based on deep learning, which is applied to the audio noise reduction method based on deep learning described in any one of the above, and includes: An acquisition module, which is used to acquire the input noisy audio signal and perform multi-scale time-frequency decomposition to obtain mixed time-frequency features and a noise fingerprint map; An analysis module, which is used to perform parameter parallel processing on the noise fingerprint map through a preset dynamic kernel generation network, and perform preliminary noise reduction processing on the mixed time-frequency features to obtain noise-reduced mixed data; An association module, which is used to construct a dual-path processing structure for the noise-reduced mixed data to obtain amplitude-optimized data and phase-optimized data; A processing module, which is used to perform dynamic time-frequency domain cross-fusion on the amplitude-optimized data and the phase-optimized data to obtain fused audio data; A control module, which is used to perform differentiable acoustic equation-constrained adversarial training on the fused audio data and perform inverse time-frequency transform processing to obtain a target noise-reduced audio signal.
[0013] The present invention also provides an audio noise reduction system based on deep learning, including: A memory, which is used to store programs; A processor, which is used to execute the programs to implement each step of the audio noise reduction method based on deep learning described in any one of the above.
[0014] The audio noise reduction method, device and system provided by the present invention have the following beneficial effects: Through the multi-scale time-frequency decomposition technology, it can more comprehensively capture the characteristic distributions of different noise types at multiple scales, thereby improving the noise reduction effect and providing a more reliable basis for the high-quality restoration of audio signals. By using a preset dynamic kernel generation network to perform parameter parallel processing on the noise fingerprint map, it realizes the adaptive adjustment of the dynamic characteristics of the noise and can more effectively cope with the complex and changing noise environment. By respectively optimizing the amplitude and phase information through a dual-path processing structure and performing dynamic time-frequency domain cross-fusion, it ensures the acoustic consistency of the reconstructed audio and improves the overall quality of the audio signal. By using differentiable acoustic equation-constrained adversarial training and utilizing the basic laws of acoustic propagation, strict physical constraints are imposed on the fused audio data, reducing the distortion or artifact phenomenon in the noise reduction result and further improving the authenticity and reliability of the noise reduction effect. The present invention not only overcomes the limitations of traditional noise reduction methods, but also significantly improves the actual performance of audio noise reduction technology in high-noise environments and high-quality audio application scenarios, effectively improving the overall efficiency and effect of audio signal processing. Description of the Drawings
[0015] Figure 1 It is a flowchart of an audio noise reduction method based on deep learning provided by the present invention; Figure 2 It is a structural diagram of an audio noise reduction device based on deep learning provided by the present invention; Figure 3 It is a structural diagram of an audio noise reduction system based on deep learning provided by the present invention.
[0016] The realization of the purpose of the present invention, functional characteristics and advantages will be further described in conjunction with embodiments with reference to the accompanying drawings. Specific embodiments
[0017] In order to make the purpose, technical solution and advantages of the present invention clearer, the present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described here are only used to explain the present invention and are not used to limit the present invention.
[0018] Next, in conjunction with the accompanying drawings and specific embodiments, the present invention will be further described.
[0019] Referring to Figure 1 As shown, the present invention also provides an audio noise reduction method based on deep learning, including: Step S1: Obtain the input noisy audio signal and perform multi-scale time-frequency decomposition to obtain the mixed time-frequency features and noise fingerprint map; Step S2: Perform parameter parallel processing on the noise fingerprint map through a preset dynamic kernel generation network, and perform preliminary noise reduction processing on the mixed time-frequency features to obtain noise-reduced mixed data; Step S3: Construct a dual-path processing structure for the noise-reduced mixed data to obtain amplitude-optimized data and phase-optimized data; Step S4: Perform dynamic time-frequency domain cross-fusion on the amplitude-optimized data and phase-optimized data to obtain fused audio data; Step S5: Perform differentiable acoustic equation constrained adversarial training on the fused audio data and perform inverse time-frequency transform processing to obtain the target noise-reduced audio signal.
[0020] Based on the above steps, the detailed step process is as follows: Step S1: Obtain the input noisy audio signal and perform multi-scale time-frequency decomposition on it. Audio signals are usually continuous, but noise makes the signals complex and it is impossible to directly extract effective audio features. Therefore, it is necessary to transform the signal into a time-frequency domain representation through time-frequency decomposition methods. Time-frequency decomposition can help capture the frequency and time-domain features of the signal, especially in the presence of noise, so that noise and useful audio information can be clearly separated. Multi-scale time-frequency decomposition performs multi-level decomposition of the audio signal through different time windows and frequency bandwidths, so as to obtain time-frequency features at different scales. This method can better adapt to the diversity of audio signals, especially for noise signals with complex spectral characteristics.
[0021] Through time-frequency decomposition, two main outputs can be obtained: mixed time-frequency features and noise fingerprint maps. The mixed time-frequency features contain the composite information of the original audio signal and noise, which provides a basis for subsequent noise reduction processing. The noise fingerprint map is specifically used to represent the characteristics of the noise in the signal. Its construction process is to estimate the frequency distribution and its time-varying characteristics of the noise by analyzing the spectral characteristics in the signal. The accurate generation of the noise fingerprint map is the key to audio noise reduction, because it can provide an accurate identification basis of the noise for the subsequent noise reduction network, helping the model to more effectively distinguish the noise from the signal, so as to achieve precise noise reduction.
[0022] Step S2: Process the noise fingerprint map through a Dynamic Kernel Generation Network and perform preliminary noise reduction processing on the mixed time-frequency features. The Dynamic Kernel Generation Network dynamically generates noise reduction kernel functions that adapt to different noise types and noise levels by designing specific neural network models. These kernel functions can be adjusted in the time-frequency domain. According to the noise characteristics in the noise fingerprint map, corresponding noise reduction filters are generated to effectively reduce the interference of noise. The key advantage of the Dynamic Kernel Generation Network is that it can flexibly adjust the filtering parameters according to different noise environments, so as to adapt to different noise reduction requirements.
[0023] When performing preliminary noise reduction processing on the mixed time-frequency features, the main goal is to reduce the noise component in the mixed signal without damaging the useful audio signal. Through the filtering effect of the dynamic kernel, the noise and the signal can be optimized respectively in different frequency bands and time windows. This step usually adopts a parallel processing method, that is, the noise fingerprint map and the mixed time-frequency features are processed separately, so as to ensure the efficiency and accuracy of the noise reduction process. Finally, through this process, the preliminarily noise-reduced mixed data is obtained, laying a foundation for subsequent further processing and optimization.
[0024] Step S3: Construct a dual-path processing structure for the denoised mixed data to obtain amplitude-optimized data and phase-optimized data. The dual-path processing structure is a deep learning model architecture where the input data is processed through two independent paths, each responsible for different audio features. Specifically, one path is responsible for optimizing the amplitude information of the signal, and the other path focuses on optimizing the phase information of the signal. The amplitude and phase of an audio signal are important components of its spectrum, and they play different roles in audio reconstruction. The amplitude reflects the energy distribution of the signal, while the phase determines the waveform of the signal.
[0025] The amplitude optimization path further removes the noise components in the signal, making the amplitude of the signal purer and reducing the interference of noise on the audio. The phase optimization path focuses on restoring the time-domain waveform of the signal to ensure that the timing characteristics of the audio signal are not damaged. This process is usually achieved through a specific network design to ensure effective cooperation between the two paths, which can not only retain the key features of the audio signal but also remove noise to the greatest extent. Finally, after dual-path processing, the obtained amplitude-optimized data and phase-optimized data can provide a more accurate basis for subsequent audio fusion and reconstruction, helping to achieve high-quality audio denoising effects.
[0026] Step S4: The amplitude-optimized data and phase-optimized data are integrated through dynamic time-frequency domain cross-fusion to finally obtain fused audio data. The core purpose of this process is to combine the amplitude and phase-optimized data into a complete audio signal. The amplitude and phase play different roles in an audio signal. The amplitude determines the intensity of the sound, while the phase determines the waveform shape of the sound. Both must be effectively fused during time-domain reconstruction. Through dynamic time-frequency domain cross-fusion, the amplitude and phase can be more finely combined in the time-frequency domain, enabling the audio signal to retain both the signal quality optimized by the amplitude and the time-domain characteristics optimized by the phase during the recovery process.
[0027] The cross-fusion process first needs to consider the matching problem between the amplitude and phase, that is, how to combine their features in the time-frequency domain to reduce information loss during the fusion process. Dynamic time-frequency domain cross-fusion adjusts the fusion strategy of the amplitude and phase by introducing dynamic weights, enabling it to make flexible adjustments according to different time and frequency conditions, ensuring that the fused signal has both high fidelity and can adapt to the audio characteristics of different frequency bands and time periods. This process may rely on a deep neural network for learning to optimize the parameter selection during the fusion process to achieve a balance between denoising and fidelity. Through this step, the finally generated fused audio data is significantly improved in quality, and the noise components are greatly reduced, preparing for the final audio reconstruction.
[0028] Step S5: Fusing audio data will first perform differentiable acoustic equation constrained adversarial training. The goal of this process is to further improve the noise reduction effect and ensure that the audio signal quality is restored to the maximum extent. The differentiable acoustic equation is a mathematical model used to simulate the propagation characteristics of audio signals in a physical environment. During the audio noise reduction process, applying this model can constrain the physical characteristics of the audio and ensure that the optimized audio signal conforms to the sound propagation characteristics in the actual environment. Through the method of adversarial training, the noise reduction network can consider the actual propagation characteristics of the audio during the optimization process, thereby avoiding over-noise reduction or unreasonable signal distortion.
[0029] Adversarial training is a commonly used deep learning technique. By introducing an adversarial network, the model continuously "fights" against noise during the training process, thereby enhancing its noise reduction ability. This process involves a game between the generator network and the discriminator network, ultimately making the generated noise reduction signal as similar as possible to the real audio signal while removing noise to the greatest extent. This step also includes inverse time-frequency transform processing, which converts time-frequency domain data back into a time-domain signal. The purpose of the inverse time-frequency transform is to reconstruct the optimized time-frequency data into an audio signal, enabling the output audio signal to be restored to nearly the original quality and effectively suppressing the noise component.
[0030] By combining adversarial training and inverse time-frequency transform, the finally obtained target noise reduction audio signal has high fidelity and clarity, the noise is effectively removed, and the sound quality is significantly improved.
[0031] Step S6: Output the target noise reduction audio signal. After a series of complex processing steps above, the noise reduction system has successfully converted the input noisy audio signal into a clear and noise-free audio signal. The output target noise reduction audio signal should be able to accurately restore the main features of the original audio signal while removing background noise and interference. This signal can be used in various audio applications such as speech recognition, audio playback, and voice communication.
[0032] In the output stage, further format conversion or compression processing is performed on the audio signal for playback or storage on different devices and platforms. The core goal of this step is to ensure that the noise-reduced audio signal meets the expectations in terms of quality and can meet the application requirements, providing a user-friendly audio experience. Finally, the system will generate the target noise reduction audio signal and output it in an appropriate format, providing the user with a clear and noise-free audio experience.
[0033] A method, device, and system for audio noise reduction based on deep learning provided by the present invention have the following beneficial effects: Through the multi-scale time-frequency decomposition technique, the characteristic distributions of different noise types at multiple scales can be captured more comprehensively, thus improving the noise reduction effect and providing a more reliable basis for the high-quality restoration of audio signals. By using the preset dynamic kernel generation network to perform parallel parameter processing on the noise fingerprint map, the adaptive adjustment of the dynamic characteristics of the noise is realized, and it can more effectively cope with the complex and changing noise environment. Through the dual-path processing structure, the amplitude and phase information are optimized respectively, and dynamic time-frequency domain cross-fusion is carried out to ensure the acoustic consistency of the reconstructed audio and improve the overall quality of the audio signal. Through differentiable acoustic equation-constrained adversarial training, using the basic laws of acoustic propagation, strict physical constraints are imposed on the fused audio data, reducing the distortion or artifact phenomena in the noise reduction results and further enhancing the authenticity and reliability of the noise reduction effect. The present invention not only overcomes the limitations of traditional noise reduction methods, but also significantly improves the actual performance of audio noise reduction technology in high-noise environments and high-quality audio application scenarios, effectively enhancing the overall efficiency and effect of audio signal processing.
[0034] In one embodiment, an input noisy audio signal is obtained and multi-scale time-frequency decomposition is performed to obtain mixed time-frequency features and a noise fingerprint map, including: A noisy audio signal refers to an original acquisition signal containing the target audio and background noise, usually represented as a one-dimensional array in the time domain.
[0035] Performing multi-band decomposition on the noisy audio signal is a process of dividing the noisy audio signal into multiple sub-bands in the frequency domain. This process is implemented using a filter bank, and each filter is responsible for extracting the signal components within a specific frequency range. In a specific implementation, wavelet packet transform is used to decompose the noisy audio signal, and the decomposition level is set to 4 layers, obtaining 16 frequency bands. For an audio signal with a sampling rate of 16 kHz, the first frequency band covers the range of 0 - 500 Hz, and the last frequency band covers the range of 7500 - 8000 Hz. The set of time-domain sequences obtained after decomposition contains 16 sub-signals, and each sub-signal represents the time-domain signal component corresponding to the corresponding frequency band.
[0036] A set of time-domain sequences refers to a group of time-domain signals obtained after multi-band decomposition, and each signal corresponds to the component of the original audio within a specific frequency band.
[0037] Performing spectral entropy analysis on a set of time-domain sequences is a process of performing short-time Fourier transform on each band signal and calculating its spectral entropy value. Spectral entropy measures the uncertainty of the signal spectrum distribution. The higher the value, the more uniform the spectrum distribution, usually corresponding to noise; the lower the value, the more concentrated the spectrum distribution, usually corresponding to the effective signal. During processing, each band signal is framed using a Hamming window with a frame length of 20 milliseconds and a frame shift of 10 milliseconds, and the short-time Fourier transform of each frame is calculated to obtain the spectrum. The spectral entropy of each frame is calculated, and its reciprocal is defined as the spectral activity index. Regions with a higher spectral activity index usually correspond to the effective signal, while regions with a lower index usually correspond to noise.
[0038] The spectral activity index is a quantitative index that describes the degree of spectral change of each band signal and is used to distinguish the target signal from the background noise.
[0039] Adaptive threshold filtering based on the spectral activity index is a process of enhancing or suppressing the signal based on the spectral activity index. In this process, an adaptive threshold is set. When the spectral activity exceeds the threshold, the corresponding spectral components are retained; otherwise, they are regarded as noise and suppressed. The adaptive threshold is dynamically adjusted according to the average activity within the sliding window, multiplied by a coefficient factor of 1.5, and the half-length of the sliding window is taken as 5 frames. By comparing the spectral activity with the adaptive threshold, a suppression factor is calculated, and this factor is multiplied by the original spectrum to obtain the noise feature enhancement matrix. This matrix highlights the noise components in the original signal while suppressing the effective signal components.
[0040] The noise feature enhancement matrix refers to the time-frequency representation obtained after adaptive threshold filtering, in which the noise components are enhanced while the effective signals are suppressed.
[0041] Performing Mel-frequency cepstral coefficient extraction on the noise feature enhancement matrix is to convert the noise feature enhancement matrix into a more compact and more in line with the human auditory characteristics feature representation. In this process, the Mel filter bank is applied to the noise feature enhancement matrix of each band to convert the linear frequency scale to the Mel scale, and a total of 24 Mel filters are set. After Mel filtering, logarithmic transformation and discrete cosine transformation are applied, and the first 13 coefficients are extracted as the Mel-frequency cepstral coefficient features. For each frame of each band, a 13-dimensional feature vector is obtained. The Mel-frequency cepstral coefficient features of all bands are concatenated together to form a 208-dimensional noise feature vector.
[0042] The noise feature vector is a compact feature representation obtained through Mel-frequency cepstral coefficient extraction and describes the acoustic characteristics of the noise.
[0043] The auto-encoding process of the noise feature vector uses a deep neural network to reduce the dimension and reconstruct the noise feature vector, extracting the essential features of the noise. This process adopts a convolutional auto-encoder structure, which consists of an encoder and a decoder. The encoder is composed of 3 convolutional layers, with the convolutional kernel sizes being 5×5, 3×3, and 3×3 respectively, the strides being 2 for all, and the number of channels being 32, 64, and 128 respectively; the decoder is composed of 3 transposed convolutional layers, with the parameter settings symmetric to those of the encoder. The noise feature vector is rearranged into a two-dimensional tensor and input into the auto-encoder. The feature map output by the encoder is the noise fingerprint spectrum, with a dimension of 8×13×128.
[0044] The noise fingerprint spectrum refers to the compact representation of the noise features extracted by the auto-encoder, which contains the key feature information of the noise.
[0045] Performing discrete cosine transform on the noise fingerprint spectrum is to further process the noise fingerprint spectrum and extract its main frequency components. This process applies one-dimensional discrete cosine transform to each frequency channel of the noise fingerprint spectrum along the time dimension, and extracts the low-frequency coefficients as the time-varying spectral features. For a time series of length T, the first T / 4 discrete cosine transform coefficients are retained, and the resulting time-varying spectral features have a dimension of 8×13×32.
[0046] The time-varying spectral features are the frequency-domain features obtained through discrete cosine transform, representing the main frequency components of the time-varying characteristics of the noise.
[0047] The process of tensor concatenation and normalization of the time-varying spectral features and the noise feature vector combines the time-varying spectral features and the noise feature vector to form a comprehensive representation. This process reshapes the noise feature vector into a 13×16 matrix, and then concatenates it with the time-varying spectral features in the channel dimension to obtain the concatenated features. Batch normalization is applied to the concatenated features to standardize each element to a distribution with a mean of 0 and a variance of 1, resulting in the final mixed time-frequency features.
[0048] The mixed time-frequency features are the final output of multi-scale time-frequency decomposition, which combines the time-varying spectral features and the spectral features of the noise, providing input for the subsequent deep learning noise reduction network. These features contain rich time-domain and frequency-domain information, can comprehensively characterize the characteristics of the noisy audio signal, and are beneficial for the subsequent neural network to accurately distinguish noise and valid signals.
[0049] In this embodiment, by performing multi-scale time-frequency decomposition on the noisy audio signal and using wavelet packet transform to achieve multi-band decomposition, it is possible to finely capture the noise characteristics in different frequency bands, improving the accuracy and adaptability of noise reduction. Combining spectral entropy analysis and adaptive threshold filtering technology, the system can dynamically identify the noise components in the signal, overcoming the limitation of fixed parameters in traditional noise reduction methods and having a stronger adaptability to different types of environmental noise. By using Mel-frequency cepstral coefficient extraction and auto-encoding processing, the high-dimensional noise features are compressed into low-dimensional noise fingerprint maps, which not only reduces the computational complexity but also retains the key feature information of the noise, providing a more compact and information-rich feature representation for the deep learning network. By further processing the noise fingerprint map through discrete cosine transform, the time-varying characteristics of the noise are effectively extracted, enhancing the system's ability to process non-stationary noise. The obtained hybrid time-frequency features integrate time-domain and frequency-domain information, providing a comprehensive feature representation for the deep learning noise reduction network, enabling the network to more accurately distinguish noise from valid signals and significantly improving the audio noise reduction effect, especially the noise reduction performance in low signal-to-noise ratio and complex noise environments.
[0050] In one embodiment, the noise fingerprint map is processed in parallel with parameters through a preset dynamic kernel generation network, and preliminary noise reduction processing is performed on the hybrid time-frequency features to obtain noise-reduced hybrid data, including: Noise level extraction refers to the process of hierarchical feature extraction of the noise fingerprint map. The noise fingerprint map is a time-frequency map representing the noise characteristics extracted from the original audio signal through methods such as short-time Fourier transform. In this step, a multi-level filter bank is used to decompose the noise fingerprint map into sub-bands in different frequency ranges. Adaptive threshold processing is applied to each sub-band to separate the noise components from the signal components. Through recursive hierarchical analysis, the distribution characteristics of the noise in the time domain and frequency domain are obtained. The time-domain feature sequence represents the characteristics of the noise changing with time, including information such as the time envelope and amplitude change of the noise; the frequency-domain feature sequence represents the distribution characteristics of the noise in the frequency dimension, including information such as the spectral shape and energy distribution of the noise. After processing this step, two sets of data, namely the time-domain feature sequence and the frequency-domain feature sequence, are obtained, providing a basis for subsequent dynamic kernel generation.
[0051] The construction of the time-domain dynamic convolution kernel refers to the process of performing non-linear transformation on the time-domain feature sequence to generate convolution kernel parameters with strong adaptability. This step uses a multi-layer perceptron to map the time-domain feature sequence and convert the feature vector into convolution kernel parameters. During this process, a residual connection structure is adopted to enhance feature transmission and ensure that low-level features are not lost. The importance of time-domain features is weighted through an attention mechanism to highlight the feature expression in key time periods. After non-linear transformation by the activation function, the features are mapped into an appropriate parameter space. The time-domain dynamic convolution kernel refers to the convolution kernel parameters dynamically generated according to the input noise features, whose shape and weight will be adaptively adjusted with the change of input features, having strong environmental adaptability. This convolution kernel processes the signal in the time domain and effectively suppresses time-varying noise. After this step of processing, a time-domain dynamic convolution kernel with strong adaptability is generated, providing a time-domain processing tool for subsequent noise suppression.
[0052] The construction of the frequency-domain dynamic filter bank refers to the process of performing band-pass filtering design on the frequency-domain feature sequence to generate a set of adaptive filters. This step uses the spectral energy distribution information in the frequency-domain feature sequence to design multiple sets of band-pass filters. The center frequency and bandwidth of the filters are dynamically adjusted according to the distribution of frequency-domain features to ensure targeted processing of noise in different frequency bands. A complex-valued neural network is used to map the frequency-domain features to generate the frequency response characteristics of the filters. The frequency-domain dynamic filter bank refers to a set of filter parameter sets dynamically generated according to the input frequency-domain features, where each filter is responsible for suppressing noise in a specific frequency band, and the whole forms a complete frequency-domain processing system. This filter bank processes the signal in the frequency domain and effectively suppresses frequency-varying noise. After this step of processing, a frequency-domain dynamic filter bank adapted to the noise characteristics of different frequency bands is generated, providing a frequency-domain processing tool for subsequent noise suppression.
[0053] The noise suppression mask analysis refers to the process of processing the mixed time-frequency features based on the generated time-domain dynamic convolution kernel and frequency-domain dynamic filter bank to generate a refined mask matrix. The mixed time-frequency features refer to the feature representation after the time-frequency transformation of the original audio signal containing the effective signal and noise. This step applies the time-domain dynamic convolution kernel to the time dimension of the mixed time-frequency features to suppress the noise in time; applies the frequency-domain dynamic filter bank to the frequency dimension of the mixed time-frequency features to suppress the noise in frequency. The processing results in the two dimensions are integrated through a bilinear fusion network to generate a joint representation. The joint representation is mapped to the [0, 1] interval through the activation function to form a refined mask matrix. The refined mask matrix refers to the estimation of the proportion of the effective signal component at each time-frequency point in the mixed time-frequency features. The value closer to 1 indicates that the noise component at this time-frequency point is less, and the value closer to 0 indicates that the noise component at this time-frequency point is more. After this step of processing, a refined mask matrix is obtained, providing accurate noise distribution information for the final noise reduction processing.
[0054] Noise reduction refers to the process of suppressing noise on mixed time-frequency features based on a refined mask matrix. This step multiplies the refined mask matrix by the mixed time-frequency features element by element to achieve noise suppression for each time-frequency point. Time-frequency points with mask values close to 1 basically retain the original features, while time-frequency points with mask values close to 0 are significantly suppressed. Phase retention technology is used to ensure that the phase information of the signal is not distorted during the noise reduction process. The processed time-frequency features are inversely transformed and converted back to time domain signals. The converted signal is post-processed to eliminate potential music noise and artifacts. Denoised mixed data refers to audio signal data that has been processed with noise suppression. Compared with the original mixed data, the noise component is significantly reduced, while the effective signal component is retained. After processing in this step, denoised mixed data is obtained, which achieves effective noise reduction of the original audio signal.
[0055] This embodiment uses a preset dynamic kernel generation network to perform parameter parallel processing on the noise fingerprint map, realizes the fine extraction and analysis of noise features, and significantly improves the accuracy of audio noise reduction. The dual-channel processing structure of the time domain dynamic convolution kernel and the frequency domain dynamic filter group enables the system to deal with time-varying noise and frequency-varying noise at the same time, greatly enhancing the environmental adaptability of the noise reduction method. The generation mechanism of the refined mask matrix ensures that the noise suppression of each time-frequency point is more accurate, and effectively avoids the problem of speech distortion caused by excessive noise reduction in traditional methods. The phase information of the signal is retained by the phase retention technology, which reduces the sound quality loss during the noise reduction process and improves the naturalness of the audio after noise reduction. This method does not need to be pre-trained for a specific noise environment, can adaptively adjust the processing parameters, is applicable to a variety of complex noise environments, and significantly improves the versatility and practical value of the system. Overall, this method has achieved significant improvements in noise reduction effect, signal fidelity and system adaptability.
[0056] In one embodiment, noise suppression mask analysis is performed on the mixed time-frequency features according to the time-domain dynamic convolution kernel and the frequency-domain dynamic filter bank to obtain a refined mask matrix, including: Perform a time-domain convolution operation on the mixed time-frequency features according to the time-domain dynamic convolution kernel to obtain time-domain suppression features. The time-domain dynamic convolution kernel is a convolution kernel that adaptively adjusts according to the characteristics of the input signal, and its parameters are learned by the neural network. In this step, the mixed time-frequency features are convolved with the dynamically generated convolution kernel along the time axis. The size of the convolution kernel is 5×5, the stride is 1, and zero-padding is used at the edges. The parameters of the convolution kernel are dynamically generated by the feed-forward network according to the input features, enabling the convolution operation to adaptively adjust to the characteristics of audio signals in different noise environments. The time-domain suppression features obtained after the convolution operation retain the time structure information of the original signal while suppressing the noise components in the time domain. The dimension of the time-domain suppression features is consistent with that of the input mixed time-frequency features in the time dimension, facilitating subsequent fusion with the frequency-domain features.
[0057] Perform a frequency-domain filtering operation on the mixed time-frequency features according to the frequency-domain dynamic filter bank to obtain frequency-domain suppression features. The frequency-domain dynamic filter bank is a set of filters that adaptively adjusts according to the frequency distribution characteristics of the input signal. In this step, the mixed time-frequency features are processed through multiple parallel band-pass filters in the frequency dimension. The center frequencies and bandwidths of the filters are dynamically determined by the neural network according to the input features. The frequency-domain filter bank contains 32 sub-filters, covering the audible range from 20 Hz to 20 kHz. The frequency response curves of the filters are in the form of Gaussian functions, and the filter coefficients are dynamically adjusted by the neural network according to the spectral characteristics of the input signal. The frequency-domain suppression features obtained after the frequency-domain filtering operation enhance the target speech components in the frequency domain while suppressing the noise components in the frequency domain. The frequency-domain suppression features retain the frequency structure characteristics of the original signal, providing frequency-domain information support for subsequent mask generation.
[0058] Fuse the time-domain suppression features and the frequency-domain suppression features to obtain the initial mask feature map. The feature fusion uses an attention mechanism to calculate the correlation between the time-domain suppression features and the frequency-domain suppression features, and weight-combines the two features through adaptive weights. The process of feature fusion uses a multi-head attention mechanism with 8 heads, and the dimension of each attention head is 64. The attention weights are normalized by the softmax function. The initial mask feature map obtained by the fusion contains both time-domain and frequency-domain information, with a dimension of T×F×C, where T represents the number of time frames, F represents the number of frequency points, and C represents the number of feature channels. The initial mask feature map provides a comprehensive time-frequency feature representation for subsequent mask estimation, containing the time-domain structure and frequency-domain distribution characteristics of the audio signal.
[0059] Perform cross-attention operation on the initial mask feature map and the hybrid time-frequency feature to obtain context-aware features. The cross-attention operation captures long-range dependencies by calculating the correlation between the initial mask feature map and the original hybrid time-frequency feature. The cross-attention mechanism uses a Query-Key-Value structure, where the query matrix is generated from the initial mask feature map, and the key matrix and value matrix are generated from the hybrid time-frequency feature. The attention scores are calculated by the dot product of the query and the key, and are scaled and normalized. The cross-attention operation effectively integrates global context information into local features, enhancing the feature representation ability. The context-aware feature has the same dimension as the initial mask feature map, but contains richer context correlation information, which helps to distinguish speech and noise components at similar time-frequency positions.
[0060] Perform non-linear activation mapping on the context-aware features to obtain an activation feature matrix. The non-linear activation function uses parametric PReLU. Compared with the traditional ReLU, PReLU introduces a learnable slope parameter in the negative value interval, avoiding information loss in the negative value interval. The expression of the activation function is f(x) = max(0, x) + α·min(0, x), where α is the learnable parameter. The non-linear activation processing increases the expressive power of the model, enabling the network to learn complex non-linear mapping relationships. The activation feature matrix retains the dimensional structure of the context-aware features, but enhances the discriminability of the features through non-linear transformation, providing a non-linear representation for subsequent recursive processing.
[0061] Perform bidirectional gated recursive processing on the activation feature matrix to obtain time-frequency correlation features. The bidirectional gated recurrent network (Bi-GRU) includes forward and backward processing, with 256 hidden units in each direction. The gating mechanism includes an update gate and a reset gate, which are used to control information flow and historical state reset. The update gate controls the weight ratio of the current input to the historical state, and the reset gate controls the degree to which the historical state participates in the current calculation. The forward recursion captures the left-to-right dependencies in the time dimension, and the backward recursion captures the right-to-left dependencies. The results of the bidirectional recursive processing are combined in a concatenated manner to form a feature representation containing bidirectional time information. The dimension of the time-frequency correlation feature is T×F×2H, where H is the number of hidden units in each direction. The time-frequency correlation feature captures the long-term dependencies of the speech signal in the time dimension, enhancing the model's ability to model the time structure of the speech signal.
[0062] Perform a sparse constraint transformation on the time-frequency correlation features to obtain a mask coefficient vector. The sparse constraint transformation is achieved through a fully connected layer and L1 regularization, which encourages the sparse distribution of the mask coefficients. The input of the fully connected layer is the time-frequency correlation features, and the output dimension is the mask coefficients for each time-frequency point. The weight coefficient of the L1 regularization term is 0.01, which is added to the loss function to promote sparsity. The sparse constraint helps the mask coefficients to approach zero in the noise region and one in the speech region, enhancing the discrimination ability of the mask. The dimension of the mask coefficient vector is T×F×1, and each element represents the retention ratio of the corresponding time-frequency point. The mask coefficient vector has the characteristic of sparsity, effectively distinguishing the speech-dominated region and the noise-dominated region.
[0063] Perform a frequency-band adaptive expansion on the mask coefficient vector to obtain an expanded mask matrix. The frequency-band adaptive expansion is based on the characteristics of human auditory perception and adopts different expansion strategies for different frequency bands. The low-frequency band (0 - 1 kHz) uses fine-grained expansion, the middle-frequency band (1 - 4 kHz) uses medium-grained expansion, and the high-frequency band (4 - 8 kHz) uses coarse-grained expansion. The expansion process uses a frequency-band-dependent interpolation function to perform upsampling and smoothing within each frequency band. The frequency-band adaptive expansion increases the resolution of the mask in the critical frequency bands and improves the perceptual quality of speech enhancement. The dimension of the expanded mask matrix is T×F', where F' represents the number of frequency points after expansion, usually 2 - 4 times the original number of frequency points. The expanded mask matrix provides noise suppression control with higher frequency resolution, especially in the frequency bands containing the speech harmonic structure.
[0064] Perform edge smoothing on the expanded mask matrix to obtain a refined mask matrix. The edge smoothing performs local averaging and gradient constraint on the mask in both the time and frequency dimensions, reducing the discontinuity of the mask in the time-frequency plane. A median filter with a width of 3 is used in the time dimension to eliminate isolated points, and a Gaussian smoothing kernel (σ = 1.5) is used in the frequency dimension for smoothing. The mask values are mapped to the interval [0, 1] through the sigmoid function to ensure the effectiveness of the mask coefficients. The edge smoothing reduces the abrupt changes of the mask in time and frequency, avoiding musical noise and speech distortion. The dimension of the refined mask matrix is the same as that of the expanded mask matrix, but it has smoother time-frequency characteristics. The refined mask matrix is applied to the short-time Fourier transform (STFT) coefficients of the original mixed signal, and the enhanced speech signal is reconstructed through the inverse short-time Fourier transform (ISTFT).
[0065] The refined mask matrix weights the time-frequency representation of the mixed signal, retains the speech components, suppresses the noise components, and finally reconstructs a high-quality enhanced speech signal. The entire noise suppression mask analysis process combines time-domain and frequency-domain information and achieves efficient separation of speech and noise through a deep learning network.
[0066] In this embodiment, by performing time-domain dynamic convolution and frequency-domain dynamic filtering on the mixed time-frequency features, it is possible to effectively suppress noise in different dimensions while retaining the key features of the speech signal, achieving a higher noise suppression effect. Through feature fusion and cross-attention processing, the information in the time domain and frequency domain is combined, long-range dependencies are captured, and the feature representation ability is enhanced, enabling the model to more accurately distinguish speech and noise. Nonlinear activation and bidirectional gated recurrent processing further improve the model's nonlinear mapping and time-dependency modeling capabilities, ensuring that the processed features have higher discriminability and time-frequency correlation. Sparse constraint transformation and frequency-band adaptive expansion processing, through sparse and frequency-band-dependent expansion strategies, enhance the resolution and effectiveness of the mask, improving the fidelity and perceptual quality of the speech signal. Edge smoothing processing reduces the discontinuity of the mask in the time-frequency plane through local averaging and gradient constraints, avoiding musical noise and speech distortion, thereby achieving a smoother and more natural audio enhancement effect.
[0067] In one embodiment, a dual-path processing structure is constructed for the noise-reduced mixed data to obtain amplitude-optimized data and phase-optimized data. In the audio noise reduction method based on deep learning, the construction of the dual-path processing structure refers to a processing architecture that optimizes the noise-reduced mixed data in the amplitude domain and the phase domain respectively. The noise-reduced mixed data is a mixed audio signal containing the target sound and noise, and this mixed data is transformed into the time-frequency domain through the short-time Fourier transform (STFT). In the dual-path processing structure, one path focuses on optimizing the amplitude component of the signal, and the other path focuses on optimizing the phase component of the signal, finally obtaining the amplitude-optimized data and the phase-optimized data.
[0068] The complex-domain separation of the noise-reduced mixed data means decomposing the time-frequency representation after the short-time Fourier transform into an amplitude component and a phase component. The complex-domain separation adopts the polar coordinate representation method, decomposing the complex value of each time-frequency point into two parts: amplitude (magnitude) and phase (phase). The amplitude component represents the energy size of the signal at a specific time-frequency point, and the phase component represents the phase angle of the signal at a specific time-frequency point. The amplitude component is obtained by calculating the modulus of the complex number, and the phase component is obtained by calculating the argument of the complex number. The result of the complex-domain separation is an amplitude component matrix and a phase component matrix, and these two matrices have the same time-frequency dimension.
[0069] The multi - layer residual mapping process for the amplitude component is implemented through a multi - layer convolutional neural network structure. This process uses the Residual Connections mechanism to directly add the input of each layer to the output of that layer, forming a residual learning pattern. The multi - layer residual mapping structure contains multiple Residual Blocks, and each Residual Block consists of two convolutional layers, a batch normalization layer, and an activation function. The convolutional layers are responsible for extracting local features of the amplitude information, the batch normalization layer accelerates network convergence and enhances generalization ability, and the activation function introduces non - linear transformations. The result of the multi - layer residual mapping process is an amplitude feature map, which retains the useful information in the original amplitude component while suppressing the noise components.
[0070] The adaptive threshold constraint process for the amplitude feature map filters out noise interference by setting a dynamic threshold mechanism. The adaptive threshold constraint process is based on the energy distribution difference between the signal and the noise, calculates the estimated noise level value for each frequency band, and sets the adaptive threshold accordingly. When the value in the amplitude feature map is lower than the corresponding threshold, it is attenuated or set to zero; when it is higher than the threshold, the value is retained or enhanced. The threshold is set using the frequency - band correlation analysis method, and different threshold calculation strategies are adopted for the low - frequency region and the high - frequency region considering the energy distribution characteristics of the speech signal in different frequency bands. The result of the adaptive threshold constraint process is amplitude - optimized data, in which the target sound is enhanced while the noise components are suppressed.
[0071] The extraction of the phase component in the complex domain is a method of processing by remapping the phase information onto the complex plane. The phase component is represented as angular information, and there are problems of periodicity and discontinuity in direct processing. The extraction in the complex domain converts the phase angle into a unit vector representation on the complex plane, that is, uses the cosine and sine functions to map the phase angle into two - dimensional coordinate points. This representation method avoids the periodicity problem in phase processing and makes the phase information easier to learn in the neural network. The extraction in the complex domain uses a complex convolutional network structure to extract features for the real and imaginary parts respectively, while maintaining the integrity and coherence of the complex - domain information. The result of the extraction in the complex domain is a phase feature representation, which contains the key features and patterns of the phase information.
[0072] The cyclic consistency constraint processing of the phase feature representation is a method to optimize the phase information by introducing a cyclic consistency loss function. The cyclic consistency constraint processing is based on the continuity characteristic of the phase, ensuring the consistency before and after phase correction. This processing uses a recurrent neural network structure to capture the temporal dependence of the phase information and establish the connection between the phases of the front and back frames. The cyclic consistency loss function ensures that the corrected phase is consistent with the input information after being transformed back to the original domain, avoiding phase distortion. The cyclic consistency constraint processing simultaneously considers the local phase coherence and the global phase structure, balancing the accuracy and stability of the phase correction. The result of the cyclic consistency constraint processing is the preliminary phase correction data, in which the phase noise is preliminarily suppressed.
[0073] The complementary information extraction from the preliminary phase correction data and the amplitude optimization data is achieved through a fusion module. The complementary information extraction is based on the mutual dependence of the amplitude and phase information, and extracts the auxiliary information from the amplitude optimization data that is helpful for phase correction. This process uses an attention mechanism to calculate the correlation weights between the amplitude optimization data and the preliminary phase correction data, highlighting the common time-frequency patterns of the two. The attention weights are applied to the amplitude features to generate the key information related to phase correction. The complementary information extraction considers the characteristic differences of different frequency bands and adopts different fusion strategies for the speech-dominated frequency band and the noise-dominated frequency band. The result of the complementary information extraction is the complementary phase feature, which contains the information helpful for phase optimization extracted from the amplitude optimization data.
[0074] The self-calibration adjustment of the preliminary phase correction data according to the complementary phase feature is achieved through an adaptive fusion network. The self-calibration adjustment is based on the guiding information in the complementary phase feature and finely adjusts the preliminary phase correction data. This process uses a gating mechanism to control the influence degree of the complementary phase feature on the preliminary phase correction and achieve adaptive fusion. The gating unit calculates the fusion weights of each time-frequency point to ensure the optimal fusion effect under different conditions. The self-calibration adjustment simultaneously maintains the temporal coherence and frequency band consistency of the phase, avoiding phase jumps and distortions. The result of the self-calibration adjustment is the phase optimization data, which together with the amplitude optimization data constitutes the complete noise reduction output.
[0075] The final step of the deep learning-based audio noise reduction method is to recombine the amplitude optimization data and the phase optimization data and reconstruct the time-domain audio signal through the inverse short-time Fourier transform (ISTFT). The amplitude optimization data and the phase optimization data jointly determine the complex value of each time-frequency point, thus reconstructing the complete time-frequency representation. The reconstruction process maintains the coordination of the amplitude and the phase, ensuring the naturalness and clarity of the output audio. The noise reduction output has enhanced target sounds and suppressed background noise, improving the overall quality and intelligibility of the audio.
[0076] Through complex domain separation and multi-layer residual mapping processing, this embodiment can fully extract useful features in the amplitude component and suppress noise interference. The adaptive threshold constraint processing mechanism dynamically adjusts the threshold according to the characteristics of different frequency bands, making the noise reduction process more accurate and effective. In terms of phase processing, the present invention solves the problems of periodicity and discontinuity in direct phase processing through a complex domain extraction method, and the cyclic consistency constraint processing ensures the stability of phase correction. An innovative complementary information extraction mechanism is introduced to achieve the collaborative optimization of amplitude and phase information, avoiding the defect of insufficient phase information processing in traditional methods. The self-calibration adjustment technology realizes the adaptive fusion of complementary phase features and preliminary phase correction data through a gating mechanism, ensuring the optimal noise reduction effect in different noise environments.
[0077] In one embodiment, dynamic time-frequency domain cross-fusion is performed on the amplitude optimization data and the phase optimization data to obtain fused audio data, including: The amplitude optimization data refers to the amplitude component of the audio signal after preliminary processing. The amplitude weight feature map is a matrix representing the importance of each time-frequency point in the amplitude optimization data. The calculation of the amplitude weight feature map uses a convolutional neural network structure, which includes three convolutional layers. Each convolutional layer uses a 3×3 convolutional kernel, with a stride of 1 and a padding of 1. After the first convolution, a BatchNorm and a ReLU activation function are connected, and 64 feature channels are output; after the second convolution, a BatchNorm and a ReLU activation function are also connected, and 32 feature channels are output; after the third convolution, a Sigmoid activation function is connected to compress the output to the range of 0-1, forming a weight matrix. This weight matrix is the amplitude weight feature map, whose dimension is the same as that of the input amplitude optimization data, and each element value represents the importance degree of the corresponding time-frequency point. The calculated amplitude weight feature map retains the structural features of the amplitude information and provides a weight basis for subsequent fusion.
[0078] The phase optimization data refers to the phase component of the audio signal after preliminary processing. The phase enhancement feature is the transformation result representing the characteristics of the phase information. The phase transformation uses the Fourier transform method in the complex domain. The phase optimization data is first converted into a complex form and then subjected to a short-time Fourier transform to obtain the phase spectrum feature. The phase spectrum feature is then processed through a convolutional network in the complex domain. This network uses complex convolution operations and includes two complex convolutional layers. The real and imaginary parts of each convolution are processed using independent convolutional kernels respectively. The convolutional kernel size is 5×5, the stride is 1, and the padding is 2. After the complex convolution, normalization and non-linear transformation are performed through a complex BatchNorm and a complex ReLU activation function. The finally obtained complex domain feature is converted into a phase enhancement feature through a phase renormalization operation. The phase enhancement feature retains the key time-frequency structure of the original phase data and at the same time enhances the coherence and consistency of the phase change.
[0079] The interaction weight matrix is a matrix that measures the correlation between amplitude and phase data. In the interaction construction process, the attention mechanism is adopted, with the amplitude weight feature map as the query matrix (Query), and the phase enhancement feature as the key matrix (Key) and value matrix (Value). The dot product of the query matrix and the key matrix is calculated through matrix multiplication, and is normalized using a scaling factor (usually the square root of the feature dimension), and then the Softmax function is applied to obtain the attention weights. The attention weights are multiplied by the value matrix to obtain the initial interaction representation. The initial interaction representation is processed through residual connection and layer normalization to form the interaction weight matrix. The interaction weight matrix captures the structural association between amplitude and phase, providing a basis for subsequent attention decomposition.
[0080] The time-frequency attention coefficient is a weight factor that guides the subsequent reweighting process. The multi-head attention decomposition uses 8 attention heads, and each attention head independently processes different subspace mappings of the interaction weight matrix. Specifically, the interaction weight matrix is first divided into 8 sub-matrices through linear projection, and self-attention calculations are performed on each sub-matrix respectively to generate corresponding attention mappings. These 8 attention mappings are concatenated along the last dimension and projected back to the original dimension space through a linear layer. The concatenated multi-head attention output is then processed through a feed-forward neural network, which contains two linear layers with the GELU activation function in the middle. Finally, through residual connection and layer normalization processing, the time-frequency attention coefficient is obtained. The time-frequency attention coefficient reflects the correlation strength between different time-frequency points, providing fine-grained weight guidance for the fusion of amplitude and phase data.
[0081] The time-frequency enhanced feature is a combined representation of the amplitude and phase information after reweighting processing. During the reweighting process, the time-frequency attention coefficient is decomposed into an amplitude weight and a phase weight. The amplitude weight is obtained by applying the Sigmoid function to the time-frequency attention coefficient, and the phase weight is obtained by applying the Tanh function to the time-frequency attention coefficient. The amplitude-optimized data is multiplied by the amplitude weight, and the phase-optimized data is multiplied by the phase weight to obtain the weighted amplitude data and phase data. The weighted amplitude data and phase data are combined through an addition operation in the complex domain to form a complex representation. This complex representation is subjected to feature extraction through a convolutional layer in the complex domain, using a 3×3 complex convolutional kernel, a stride of 1, and a padding of 1, with the number of output channels being the same as the input. The convolved features are processed through complex BatchNorm and complex PReLU activation functions to obtain the time-frequency enhanced feature. The time-frequency enhanced feature integrates the key information of amplitude and phase in the time-frequency domain, highlighting the main components of the signal.
[0082] Context-dependent representation is a feature representation that integrates global context information. The residual connection process performs a weighted combination of the time-frequency enhanced features with the original amplitude-optimized data and phase-optimized data. The combination weights are determined by a gating mechanism that uses a 1×1 convolutional layer followed by a Sigmoid activation function to generate weight factors. The weighted sum of the time-frequency enhanced features and the original data is processed by a bidirectional long short-term memory network (BiLSTM) with two layers and 256 hidden units to capture long-range dependencies in the time dimension. The output of the BiLSTM is further processed by a temporal convolutional network (TCN) with eight dilated convolutional layers where the dilation rate exponentially increases from 1 to 128 and each layer uses a 1×3 convolutional kernel to capture time patterns at different scales. After the TCN output, a self-attention layer is used to capture global context dependencies, forming the context-dependent representation. The context-dependent representation integrates rich context information in both the time and frequency dimensions, making the audio features more coherent and consistent.
[0083] The time-frequency alignment fusion feature is a feature representation that has undergone complex domain reconstruction and consistency optimization. The complex domain reconstruction process converts the context-dependent representation into a complex form where the real part corresponds to the amplitude information and the imaginary part corresponds to the phase information. The features after complex domain reconstruction are processed by a complex gating convolutional network (Complex GCN) with three layers of complex convolutional layers, each using different-sized convolutional kernels (3×3, 5×5, 7×7 respectively) to capture features with different receptive fields. The amplitude-phase consistency optimization uses a cyclic consistency loss function that calculates the similarity between the original complex representation and the reconstructed complex representation and minimizes the difference between them. The optimization process also introduces a phase constraint term to ensure the consistency of the reconstructed phase information with the original phase information. After complex domain reconstruction and amplitude-phase consistency optimization, the resulting features achieve alignment of amplitude and phase in the time-frequency domain, forming the time-frequency alignment fusion feature. The time-frequency alignment fusion feature enhances the coordination between amplitude and phase while maintaining the integrity of the signal structure.
[0084] The fused audio data is the finally synthesized noise-reduced audio signal. The segmented overlapping process adopts the overlapping add (OLA) method with windowing, dividing the time-frequency aligned fusion features into multiple overlapping time-frequency blocks, with an overlapping rate of 50% for each block. The inverse short-time Fourier transform (ISTFT) is applied to each time-frequency block to convert the time-frequency domain features into a time-domain signal. A Hanning window function is used for windowing during the conversion, with a window length of 1024 sampling points and a frame shift of 512 sampling points. The time-domain signals of adjacent time-frequency blocks are synthesized by the overlapping add method, and the overlapping part is fused by weighted averaging, where the weights are related to the shape of the window function. The finally synthesized time-domain signal undergoes energy normalization to ensure that the energy of the output signal is equivalent to the energy of the input noise signal, obtaining the fused audio data. The fused audio data retains the semantic content of the original audio, while significantly reducing background noise and improving audio quality and intelligibility.
[0085] In this embodiment, by performing dynamic time-frequency domain cross-fusion on the amplitude-optimized data and phase-optimized data, the limitation of traditional audio noise reduction methods that only focus on amplitude information and ignore phase information is overcome, achieving amplitude-phase collaborative optimization and effectively improving the quality and naturalness of the noise-reduced audio. The introduction of the time-frequency attention coefficient enables the system to adaptively capture the noise characteristics in different time-frequency regions and has stronger robustness to different types of noise. The construction of residual connections and context-dependent representations alleviates the vanishing gradient problem in the training of deep networks, retains the key features of the original signal, and improves the convergence speed and stability of the model. The complex-domain reconstruction and amplitude-phase consistency optimization ensure the coordination of amplitude and phase information and avoid the audio distortion problem caused by amplitude-phase mismatch in traditional methods.
[0086] In one embodiment, differential acoustic equation constrained adversarial training is performed on the fused audio data, and inverse time-frequency transformation processing is carried out to obtain the target noise-reduced audio signal, including: When constructing the differential acoustic equation constraint for the fused audio data, the system imposes mathematical constraints on the fused audio data according to a preset acoustic propagation equation. The preset acoustic propagation equation is a mathematical expression formulated based on the physical propagation characteristics of sound waves in the air medium and includes core physical models such as the wave equation and the Helmholtz equation. The system discretizes these acoustic propagation equations and converts them into a differentiable form to make them suitable for gradient calculation in the deep learning framework. During the construction process, the system establishes an analytical expression for acoustic propagation and realizes continuous optimization and adjustment of parameters through automatic differentiation technology, finally forming a differential constraint structure.
[0087] The differential constraint structure refers to a hybrid model structure that combines the physical laws of acoustics with the characteristics of deep learning. This structure not only conforms to the physical laws of acoustic propagation but also has the optimization characteristics of deep learning models.
[0088] When constructing the gradient penalty term for the differentiable constraint structure based on the preset clear audio reference samples, the system selects multiple high-quality clear audio samples as the reference standard. The preset clear audio reference samples refer to high-quality audio segments that are professionally recorded and have no obvious noise interference. These samples are usually collected in professional recording studios and have high signal-to-noise ratios and clear acoustic characteristics. The system calculates the difference between the fused audio data and the reference samples in the frequency domain representation and introduces a gradient penalty mechanism to impose a greater penalty weight on the spectral components that do not conform to the acoustic characteristics of the reference samples. In this way, the spectral distance loss function is constructed.
[0089] The spectral distance loss function is a mathematical expression that measures the spectral difference between the processed audio and the ideal clear audio. This function not only considers the difference in spectral amplitudes but also includes the difference metrics of acoustic characteristics such as phase information and harmonic structure.
[0090] When evaluating the authenticity of the fused audio data according to the spectral distance loss function, the system constructs a discriminator network. This network receives the fused audio data as input and performs time-frequency analysis on it. The discriminator network calculates the similarity between the spectral characteristics of the fused audio data and the preset clear samples based on the spectral distance loss function and comprehensively evaluates its authenticity. During the evaluation process, the system detects features such as unnatural fluctuations, signal discontinuities, and abnormal harmonic structures in the spectrum and gives a quantitative discrimination score.
[0091] The discrimination score refers to the quantitative evaluation of the authenticity of the fused audio data by the discriminator network. The score range is usually between 0 and 1, where 1 indicates a perfect match with the clear audio characteristics and 0 indicates a complete mismatch with the natural audio characteristics.
[0092] When reconstructing the waveform of the fused audio data according to the discrimination score and the spectral distance loss function, the system applies the generative adversarial network architecture and gradually improves the discrimination score of the reconstructed audio through iterative optimization. During the reconstruction process, the system adjusts the spectral characteristics according to the feedback of the discrimination score, preferentially repairs the frequency bands with low discrimination scores, and maintains spectral continuity. At the same time, the system uses the spectral distance loss function to guide the waveform reconstruction direction, making the reconstruction result conform to the spectral characteristics of the clear audio while retaining the semantic information of the original audio. After multiple rounds of iterative optimization, the system generates the corrected audio data.
[0093] The corrected audio data is an intermediate result after waveform reconstruction. At this time, the main noise of the audio data has been suppressed, but there may still be subtle residual noise components.
[0094] When processing the residual noise components of the corrected audio data, the system applies a method that combines spectral subtraction and adaptive filtering. The system makes a fine frequency band division of the corrected audio data and identifies the residual noise characteristics in each frequency band. For the frequency bands with significant noise characteristics, the system adopts a selective attenuation strategy, dynamically adjusting the attenuation coefficient according to the intensity of the noise characteristics to ensure that while reducing the noise, the useful signal components are retained to the greatest extent. Through this refined residual noise processing, the system obtains a clean time-frequency representation.
[0095] A clean time-frequency representation refers to the representation form of the audio data in the time-frequency domain after the residual noise components have been reduced. This representation contains almost no noise interference and retains the key acoustic information of the original audio.
[0096] When performing a segmented inverse short-time Fourier transform on the clean time-frequency representation, the system divides the clean time-frequency representation into multiple time windows and performs the inverse short-time Fourier transform on each window separately. The segmented processing uses the overlap-and-add method, with a 50% overlap rate between adjacent windows to ensure a smooth transition of the transformed signal. During the transformation process, the system applies window functions such as the Hann window to adjust the boundary effect and reduce spectral leakage. Subsequently, the system recombines the transformation results of each window in chronological order and performs a normalization process to obtain the initial target audio.
[0097] The initial target audio refers to the time-domain audio signal restored through the inverse short-time Fourier transform. This signal has basically restored clarity but may have minor artifacts introduced during the transformation process.
[0098] When performing artifact filtering on the initial target audio, the system uses a method that combines median filtering and bilateral filtering to process the transformation artifacts. The system identifies the high-frequency spikes and unnatural fluctuations in the initial target audio and uses an adaptive threshold to distinguish the artifacts from the normal audio components. For the detected artifact regions, the system applies local smoothing processing to maintain the overall coherence of the signal. During the filtering process, the system maintains the harmonic structure and transient characteristics of the original audio, avoiding a reduction in sound quality caused by over-smoothing. After the artifact filtering process, the system finally obtains the target noise-reduced audio signal.
[0099] The target noise-reduced audio signal is the final audio output obtained after the complete processing flow. This signal not only removes the noise interference in the original audio but also retains the natural characteristics and semantic information of the sound, achieving a high-quality audio noise reduction effect.
[0100] In this embodiment, by performing differentiable acoustic equation-constrained adversarial training on the fused audio data, the physical acoustic laws and deep learning techniques are organically combined, enabling the noise reduction process to not only conform to the physical characteristics of sound wave propagation but also possess the optimization ability of deep learning, significantly improving the accuracy and physical rationality of the noise reduction process. The authenticity evaluation mechanism based on the spectral distance loss function can accurately identify the unnatural components in the audio and guide the waveform reconstruction through the discrimination score, ensuring the complete retention of semantic information during the noise reduction process. The processing method of segmented inverse short-time Fourier transform combined with the overlap-add method effectively reduces spectral leakage and ensures the coherence of the time-domain signal reconstruction. The dual processing mechanism of residual noise component reduction and artifact filtering removes weak noise while avoiding the sound quality loss caused by over-smoothing, maximizing the retention of the natural acoustic characteristics of the original audio. The overall method can still extract clear audio signals in complex noise environments and has significant value for practical applications in fields such as voice communication and audio processing.
[0101] In one embodiment, a differentiable constraint structure is constructed by performing differentiable constraints on the fused audio data according to a preset acoustic propagation equation, including: When extracting the acoustic propagation characteristic parameters of the fused audio data, the time-domain fused audio signal is converted to the frequency domain through Fourier transform, and parameters such as sound wave frequency, amplitude, phase, and propagation speed are extracted. The acoustic characteristic matrix is a multi-dimensional data structure that contains the distribution information of the above parameters at different frequencies and spatial positions. This matrix records the attenuation characteristics, reflection characteristics, and absorption characteristics of the audio signal during propagation, providing basic data for subsequent sound field modeling.
[0102] During the process of constructing the expression of the acoustic characteristic matrix according to the acoustic propagation equation, a sound field propagation model is established based on the Helmholtz equation. The sound field propagation expression describes the propagation law of sound waves in space, integrating the frequency dependence and spatial distribution characteristics in the acoustic characteristic matrix. This expression mathematizes the physical phenomena such as sound wave reflection, attenuation, and interference, establishing a mathematical relationship between the sound source and the receiving point, providing a theoretical basis for subsequent sound field analysis.
[0103] When discretizing the sound field propagation expression, the finite difference method is used to divide the continuous sound field space into discrete grid points, and the sound pressure value is calculated at each grid point. The grid sound field distribution map is a discrete data structure representing the spatial distribution of the sound field, recording the sound pressure distribution at different spatial positions and time points, providing a spatial reference for boundary condition construction.
[0104] When constructing differentiable boundary conditions based on the grid sound field distribution map, reflection, absorption, or transmission conditions are set at the sound field boundary positions. The acoustic boundary constraint expression defines the behavior rules of sound waves when encountering boundaries, including the total reflection condition of a rigid boundary, the partial reflection condition of an impedance boundary, and the radiation condition of an open boundary, ensuring that the sound field calculation conforms to physical laws.
[0105] When calculating the gradient of the acoustic boundary constraint expression, the impact of the boundary constraint on the sound field is calculated through automatic differentiation technology. The acoustic gradient field represents the rate of change of acoustic parameters in space, describes the change trend of the spatial distribution of the sound field, provides directional information for the constraint structure, and guides the intensity and direction of constraint application.
[0106] When performing tensor fusion of the acoustic gradient field and the acoustic property matrix, the tensor product operation is used to combine the two in the same dimension. The initial constraint structure is a fused multi-dimensional data structure that contains both acoustic properties and gradient information, provides a basis for subsequent constraint construction, and enhances the expressive power and physical meaning of the constraints.
[0107] The process of obtaining second-order partial differential constraints by applying operator constraints to the initial constraint structure involves multiple delicate operations. In this process, the initial constraint structure is first represented in the form of a multi-dimensional tensor, containing time-domain, frequency-domain, and spatial-domain information. When the Laplace operator is applied to the initial constraint structure, second-order partial derivatives are calculated for each spatial dimension. For a point (x, y, z) in three-dimensional space, the second-order partial derivatives along the x, y, and z directions are calculated separately, and the results are superimposed. The central difference method is used for the calculation. For each grid point, the values of the surrounding adjacent points are used to calculate the approximate value of the second derivative.
[0108] The second-order partial differential constraints capture the abrupt, discontinuous, and abnormal points in the sound field by considering the acceleration change characteristics in sound wave propagation. These characteristics play an important role in identifying the positions and characteristics of noise sources. The calculation results are regularized to eliminate numerical errors and instabilities caused by discretization. During the regularization process, low-pass filtering technology is used to filter out high-frequency noise components and retain the main acoustic characteristics.
[0109] The mathematical expression form of the second-order partial differential constraints ensures the differentiability of the constraints in the deep learning model, enabling the gradient to flow smoothly during the backpropagation process. After the calculation is completed, the second-order partial differential constraints are stored as a tensor with the same dimension as the initial constraint structure, retaining the complete spatial distribution information of the second derivative of the sound field.
[0110] The process of calculating constraint parameters for the second-order partial differential constraints to obtain dynamic constraint parameters is based on sound field characteristic analysis and an adaptive mechanism. First, spectral analysis is performed on the second-order partial differential constraints to extract the energy distribution characteristics of different frequency bands. The spectral analysis uses the short-time Fourier transform method to decompose the time-domain signal into multiple frequency bands.
[0111] Based on the spectrum analysis results, a frequency-dependent parameter model is constructed, and different weight coefficients are set for different frequency bands. Low-frequency sounds usually require a smaller constraint intensity, while high-frequency noises require a larger constraint intensity. At the same time, the time-varying characteristics of the sound field are evaluated, the signal-to-noise ratio of the noise and the target audio signal is calculated, and the constraint parameters are dynamically adjusted according to the signal-to-noise ratio value.
[0112] The nonlinear mapping function is introduced in the dynamic constraint parameter calculation process to convert the acoustic characteristic index into the constraint intensity parameter. The S-shaped function is used for the nonlinear mapping, providing strong constraints in the low signal-to-noise ratio region and weak constraints in the high signal-to-noise ratio region. The position-dependent parameter is set in the spatial domain, and the constraint intensity is adjusted according to the source distance and directional characteristics.
[0113] The time smoothing strategy is introduced in the calculation process to avoid the problem of incoherence in the listening experience caused by the drastic fluctuation of the constraint parameters. The exponential weighted moving average method is used for the smoothing strategy, comprehensively considering the parameter values of the historical frames and the current frame. After the calculation is completed, the dynamic constraint parameters form a multi-dimensional parameter set, including the constraint intensity modulation factors that depend on frequency, time, and space.
[0114] When modulating the initial constraint structure according to the dynamic constraint parameters, the parameters are incorporated into the constraint structure through the weighted fusion method. The enhanced constraint structure is a constraint expression with adaptive characteristics, integrating the initial constraint and the dynamic parameters, improving the flexibility and pertinence of the constraint, and better adapting to the complex sound field environment.
[0115] When performing the Taylor series expansion on the enhanced constraint structure, the first-order or second-order expansion is carried out near the working point, and the main terms are retained. The linear approximation representation is a simplified form of the enhanced constraint structure, reducing the computational complexity, retaining the main characteristics of the constraint, and facilitating the efficient implementation in the deep learning model.
[0116] The process of constructing the differentiable constraint structure by constraining the fused audio data based on the linear approximation representation and the enhanced constraint structure is the core link of the whole method. First, the linear approximation representation is converted into the form of a loss function, defining the balance relationship between the audio reconstruction target and the physical constraint. The loss function is designed in the form of a weighted sum, combining the data reconstruction error term and the physical constraint violation term.
[0117] The enhanced constraint structure is transformed into a custom layer in the neural network. This layer receives the audio feature map and outputs the constraint satisfaction score. The implementation of the custom layer involves the forward propagation calculation and the gradient backpropagation mechanism, ensuring that the constraint can correctly guide the parameter update during the model training. In the forward propagation calculation, the physical constraint test is applied to the input feature map to calculate the degree of constraint satisfaction.
[0118] During the construction process, a gradient scaling mechanism is designed to prevent the constraint gradient from being too large or too small, which may affect the model convergence. The gradient scaling adopts the adaptive learning rate method to dynamically adjust the constraint gradient intensity according to the training process. Batch processing mode calculation is implemented to support the constraint application of multiple audio samples simultaneously, improving the training efficiency.
[0119] The final form of the differentiable constraint structure is a composite function, which includes a linear approximation part and a non - linear modulation part. The linear approximation part provides the basic constraint form to ensure the computational efficiency. The non - linear modulation part introduces adaptive characteristics to improve the expressive ability and flexibility of the constraint. Numerical stability verification is carried out on the constructed constraint structure to ensure that it will not cause gradient explosion or vanishing problems in extreme cases.
[0120] The constructed differentiable constraint structure is used as a regularization term in the training process of the deep - learning model, guiding the model optimization direction together with the traditional loss function, so that the noise reduction result can meet both the data fitting goal and the physical acoustic law.
[0121] In this embodiment, by introducing the construction process of differentiable constraints into the deep - learning - based audio noise reduction method, the acoustic physical law is organically combined with the deep - learning model, so that the noise reduction result can meet both the data fitting goal and the physical acoustic law. This method uses acoustic propagation characteristic parameter extraction and sound field modeling to establish constraint conditions consistent with the actual acoustic environment, improving the physical rationality and auditory realism of the noise reduction result. Through the second - order partial differential constraint acquisition process, the system can accurately capture the mutations, discontinuities and abnormal points in the sound field, providing important support for identifying the position and characteristics of noise sources and improving the accuracy of noise reduction. The dynamic constraint parameter calculation process introduces a frequency - dependent parameter model and a non - linear mapping mechanism, enabling the system to adaptively adjust according to the sound characteristics of different frequency bands and improving the adaptability to complex sound field environments. The final form of the differentiable constraint structure is a composite function, including a linear approximation part and a non - linear modulation part, which not only ensures the computational efficiency but also improves the expressive ability and flexibility of the constraint. This method effectively prevents the constraint gradient from being too large or too small from affecting the model convergence through the gradient scaling mechanism and batch processing mode calculation, and at the same time improves the training efficiency, providing reliable technical support for practical applications.
[0122] Refer to Figure 2 As shown in The present invention also provides a deep - learning - based audio noise reduction device applied to the deep - learning - based audio noise reduction method of any one of the above, including: An acquisition module, which is used to obtain the input noisy audio signal and perform multi - scale time - frequency decomposition to obtain the mixed time - frequency features and the noise fingerprint map;An analysis module, which is used to perform parameter parallel processing on the noise fingerprint spectrum through a preset dynamic kernel generation network, and perform preliminary noise reduction processing on the mixed time-frequency features to obtain noise-reduced mixed data; An association module, which is used to construct a dual-path processing structure for the noise-reduced mixed data to obtain amplitude-optimized data and phase-optimized data; A processing module, which is used to perform dynamic time-frequency domain cross-fusion on the amplitude-optimized data and the phase-optimized data to obtain fused audio data; A control module, which is used to perform differentiable acoustic equation-constrained adversarial training on the fused audio data and perform inverse time-frequency transform processing to obtain the target noise-reduced audio signal.
[0123] An audio noise reduction device based on deep learning provided by the present invention can more comprehensively capture the feature distributions of different noise types at multiple scales through multi-scale time-frequency decomposition technology, thereby improving the noise reduction effect and providing a more reliable basis for the high-quality restoration of audio signals. By using a preset dynamic kernel generation network to perform parameter parallel processing on the noise fingerprint spectrum, the adaptive adjustment of the dynamic characteristics of the noise is realized, and it can more effectively cope with the complex and changing noise environment. By respectively optimizing the amplitude and phase information through a dual-path processing structure and performing dynamic time-frequency domain cross-fusion, the acoustic consistency of the reconstructed audio is ensured, and the overall quality of the audio signal is improved. Through differentiable acoustic equation-constrained adversarial training, using the basic laws of acoustic propagation, strict physical constraints are imposed on the fused audio data, reducing the distortion or artifact phenomenon in the noise reduction result, and further improving the authenticity and reliability of the noise reduction effect. The present invention not only overcomes the limitations of traditional noise reduction methods, but also significantly improves the actual performance of audio noise reduction technology in high-noise environments and high-quality audio application scenarios, effectively improving the overall efficiency and effect of audio signal processing.
[0124] Refer to Figure 3 As shown, the present invention also provides an audio noise reduction system based on deep learning, including: A memory, which is used to store programs; A processor, which is used to execute programs to implement the steps of an audio noise reduction method based on deep learning according to any one of claims 1-8.
[0125] It should be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described system and each module can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.
[0126] In this embodiment, the processor and the memory can be connected via a bus or other means. The memory may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as read-only memory, flash memory, hard disk, or solid state drive. The processor may be a general-purpose processor, such as a central processing unit, a digital signal processor, an application specific integrated circuit, or one or more integrated circuits configured to implement the embodiments of the present invention.
[0127] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. An audio noise reduction method based on deep learning, characterized in that: include: Get the input noisy frequency signal and perform multi-scale time-frequency decomposition to obtain mixed time-frequency features and noise fingerprints; The noise fingerprint spectrum is processed in parallel by using a preset dynamic kernel generation network, and the mixed time-frequency features are subjected to preliminary noise reduction processing to obtain noise-reduced mixed data; Performing a dual-path processing structure on the noise reduction mixed data to obtain amplitude optimized data and phase optimized data; Performing dynamic time-frequency domain cross-fusion on the amplitude optimized data and the phase optimized data to obtain fused audio data; The fused audio data is subjected to differentiable acoustic equation constrained adversarial training and inverse time-frequency transform processing to obtain a target noise-reduced audio signal.
2. The audio noise reduction method based on deep learning according to claim 1, characterized in that: The step of obtaining an input noisy frequency signal and performing multi-scale time-frequency decomposition to obtain a mixed time-frequency feature and a noise fingerprint spectrum includes: Performing multi-band decomposition on the noisy frequency signal to obtain a set of time domain sequences; Performing spectral entropy analysis on the time domain sequence set to obtain a spectrum activity index; Performing adaptive threshold filtering according to the spectrum activity index to obtain a noise feature enhancement matrix; Extracting Mel-frequency cepstral coefficients from the noise feature enhancement matrix to obtain a noise feature vector; Performing self-encoding processing on the noise feature vector to obtain the noise fingerprint spectrum; Performing discrete cosine transform on the noise fingerprint spectrum to obtain time-varying spectrum features; The time-varying spectrum feature and the noise feature vector are subjected to tensor concatenation and normalization processing to obtain the mixed time-frequency feature.
3. The audio noise reduction method based on deep learning according to claim 1, characterized in that: The method of performing parameter parallel processing on the noise fingerprint spectrum through a preset dynamic kernel generation network and performing preliminary noise reduction processing on the mixed time-frequency features to obtain noise-reduced mixed data includes: Extracting the noise level of the noise fingerprint to obtain a time domain feature sequence and a frequency domain feature sequence; Performing nonlinear transformation on the time domain feature sequence to obtain the time domain dynamic convolution kernel; Performing frequency band filtering on the frequency domain feature sequence to obtain the frequency domain dynamic filter bank; Performing noise suppression mask analysis on the mixed time-frequency features according to the time-domain dynamic convolution kernel and the frequency-domain dynamic filter bank to obtain a refined mask matrix; The mixed time-frequency features are subjected to denoising according to the refined mask matrix to obtain the denoised mixed data.
4. The audio noise reduction method based on deep learning according to claim 3, characterized in that: The noise suppression mask analysis is performed on the mixed time-frequency features according to the time-domain dynamic convolution kernel and the frequency-domain dynamic filter group to obtain a refined mask matrix, including: Performing a time-domain convolution operation on the mixed time-frequency feature according to the time-domain dynamic convolution kernel to obtain a time-domain suppression feature; Performing frequency domain filtering operation on the mixed time-frequency features according to the frequency domain dynamic filter group to obtain frequency domain suppression features; Performing feature fusion on the time domain suppression feature and the frequency domain suppression feature to obtain an initial mask feature map; Performing a cross attention operation on the initial mask feature map and the mixed time-frequency feature to obtain a context-aware feature; Performing nonlinear activation mapping on the context-aware features to obtain an activation feature matrix; Performing bidirectional gated recursive processing on the activation feature matrix to obtain time-frequency correlation features; Performing a sparse constraint transformation on the time-frequency correlation feature to obtain a mask coefficient vector; Performing frequency band adaptive expansion on the mask coefficient vector to obtain an expanded mask matrix; The extended mask matrix is subjected to edge smoothing processing to obtain the refined mask matrix.
5. The audio noise reduction method based on deep learning according to claim 1, characterized in that: A dual-path processing structure is constructed for the noise reduction mixed data to obtain amplitude optimized data and phase optimized data. Performing complex domain separation on the noise reduction mixed data to obtain an amplitude component and a phase component; Performing multi-layer residual mapping processing on the amplitude component to obtain an amplitude feature map; Performing adaptive threshold constraint processing on the amplitude characteristic graph to obtain the amplitude optimization data; Performing complex domain extraction on the phase component to obtain a phase feature representation; Performing cyclic consistency constraint processing on the phase feature representation to obtain preliminary phase correction data; Extracting complementary information from the preliminary phase correction data and the amplitude optimization data to obtain complementary phase features; The preliminary phase correction data is self-calibrated and adjusted according to the complementary phase characteristics to obtain the phase optimization data.
6. The audio noise reduction method based on deep learning according to claim 1, characterized in that: The step of performing dynamic time-frequency domain cross-fusion on the amplitude optimized data and the phase optimized data to obtain fused audio data includes: Performing weight matrix calculation on the amplitude optimization data to obtain an amplitude weight characteristic graph; Performing phase transformation on the phase optimization data to obtain a phase enhancement feature; Interactively constructing according to the amplitude weight feature map and the phase enhancement feature to obtain an interactive weight matrix; Performing multi-head attention decomposition on the interaction weight matrix to obtain a time-frequency attention coefficient; Re-weighting the amplitude optimized data and the phase optimized data according to the time-frequency attention coefficient to obtain a time-frequency enhanced feature; Performing residual connection on the time-frequency enhancement features to obtain context-dependent representation; Performing complex domain reconstruction and amplitude-phase consistency optimization on the context-dependent representation to obtain time-frequency aligned fusion features; The time-frequency alignment fusion features are subjected to segmented overlapping processing to obtain fused audio data.
7. The audio noise reduction method based on deep learning according to claim 1, characterized in that: The step of performing differentiable acoustic equation constrained adversarial training on the fused audio data and performing inverse time-frequency transform processing to obtain a target noise reduction audio signal includes: Performing differentiable constraint construction on the fused audio data according to a preset acoustic propagation equation to obtain a differentiable constraint structure; Constructing a gradient penalty term for the differentiable constraint structure according to a preset clear audio reference sample to obtain a spectrum distance loss function; Performing authenticity evaluation on the fused audio data according to the spectral distance loss function to obtain a discrimination score; Reconstructing the waveform of the fused audio data according to the discrimination score and the spectrum distance loss function to obtain corrected audio data; Performing residual noise component reduction processing on the modified audio data to obtain a pure time-frequency representation; Performing a piecewise inverse short-time Fourier transform on the pure time-frequency representation and performing time-domain signal synthesis to obtain an initial target audio; Performing artifact filtering on the initial target audio to obtain a target noise-reduced audio signal.
8. The audio noise reduction method based on deep learning according to claim 7, characterized in that: The step of constructing a differentiable constraint on the fused audio data according to a preset acoustic propagation equation to obtain a differentiable constraint structure includes: Extracting acoustic propagation characteristic parameters from the fused audio data to obtain an acoustic characteristic matrix; Constructing an expression for the acoustic characteristic matrix according to the acoustic propagation equation to obtain an expression for sound field propagation; Discretizing the sound field propagation expression to obtain a grid sound field distribution diagram; Constructing a differentiable boundary condition according to the grid acoustic field distribution diagram to obtain an acoustic boundary constraint expression; Performing gradient calculation on the acoustic boundary constraint expression to obtain an acoustic gradient field; Performing tensor fusion on the acoustic gradient field and the acoustic characteristic matrix to obtain an initial constraint structure; Performing operator constraints on the initial constraint structure to obtain a second-order partial differential constraint; Calculating constraint parameters of the second-order partial differential constraint to obtain dynamic constraint parameters; Modulating the initial constraint structure according to the dynamic constraint parameters to obtain an enhanced constraint structure; Performing Taylor series expansion on the enhanced constraint structure to obtain a linear approximate representation; The fused audio data is constrained and constructed according to the linear approximation representation and the enhanced constraint structure to obtain a differentiable constraint structure.
9. An audio noise reduction device based on deep learning, characterized in that: The deep learning-based audio noise reduction method applied to any one of claims 1 to 8, comprising: An acquisition module, which is used to acquire an input noisy frequency signal and perform multi-scale time-frequency decomposition to obtain a mixed time-frequency feature and a noise fingerprint spectrum; An analysis module, the analysis module is used to perform parameter parallel processing on the noise fingerprint spectrum through a preset dynamic kernel generation network, and perform preliminary noise reduction processing on the mixed time-frequency features to obtain noise-reduced mixed data; An association module, the association module is used to construct a dual-path processing structure for the noise reduction mixed data to obtain amplitude optimized data and phase optimized data; A processing module, the processing module is used to perform dynamic time-frequency domain cross-fusion on the amplitude optimization data and the phase optimization data to obtain fused audio data; A control module is used to perform differentiable acoustic equation constrained adversarial training on the fused audio data and perform inverse time-frequency transform processing to obtain a target noise reduction audio signal.
10. An audio noise reduction system based on deep learning, characterized in that: include: Memory, used to store programs; A processor is used to execute the program to implement the various steps of the audio noise reduction method based on deep learning as described in any one of claims 1 to 8.
Citation Information
Cited By
Audio noise suppression method based on deep learning and intelligent sound equipment
CN120412619A
Analog-digital hybrid SRAM memory audio noise reduction method and system
CN120708645A
Rare earth ion chromatography online analysis and detection method and system
CN120721908A
Passive detection signal enhancement method and device under low signal-to-noise ratio, equipment and medium
CN120741957A
Multitask speech enhancement method and device
CN120766701A