Multi-domain fusion voice noise reduction method and system and computer readable storage medium

By employing a multi-domain fusion speech denoising method that combines adaptive notch filtering, wavelet decomposition, frequency domain Kalman filtering, and lightweight convolutional networks, the problem of simultaneously suppressing narrowband line spectrum and broadband noise in complex noise environments is solved, achieving high-gain denoising and low-distortion effects, and is applicable to a variety of complex acoustic scenarios.

CN121983017APending Publication Date: 2026-05-05NANJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANJING UNIV OF POSTS & TELECOMM
Filing Date
2026-02-09
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively handle complex noise in scenarios such as industrial sites, vehicle/airborne environments, and conference data acquisition, especially the simultaneous suppression of narrowband line spectrum and broadband noise. Furthermore, deep learning models suffer from high computational resource requirements and poor interpretability.

Method used

A multi-domain fusion speech denoising method is adopted, which simultaneously suppresses narrowband line spectrum and broadband noise through adaptive notch filtering, wavelet decomposition, improved logarithmic spectrum amplitude estimation, frequency domain Kalman filtering and lightweight convolutional network, and performs signal reconstruction and spectrum equalization.

Benefits of technology

It achieves high-gain noise reduction and low distortion in complex noise environments, supports real-time deployment, has good interpretability, is suitable for speech enhancement in a variety of complex acoustic scenarios, and is lightweight to deploy, with multi-dimensional visualization analysis tools.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121983017A_ABST
    Figure CN121983017A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-domain fusion voice noise reduction method and system and a computer readable storage medium, and relates to the technical field of voice enhancement and signal processing, and the method comprises the following steps: recognizing a candidate line spectrum through constructing a power spectrum peak value comprehensive score, and carrying out the adaptive notch; performing compensation in combination with wavelet adaptive threshold processing and improved logarithmic spectrum amplitude estimation; and weighted fusion is further carried out on the signals after wave trapping and compensation through spectrum energy entropy. Rapid suppression of narrowband line spectrum interference is realized through spectrum line peak detection and an adaptive notch technology, broadband background noise is effectively eliminated by means of a wavelet coefficient energy ratio eta adaptive threshold algorithm and an OM-LSA spectrum compensation technology, and recursive robust estimation in a time-frequency domain is completed by adopting frequency domain three-dimensional Kalman filtering. And finally, signal detail reconstruction and spectrum equalization are realized by using a lightweight residual CNN network without training weight, so that the optimal balance of high gain and low distortion is achieved in a complex noise environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech enhancement and signal processing technology, specifically to a multi-domain fusion speech noise reduction method, system, and computer-readable storage medium. Background Technology

[0002] In current industrial sites, vehicle / airborne environments, conference data acquisition, and communication links, voice signals are generally subject to a combination of two types of noise interference: one is the narrowband line spectrum and its harmonic interference caused by equipment characteristics such as motor operation, inverter operation, and rectifier residue; the other is broadband Gaussian or near-Gaussian noise from the environmental background, sensor background, and transmission channel.

[0003] Faced with such complex noise challenges, existing technical solutions have obvious limitations: traditional adaptive filtering (such as NLMS) can effectively suppress line spectrum components, but its effect on broadband noise is limited; methods based on spectral subtraction or improved logarithmic spectral amplitude estimation can handle broadband noise, but often leave obvious line spectrum spikes; statistical modeling methods and single Kalman filtering techniques are not adaptable to non-stationary line spectra; and while end-to-end deep learning models have shown some effectiveness, they have inherent defects such as model training relying on a large amount of data, high computational resource requirements, weak system interpretability, and insufficient engineering transferability.

[0004] Therefore, the industry urgently needs a composite noise reduction technology solution that does not require training, has good interpretability, supports real-time deployment, and can effectively handle both narrowband line spectrum and broadband noise. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides a multi-domain fusion speech denoising method, system, and computer-readable storage medium, which solves the problems mentioned in the background art.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a multi-domain fusion speech denoising method, comprising the following steps: S1: Perform power spectral density estimation on the input signal, and construct a comprehensive score by combining peak amplitude ratio, adjacent peak contrast, noise peak ratio and local power spectrum smoothness to identify candidate line spectrum frequencies and their corresponding bandwidths; S2: Perform adaptive notch filtering based on the identified candidate frequencies, and use the correlation coefficient to root mean square amplitude ratio as the iteration termination condition; S3: Perform wavelet decomposition on the notch-filtered signal, adaptively determine the threshold and retention coefficients based on the energy distribution of each decomposition layer, and perform threshold processing and reconstruction on the detail coefficients and approximation coefficients respectively. S4: The spectrum compensation of the reconstructed signal after S3 is performed using an improved logarithmic spectrum amplitude estimation method; S5: Based on the spectral energy entropy of the notch-filtered signal in S2 and the spectrum-compensated signal in S4, perform adaptive weighted fusion on the two; S6: In the short-time Fourier transform domain, Kalman filtering is performed on the signal fused by S5 at frequency points and frame by frame. The process noise variance and measurement noise variance are dynamically adjusted according to wavelet energy interpolation, spectral gradient and instantaneous signal-to-noise ratio. S7: Calculate the residual between the output signal in S6 and the output signal in S5, and use a lightweight convolutional network with fixed weights to perform detail restoration and signal equalization on the residual to obtain the final denoised signal.

[0007] Preferably, the comprehensive score in S1 is calculated using the following formula:

[0008] in, This indicates the overall score. This represents the standardized peak amplitude ratio. This represents the ratio of adjacent peak values ​​after standardization. This represents the normalized peak noise ratio. This represents the smoothness of the local power after standardization.

[0009] During the online spectrum screening process, a peak height threshold and a minimum frequency interval threshold are set simultaneously to determine the effective line spectrum components from the candidate frequencies.

[0010] Preferably, the notch filtering in S2 employs an infinite impulse response notch filter, and the bandwidth scaling factor of the notch filter is dynamically adjusted within the range of 1.0 to 1.3.

[0011] Preferably, the number of wavelet decomposition layers in S3 is determined by the following formula:

[0012] Where L is the wavelet decomposition level, sam is the sampling rate of the input signal, and the wavelet basis function used in the wavelet decomposition is coif4 or sym4.

[0013] Preferably, in step S3, a threshold is set for the wavelet coefficients of the j-th layer. The calculation method is as follows

[0014] in, This is the threshold scaling factor, used to adjust the overall threshold strength. A value greater than 1 indicates enhanced noise reduction strength, while a value less than 1 indicates weakened noise reduction strength. The noise standard deviation of the wavelet coefficients at the j-th level is estimated based on the median absolute deviation of the wavelet coefficients. denoted as the number of wavelet coefficients in the j-th layer.

[0015] At the same time, the retention coefficient of the j-th layer is set. The calculation method is as follows:

[0016] in, To preserve the upper limit of the retention coefficient and protect the integrity of the signal structure, the value ranges from 0.95 to 0.98; The base retention factor is used to avoid excessive distortion of the signal in the high-noise layer or the introduction of artifacts, and its value ranges from 0.1 to 0.3. This is the energy gain factor, used to adaptively adjust the noise reduction intensity according to the energy ratio of each layer, with a value range of 8 to 12. The energy proportion of the wavelet coefficients at the j-th level is calculated as follows:

[0017] in, The total energy of the wavelet coefficients at the j-th level is calculated using the following formula:

[0018] The total energy of all wavelet coefficients is calculated using the following formula:

[0019] Preferably, the fusion weights in S5 The calculation method is as follows:

[0020] Where H represents the Shannon entropy function of the average spectral energy, This is the output signal after notch filtering in S2. This is the output signal after spectral compensation in S4.

[0021] Preferably, the process noise variance Q and measurement noise variance R in S6 are set as follows:

[0022] in,

[0023] In the formula, This is the baseline value for process noise. To measure the noise reference value; The parameter is estimated through the spectral energy distribution, representing the proportion of energy in each frequency band to the total signal energy; The signal energy of the current frame is given by `prevE`, and the signal energy of the previous frame is given by `prevE`. P represents the process noise power; The coefficients, determined by the optimization algorithm, are used to adjust the intensity of the influence of each physical quantity on Q and R.

[0024] Preferably, the residual convolutional network in S7 consists of 2 to 5 layers of one-dimensional convolutions, with each convolutional kernel having a length of 3 to 7. The network structure includes a residual gating mechanism, and all network parameters remain fixed during deployment and use, requiring no training optimization.

[0025] A multi-domain fusion speech denoising system includes a line spectrum recognition unit, a notch filtering unit, a wavelet thresholding unit, a spectrum compensation unit, an entropy fusion unit, a frequency domain Kalman filtering unit, a residual convolution unit, and a control module. The line spectrum recognition unit is used to detect candidate line spectrum frequencies and their bandwidths in the input signal; The notch processing unit is connected to the line spectrum recognition unit and is used to perform adaptive notch filtering on the candidate line spectrum frequencies. The wavelet thresholding unit is connected to the notch processing unit and is used to perform wavelet decomposition and adaptive thresholding on the notch-filtered signal. The spectrum compensation unit is connected to the wavelet threshold processing unit and is used to perform logarithmic spectrum amplitude compensation on the reconstructed signal; The entropy fusion unit is connected to the notch processing unit and the spectrum compensation unit respectively, and is used to perform adaptive weighted fusion of the two signals based on the spectral energy entropy. The frequency domain Kalman filter unit is connected to the entropy fusion unit to perform frequency-by-frequency Kalman filtering in the short-time Fourier transform domain; The residual convolution unit is connected to the frequency domain Kalman filter unit and the entropy fusion unit to achieve detail recovery and signal equalization through a convolutional network with fixed weights; The control module is connected to the line spectrum recognition unit, notch filtering unit, wavelet thresholding unit, spectrum compensation unit, entropy fusion unit, frequency domain Kalman filtering unit, and residual convolution unit, respectively, and is used to coordinate the timing processing and dynamically adjust system parameters.

[0026] The proposed multi-domain fusion speech denoising method is based on a cascaded processing system constructed using adaptive notch filtering, wavelet thresholding, OM-LSA spectral compensation, frequency-domain Kalman filtering, and residual CNN. This multi-domain fusion speech denoising system comprises three core parts: a line spectrum suppression module, a broadband noise suppression module, and a detail restoration module. First, adaptive notch filtering effectively suppresses line spectrum interference. Then, an adaptive thresholding method based on the wavelet coefficient energy ratio η, combined with OM-LSA spectral compensation, deeply suppresses broadband noise. Next, a frequency-domain three-dimensional Kalman filter is used to achieve robust recursive estimation in the time and frequency domains. Finally, a lightweight residual CNN is used to accurately restore speech details and achieve spectral equalization.

[0027] This invention combines high-gain noise reduction with low-distortion fidelity in complex noise environments, making it suitable for speech enhancement needs in various complex acoustic scenarios.

[0028] A computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of a multi-domain fusion speech denoising method, system, and computer-readable storage medium as described in any one of claims 1 to 8.

[0029] This invention provides a multi-domain fusion speech denoising method, system, and computer-readable storage medium, which have the following beneficial effects: 1. This multi-domain fusion speech denoising method, system, and computer-readable storage medium possess five core advantages compared to traditional technologies. It achieves effective suppression of multi-frequency / converted-frequency line spectra and broadband noise across the entire frequency band through multi-mechanism collaboration; validated with 40 sets of random samples, it achieves an objective performance improvement of an average ΔSNR of 12.39±2.48dB, significantly surpassing traditional single-strategy methods; it adopts a completely training-free design concept, offering superior deployment lightweightness and engineering interpretability compared to deep learning solutions such as DCCRN / Conv-TasNet; it supports real-time processing rhythms of 32ms frame length and 8ms step size, meeting the demands of latency-sensitive application scenarios; and it provides... Multi-dimensional visualization analysis tools, including adaptive heatmaps, Δ spectra, and line spectrum energy slices, ensure that the technical process is completely transparent and analyzable.

[0030] 2. This multi-domain fusion speech denoising method, system, and computer-readable storage medium achieve rapid suppression of narrowband line spectrum interference through spectral peak detection and adaptive notch filtering. It effectively eliminates broadband background noise by using wavelet coefficient energy ratio η adaptive threshold algorithm and OM-LSA spectral compensation technology. It completes recursive robust estimation in the time-frequency domain by using frequency domain three-dimensional Kalman filtering. Finally, it uses a lightweight residual CNN network that does not require training weights to achieve signal detail reconstruction and spectral equalization, thereby achieving the best balance between high gain and low distortion in complex noise environments. Attached Figure Description

[0031] Figure 1 This is a schematic diagram of the original signal input to this invention; Figure 2 This is a signal input diagram (with noise added) of the signal receiver of the present invention. Figure 3 This is a schematic diagram of the waveform after the first stage of signal processing in this invention; Figure 4 This is a schematic diagram of the waveform after the second stage of signal processing in this invention; Figure 5This is a schematic diagram of the waveforms after the third and fourth stages of signal processing in this invention; Figure 6 This is a schematic diagram of the energy attenuation spectrum of the system signal of the present invention; Figure 7 This is a schematic diagram of the frequency band through which the signal of this invention obtains the main signal-to-noise ratio gain after passing through a noise reduction system; Figure 8 This is a schematic diagram of the Q / R adaptive dynamic adjustment of the Kalman filter of the present invention; Figure 9 This is a visualization diagram of the signal-to-noise ratio of the present invention under 40 sets of signal samples; Figure 10 This is a schematic diagram of the method flow of the present invention. Detailed Implementation

[0032] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0033] Please see Figures 1 to 10 The present invention provides a technical solution: a multi-domain fusion speech noise reduction method, comprising the following steps: S1: Perform power spectral density estimation on the input signal, and construct a comprehensive score by combining peak amplitude ratio, adjacent peak contrast, noise peak ratio and local power spectrum smoothness to identify candidate line spectrum frequencies and their corresponding bandwidths; S2: Perform adaptive notch filtering based on the identified candidate frequencies, and use the correlation coefficient to root mean square amplitude ratio as the iteration termination condition; S3: Perform wavelet decomposition on the notch-filtered signal, adaptively determine the threshold and retention coefficients based on the energy distribution of each decomposition layer, and perform threshold processing and reconstruction on the detail coefficients and approximation coefficients respectively. S4: The spectrum compensation of the reconstructed signal after S3 is performed using an improved logarithmic spectrum amplitude estimation method; S5: Based on the spectral energy entropy of the notch-filtered signal in S2 and the spectrum-compensated signal in S4, perform adaptive weighted fusion on the two; S6: In the short-time Fourier transform domain, Kalman filtering is performed on the signal fused by S5 at frequency points and frame by frame. The process noise variance and measurement noise variance are dynamically adjusted according to wavelet energy interpolation, spectral gradient and instantaneous signal-to-noise ratio. S7: Calculate the residual between the output signal in S6 and the output signal in S5, and use a lightweight convolutional network with fixed weights to perform detail restoration and signal equalization on the residual to obtain the final denoised signal.

[0034] The overall score in S1 is calculated using the following formula:

[0035] in, This indicates the overall score. This represents the standardized peak amplitude ratio. This represents the ratio of adjacent peak values ​​after standardization. This represents the normalized peak noise ratio. This represents the smoothness of the local power after standardization.

[0036] During the online spectrum screening process, a peak height threshold and a minimum frequency interval threshold are set simultaneously to determine the effective line spectrum components from the candidate frequencies.

[0037] The notch filtering in S2 uses an infinite impulse response notch filter, and the bandwidth scaling factor of the notch filter is dynamically adjusted in the range of 1.0 to 1.3.

[0038] The number of wavelet decomposition layers in S3 is determined by the following formula:

[0039] Where L is the wavelet decomposition level, sam is the sampling rate of the input signal, and the wavelet basis function used in the wavelet decomposition is coif4 or sym4.

[0040] In S3, a threshold is set for the wavelet coefficients of the j-th layer. The calculation method is as follows

[0041] in, This is the threshold scaling factor, used to adjust the overall threshold strength. A value greater than 1 indicates enhanced noise reduction strength, while a value less than 1 indicates weakened noise reduction strength. The noise standard deviation of the wavelet coefficients at the j-th level is estimated based on the median absolute deviation of the wavelet coefficients. denoted as the number of wavelet coefficients in the j-th layer.

[0042] At the same time, the retention coefficient of the j-th layer is set. The calculation method is as follows:

[0043] in, To preserve the upper limit of the retention coefficient and protect the integrity of the signal structure, the value ranges from 0.95 to 0.98; The base retention factor is used to avoid excessive distortion of the signal in the high-noise layer or the introduction of artifacts, and its value ranges from 0.1 to 0.3. This is the energy gain factor, used to adaptively adjust the noise reduction intensity according to the energy ratio of each layer, with a value range of 8 to 12. The energy proportion of the wavelet coefficients at the j-th level is calculated as follows:

[0044] in, The total energy of the wavelet coefficients at the j-th level is calculated using the following formula:

[0045] The total energy of all wavelet coefficients is calculated using the following formula:

[0046] The fusion weight in S5 The calculation method is as follows:

[0047] Where H represents the Shannon entropy function of the average spectral energy, This is the output signal after notch filtering in S2. This is the output signal after spectral compensation in S4.

[0048] The process noise variance Q and measurement noise variance R in S6 are set as follows:

[0049] in,

[0050] In the formula, This is the baseline value for process noise. To measure the noise reference value; The parameter is estimated through the spectral energy distribution, representing the proportion of energy in each frequency band to the total signal energy; The signal energy of the current frame is given by `prevE`, and the signal energy of the previous frame is given by `prevE`. P represents the process noise power; The coefficients, determined by the optimization algorithm, are used to adjust the intensity of the influence of each physical quantity on Q and R.

[0051] The residual convolutional network in S7 consists of 2 to 5 layers of one-dimensional convolutions, with each layer having a kernel length of 3 to 7. The network structure includes a residual gating mechanism, and all network parameters remain fixed during deployment and use, requiring no training optimization.

[0052] This multi-domain fusion speech denoising method is applicable to speech enhancement processing in the following application scenarios: vehicle-mounted speech acquisition systems, aviation communication equipment, industrial human-machine interfaces, remote conferencing systems, intercom equipment, and embedded speech processing front-ends.

[0053] The multi-domain fusion speech denoising method sets the frame length to 32 milliseconds and the frame shift to 8 milliseconds during processing, and the end-to-end processing delay of a single frame of data does not exceed 20 milliseconds.

[0054] This multi-domain fusion speech denoising method achieves rapid suppression of narrowband line spectrum interference through spectral peak detection and adaptive notch filtering. It effectively eliminates broadband background noise by using wavelet coefficient energy ratio η adaptive threshold algorithm and OM-LSA spectral compensation technology. It uses frequency domain three-dimensional Kalman filtering to complete recursive robust estimation in the time and frequency domains. Finally, it uses a lightweight residual CNN network that does not require training weights to achieve signal detail reconstruction and spectral equalization, thereby achieving the best balance between high gain and low distortion in complex noise environments.

[0055] A multi-domain fusion speech denoising system is provided for use in a multi-domain fusion speech denoising method. The system consists of a line spectrum recognition unit, a notch filtering unit, a wavelet thresholding unit, a spectrum compensation unit, an entropy fusion unit, a frequency domain Kalman filtering unit, a residual convolution unit, and a control module. The line spectrum recognition unit is used to detect the candidate line spectrum frequencies and their bandwidths in the input signal. The notch processing unit is connected to the line spectrum recognition unit and is used to perform adaptive notch filtering on the candidate line spectrum frequencies. The wavelet thresholding unit is connected to the notch processing unit and is used to perform wavelet decomposition and adaptive thresholding on the notch-filtered signal. The spectrum compensation unit is connected to the wavelet threshold processing unit and is used to perform logarithmic spectrum amplitude compensation on the reconstructed signal. The entropy fusion unit is connected to the notch processing unit and the spectrum compensation unit respectively, and is used to perform adaptive weighted fusion of the two signals based on the spectral energy entropy. The frequency domain Kalman filter unit is connected to the entropy fusion unit and is used to perform frequency-by-frequency Kalman filtering in the short-time Fourier transform domain. The residual convolution unit is connected to the frequency domain Kalman filter unit and the entropy fusion unit, and is used to achieve detail recovery and signal equalization through a convolutional network with fixed weights. The control module is connected to the line spectrum recognition unit, notch filtering unit, wavelet thresholding unit, spectrum compensation unit, entropy fusion unit, frequency domain Kalman filtering unit, and residual convolution unit, respectively, and is used to coordinate the timing processing and dynamically adjust system parameters.

[0056] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is used to implement a multi-domain fusion speech noise reduction method.

[0057] like Figure 1The diagram illustrates the time-frequency distribution of the input speech signal to the noise reduction system of this invention. This signal consists of speech components superimposed with narrowband line spectrum interference and broadband background noise. The horizontal axis represents time, the vertical axis represents frequency, and the color intensity reflects the power spectral density. Figure 1 As can be seen, there are obvious and stable bright stripes in the frequency bands of 1kHz, 2kHz, 4kHz and 7kHz, which are typical line spectrum interference characteristics; at the same time, the background part shows a continuous random distribution, indicating the presence of broadband noise components. The spectrum intuitively shows the complex noise scene that needs to be processed by the present invention.

[0058] like Figure 2 The image shown is the spectrum after Stage A adaptive notch filtering. In this stage, a multi-feature joint detection mechanism, including peak amplitude ratio, harmonic ratio, neighborhood spectrum ratio, and local spectrum smoothness, is used to accurately identify the line spectrum frequency and adaptively set the notch filter parameters. like Figure 2 As shown, the original high-brightness narrowband interference was significantly suppressed, while the main speech structure remained intact, proving that the main line spectrum components were effectively filtered out at this stage.

[0059] like Figure 3 As shown, the spectral effect after StageB wavelet adaptive threshold denoising is demonstrated, achieved through multi-scale wavelet decomposition combined with energy proportion-based denoising. Adaptive threshold With retention factor Through the adjustment mechanism, random noise and residual line spectrum were further suppressed in this stage. It can be observed that the fine-grained noise in the spectrum is significantly reduced and the speech formant structure is clearer, which reflects the balanced advantage of energy adaptive threshold in terms of detail preservation and noise suppression.

[0060] like Figure 4 The image shown is a spectrum of the intermediate result after OM-LSA spectral compensation and entropy-guided Wiener fusion processing in Stage B. and Figure 3 In comparison, the speech formant features are more prominent, the turbidity in the low-frequency region is significantly reduced, and the overall speech clarity is improved. This indicates that the strategy of using least mean square logarithmic spectral amplitude estimation and information entropy weighted fusion effectively enhances the smoothness and clarity of the signal at the power spectral density level.

[0061] like Figure 5 As shown, the time-spectral graph of ideal clean speech is presented as a benchmark reference for evaluating the noise reduction effect. Figure 5 There are no obvious fixed frequency line spectrum components, and the speech harmonic structure is complete and clear, providing an intuitive basis for comparison of the effects of subsequent processing at each level.

[0062] like Figure 6As shown, the spectrum difference diagram of the noise reduction system of the present invention between the output and input signals is displayed. The red area represents energy enhancement and the blue area represents energy reduction. It can be seen that continuous blue bands appear in the line spectrum interference frequency bands such as 1kHz, 2kHz, 4kHz and 7kHz, indicating that the interference energy in these frequency bands is effectively attenuated. However, no significant negative gain appears in the main speech energy region, proving that the system achieves directional suppression of narrowband interference while maintaining speech integrity.

[0063] like Figure 7 As shown in the figure, the average gain variation at each line spectrum frequency point is displayed in the form of a bar chart. The data shows that energy reduction of -30dB to -40dB was achieved in the main interference frequency bands such as 1kHz, 2kHz, 4kHz, and 7kHz. This indicates that the present invention has achieved stable and significant suppression effects at key interference frequencies, providing a reliable quantitative basis for the technical performance.

[0064] like Figure 8 As shown, the process noise during the Stage C frequency domain Kalman filtering process is displayed. With observation noise The adaptive parameter distribution heatmap is shown in the figure. In regions with abrupt changes in signal energy or large spectral gradients, The corresponding increase is to improve the tracking response speed; while in the high signal-to-noise ratio region, The value is then appropriately reduced to enhance the reliability of the observed data. This dynamic adjustment characteristic aligns with the design expectation of the following formula, ensuring the filter's rapid response and stable tracking capability in time-varying noise environments:

[0065] like Figure 9 The figure shows a scatter plot of batch test results on 40 randomly generated samples of "narrowband line spectrum interference + broadband background noise". The horizontal axis represents the center frequency of the line spectrum, the vertical axis represents the signal-to-noise ratio (SNR) gain, and the color intensity represents the bandwidth. The test results show that under different frequency and bandwidth combinations, this invention can achieve an SNR improvement of 8dB to 18dB, with an average gain of approximately 12dB, demonstrating the algorithm's good robustness and adaptability in diverse noise scenarios.

[0066] This invention relates to a multi-domain fusion speech denoising method, system, and computer-readable storage medium. When used, taking a 16kHz sampling rate as an example... Step 1: Line spectrum candidate detection The signal power spectral density is calculated, and a multi-feature evaluation system including peak amplitude ratio, harmonic ratio, neighborhood spectral ratio, and local spectral smoothness is constructed. A weighted comprehensive scoring model is then adopted.

[0067] Candidate line spectrum frequencies are screened by combining peak height and minimum frequency spacing threshold, and the local bandwidth of each candidate frequency is estimated.

[0068] Step 2: Adaptive Notch Filtering An IIR notch filter is designed for each candidate frequency to support adaptive bandwidth adjustment. An enhanced iterative mechanism is introduced, with the bandwidth adjustment range being 1.0-1.3 times. Finally, the correlation coefficient and the root mean square amplitude ratio are used as fidelity constraints to prevent excessive suppression of speech components.

[0069] Step 3: Wavelet adaptive thresholding denoising First, the number of decomposition levels is adaptively determined:

[0070] Then, the coif4 or sym4 wavelet basis function is selected, and the noise estimate is calculated based on the statistical characteristics of the coefficients at each level:

[0071] Next, we construct an adaptive threshold model:

[0072] The system uses energy percentage Dynamically adjust retention factor This achieves a trade-off between soft and hard thresholds:

[0073] Step 4: Improve spectral compensation The noise spectral density and posterior or prior signal-to-noise ratio are estimated, and then the frequency domain gain is calculated based on the minimum distortion criterion and the probability of speech presence, while setting a lower limit for the gain. Maintain the integrity of the audio.

[0074] Step 5: Entropy-guided adaptive fusion Calculate the spectral energy distribution entropy value of each path output, and use an entropy weight allocation strategy:

[0075] Use the above formula to achieve optimal fusion.

[0076] Step 6: Frequency Domain Kalman Filtering Implement frequency-by-frequency Kalman recursive estimation in the STFT domain and establish an adaptive noise variance model:

[0077] Using the standard Kalman update equation with the above formula, a visualization of the parameter adaptation process is finally output.

[0078] Step 7: Residual Convolution Refinement First, the residual signals of the Kalman output and the fused output are constructed. Then, a first-order high-pass preprocessing is used to enhance the detail response. Next, a fixed-weight lightweight convolutional network combined with residual gating is used for detail recovery. Finally, high-frequency equalization is introduced to improve speech intelligibility, and the final output is given.

[0079] This invention establishes a complete parameter optimization system, in which the line spectrum detection module includes constraints such as a peak height threshold of 1.8-3.0 and a minimum peak distance of 100-400Hz; the wavelet processing module sets a threshold scaling factor of 20-60, a base value of 0.05-0.25 for the retention factor, a gain factor of 5-15, and an upper limit of 0.90-0.995 for the dynamic range; the spectrum compensation module limits the working interval of the minimum gain to 0.08-0.20; the Kalman filter module configures core parameters of a base process noise of 0.002-0.01 and a base measurement noise of 0.8-1.6; the residual fusion module adopts a weight coefficient of 0.1-0.6; and the network structure module designs a compact architecture with a convolution kernel length of 3-7 and a network layer number of 2-5.

[0080] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A multi-domain fusion speech denoising method, characterized in that: Includes the following steps: S1: Perform power spectral density estimation on the input signal, and construct a comprehensive score by combining peak amplitude ratio, adjacent peak contrast, noise peak ratio and local power spectrum smoothness to identify candidate line spectrum frequencies and their corresponding bandwidths; S2: Perform adaptive notch filtering based on the identified candidate frequencies, and use the correlation coefficient to root mean square amplitude ratio as the iteration termination condition; S3: Perform wavelet decomposition on the notch-filtered signal, adaptively determine the threshold and retention coefficients based on the energy distribution of each decomposition layer, and perform threshold processing and reconstruction on the detail coefficients and approximation coefficients respectively. S4: The spectrum compensation of the reconstructed signal after S3 is performed using an improved logarithmic spectrum amplitude estimation method; S5: Based on the spectral energy entropy of the notch-filtered signal in S2 and the spectrum-compensated signal in S4, perform adaptive weighted fusion on the two; S6: In the short-time Fourier transform domain, Kalman filtering is performed on the signal fused by S5 at frequency points and frame by frame. The process noise variance and measurement noise variance are dynamically adjusted according to wavelet energy interpolation, spectral gradient and instantaneous signal-to-noise ratio. S7: Calculate the residual between the output signal in S6 and the output signal in S5, and use a lightweight convolutional network with fixed weights to perform detail restoration and signal equalization on the residual to obtain the final denoised signal.

2. The multi-domain fusion speech denoising method according to claim 1, characterized in that: The overall score in S1 is calculated using the following formula: ; in, This indicates the overall score. This represents the standardized peak amplitude ratio. This represents the ratio of adjacent peak values ​​after standardization. This represents the normalized peak noise ratio. This represents the smoothness of the local power after standardization. During the online spectrum screening process, a peak height threshold and a minimum frequency interval threshold are set simultaneously to determine the effective line spectrum components from the candidate frequencies.

3. The multi-domain fusion speech denoising method according to claim 1, characterized in that: The notch filtering in S2 uses an infinite impulse response notch filter, and the bandwidth scaling factor of the notch filter is dynamically adjusted in the range of 1.0 to 1.

3.

4. The multi-domain fusion speech denoising method according to claim 1, characterized in that: The number of wavelet decomposition layers in S3 is determined by the following formula: ; Where L is the wavelet decomposition level, sam is the sampling rate of the input signal, and the wavelet basis function used in the wavelet decomposition is coif4 or sym4.

5. The multi-domain fusion speech denoising method according to claim 1, characterized in that: In S3, a threshold is set for the wavelet coefficients of the j-th layer. The calculation method is as follows ; in, This is the threshold scaling factor, used to adjust the overall threshold strength. A value greater than 1 indicates enhanced noise reduction strength, while a value less than 1 indicates weakened noise reduction strength. The noise standard deviation of the wavelet coefficients at the j-th level is estimated based on the median absolute deviation of the wavelet coefficients. The number of wavelet coefficients in the j-th layer. At the same time, the retention coefficient of the j-th layer is set. The calculation method is as follows: ; in, To preserve the upper limit of the retention coefficient and protect the integrity of the signal structure, the value ranges from 0.95 to 0.98; The base retention factor is used to avoid excessive distortion of the signal in the high-noise layer or the introduction of artifacts, and its value ranges from 0.1 to 0.

3. This is the energy gain factor, used to adaptively adjust the noise reduction intensity according to the energy ratio of each layer, with a value range of 8 to 12. The energy proportion of the wavelet coefficients at the j-th level is calculated as follows: ; in, The total energy of the wavelet coefficients at the j-th level is calculated using the following formula: ; The total energy of all wavelet coefficients is calculated using the following formula: 。 6. The multi-domain fusion speech denoising method according to claim 1, characterized in that: The fusion weight in S5 The calculation method is as follows: ; Where H represents the Shannon entropy function of the average spectral energy, This is the output signal after notch filtering in S2. This is the output signal after spectral compensation in S4.

7. The multi-domain fusion speech denoising method according to claim 1, characterized in that: The process noise variance Q and measurement noise variance R in S6 are set as follows: ; in, ; In the formula, This is the baseline value for process noise. To measure the noise reference value; The parameter is estimated through the spectral energy distribution, representing the proportion of energy in each frequency band to the total signal energy; The signal energy of the current frame is given by `prevE`, and the signal energy of the previous frame is given by `prevE`. P represents the process noise power; The coefficients, determined by the optimization algorithm, are used to adjust the intensity of the influence of each physical quantity on Q and R.

8. The multi-domain fusion speech denoising method according to claim 1, characterized in that: The residual convolutional network in S7 consists of 2 to 5 layers of one-dimensional convolutions, with each layer having a kernel length of 3 to 7. The network structure includes a residual gating mechanism, and all network parameters remain fixed during deployment and use, requiring no training optimization.

9. A multi-domain fusion speech denoising system, used to implement the multi-domain fusion speech denoising method according to any one of claims 1 to 8, characterized in that: It includes a line spectrum recognition unit, a notch filtering unit, a wavelet thresholding unit, a spectrum compensation unit, an entropy fusion unit, a frequency domain Kalman filter unit, a residual convolution unit, and a control module; The line spectrum recognition unit is used to detect candidate line spectrum frequencies and their bandwidths in the input signal; The notch processing unit is connected to the line spectrum recognition unit and is used to perform adaptive notch filtering on the candidate line spectrum frequencies. The wavelet thresholding unit is connected to the notch processing unit and is used to perform wavelet decomposition and adaptive thresholding on the notch-filtered signal. The spectrum compensation unit is connected to the wavelet threshold processing unit and is used to perform logarithmic spectrum amplitude compensation on the reconstructed signal; The entropy fusion unit is connected to the notch processing unit and the spectrum compensation unit respectively, and is used to perform adaptive weighted fusion of the two signals based on the spectral energy entropy. The frequency domain Kalman filter unit is connected to the entropy fusion unit to perform frequency-by-frequency Kalman filtering in the short-time Fourier transform domain; The residual convolution unit is connected to the frequency domain Kalman filter unit and the entropy fusion unit to achieve detail recovery and signal equalization through a convolutional network with fixed weights; The control module is connected to the line spectrum recognition unit, notch filtering unit, wavelet thresholding unit, spectrum compensation unit, entropy fusion unit, frequency domain Kalman filtering unit, and residual convolution unit, respectively, and is used to coordinate the timing processing and dynamically adjust system parameters.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of a multi-domain fusion speech denoising method as described in any one of claims 1 to 8.