A photoacoustic signal enhancement method and device based on adaptive multi-band spectral subtraction and wiener filtering
By using adaptive multi-band spectral subtraction and Wiener filtering, the sub-band over-subtraction factor and lower limit factor are dynamically adjusted. Combined with inter-sub-band smoothing and Wiener filtering, the problems of noise residue and speech distortion in photoacoustic signal processing are solved, achieving a higher signal-to-noise ratio and clearer speech signal recovery.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-24
- Publication Date
- 2026-07-24
AI Technical Summary
Existing technologies in photoacoustic signal processing suffer from severe musical noise, poor adaptability to non-stationary noise, and easy damage to speech under high signal-to-noise ratios. In particular, in fixed-parameter multi-band subtraction, it is impossible to make differentiated adjustments according to changes in signal-to-noise ratio, resulting in noise residue and speech distortion.
An adaptive multi-band spectral subtraction and Wiener filtering method is adopted. After framing, windowing and short-time Fourier transform, subbands are divided according to the Mel scale. Combined with speech activity detection and recursive averaging to update the noise spectrum estimation, the subband over-subtraction factor and lower limit factor are dynamically adjusted to perform adaptive spectral subtraction. Noise is suppressed by inter-subband smoothing and Wiener filtering to protect speech components.
It effectively solves the contradiction between over-subtraction distortion and residual noise caused by noise estimation bias in traditional methods, and achieves a balance between denoising and speech component preservation under different signal-to-noise ratio conditions, thereby improving the signal-to-noise ratio and auditory clarity.
Smart Images

Figure CN122454993A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech signal enhancement technology, and in particular to a photoacoustic signal enhancement method and apparatus based on adaptive multi-band spectral subtraction and Wiener filtering. Background Technology
[0002] When acquiring speech signals based on the principle of optical interference, the signal-to-noise ratio of the demodulated speech signal is severely reduced due to the superposition of various distortion factors such as system noise, background light noise, and environmental interference during the demodulation process. How to remove noise while avoiding the introduction of musical noise and protecting the speech components from damage has always been the core technical bottleneck restricting the practical application of photoacoustic speech acquisition.
[0003] Currently, the industry generally adopts a multi-band spectral subtraction-based technical solution when handling the aforementioned speech enhancement tasks. The specific workflow of this solution is as follows: First, the noisy time-domain speech signal is framed, windowed, and converted to a frequency-domain signal using a short-time Fourier transform. Then, the entire frequency-domain signal is divided into several sub-bands, and the amplitude spectrum of the noise is independently estimated within each sub-band to obtain the noisy speech amplitude spectrum for that sub-band. Next, based on a preset fixed over-subtraction factor, the estimated noise amplitude spectrum is subtracted from the noisy speech amplitude spectrum to obtain the enhanced amplitude spectrum of each sub-band. For spectral components with negative amplitudes after subtraction, they are directly set to zero, i.e., half-wave rectification is applied. Finally, the enhanced amplitude spectrum is combined with the original noisy speech phase, and the time-domain speech signal is reconstructed through inverse short-time Fourier transform and overlap addition. The core of this solution lies in using a fixed over-subtraction factor to control the noise subtraction intensity of each sub-band and relying on half-wave rectification to process the negative spectrum generated after spectral subtraction.
[0004] However, the above-mentioned scheme has insurmountable technical shortcomings in practical applications. First, since the over-subtraction factor remains constant throughout the processing, when the input signal-to-noise ratio (SNR) is low, the preset over-subtraction strength is often insufficient, resulting in a large amount of broadband noise remaining in the output speech. Conversely, when the input signal-to-noise ratio is high, the fixed over-subtraction strength is too large, which will subtract the weak energy components of the speech along with the noise, causing speech distortion. This contradiction stems from the fact that the fixed parameters cannot be adjusted differently according to the real-time SNR changes of each subband. Second, directly applying half-wave rectification to zero the negative value part after spectrum subtraction will randomly generate isolated spectral peaks and valleys on each subband spectrum. These unnaturally generated spectral structures appear as perceptible musical noise after reconstructing the time-domain signal, severely impairing the listening experience of the speech. Half-wave rectification is the only means for this scheme to process negative spectrum values, and its mathematical core determines that such musical noise cannot be eliminated by simply adjusting parameters. When noise characteristics change rapidly or the signal-to-noise ratio fluctuates drastically, the noise spectrum estimation of this scheme uses a recursive average with a fixed update rate, which is prone to lagging behind the actual noise changes and is also prone to being contaminated by speech components during speech activity segments. This leads to inaccuracy of the basic reference quantity of the entire multi-band spectrum reduction framework, further amplifying the aforementioned noise residue and speech distortion problems. Summary of the Invention
[0005] To address the aforementioned problems, this invention provides a photoacoustic signal enhancement method and apparatus based on adaptive multi-band spectral subtraction and Wiener filtering. The invention aims to solve the problems of severe music noise, poor adaptability to non-stationary noise, and easy damage to speech under high signal-to-noise ratio conditions in existing fixed-parameter multi-band spectral subtraction methods during photoacoustic signal processing.
[0006] To achieve the above objectives, the technical solution adopted by the present invention is: a photoacoustic signal enhancement method based on adaptive multi-band spectral subtraction and Wiener filtering, comprising: acquiring the original speech signal, framing, windowing and short-time Fourier transform to obtain the original frequency domain signal; The original frequency domain signal is divided into multiple continuous sub-bands according to the Mel scale to obtain sub-band signals; For each sub-band speech activity detection, the noise amplitude spectrum estimate is updated recursively by averaging in the speech inactive segment, and the noise power spectrum estimate is obtained by squaring. The sub-band signal-to-noise ratio is calculated based on the sub-band signal power spectrum and the noise power spectrum estimate. Based on the subband signal-to-noise ratio, the subband over-decrease factor and lower limit factor are determined by a monotonically decreasing function, both of which are negatively correlated with the subband signal-to-noise ratio. For the current sub-band and all sub-bands before and after it M A smoothed power spectrum is obtained by weighted moving average of the power spectra of adjacent sub-bands. M To smooth the neighborhood radius; An adaptive spectral base function containing the lower limit factor is used to perform spectral subtraction on the smooth power spectrum and the noise power spectrum estimation, and adaptively fill the spectral valleys to obtain the enhanced speech power spectrum; Calculate the maximum noise residual threshold, and limit the power spectrum of the enhanced speech to obtain pre-enhanced frequency domain data; Using the pre-enhanced frequency domain data as a clean signal power spectrum estimate, and combining it with the noise power spectrum estimate Wiener filter, the filtered frequency domain signal is obtained. The filtered frequency domain signal is then subjected to inverse short-time Fourier transform and overlap addition to synthesize an enhanced time-domain speech signal.
[0007] Preferably, the sub-band over-subtraction factor α i Determined by the following piecewise function: in, SNR i For the first i Subband signal-to-noise ratio of each subband.
[0008] Preferably, the lower limit factor β i Determined by the following piecewise function: in, SNR i For the first i Subband signal-to-noise ratio of each subband.
[0009] Preferably, the enhanced speech power spectrum | X i ( oh k )| 2 Determined by the following adaptive spectral base function: in, For the first i Subband at frequency oh k The smoothed power spectrum at that location, For the noise power spectrum estimation, α i The subband over-subtraction factor, d i This is the subband frequency fine-tuning factor. β i The lower limit factor is... To predict the power spectrum of the enhanced signal.
[0010] Preferably, the noise amplitude spectrum estimation The update formula is: in, For the current frame t No. i Subband at frequency oh k Noise amplitude spectrum estimation at the location, For noise amplitude spectrum estimation of the previous frame, | Y i ( oh k , t | represents the subband signal amplitude spectrum of the current frame. c The forgetting factor ranges from 0 to 1; the square of the noise amplitude spectrum estimation That is, the noise power spectrum estimation. .
[0011] Preferably, the sub-band frequency fine-tuning factor is determined by the following piecewise function: in, f For the first i The center frequency of the sub-band fs Sampling rate, d i This is the subband frequency fine-tuning factor.
[0012] Preferably, the smooth neighborhood radius M The value is 2, which means that the power spectrum of the current sub-band and the two adjacent sub-bands before and after it, a total of 5 sub-bands, is weighted moving average.
[0013] A photoacoustic signal enhancement device, employing the aforementioned photoacoustic signal enhancement method, includes: a signal acquisition module for acquiring raw speech signals; The time-frequency conversion module is used for framing, windowing, and short-time Fourier transform to obtain the original frequency domain signal; The sub-band division module is used to divide the signal into multiple continuous sub-bands according to the Mel scale to obtain the sub-band signal. The noise estimation and signal-to-noise ratio calculation module is used for speech activity detection. It recursively averages and adaptively updates the noise amplitude spectrum estimate in the speech inactive segment, and squares it to obtain the noise power spectrum estimate. It calculates the subband signal-to-noise ratio based on the subband signal power spectrum and the noise power spectrum estimate. The parameter determination module is used to determine the subband over-reduction factor and the lower limit factor based on the subband signal-to-noise ratio using a monotonically decreasing function. Both factors are negatively correlated with the subband signal-to-noise ratio. The sub-band smoothing module is used to smooth the current sub-band and the sub-bands before and after it. M A smoothed power spectrum is obtained by weighted moving average of the power spectra of adjacent sub-bands. MTo smooth the neighborhood radius; An adaptive enhancement module is used to perform spectral subtraction on the smooth power spectrum and the noise power spectrum estimate using an adaptive spectral base function that includes the lower limit factor, and adaptively fill the spectral valleys to obtain an enhanced speech power spectrum. The residual limiting module is used to calculate the maximum noise residual threshold and limit the power spectrum of the enhanced speech to obtain pre-enhanced frequency domain data; The post-filtering processing module is used to use the pre-enhanced frequency domain data as a clean signal power spectrum estimate, and combine it with the noise power spectrum estimate to perform Wiener filtering to obtain the filtered frequency domain signal. The signal reconstruction module is used to synthesize an enhanced time-domain speech signal by performing inverse short-time Fourier transform and overlap addition on the filtered frequency domain signal.
[0014] Preferably, the adaptive spectral base function used by the adaptive enhancement module is: in, For the first i Subband at frequency oh k The smoothed power spectrum at that location, For the noise power spectrum estimation, α i The subband over-subtraction factor, d i This is the subband frequency fine-tuning factor. β i The lower limit factor is... To predict the power spectrum of the enhanced signal, | X i ( oh k )| 2 To enhance the speech power spectrum.
[0015] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described photoacoustic signal enhancement method.
[0016] By adopting the above technical solution, the present invention has the following beneficial effects.
[0017] (1) This invention addresses the problems of severe music noise, poor adaptability to non-stationary noise, and easy damage to speech under high signal-to-noise ratio in traditional spectral subtraction for photoacoustic signal enhancement. By constructing an adaptive spectral subtraction framework in which the subband oversubtraction factor and the lower limit factor are dynamically linked with the signal-to-noise ratio, and combining inter-subband smoothing and two-stage Wiener filtering, the contradiction between oversubtraction distortion and residual noise caused by noise estimation bias is fundamentally solved.
[0018] (2) After dividing the signal into sub-bands according to the Mel scale, this invention determines the over-subtraction factor and the lower limit factor of each sub-band simultaneously through a monotonically decreasing function based on the real-time signal-to-noise ratio (SNR) of each sub-band. Both factors are negatively correlated with the SNR. The lower the SNR, the greater the noise reduction and the more fully the spectral valleys are filled. The higher the SNR, the lower the values of both factors automatically decrease to avoid damaging the speech components. This linkage mechanism enables the noise reduction intensity and the filling protection intensity to be matched in real time under each sub-band and each SNR condition. The reason for this is that the over-subtraction factor controls the amount of noise reduction, and the lower limit factor controls the filling depth of the negative spectral valleys after subtraction. The two factors work together back-to-back, solving the dilemma of residual noise at low SNR and speech distortion at high SNR in traditional methods.
[0019] (3) This invention performs a weighted moving average of the power spectra of adjacent sub-bands. This smoothing operation fuses the current sub-band processing result with the information of the two adjacent sub-bands before and after it along the frequency axis, smoothing out the isolated spectral peaks and random spectral valleys left after the independent spectrum subtraction of each sub-band. Isolated spectral peaks and valleys are the frequency domain roots of musical noise. The smoothing operation utilizes the power spectrum processed by an adaptive factor. The smoothing intensity is adapted to the current noise reduction state, suppressing residual noise while preserving the speech details at the sub-band boundaries.
[0020] (4) This invention adds a residual limiting step before Wiener filtering. First, the maximum noise residual threshold is calculated based on the cumulative frequency point of the enhanced speech power spectrum of the current frame and several frames before and after it. Then, the enhanced speech power spectrum is limited using this threshold. The limiting forces the frequency point power to be constrained below the statistical upper limit of the residual noise energy, thus removing the random spikes remaining after spectral subtraction. The pre-enhanced frequency domain data after limiting is used as the clean signal power spectrum estimate for Wiener filtering, so that the filter can more accurately fit the clean speech spectrum envelope, enhance speech details in the high frequency band, and suppress residual broadband noise in the low frequency band.
[0021] (5) This invention introduces an adaptive recursive update linked to the signal-to-noise ratio (SNR) in the noise spectrum estimation stage. During inactive speech segments, the forgetting factor is dynamically adjusted based on the current frame's SNR. When the SNR is extremely low, the forgetting factor approaches its upper limit, slowing down the update speed and preventing weak speech components from being misjudged as noise. When the SNR is high, the forgetting factor approaches 1, and the noise spectrum is hardly updated, fully protecting the speech harmonic structure. This mechanism ensures that the noise spectrum estimation always follows the fluctuations of the real noise floor, providing an accurate reference for adaptive oversubtraction and spectral floor filling. Attached Figure Description
[0022] The following provides a detailed discussion of the manufacture and application of preferred embodiments of the present invention. However, it should be understood that the present invention provides many applicable inventive concepts that can be embodied in various specific environments. The specific embodiments discussed are merely illustrative of specific ways of manufacturing and using the present invention and do not limit the scope of the invention. For those skilled in the art, other drawings can be obtained from these drawings without any creative effort.
[0023] Figure 1 This is a flowchart of the method of the present invention.
[0024] Figure 2 This is a comparison diagram of the time-domain waveforms of the present invention and the original photoacoustic signal.
[0025] Figure 3 This is a comparison chart showing the signal-to-noise ratio of the enhanced signal obtained by the present invention and the traditional spectral subtraction method. Detailed Implementation
[0026] The following provides a detailed discussion of the manufacture and application of preferred embodiments of the present invention. However, it should be understood that the present invention provides many applicable inventive concepts that can be embodied in various specific environments. The specific embodiments discussed are merely illustrative of specific ways of manufacturing and using the invention and do not limit the scope of the invention.
[0027] This invention provides a photoacoustic signal enhancement method based on adaptive multi-band spectral subtraction and Wiener filtering. The method acquires the original speech signal using an optical sensing device. The signal is then framed, windowed, and subjected to short-time Fourier transform. In the frequency domain, the signal is divided into multiple continuous sub-bands according to the Mel scale, which approximates the characteristics of human auditory perception. Within each sub-band, based on the real-time estimated sub-band signal-to-noise ratio, a sub-band over-subtraction factor and a lower limit factor are dynamically determined through a monotonically decreasing function. Simultaneously, a sub-band frequency fine-tuning factor is introduced, and adaptive multi-band spectral subtraction is performed. During the spectral subtraction process, the generated spectral valleys are adaptively filled using a lower limit factor, rather than traditional half-wave rectification. A weighted moving average is applied to the power spectra of adjacent sub-bands to suppress isolated spectral peaks and valleys. Then, the maximum noise residual threshold is calculated, and the enhanced speech power spectrum is clipped. The clipped pre-enhanced frequency domain data is then used as a clean signal power spectrum estimate for Wiener filtering. Finally, the enhanced time-domain speech signal is synthesized through inverse short-time Fourier transform and overlapping addition.
[0028] Figure 1 This is a flowchart illustrating a photoacoustic signal enhancement method based on adaptive multi-band spectral subtraction and Wiener filtering, provided as an embodiment of the present invention. Figure 1 As shown, the method specifically includes the following steps: Step S101, acquiring the original speech signal.
[0029] In this step, an optical sensing device is used to detect the minute vibrations on the surface of the target object caused by sound waves based on the principle of optical interference. The optical signal is then converted into an electrical signal to obtain the "raw speech signal" in the time domain, denoted as... y ( n The raw speech signal contains clean speech components and superimposed noise components. The noise sources include system noise, background light noise, and environmental interference. The signal acquired in this step is the basis input for all subsequent frequency domain processing.
[0030] Step S102: The original speech signal is framed, windowed, and subjected to short-time Fourier transform to obtain the original frequency domain signal.
[0031] In this step, the original time-domain speech signal obtained in step S101 is processed. y ( n Preprocessing is performed first. Each frame is segmented, with a length of 20 to 40 milliseconds and a frame shift of 10 to 20 milliseconds. For example, at a sampling rate of 16 kHz, each frame takes 512 samples, and the frame shift takes 256 samples. Then, a Hamming window is applied to each frame to reduce spectral leakage. Afterward, a short-time Fourier transform is performed on each windowed frame to convert the time-domain signal to the frequency domain, obtaining the "original frequency-domain signal". Y ( oh The result of the short-time Fourier transform is a complex number, containing the amplitude spectrum. Y ( oh )| and phase spectrum ∠ Y ( oh ).
[0032] Step S103: Divide the original frequency domain signal into multiple continuous sub-bands according to the Mel scale to obtain the sub-band signal of each sub-band.
[0033] In this step, the Mel scale, which approximates the auditory perception characteristics of the human ear, is used to perform non-uniform sub-band division on the original frequency domain signal obtained in step S102. The Mel scale has high frequency resolution in the low-frequency band and low frequency resolution in the high-frequency band, a characteristic that matches the fact that the energy of speech signals is mainly concentrated in the low-frequency band. Specifically, the entire frequency domain is uniformly segmented in the Mel domain and then mapped back to the linear frequency domain, resulting in N non-overlapping continuous sub-bands. The value of N is determined based on the sampling rate; for example, at a sampling rate of 16kHz, N The typical value range is 20 to 30. For each sub-band i Determine its starting frequency point b i and termination frequency e i Then the subband signal of that subband Yi ( oh k ) represents the original frequency domain signal in the frequency range [ b i , e i The part within ] . Among them, oh k For discrete frequency points, k This is the frequency index. After division, the sub-band signal amplitude spectrum of each sub-band is obtained. Y i ( oh k )|.
[0034] Step S104: Perform speech activity detection on the sub-band signal, and adaptively update the noise amplitude spectrum estimate of each sub-band when a non-active speech segment is detected. Squaring the estimate yields the noise power spectrum estimate, and the sub-band signal-to-noise ratio of each sub-band is calculated.
[0035] This step is the core of adaptive noise estimation. First, "speech activity detection" is performed on each sub-band signal obtained in step S103. The detection method is as follows: if the energy of a sub-band in the current frame is lower than a preset speech presence threshold, the sub-band is determined to be in a speech inactivity segment in the current frame; otherwise, it is determined to be a speech activity segment. This speech presence threshold can be set based on the noise energy statistical characteristics of the initial few frames.
[0036] When a non-active segment of speech is detected in the current frame, an adaptive recursive update of the noise amplitude spectrum estimation is triggered. The update formula is: in, For the current frame t No. i Subband at frequency oh k The noise amplitude spectrum estimate at point is a non-negative real number; The previous frame t -1 corresponds to the noise amplitude spectrum estimation of the sub-band and frequency, which is a recursive term that provides historical noise statistics; | Y i ( oh k , t )| is the current frame t No. i Subband at frequency oh k The sub-band signal amplitude spectrum at that location is the current observation value; c The forgetting factor is a real number ranging from 0 to 1, used to control the update speed of the noise spectrum estimation.
[0037] Forgetting factor c The subband signal-to-noise ratio of the current frame is dynamically determined. When SNR i When less than 0dB, c The value range is from 0.85 to 0.90. When SNR i Between 0dB and 10dB c The value range is from 0.90 to 0.95. When SNR i When it is greater than 10dB, c The value range is from 0.95 to 0.98. In this embodiment, the above three intervals are respectively taken as... c The values are 0.87, 0.93, and 0.97. The mechanism of this dynamic adjustment is: when the signal-to-noise ratio is extremely low, the current frame is dominated by noise, so increasing... c This can reduce the contribution of the current frame to noise estimation, avoiding mistaking weak speech for noise. When the signal-to-noise ratio is high, the current frame is dominated by speech, significantly increasing... c This ensures that the noise spectrum hardly updates, thus protecting the speech harmonic structure.
[0038] When a speech activity segment is detected in the current frame, the noise amplitude spectrum estimation remains unchanged, that is, the noise amplitude spectrum estimation of the current frame uses the estimation value of the previous frame.
[0039] After obtaining the updated noise amplitude spectrum estimate Then, square it to obtain the noise power spectrum estimate. .Right now = .
[0040] Then, the "subband signal-to-noise ratio" is calculated for each subband. SNR i The subband signal-to-noise ratio (SNR) is defined as the logarithm of the ratio of the total power of the subband signal to the total power of the noise within that subband. Where the total power of the subband signal is the power spectrum of the subband signal. Y i ( oh k )| 2 The summation within this sub-band yields the total noise power as an estimate of the noise power spectrum. The summation within this subband. The signal-to-noise ratio of this subband is the direct basis for determining the subband over-subtraction factor and lower limit factor in subsequent steps.
[0041] Step S105: Determine the subband over-attenuation factor and lower limit factor based on the subband signal-to-noise ratio, and determine the subband frequency fine-tuning factor based on the subband frequency range.
[0042] In this step, the signal-to-noise ratio of each sub-band calculated in step S104 is used. SNR iBy using a preset monotonically decreasing function, the two core parameters of the linkage are dynamically determined: the "sub-band over-reduction factor". α i And "lower limit factor" β i Both are negatively correlated with the subband signal-to-noise ratio.
[0043] The subband over-reduction factor α i Determined by the following piecewise function: in, SNR i For the first i The subband signal-to-noise ratio (SNR) is expressed in dB. This function provides the specific principles for determining the subband over-subtraction factor: at extremely low SNR, the subband over-subtraction factor takes a large value of 5 to strongly subtract noise; as the SNR increases, the subband over-subtraction factor decreases linearly; at high SNR, the subband over-subtraction factor takes a small value of 1 to perform only mild subtraction and preserve speech components.
[0044] The lower limit factor β i Determined by the following piecewise function:
[0045] The function provides specific principles for determining the lower limit factor: at extremely low signal-to-noise ratios (SNR), the lower limit factor takes a relatively large value of 0.05 to provide deep spectral filling and avoid generating musical noise; as the SNR increases, the lower limit factor decreases linearly; at high SNRs, the lower limit factor takes a very small value of 0.002, with almost no filling, to avoid polluting the speech spectrum.
[0046] The subband over-attenuation factor and lower limit factor are determined by the signal-to-noise ratio (SNR) of the same subband through their respective monotonically decreasing functions. They work synergistically: at low SNR, the over-attenuation is strong, and the spectral fill is deep, preventing the spectral valleys caused by over-attenuation from being set to zero and introducing musical noise; at high SNR, the over-attenuation is weak, and the spectral fill is shallow, avoiding unnecessary modifications to the speech spectrum. This linkage mechanism fundamentally solves the contradiction of balancing high and low SNR that traditional fixed over-attenuation factors and half-wave rectification cannot overcome.
[0047] In addition, this step also determines the "sub-band frequency fine-tuning factor". d i This factor is based on the center frequency of each sub-band. f The preset frequency range in which it is located has been determined, specifically as follows: in, fsis the sampling rate of the original speech signal. The sub-band frequency fine-tuning factor is used in the subsequent spectral subtraction operation to differentially control the degree of noise subtraction according to the frequency range of the sub-band, giving a smaller subtraction strength to the low-frequency band to protect the speech fundamental frequency and harmonics, and a larger subtraction strength to the mid-high frequency band to fully remove noise.
[0048] Step S106: Perform a weighted moving average on the sub-band signal power spectra of adjacent sub-bands to obtain the smoothed power spectra of each sub-band.
[0049] In this step, an inter-sub-band smoothing operation is performed to suppress musical noise. The smoothing is performed on the frequency axis, for the current sub-band i and its M adjacent sub-bands before and after it M perform a weighted moving average on the sub-band signal power spectra. The smoothing neighborhood radius takes a value of 2, that is, perform a weighted moving average on the power spectra of the current sub-band and its 2 adjacent sub-bands before and after it, a total of 5 sub-bands. The formula for the smoothing process is: where i is the smoothed power spectrum of the oh k sub-band at frequency is the output after smoothing; i - j is the smoothed power spectrum of the i sub-band (i.e., the current sub-band j and its oh k adjacent sub-bands before and after it) at frequency is the predicted enhanced signal amplitude of the i sub-band at frequency oh k .
[0050] The above smoothed power formula describes the inter-sub-band weighted smoothing operation. Centered on the current i sub-band, sum the smoothed power spectra of its M adjacent sub-bands before and after it after weighting. The weighting coefficient is the predicted enhanced signal amplitude of the current sub-band. The summation range covers from i - M to i +[[ID=6[3]]<000038a total of 2 M +1 sub-bands. When M =2, then a total of 5 sub-bands participate in the smoothing. This smoothing operation expands the noise statistical samples on the frequency axis, and can smooth out the isolated spectral peaks and random spectral valleys left after independent spectral subtraction of each sub-band, suppressing musical noise from the frequency domain root cause.
[0051] Step S107: An adaptive spectral base function is used to perform spectral subtraction on the smooth power spectrum and noise power spectrum estimation, and adaptive filling is performed on the spectral valleys to obtain the enhanced speech power spectrum.
[0052] In this step, the subband over-subtraction factor determined in step S105 is used. α i Lower limit factor β i Subband frequency fine-tuning factor d i And the noise power spectrum estimate obtained in step S104 and the smoothed power spectrum obtained in step S106 This involves performing an adaptive spectral subtraction operation. This operation is implemented using an "adaptive spectral base function," specifically: Among them, | X i ( oh k )| 2 For the first i Subband at frequency oh k The enhanced speech power spectrum at that location is the output of this step; The smoothed power spectrum obtained in step S106; For the noise power spectrum estimation obtained in step S104; α i The sub-band has an over-subtraction factor; d i This is the subband frequency fine-tuning factor; β i The lower limit factor; To predict the power spectrum of the enhanced signal.
[0053] The adaptive spectral base function works as follows: when the predicted enhanced signal power spectrum is greater than zero, this non-negative value is directly used as the enhanced speech power spectrum to complete spectral subtraction and noise reduction at that frequency point; when the predicted enhanced signal power spectrum is less than or equal to zero, i.e., a "spectral valley" appears, instead of simply setting the valley to zero as in traditional methods, a lower limit factor is used. β i Multiply by smoothed power spectrum This serves as the enhanced speech power spectrum at that frequency. This adaptive padding operation ensures that the spectral valleys are filled with a small value proportional to the original speech power, rather than zero, thus eliminating the physical conditions for isolated spectral peaks and valleys caused by half-wave rectification, thereby effectively suppressing musical noise.
[0054] Step S108: Calculate the maximum noise residual threshold and use the threshold to limit the power spectrum of the enhanced speech to obtain pre-enhanced frequency domain data.
[0055] In this step, the enhanced speech power spectrum output in step S107 is... X i ( oh k )| 2 Perform residual limiting processing. First, calculate the "maximum noise residual threshold", denoted as max( oh k The threshold is calculated as follows: taking the current frame as the center, and taking the values before and after it... T Frames, 2 in total T +1 frame of enhanced speech power spectrum, at each frequency point oh k Calculate these 2 T The minimum sum of the power spectrum values of +1 frames. T The cumulative frame half-window length ranges from 2 to 4. This calculation reflects the statistically permissible upper limit of residual noise energy and can track the slow drift of noise.
[0056] Then, the enhanced speech power spectrum is limited using the maximum noise residual threshold to obtain the "pre-enhanced frequency domain data", denoted as | X pre ( oh k The amplitude limit rule is: when | X i ( oh k )| 2 Less than max( oh k When the threshold is reached, the value is the minimum value selected from adjacent frames; otherwise, the original value remains unchanged. After this limiting process, the random spikes remaining after spectral subtraction are forcibly removed, and the power spectrum is constrained below the statistical upper limit of residual noise, making the input for subsequent Wiener filtering cleaner.
[0057] Step S109: Using the pre-enhanced frequency domain data as a clean signal power spectrum estimate, combined with the noise power spectrum estimate, Wiener filtering is performed to obtain the filtered frequency domain signal.
[0058] In this step, the pre-enhanced frequency domain data output in step S108 is used. X pre ( oh k )|² as “clean signal power spectrum estimation” P XX ( oh k The noise power spectrum obtained in step S104 is used to estimate... As "noise power spectrum estimation" P VV ( oh k Design a Wiener filter. The frequency response H( of the Wiener filter) oh k The following formula is given: Among them, H( oh k ) represents frequency oh k The Wiener filter coefficient at point P is a real number between 0 and 1; XX ( oh k P is the power spectrum estimate for a clean signal; VV ( oh k () is used for noise power spectrum estimation.
[0059] After calculating the filter coefficients, multiply them by the amplitude spectrum of the pre-enhanced frequency domain data output in step S108 to obtain the amplitude spectrum of the filtered frequency domain signal. Combine the filtered amplitude spectrum with the original phase spectrum retained in step S102 to obtain the complete "filtered frequency domain signal".
[0060] The Wiener filter in this step, together with the aforementioned residual limiting stage, forms a two-tiered post-processing architecture. The first-stage residual limiting removes the random spikes remaining after spectral subtraction, ensuring that the input data used to estimate the Wiener filter coefficients has been pre-processed to remove large residual biases. The second-stage Wiener filter further smooths the spectral envelope, enhancing speech details in the high-frequency band and suppressing residual broadband noise in the low-frequency band. The two stages work together to improve the signal-to-noise ratio and auditory clarity of the final output speech.
[0061] Step S110: The filtered frequency domain signal is synthesized into an enhanced time domain speech signal through inverse short-time Fourier transform and overlap addition.
[0062] In this step, an inverse short-time Fourier transform is performed on the filtered frequency domain signal obtained in step S109 to convert the frequency domain data of each frame back to the time domain. Then, an overlap addition method is used to overlap and add the time domain signals of adjacent frames according to the frame shift, eliminating the discontinuity between frames, and finally synthesizing a continuous "enhanced time domain speech signal". This signal is the final output of this embodiment.
[0063] To more clearly illustrate the data flow and collaborative relationship between the steps in this embodiment, a specific example throughout the process will be used for explanation below.
[0064] The sampling rate is set to 16kHz, with 512 samples per frame, a frame shift of 256 samples, and the number of Mel subbands. N Select 24. Input a noisy speech signal with a signal-to-noise ratio of 0dB as the original speech signal.
[0065] First, steps S101 and S102 are executed to obtain the frequency domain signal of each frame. Then, step S103 is executed to divide the frequency domain signal of each frame into 24 Mel sub-bands. Assume the current frame is the [missing information - likely a specific frame number]. t For the fifth sub-band, whose frequency range is approximately 300Hz to 400Hz, the initial noise amplitude spectrum estimate obtained by step S104 in the initial silent segment of this sub-band is approximately 0.15, and the corresponding noise power spectrum estimate is... Approximately 0.0225. After the speech begins, the... t The amplitude spectrum of the subband signal in this subband | Y 5| is 0.50, subband signal power spectrum| Y 5| 2 The value is 0.25. Based on speech activity detection, this segment is determined to be inactive. The sub-band signal-to-noise ratio is then calculated. SNR 5 is 10dB.
[0066] Proceed to step S105, according to SNR 5 = 10dB, calculate the subband over-subtraction factor. α 5 = 4 - (3 / 20) × 10 = 2.5; lower limit factor β 5 = 0.06 - (1 / 625) × 10 = 0.044. The center frequency of this sub-band is approximately 350 Hz, which is less than 1 kHz. Therefore, the sub-band frequency adjustment factor is... d 5 = 1.
[0067] Proceed to step S106, take the power spectrum of the current 5th sub-band and the 2 sub-bands before and after it, for a total of 5 sub-bands, and perform a weighted moving average to obtain a smoothed power spectrum. After smoothing, It is 0.24.
[0068] Proceed to step S107 to calculate the power spectrum of the predicted enhanced signal. =0.24 - 2.5 × 1 × 0.0225 = 0.24 - 0.05625 = 0.18375. Since the result is greater than zero, the speech power spectrum is enhanced. The value is directly taken as 0.18375. If at another frequency point in this sub-band, If the calculation result is -0.01, adaptive spectral padding is triggered, and the enhanced speech power spectrum is set to 0.044 × 0.24 = 0.01056. The introduction of this non-zero padding value directly avoids the music noise caused by setting this frequency point to zero.
[0069] After residual limiting in step S108 and Wiener filtering in step S109, the residual noise in the subband is further suppressed. Finally, after synthesizing the time-domain signal in step S110, the overall signal-to-noise ratio of the output speech is improved from 0dB to approximately 5.7dB.
[0070] In contrast, under the same input conditions, the traditional multi-band spectral subtraction method uses a fixed over-subtraction factor of 4 and half-wave rectification to process the negative spectrum, and the noise estimation uses a recursive average with a fixed forgetting factor of 0.95, resulting in an output signal-to-noise ratio of approximately 3.6 dB.
[0071] Figure 2 shows a comparison of the time-domain waveforms of the original photoacoustic signal and the signal obtained by the present invention, used to visually verify the enhancement effect of the method of the present invention on the photoacoustic speech signal from the time domain dimension. The horizontal axis of the figure represents time t, in seconds (s); the vertical axis represents signal amplitude, in decibels (dB). The upper sub-figure shows the time-domain waveform of the original photoacoustic signal, with the vertical axis ranging from -0.5dB to 1.0dB; the lower sub-figure shows the time-domain waveform of the signal enhanced by the method of the present invention, with the vertical axis ranging from -1.0dB to 1.0dB.
[0072] From the waveform feature comparison: In the non-active speech segments (such as 0~0.5s, 1~2s, 3~4s, 5~6s, 7~8s, 9~10s, etc., where there are no obvious speech pulses), the original photoacoustic signal has a significant and continuous background noise floor, the waveform is chaotic, and system noise, background light noise, and environmental interference are superimposed on the signal throughout, making it impossible to visually distinguish noise from weak speech components; the enhanced signal of this invention almost completely suppresses the noise floor in the corresponding segments, the waveform approaches zero level, and there is no obvious random noise fluctuation. In the active speech segments (such as 0.5~1s, 2~3s, 4~5s, 6~7s, 8~9s, etc., where there are speech pulses), the speech waveform of the original photoacoustic signal is severely submerged by noise, the outline is blurred, and the peak, rising edge, and falling edge characteristics of the speech are lost; the enhanced signal of this invention completely preserves the temporal structure of the speech, and the amplitude, duration, and envelope characteristics of the speech pulses are clear and sharp, without speech clipping, distortion, or loss of weak speech components caused by over-subtraction. Figure 2 visually demonstrates that the method of the present invention can accurately separate speech components from background noise in photoacoustic signals. While achieving deep denoising, it has excellent ability to protect speech temporal features, solving the core problems of traditional spectral subtraction, such as severe noise residue at low signal-to-noise ratio and easy damage to speech at high signal-to-noise ratio.
[0073] This embodiment achieves a more significant enhancement effect at extremely low signal-to-noise ratios. Table 1 shows the comparison of the enhanced signal and the traditional spectral subtraction method in terms of speech quality signal-to-noise ratio (SNR).
[0074] Table 1: Comparison of the enhanced signals obtained by the present invention and the traditional spectral subtraction method in terms of speech quality signal-to-noise ratio (SNR).
[0075] .
[0076] A comparison of the enhanced signal obtained by this application and the traditional spectral subtraction method shows that, for example, at 0dB... SNR In scenarios where the input signal is low, the signal-to-noise ratio of traditional spectral subtraction is... SNR The method of this invention is improved to 3.6dB. SNR It can be boosted to 5.7dB, output SNR 58% improvement over traditional spectral subtraction, at 5dB SNR In scenarios with high-quality input signals, such as signal-to-noise ratios (SNRs) of 10dB and 15dB, the traditional spectral subtraction method achieves a 126% improvement over conventional spectral subtraction. SNR Compared to the original input signal SNR The output of the method of this invention is slightly reduced in high signal-to-noise ratio scenarios. SNR It did not decrease compared to the original input signal. SNR It also provides a stable enhancement of approximately 8%, demonstrating that the solution of this invention can improve the output speech signal in scenarios with low signal-to-noise ratio and poor signal quality. SNR Significant improvement in signal-to-noise ratio, such as Figure 3 As shown.
[0077] The core of this embodiment's achievement lies in the linkage mechanism between the subband over-subtraction factor and the lower limit factor, which varies with the signal-to-noise ratio (SNR). Unlike traditional methods where the over-subtraction factor is fixed and negative values are simply set to zero, in this embodiment, both factors are synchronously driven by the same subband SNR, with denoising intensity and spectral fill depth matched in real time. At low SNR, strong denoising is combined with deep fill. At high SNR, light denoising is combined with shallow fill. This collaborative mechanism enables the entire adaptive spectral subtraction framework to maintain a balance between denoising and fidelity across all subbands and SNR conditions, resolving the contradiction of traditional methods struggling to balance high and low SNR.
[0078] Furthermore, the inter-subband smoothing operation fuses the spectral information of the current subband with that of adjacent subbands on the frequency axis, smoothing out isolated spectral peaks and random spectral valleys left over from independent subband operations. These isolated peaks and valleys are the frequency domain root of musical noise. The smoothing operation utilizes the power spectrum processed by an adaptive factor. The smoothing intensity is adapted to the current noise reduction state, suppressing residual noise while preserving speech details at subband boundaries.
[0079] In a preferred embodiment, the forgetting factor in step S104 above c Dynamically adjusted based on subband signal-to-noise ratio. When the subband signal-to-noise ratio is less than 0dB... cTake 0.90; when the subband signal-to-noise ratio is between 0dB and 10dB, c Take 0.95; when the subband signal-to-noise ratio is greater than 10dB, c We set it to 0.98. This dynamic adjustment allows the noise spectrum estimation to keep up with the slow drift of noise without being contaminated by speech components during speech activity segments, providing an accurate noise reference for the entire adaptive framework.
[0080] In another preferred embodiment, the smoothed neighborhood radius in step S106 above M The value can be 1 or 3. When the noise spectrum fluctuation is small, M Setting it to 1 satisfies the smoothing requirement while further reducing computational complexity; when the noise spectrum fluctuates significantly, M Choosing 3 allows for wider frequency coverage and provides a stronger smoothing effect.
[0081] In another preferred embodiment, the cumulative frame half-window length in step S108 above T The value is set to 2, meaning that 5 frames (2 frames before and 2 frames after) are used in the calculation of the maximum noise residual threshold. This value corresponds to a time window of approximately 60 milliseconds, which ensures sufficient statistical samples while maintaining temporal resolution for noise changes.
[0082] In another alternative embodiment, the subband division is not limited to the Mel scale; other non-uniform division methods that approximate the characteristics of human auditory perception, such as the Bark scale or the equivalent rectangular bandwidth scale, can also be used. These scales also have the characteristics of high low-frequency resolution and low high-frequency resolution, and can replace the Mel scale to achieve similar technical effects.
[0083] In another optional embodiment, the speech activity detection in step S104 can employ a deep learning-based speech activity detection model instead of the energy threshold method. This model infers from the sub-band signal, outputting the probability of speech presence; when the probability falls below a preset threshold, a noise spectrum update is triggered. This alternative approach offers higher detection accuracy in non-stationary noise scenarios.
[0084] Those skilled in the art will understand that the specific parameter values given in the above embodiments, for example... N For 24, M For 2, T The values of the forgetting factor, the piecewise function expressions of the sub-band over-subtraction factor and the lower limit factor are illustrative and not intended to limit the scope of protection. In practical applications, adjustments can be made around the above values according to the specific sampling rate, noise environment, and application scenario. The adjusted solution can still achieve the technical effects of this invention.
[0085] Although the specification has provided a detailed description, it should be understood that various changes, substitutions, and modifications can be made without departing from the spirit and scope of the invention as defined by the appended claims. Furthermore, the specific embodiments described are not intended to limit the scope of the invention, and those skilled in the art will readily understand based on this invention that existing or future-developed processes, machines, manufactures, compositions of matter, means, methods, or steps can perform substantially the same functions or achieve substantially the same results as the embodiments of the invention. Therefore, the appended claims are intended to include such processes, machines, manufactures, compositions of matter, means, methods, or steps within their scope.
Claims
1. A photoacoustic signal enhancement method based on adaptive multi-band spectral subtraction and Wiener filtering, characterized in that, include: The original speech signal is acquired, framed, windowed, and subjected to short-time Fourier transform to obtain the original frequency domain signal; The original frequency domain signal is divided into multiple continuous sub-bands according to the Mel scale to obtain sub-band signals; For each sub-band speech activity detection, the noise amplitude spectrum estimate is updated recursively by averaging in the speech inactive segment, and the noise power spectrum estimate is obtained by squaring. The sub-band signal-to-noise ratio is calculated based on the sub-band signal power spectrum and the noise power spectrum estimate. Based on the subband signal-to-noise ratio, the subband over-decrease factor and lower limit factor are determined by a monotonically decreasing function, both of which are negatively correlated with the subband signal-to-noise ratio. For the current sub-band and all sub-bands before and after it M A smoothed power spectrum is obtained by weighted moving average of the power spectra of adjacent sub-bands. M To smooth the neighborhood radius; An adaptive spectral base function containing the lower limit factor is used to perform spectral subtraction on the smooth power spectrum and the noise power spectrum estimation, and adaptively fill the spectral valleys to obtain the enhanced speech power spectrum; Calculate the maximum noise residual threshold, and limit the power spectrum of the enhanced speech to obtain pre-enhanced frequency domain data; Using the pre-enhanced frequency domain data as a clean signal power spectrum estimate, and combining it with the noise power spectrum estimate Wiener filter, the filtered frequency domain signal is obtained. The filtered frequency domain signal is then subjected to inverse short-time Fourier transform and overlap addition to synthesize an enhanced time-domain speech signal.
2. The method according to claim 1, characterized in that, The subband over-reduction factor α i Determined by the following piecewise function: in, SNR i For the first i Subband signal-to-noise ratio of each subband.
3. The method according to claim 1, characterized in that, The lower limit factor β i Determined by the following piecewise function: in, SNR i For the first i Subband signal-to-noise ratio of each subband.
4. The method according to claim 1, characterized in that, The enhanced speech power spectrum | X i ( ω k )| 2 Determined by the following adaptive spectral base function: in, For the first i Subband at frequency ω k The smoothed power spectrum at that location, For the noise power spectrum estimation, α i The subband over-subtraction factor, δ i This is the subband frequency fine-tuning factor. β i The lower limit factor is... To predict the power spectrum of the enhanced signal.
5. The method according to claim 1, characterized in that, The noise amplitude spectrum estimation The update formula is: in, For the current frame t No. i Subband at frequency ω k Noise amplitude spectrum estimation at the location, For noise amplitude spectrum estimation of the previous frame, | Y i ( ω k , t | represents the subband signal amplitude spectrum of the current frame. γ The forgetting factor ranges from 0 to 1; the square of the noise amplitude spectrum estimation That is, the noise power spectrum estimation. .
6. The method according to claim 4, characterized in that, The sub-band frequency fine-tuning factor is determined by the following piecewise function: in, f For the first i The center frequency of the sub-band fs Sampling rate, δ i This is the subband frequency fine-tuning factor.
7. The method according to claim 1, characterized in that, The smooth neighborhood radius M The value is 2, which means that the power spectrum of the current sub-band and the two adjacent sub-bands before and after it, a total of 5 sub-bands, is weighted moving average.
8. A photoacoustic signal enhancement device, employing the method described in any one of claims 1-7, characterized in that, include: The signal acquisition module is used to acquire raw speech signals; The time-frequency conversion module is used for framing, windowing, and short-time Fourier transform to obtain the original frequency domain signal; The sub-band division module is used to divide the signal into multiple continuous sub-bands according to the Mel scale to obtain the sub-band signal. The noise estimation and signal-to-noise ratio calculation module is used for speech activity detection. It recursively averages and adaptively updates the noise amplitude spectrum estimate in the speech inactive segment, and squares it to obtain the noise power spectrum estimate. It calculates the subband signal-to-noise ratio based on the subband signal power spectrum and the noise power spectrum estimate. The parameter determination module is used to determine the subband over-reduction factor and the lower limit factor based on the subband signal-to-noise ratio using a monotonically decreasing function. Both factors are negatively correlated with the subband signal-to-noise ratio. The sub-band smoothing module is used to smooth the current sub-band and the sub-bands before and after it. M A smoothed power spectrum is obtained by weighted moving average of the power spectra of adjacent sub-bands. M To smooth the neighborhood radius; An adaptive enhancement module is used to perform spectral subtraction on the smooth power spectrum and the noise power spectrum estimate using an adaptive spectral base function that includes the lower limit factor, and adaptively fill the spectral valleys to obtain an enhanced speech power spectrum. The residual limiting module is used to calculate the maximum noise residual threshold and limit the power spectrum of the enhanced speech to obtain pre-enhanced frequency domain data; The post-filtering processing module is used to use the pre-enhanced frequency domain data as a clean signal power spectrum estimate, and combine it with the noise power spectrum estimate to perform Wiener filtering to obtain the filtered frequency domain signal. The signal reconstruction module is used to synthesize an enhanced time-domain speech signal by performing inverse short-time Fourier transform and overlap addition on the filtered frequency domain signal.
9. The photoacoustic signal enhancement device according to claim 8, characterized in that, The adaptive spectral base function used by the adaptive enhancement module is: in, For the first i Subband at frequency ω k The smoothed power spectrum at that location, For the noise power spectrum estimation, α i The subband over-subtraction factor, δ i This is the subband frequency fine-tuning factor. β i The lower limit factor is... To predict the power spectrum of the enhanced signal, | X i ( ω k )| 2 To enhance the speech power spectrum.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 7.