Adaptive Spectral Subtraction of Speech Spectrogram Based on Non-uniform Graph Sub-band Partitioning
Through the speech graph topology based on forgetting factor and dichotomy, adaptive spectral subtraction of non-uniform graph subbands is realized, which solves the problem of insufficient noise reduction performance of graph subtraction in high-dimensional speech signal processing in the prior art, and significantly improves the signal-to-noise ratio and PESQ value.
Patent Information
- Application Number
- CN202111362013.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-17
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-11-17
AI Technical Summary
When processing irregular high-dimensional speech signals, the existing graph subtraction fails to effectively consider the differences in the distribution characteristics of speech signals and noise signals in the graph frequency domain, resulting in insufficient noise reduction performance.
The speech graph topology based on forgetting factors is used to divide the frequency domain of the non-uniform graph to form a graph subband signal, and an adaptive spectrum subtraction strategy is performed according to the speech distribution characteristics. The graph frequency is divided into 12 non-uniform graph subbands through the forgetting graph topology and dichotomy method, and the effective graph subband is filtered and different spectrum subtraction strategies are adopted.
The output signal-to-noise ratio is improved by 3-5dB under low signal-to-noise ratio, the PESQ value is increased by 0.2-0.5 under high signal-to-noise ratio, 5-7dB in white noise environment, and the PESQ value is increased by 0.1-0.3, which significantly improves noise reduction performance in factory noise environment.
Smart Images

Figure CN114220445B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of graph voice signal denoising, and specifically designs a voice graph adaptive spectral subtraction method based on non-uniform graph subband division. Background Art
[0002] With the advent of the big data era, the scenarios of signals are gradually becoming complex and diverse, such as large-scale sensor networks, social networks, and biological networks. Therefore, signals are gradually showing high-dimensional and irregular characteristics. Traditional signal processing methods rely on digital signal processing theory and have good processing effects on regular low-dimensional signals, but are helpless for irregular high-dimensional signals. In recent years, the proposal of Graph Signal Processing (GSP) has attracted extensive attention from researchers. It mainly describes the signal model in depth with three elements: vertices, edges, and weights by establishing a graph topology structure, and further establishes a new graph frequency domain, enabling the signal to obtain different performances in the new graph frequency domain compared with the traditional frequency domain.
[0003] As a new signal processing theory, the theory of graph signal processing enriches the original signal processing methods. As a short-term stationary non-linear signal, voice signals can also be processed using graph signal theory. Based on the forward and backward correlation and periodicity of voice signals, a new voice graph topology structure can be constructed. Based on this voice graph topology structure, a new graph frequency domain is obtained, and signal processing is performed in the new graph frequency domain.
[0004] After a literature search of the existing technologies, it is found that Yan Xue et al. published an article "An Iterative Graph Spectral Subtraction Method for Speech Enhancement" in "Speech Communication" June 2020 pp. 35 - 42. This article proposed a k-shift graph as the graph topology structure of voice signals, obtained eigenvectors as the graph Fourier transform basis by performing eigenvalue decomposition on the graph adjacency matrix of this topology structure, and obtained a new graph frequency domain space with the graph Fourier transform basis as the spatial basis vectors. Project the voice signal onto this graph frequency domain and use spectral subtraction for overall voice denoising. In this scheme, the spectral subtraction uniformly adopts the same spectral subtraction strategy for the entire graph frequency domain, without considering the different distribution characteristics of voice signals and noise signals in the graph frequency domain. Therefore, there is still room for further improvement in its noise reduction technology and performance. Summary of the Invention
[0005] In view of the deficiencies of the above-mentioned existing technologies, the present invention proposes a voice graph adaptive spectral subtraction method based on non-uniform graph subbands. On the basis of the original spectral subtraction method, the graph frequencies are divided into non-uniform graph frequency domains according to the graph frequency domain characteristics of the voice, forming so-called graph subband signals, such that the widths of the graph subbands in different graph frequency ranges are different. Then, according to the voice distribution on the graph subbands, the subband signal-to-noise ratio of the voice signal is estimated, and based on this, graph subbands are screened, and an adaptive spectral subtraction strategy is adopted for different graph subbands to improve the voice enhancement performance of the algorithm.
[0006] The present invention is implemented through the following technical solutions:
[0007] The present invention introduces a voice graph topology structure based on a forgetting factor, and uses the dichotomy method to perform non-uniform graph subband division on the graph frequencies of the graph topology structure, dividing the graph frequencies into 12 non-uniform graph subbands. The number of effective graph subbands is calculated according to the estimated signal-to-noise ratio of the voice signal, and the graph subbands are classified into two types: effective graph subbands and non-effective graph subbands. Furthermore, different spectral subtraction strategies are adopted for effective graph subbands and non-effective graph subbands, that is, the spectral subtraction coefficient under the effective graph subbands is determined according to the number of effective graph subbands, so as to implement an adaptive spectral subtraction strategy for different graph frequency points.
[0008] The present invention includes the following steps:
[0009] Step 1: According to the correlation between voice signal samples, a forgetting graph topology structure based on a forgetting factor is used to describe the noisy voice signal. According to the forgetting graph topology structure, the voice graph adjacency matrix is obtained, and further the graph Fourier transform basis matrix and the graph frequencies in the corresponding graph frequency domain are obtained.
[0010] Step 2: According to the graph frequencies obtained in Step 1, non-uniform graph subband division can be performed. Since the graph frequencies obtained from the graph adjacency matrix of the forgetting graph show a characteristic of being denser in the low frequency and sparser in the high frequency after normalization, the present invention adopts a non-uniform graph subband division algorithm based on the dichotomy method to divide the graph frequencies into a group of graph subbands with wide high frequencies and narrow low frequencies.
[0011] Step 3: Map the noisy voice signal to the graph vertex domain to obtain the corresponding noisy voice graph signal. According to the graph Fourier transform basis matrix obtained in Step 1, the graph spectrum of the noisy voice signal can be obtained.
[0012] Step 4: Use the first five frames of the noisy voice signal as silent frames, and use the average graph energy spectrum of the silent frames for noise graph energy spectrum estimation. According to the graph energy spectrum of the noisy voice signal and the estimated noise graph energy spectrum, the estimated signal-to-noise ratio of the noisy voice signal can be obtained.
[0013] Step 5: Using the estimated signal-to-noise ratio obtained in Step 4, the number of useful sub-bands of a frame of signal map can be obtained. According to the map sub-band division result obtained in Step 2, the total energy of each map sub-band is calculated, and the total energies of the map sub-bands are arranged in descending order. Finally, the map sub-bands with the top several total energies are marked as useful map sub-bands, and the remaining map sub-bands are marked as useless map sub-bands.
[0014] Step 6: According to the map sub-band marking result obtained in Step 5, different spectral subtraction strategies are adopted for useful map sub-bands and useless map sub-bands. For useful map sub-bands, the energy spectrum of the noisy speech signal is directly subtracted from the energy spectrum of the noise speech to obtain the denoised speech signal. For useless map sub-bands, a penalty coefficient is multiplied in front of the energy spectrum of the noise speech signal. The penalty coefficient is related to the number of useful map sub-bands. Then, the energy spectrum of the noisy speech signal is subtracted from it to obtain the denoised speech map spectrum.
[0015] Step 7: The denoised speech map spectrum is inverse Fourier-transformed through the map Fourier transform to obtain a new speech map signal, which is finally mapped to the time domain to obtain the denoised speech signal.
[0016] Compared with the prior art, the present invention adopts the above technical solutions and has the following technical effects:
[0017] 1. In a white noise environment, compared with the spectral subtraction of Reference Solution 1 and the sub-band spectral subtraction of Reference Solution 2, the output signal-to-noise ratio of the speech map adaptive spectral subtraction based on non-uniform map sub-band division is increased by 3 - 5 dB under low signal-to-noise ratio conditions, and the PESQ value is increased by 0.2 - 0.5 under high signal-to-noise ratio conditions;
[0018] 2. In a factory noise environment, compared with the spectral subtraction of Reference Solution 1 and the sub-band spectral subtraction of Reference Solution 2, the output signal-to-noise ratio of the speech map adaptive spectral subtraction based on non-uniform map sub-band division is increased by 5 - 7 dB under low signal-to-noise ratio conditions, and the PESQ value is increased by 0.1 - 0.3 under high signal-to-noise ratio conditions. Description of the Drawings
[0019] Figure 1 is the map frequency distribution diagram of the forgetting map structure
[0020] Figure 2 is the algorithm flow chart;
[0021] Figure 3 is the graph of the output signal-to-noise ratio of the example solution and the reference solution varying with the input signal-to-noise ratio under white noise;
[0022] Figure 4 is the graph of the PESQ of the example solution and the reference solution varying with the input signal-to-noise ratio under white noise;
[0023] Figure 5It is a graph showing the variation of the output signal-to-noise ratio of the example scheme and the reference scheme with the input signal-to-noise ratio under factory noise;
[0024] Figure 6 It is a graph showing the variation of PESQ of the example scheme and the reference scheme with the input signal-to-noise ratio under factory noise. Detailed implementation manners
[0025] The following is a detailed description of the example of the present invention. This example is implemented on the premise of the technical solution of the present invention, and detailed implementation manners and specific operation processes are given.
[0026] As Figure 2 shown, this example is implemented through the following steps:
[0027] Step 1: Use the forgetting graph based on the forgetting factor to construct the graph topology structure of the speech signal. For the element in the i-th row and j-th column of the N-th order speech graph adjacency matrix A, it satisfies
[0028]
[0029] where λ∈(0, 1) is the forgetting factor, indicating that the farther the time distance between sampling points, the weaker the correlation between the two. At the same time, Ψ is the forgetting threshold. When the time correlation between the two is lower than the threshold, it is considered that there is no correlation.
[0030] According to the theory of graph signal processing, perform diagonalization decomposition on the graph adjacency matrix A of the forgetting graph
[0031] A = VΛV T (2)
[0032] where V = [v0,..., v N-1 is the eigenvector matrix, and Λ = diag[λ0, λ1,..., λ N-1 is the eigenvalue matrix. Thus, the graph Fourier transform matrix of the speech signal and the graph frequencies in the graph frequency domain can be obtained.
[0033] Step 2: Perform non-uniform sub-band division on the graph frequency domain. From Figure 1 it can be seen that the newly designed speech graph frequencies show non-uniform distribution characteristics after normalization. Therefore, a non-uniform graph sub-band division algorithm is used to divide them.
[0034] The specific steps of the non-uniform graph sub-band division algorithm are as follows:
[0035] Step1. Select the maximum graph frequency λ max and the minimum graph frequency λ min from the normalized graph frequencies in (2), and let c = 1;
[0036] Step 2. Take the average of λ max and λ min to obtain the intermediate graph frequency λ mid . Set λ mid to λ max as the c-th sub-band;
[0037] Step 3. Determine whether the number of graph frequencies between λ min and λ mid is 1. If not, go to Step 4; otherwise, go to Step 5;
[0038] Step 4. Set λ mid as the new λ max , increment c by 1, and return to Step 2;
[0039] Step 5. Set the last frequency point as the (c + 1)-th sub-band and end the algorithm.
[0040] Using this algorithm, the finally divided m graph sub-bands B = {b1, b2,..., b m}, where b i = {λ j ,..., λ k} represents the i-th graph sub-band, λ j represents the starting graph frequency of the i-th graph sub-band, and λ k represents the ending graph frequency of the i-th graph sub-band.
[0041] Step 3: Map a frame of noisy speech signal s with a length of N, s = {s1,..., s N} T to a speech graph signal Perform graph Fourier transform on the speech graph signal S G using the graph Fourier transform basis obtained in Step 1 to obtain the graph spectrum of this frame of speech signal
[0042]
[0043] where represents the graph spectrum.
[0044] Step 4: Estimate the signal-to-noise ratio of a segment of speech signal. First, assume that the first five frames of this segment of speech signal are silent frames, i.e., noise frames, and use the first five frames for noise graph energy spectrum estimation. For the first five frames of noise frames, use the following method. For a frame of noisy frame n = {n1,..., n N} T Perform graph Fourier transform using the method in Step 3 to obtain the graph spectrum of the noise By taking the average of the squares of the magnitudes of the graph spectra of the first five frames of noise, the noise graph energy spectrum estimate
[0045]
[0046] Among them represents the spectrum of the noise map of the z-th frame, represents the spectrum of the average noise map of the first five frames.
[0047] According to the noise estimation, the estimated signal-to-noise ratio of the noisy speech can be obtained
[0048]
[0049] Step Five: Using the estimated signal-to-noise ratio obtained in Step Four, filter out the useful map sub-bands. The number of useful map sub-bands is determined by the estimated signal-to-noise ratio
[0050]
[0051] where B use represents the number of useful map sub-bands.
[0052] According to the map spectrum of the signal, the energy values E of m map sub-bands of each frame of the signal can be obtained band ={E1,..., E m}, where E i is the energy value of the i-th map sub-band. The specific method for obtaining the energy value of a single map sub-band is
[0053]
[0054] where p is the serial number of the starting map frequency point of the d-th map sub-band, q is the serial number of the ending map frequency point of the d-th map sub-band, is the amplitude spectrum value corresponding to the k-th map frequency point of the noisy speech signal. After obtaining the energy values of all map sub-bands, sort them in descending order according to the energy magnitude, and refer to the value of B use to mark the first B use map sub-bands as useful map sub-bands, and mark the remaining map sub-bands as useless map sub-bands.
[0055] Step Six: For the useful map sub-bands and the useless map sub-bands, perform map subtraction with different strategies respectively. The specific spectral subtraction strategy is as follows
[0056]
[0057] where is the amplitude spectrum value corresponding to the r-th map frequency point of a frame of denoised speech signal, is the amplitude spectrum value corresponding to the r-th map frequency point of the noisy speech signal, is the amplitude spectrum value corresponding to the r-th map frequency point of the noise signal.
[0058] For the useful graph sub-bands, which contain more useful information, the energy spectrum of the speech graph can be directly subtracted from the energy spectrum of the noise graph to obtain the processed speech graph spectrum. For the useless graph sub-bands, which contain more useless information, log2(B use ) is added as a penalty factor. When the number of useful graph sub-bands is larger, it is considered that the main noise information is more concentrated in the remaining useless graph sub-bands, and the larger the penalty factor needs to be applied.
[0059] Step 7: According to the graph inverse Fourier transform method in Step 1, the speech graph signal in the vertex domain is obtained from the processed speech graph spectrum obtained in Step 6 according to formulas (2) and (8).
[0060]
[0061] Among them, Finally, the speech graph signal is mapped to the time domain to obtain the denoised speech signal s clean .
[0062] To test the performance of the present invention, the DARPA TIMIT Acoustic-phonetic Continuous Speech Corpus speech library is used as the speech material in the experiment, and the Standard noise NOISEX-92 library database is selected as the noise material. The sampling rate of the speech in the experiment is 8000 Hz, the frame length is selected as 128, and the duration of each selected speech material is 3 s. The method for generating the noisy signal in this article is to calculate the noise variance according to the preset signal-to-noise ratio and the signal power, and then scale the noise amplitude and superimpose it on the clean speech. To test the denoising performance of the method in this example, 100 pieces of speech are selected from the TIMIT speech library, including 50 sentences of male voices and 50 sentences of female voices, with a cumulative speech duration of 5 minutes. In addition, white noise, pink noise, and factory noise are selected as the three noise sources, and the signal-to-noise ratio of the noisy speech signal is set to -15 to 15 dB, with an interval of 5 dB, and the signal-to-noise ratio and PESQ performance of the speech signal before and after denoising in different scenarios are tested respectively.
[0063] From Figure 3 and Figure 4It can be seen that in a white noise environment, compared with the spectral subtraction of the reference solution 1 and the sub-band spectral subtraction of the reference solution 2, the output signal-to-noise ratio of the speech spectrogram adaptive spectral subtraction based on non-uniform spectrogram sub-band division is increased by 3-5 dB under low signal-to-noise ratio conditions, while the PESQ value is increased by 0.2-0.5 under high signal-to-noise ratio conditions. Since white noise is approximately uniformly distributed in the entire spectrogram frequency domain, applying a penalty factor to the useless spectrogram sub-bands in the spectrogram sub-band spectral subtraction will cause the loss of useful information. Therefore, the phenomenon of the decline in the signal-to-noise ratio improvement effect will occur under high signal-to-noise ratio conditions. The experimental results show that the noise reduction performance of the proposed solution in this example is significantly better than that of the reference solution 1 and the reference solution 2 under stationary noise.
[0064] From Figure 5 and Figure 6 It can be seen that in a factory noise environment, compared with the spectral subtraction of the reference solution 1 and the sub-band spectral subtraction of the reference solution 2, the output signal-to-noise ratio of the speech spectrogram adaptive spectral subtraction based on non-uniform spectrogram sub-band division is increased by 5-7 dB under low signal-to-noise ratio conditions, while the PESQ value is increased by 0.1-0.3 under high signal-to-noise ratio conditions. Since the distribution of non-stationary noise in the entire spectrogram frequency domain is non-uniform, which is consistent with the segmented noise reduction idea of the spectrogram sub-band spectral subtraction, the noise reduction effect will be more significant. The experimental results show that the noise reduction performance of the proposed solution in this example is also significantly better than that of the reference solution 1 and the reference solution 2 under non-stationary noise.
[0065] The present invention considers the characteristic that the frequency distribution of the speech forgetting spectrogram topology based on the forgetting factor is non-uniform, and designs a non-uniform spectrogram sub-band division method based on the dichotomy method. On this basis, useful spectrogram sub-bands and useless spectrogram sub-bands are screened according to the estimated signal-to-noise ratio, and then different spectral subtraction strategies are adopted for different spectrogram sub-bands. The present invention also includes the performance analysis of the system under stationary noise and non-stationary noise environments, mainly the analysis of the output signal-to-noise ratio and PESQ results of the system under different signal-to-noise ratio conditions. After multiple experimental verifications, the speech spectrogram adaptive spectral subtraction based on non-uniform spectrogram sub-band division proposed by the present invention has significantly improved the noise reduction performance compared with the spectral subtraction and the sub-band spectral subtraction.
[0066] It should be noted that the description of the above embodiments is only used to help understand the method and its core idea of the present application. For those of ordinary skill in the art, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications are also within the protection scope of the claims of the present application.
Claims
1. Speech graph adaptive spectral subtraction based on non-uniform graph sub-band division, characterized in that It includes the following steps: Step 1: According to the correlation between speech signal samples, use a forgetting graph topology based on a forgetting factor to describe the noisy speech signal, obtain the speech graph adjacency matrix according to the forgetting graph topology, and further obtain the graph Fourier transform basis matrix and the graph frequencies in the corresponding graph frequency domain; Step 2: Perform non-uniform graph sub-band division according to the graph frequencies obtained in Step 1; Step 3: Map the noisy speech signal to the graph vertex domain to obtain the corresponding speech graph signal; according to the graph Fourier transform basis matrix obtained in Step 1, obtain the graph spectrum of the noisy speech signal; Step 4: Take the first five frames of the noisy speech signal as silent frames, and take the average graph energy spectrum of the silent frames as the estimated noise graph energy spectrum; according to the graph energy spectrum of the noisy speech signal and the estimated noise graph energy spectrum, obtain the estimated signal-to-noise ratio of the noisy speech signal; Step 5: According to the sub-band division result of the graph obtained in Step 2, calculate the energy value of each graph sub-band, and sort the total energy in descending order according to the energy value. Mark the first B use graph sub-bands as useful graph sub-bands, and mark the remaining graph sub-bands as useless graph sub-bands. The number of useful graph sub-bands is obtained from the estimated signal-to-noise ratio obtained in Step 4; Step 6: According to the graph sub-band marking results obtained in Step 5, adopt different spectral subtraction strategies for useful graph sub-bands and useless graph sub-bands: for useful graph sub-bands, directly subtract the energy spectrum of the noise speech signal from the energy spectrum of the noisy speech signal to obtain the denoised speech graph spectrum; for useless graph sub-bands, multiply a penalty coefficient in front of the energy spectrum of the noise speech signal, and then subtract it from the energy spectrum of the noisy speech signal to obtain the denoised speech graph spectrum; Step 7: Obtain a new speech graph signal through the inverse graph Fourier transform of the denoised speech graph spectrum, and finally map it to the time domain to obtain the denoised speech signal; In Step 1, for the element in the \(i\)-th row and \(j\)-th column of an \(N\)-order speech graph adjacency matrix \(A\), it satisfies: where \(\lambda\in(0,1)\) is the forgetting factor and \(\Psi\) is the forgetting threshold; According to the theory of graph signal processing, perform diagonalization decomposition on the speech graph adjacency matrix \(A\); A = V Λ V T (2) where \(V = [v_0,\ldots,v N-1 \) is the eigenvector matrix, \(v t \) is the \(t\)-th eigenvector, \(\Lambda=\text{diag}[\lambda_0,\lambda_1,\ldots,\lambda N-1 \) is the eigenvalue matrix, \(\lambda t \) is the \(t\)-th eigenvalue, \(t = 1,2,\ldots,N - 1\), \(N\) is the frame length; thus, the graph Fourier transform basis matrix of the speech signal and the graph frequencies in the graph frequency domain are obtained; The specific steps of the non-uniform graph sub-band division algorithm are: Step 1. After normalizing the graph frequencies, select the largest graph frequency λ max and the smallest graph frequency λ min , and let c = 1; Step 2. Take the average of λ max and λ min to obtain the intermediate graph frequency λ mid . Set λ mid to λ max as the c-th sub-band; Step3. Determine λ min to λ mid Whether the number of graph frequencies between them is 1. If not, go to Step4; otherwise, go to Step5; Step4. Let λ max = λ mid , c = c + 1, and return to Step2; Step 5. Set this graph frequency as the \((c + 1)\)-th sub-band and end the algorithm.
2. The voice spectrogram adaptive spectral subtraction method based on non-uniform spectrogram sub-band division according to claim 1, characterized in that Specifically, Step 3 is: The noisy speech signal s = {s1, …, s N} T is mapped to a spectrogram signal where s n is the n-th frame of the speech signal, is the n-th frame of the spectrogram signal, and n = 1, 2, …, N; Use the graph Fourier transform basis obtained in Step 1 to perform a graph Fourier transform on the speech graph signal s G to obtain the graph spectrum of the noisy speech signal: Among them is s n graph spectrum of 3. The voice spectrogram adaptive spectral subtraction method based on non-uniform spectrogram sub-band division according to claim 2, wherein Specifically, in Step 4: Take the first five frames of this noisy speech signal as noise frames, and perform graph Fourier transform respectively using the graph Fourier transform basis obtained in Step 1 to obtain the graph spectra of five silent frames; Take the average of the graph spectra of the five noise frames to obtain the estimated graph energy spectrum of the noise: Among them represents the graph spectrum of the z-th frame noise frame, represents the average value of the graph spectra of five frame noise frames; According to the estimated graph energy spectrum of the noise, obtain the estimated signal-to-noise ratio of this noisy speech signal:
4. The voice spectrogram adaptive spectral subtraction method based on non-uniform spectrogram subband division according to claim 2, wherein In step five, according to the spectrogram of the noisy speech signal, the energy value of each graph sub-band in each frame of the speech signal is obtained. All the graph sub-bands are sorted in descending order according to the energy value, and the first B use graph sub-bands are marked as useful graph sub-bands.
5. The voice spectrogram adaptive spectral subtraction method based on non-uniform graph sub-band division according to claim 4, characterized in that, The calculation formula for the number of useful graph sub-bands is: where represents the number of useful graph sub-bands and SNR is the estimated signal-to-noise ratio.
6. The voice spectrogram adaptive spectral subtraction method based on non-uniform graph subband division according to claim 1, characterized in that The calculation formula for the energy value of the \(d\)-th graph sub-band is: where p is the serial number of the starting graph frequency point of the d-th graph sub-band, and q is the serial number of the ending graph frequency point of the d-th graph sub-band. is the amplitude spectrum value corresponding to the k-th graph frequency point of the noisy speech signal.
7. The voice spectrogram adaptive spectral subtraction method based on non-uniform spectrogram subband division according to claim 4, wherein The spectral subtraction strategy in Step 6 is as follows where is the amplitude spectrum value corresponding to the r-th graph frequency point of the denoised speech signal, is the amplitude spectrum value corresponding to the r-th graph frequency point of the noisy speech signal, is the amplitude spectrum value corresponding to the r-th graph frequency point of the noise signal.
8. The voice spectrogram adaptive spectral subtraction method based on non-uniform spectrogram sub-band division according to claim 7, wherein The new speech graph signal in Step 7 is: Among them,
Citation Information
Patent Citations
Speech enhancement method based on non-local mean filtering
CN103971697A
Target voice signal enhancing method based on continuous noise tracking, system and storage medium
CN109817234A