Modulation domain speech enhancement method based on group sparse representation
By using a modulation domain method based on group sparse representation, and leveraging the dynamic characteristics of the modulation spectrum and the structural characteristics of speech, a structured dictionary is trained and optimized using a Wiener filter. This solves the problems of lack of dynamic information and neglect of structural characteristics in traditional methods, and achieves better noise suppression and speech recovery results.
Patent Information
- Application Number
- CN202210615891.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-01
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2042-06-01
Smart Images

Figure CN114999516B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of single-channel speech enhancement, and more particularly to a modulation domain speech enhancement method based on group sparse representation. Background Technology
[0002] Traditional sparse representation speech enhancement methods (Supervised monaural speech enhancement using complementary joint sparse representations[J].IEEE Signal Processing Letters,2015,23(2):237-241; CN109087664A speech enhancement method) represent speech signals as acoustic spectra related to acoustic frequencies and time frames. A dictionary is trained using the acoustic amplitude spectra of speech and noise, and the mixed signal is projected onto the dictionary to recover clean speech. The acoustic spectrum can reflect the shape of the vocal tract, but it assumes that the vocal tract is stationary. Therefore, there are limitations to using acoustic domain features for speech enhancement. Studies have shown that speech enhancement in the modulation domain can achieve more noise suppression and less speech distortion than in the acoustic domain (Single-channel speech enhancement using spectral subtraction in the short-time modulation domain, Speech communication 52(2010)450–475.). As a three-dimensional feature of the modulation domain, the modulation spectrum is a function of acoustic frequency, time frame, and modulation frequency. By dynamically modeling the frequency envelope of the acoustic spectrum, the modulation spectrum can reflect the dynamic information of the vocal tract changing over time, and the characteristics of different modulation frequencies have different effects on the quality and intelligibility of speech.
[0003] Another drawback of traditional sparse representation methods is that they ignore the structured characteristics of speech. Speech signals possess inherent structured features, manifested as dependencies between features across different time frames. Traditional sparse representation methods fail to utilize this characteristic, treating atoms in the dictionary as isolated entities, which can easily lead to source confusion. Considering the sparsity of group structures, group sparse representation methods divide atoms into different groups for holistic selection, improving the robustness and accuracy of signal recovery.
[0004] As a three-dimensional feature representation, the modulation spectrum of speech can be divided into sub-band spectra with different modulation frequencies. By using the group sparse representation method, the structured characteristics of these sub-band spectra can be used as prior information to train a structured dictionary, thereby improving the performance of speech enhancement. Summary of the Invention
[0005] The technical problem solved by this invention is to overcome the lack of dynamic information caused by the use of acoustic domain features in existing technologies, as well as the shortcomings of ignoring the utilization of speech structure characteristics, and to provide a modulation domain speech enhancement method based on group sparse representation, which can greatly improve the performance of single-channel speech enhancement tasks.
[0006] The technical solution of the present invention is as follows: A modulation domain speech enhancement method based on group sparse representation, comprising:
[0007] Step 1: Training Phase
[0008] Step 11: Calculate the modulation spectrum of the mixed training signal, speech signal and noise signal, discard the acoustic phase and modulation phase in the modulation transformation process to obtain the modulation amplitude spectrum, and split the modulation amplitude spectrum into sub-band spectra according to different modulation frequencies;
[0009] Step 12: Perform frame clustering analysis based on correlation distance on each subband spectrum to obtain the signals of different clustered frames of the subband spectrum;
[0010] Step 13: Use each sub-band spectrum to train the mixed speech and noise structured dictionary for the corresponding modulation frequency. At a specific modulation frequency, first use the signals from different clustering frames of the sub-band spectrum to train the joint sub-dictionary, and then concatenate the sub-dictionaries into a structured dictionary.
[0011] Step 2, Testing Phase
[0012] Step 21: Calculate the sub-band spectrum of the mixed test signal (the mixed test signal is the noisy speech signal in the test phase, which corresponds to the mixed training signal), retain its acoustic phase and modulation phase, and use it for subsequent inverse transform.
[0013] Step 22: Calculate the projection coefficients of the subband spectrum of the hybrid test signal onto the hybrid structured dictionary using group sparse coding;
[0014] Step 23: Multiply the structured dictionary and projection coefficients to recover the modulation amplitude spectrum of speech and noise, and perform short-time Fourier inverse transform in combination with the modulation phase of the mixed test signal to obtain the acoustic amplitude spectrum of speech and noise.
[0015] Step 24: Calculate the weights of the two sets of estimation results based on the Gini coefficient, further optimize the estimation using the Wiener filter, and obtain the enhanced speech in the time domain after short-time Fourier inverse transform.
[0016] Furthermore, the calculation of the modulation spectrum of the mixed training signal, speech signal, and noise signal, discarding the acoustic phase and modulation phase during the modulation transformation process to obtain the modulation amplitude spectrum, and splitting the modulation amplitude spectrum into sub-band spectra according to different modulation frequencies includes:
[0017] Let the pure speech signal be s(t), the noise signal be n(t), and the mixed training signal x(t) be expressed as:
[0018] x(t) = s(t) + n(t)
[0019] t is the discrete-time exponent. Performing a short-time Fourier transform on x(t) yields the acoustic spectrum X(t,f) of the mixed training signal, where f is the acoustic frequency exponent. X(t,f) can be expressed in pole form:
[0020]
[0021] Where |X(t,f)| is the acoustic amplitude spectrum, X p (t,f) represents the acoustic phase spectrum, and || denotes the modulo operation; the acoustic amplitude spectrum is written in matrix form. Acoustic phase spectrum is written as F and J represent the number of acoustic frequency points and the number of acoustic time frames, respectively; similarly, the acoustic amplitude spectra S and N of the speech and noise are obtained. The modulation transform requires a second short-time Fourier transform on |X(t,f)|:
[0022]
[0023] in, Let v(t) be the modulation spectrum of the mixed training signal, v(t) be the modulation window function, r be the modulation time exponent, m be the modulation frequency exponent, and L be the modulation frame duration. It can be represented in pole form:
[0024]
[0025] For modulation amplitude spectrum, To modulate the phase spectrum, the modulation amplitude spectrum is represented in matrix form. The modulation phase spectrum is represented as F, V, and M represent the number of acoustic frequency points, the number of modulation time frames, and the number of modulation frequencies, respectively; The subband spectrum at the modulation frequency m is represented as follows: m∈{1,2,...,M}; similarly, the modulation amplitude spectra of speech and noise are obtained. and and his son's genealogy and This step separates the subband spectrum of the signal at different modulation frequencies through modulation transformation, and subsequent steps process the characteristics of different modulation frequencies respectively.
[0026] Furthermore, the frame clustering analysis based on correlation distance for each sub-band spectrum, to obtain signals of different clustered frames of the sub-band spectrum, includes:
[0027] At modulation frequency m, the K-means algorithm is used to extract the speech subband spectrum. and noise subband spectrum The frames are divided into K clusters according to the Pearson correlation distance, and the signals of the speech and noise subband spectra of different clusters are obtained:
[0028]
[0029]
[0030] For mixed training signal subband spectrum Clustering makes and and By sharing frame labels, signals from different clustering frames of the mixed training signal subband spectrum are obtained:
[0031]
[0032]
[0033] Among them, shared-Labels1 indicates that it adopts The frame label, shared-Labels2, indicates that it uses... The frame labels. This step involves performing correlation frame clustering analysis on the subband spectrum of the training signal to form a defined group pattern, which will be used to train the structured dictionary in subsequent steps.
[0034] Furthermore, a mixed speech and noise structured dictionary is trained using each sub-band spectrum for the corresponding modulation frequency. At a specific modulation frequency, a joint sub-dictionary is first trained using signals from different clustering frames of the sub-band spectrum, and then the sub-dictionaries are concatenated into a structured dictionary, including:
[0035] At modulation frequency m, a joint sub-dictionary for different clusters is trained sequentially. For the i-th cluster, i∈{1,2,...,K}, a hybrid-speech matrix is constructed. and hybrid noise matrix Training a joint sub-dictionary using a sparse constraint learning algorithm and The learning process is as follows:
[0036]
[0037]
[0038] in, The sparse coefficients of the hybrid-speech matrix, for The kth column, The sparse coefficients of the mixture-noise matrix, for The k-th column, where q is a sparse constraint; This represents the Frobenius norm, and ||||1 represents the 1-norm; after the above learning process is carried out in K clusters, the sub-dictionaries are concatenated to obtain a structured dictionary:
[0039]
[0040]
[0041]
[0042]
[0043] in, and Two sets of hybrid structured dictionaries, A structured dictionary for speech. This step trains a structured dictionary for noise using subband spectra with predefined group patterns. Compared to the disordered distribution of atoms in unstructured dictionaries of traditional methods, the atoms in the structured dictionary exhibit obvious clustering characteristics.
[0044] Furthermore, the subband spectrum of the mixed test signal is calculated, preserving its acoustic phase and modulation phase, for subsequent inverse transform, including:
[0045] Let the mixed test signal be x. test (t), calculate its subband spectrum m∈{1,2,...,M}, retain the acoustic phase in the two short-time Fourier transforms. and modulation phase This will be used in the subsequent inverse transform. This step extracts the components of the mixed test signal at different modulation frequencies and saves the noisy components and modulation phase for subsequent inverse transform.
[0046] Furthermore, the calculation of the projection coefficients of the subband spectrum of the hybrid test signal onto the hybrid structured dictionary through group sparse coding includes:
[0047] At the modulation frequency m, calculate the subband spectrum of the mixed test signal. In hybrid structured dictionaries Projection coefficients on and because and Having group structure characteristics, the sparse coding process uses the following objective function:
[0048]
[0049]
[0050] in, for Zhongyu The corresponding submatrix composed of projection coefficients, for Zhongyu The corresponding submatrix composed of projection coefficients, for The kth column, for In the k-th column, the α1 and α2 regularization coefficients adjust the group sparsity, while the β1 and β2 regularization coefficients adjust the individual sparsity within the group. This step generates sparse coefficients optimized at both the group and individual levels through group sparse coding, which can reduce source confusion and thus more robustly recover speech from noise.
[0051] Furthermore, the step of multiplying the structured dictionary and projection coefficients to recover the modulation amplitude spectrum of speech and noise, and then performing a short-time Fourier inverse transform based on the modulation phase of the mixed test signal to obtain the acoustic amplitude spectrum of speech and noise includes:
[0052] At modulation frequency m, based on the mapping relationship between the mixed signal and speech and noise, the estimated speech and noise subband spectra are recovered by multiplying the structured dictionary and projection coefficients:
[0053]
[0054]
[0055] in, The speech subband spectrum estimated based on the hybrid-speech mapping. The noise sub-band spectrum is estimated based on the hybrid-noise mapping. After the above estimation process is performed at M modulation frequencies, the sub-band spectra of speech and noise are synthesized into a complete modulation amplitude spectrum, which is then combined with the noisy modulation phase. The estimated acoustic amplitude is obtained by inverse short-time Fourier transform. and Another set of complementary estimates was calculated based on the additive mixing model:
[0056]
[0057]
[0058] Among them, X test This step involves first recovering the different modulation frequency components of the speech and noise, and then performing an inverse transform to obtain the acoustic domain estimation result.
[0059] Furthermore, the weights of the two sets of estimation results are calculated based on the Gini coefficient, and the estimation is further optimized using a Wiener filter. The enhanced speech in the time domain obtained after short-time Fourier inverse transform includes:
[0060] Because the two sets of acoustic amplitude spectrum estimation results have different recovery accuracies, the weighted average result is calculated using the Gini coefficient:
[0061]
[0062]
[0063] Where λ is the Gini coefficient, which is related to the singular value distribution of the noise signal. The optimal estimate in the least mean square sense is obtained through the Wiener filter:
[0064]
[0065] in,() 2 This represents the operation of squaring matrix elements. S represents the matrix dot product operation. test This refers to the enhanced acoustic amplitude of the speech, combined with the noisy acoustic phase. The acoustic spectrum in complex form is obtained, and the enhanced speech in the time domain is obtained after inverse short-time Fourier transform. This step involves weighting the two sets of estimated acoustic domain results and optimizing them based on Wiener filtering, which can improve the estimation accuracy.
[0066] The advantages of this invention compared to existing technologies are as follows: Compared to traditional methods, this invention processes the features of different modulation frequencies separately, and uses the features of different clustered frames in the subband spectrum to learn a structured dictionary. Based on a predetermined group pattern, it can more robustly recover speech from noise, thereby improving the performance of single-channel speech enhancement. Attached Figure Description
[0067] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0068] Figure 1 A flowchart of a modulation domain speech enhancement method based on group sparse representation provided in an embodiment of the present invention;
[0069] Figure 2The algorithm provided in this embodiment of the invention is compared with the segmented signal-to-noise ratio (segSNR) of the Complementary Joint Sparse Representation (CJSR) algorithm under the conditions of an average of 5 noise backgrounds, 20 speakers, and input signal-to-noise ratios of -5dB, 0dB, 5dB, and 10dB.
[0070] Figure 3 The algorithm provided in this embodiment of the invention is compared with the PESQ of the Complementary Joint Sparse Representation (CJSR) algorithm under the conditions of an average of 5 noise backgrounds, 20 speakers, and input signal-to-noise ratios of -5dB, 0dB, 5dB, and 10dB.
[0071] Figure 4 The source signal interference ratio (SIR) of the algorithm provided in this embodiment of the invention is compared with that of the Complementary Joint Sparse Representation (CJSR) algorithm under the conditions of an average of 5 noise backgrounds, 20 speakers, and input signal-to-noise ratios of -5dB, 0dB, 5dB, and 10dB.
[0072] Figure 5 A detailed block diagram of the method provided in the embodiments of the present invention. Detailed Implementation
[0073] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.
[0074] like Figure 1 , 5 As shown, the modulation domain speech enhancement method based on group sparse representation of the present invention includes the following steps:
[0075] Step 1: Training Phase
[0076] Step 11: Calculate the modulation spectrum of the mixed training signal, speech signal and noise signal, discard the acoustic phase and modulation phase in the modulation transformation process to obtain the modulation amplitude spectrum, and split the modulation amplitude spectrum into sub-band spectra according to different modulation frequencies;
[0077] Let the pure speech signal be s(t), the noise signal be n(t), and the mixed training signal x(t) be expressed as:
[0078] x(t) = s(t) + n(t)
[0079] t is the discrete-time exponent. Performing a short-time Fourier transform on x(t) yields the acoustic spectrum X(t,f) of the mixed training signal, where f is the acoustic frequency exponent. X(t,f) can be expressed in pole form:
[0080]
[0081] Where |X(t,f)| is the acoustic amplitude spectrum, X p (t,f) represents the acoustic phase spectrum, and || denotes the modulo operation; the acoustic amplitude spectrum is written in matrix form. Acoustic phase spectrum is written as F and J represent the number of acoustic frequency points and the number of acoustic time frames, respectively; similarly, the acoustic amplitude spectra S and N of the speech and noise are obtained. The modulation transform requires a second short-time Fourier transform on |X(t,f)|:
[0082]
[0083] in, Let v(t) be the modulation spectrum of the mixed training signal, v(t) be the modulation window function, r be the modulation time exponent, m be the modulation frequency exponent, and L be the modulation frame duration. It can be represented in pole form:
[0084]
[0085] For modulation amplitude spectrum, To modulate the phase spectrum, the modulation amplitude spectrum is represented in matrix form. The modulation phase spectrum is represented as F, V, and M represent the number of acoustic frequency points, the number of modulation time frames, and the number of modulation frequencies, respectively; The subband spectrum at the modulation frequency m is represented as follows: m∈{1,2,...,M}; similarly, the modulation amplitude spectra of speech and noise are obtained. and and his son's genealogy and This step separates the subband spectrum of the signal at different modulation frequencies through modulation transformation, and subsequent steps process the characteristics of different modulation frequencies respectively.
[0086] Step 12: Perform frame clustering analysis based on correlation distance on each sub-band spectrum to obtain the signals of different clustered frames of the sub-band spectrum; at the modulation frequency m, use the K-means algorithm to divide the speech sub-band spectrum. and noise subband spectrum The frames are divided into K clusters according to the Pearson correlation distance, and the signals of the speech and noise subband spectra of different clusters are obtained:
[0087]
[0088]
[0089] For mixed training signal subband spectrum Clustering makes and and By sharing frame labels, signals from different clustering frames of the mixed training signal subband spectrum are obtained:
[0090]
[0091]
[0092] Among them, shared-Labels1 indicates that it adopts The frame label, shared-Labels2, indicates that it uses... The frame labels. This step involves performing correlation frame clustering analysis on the subband spectrum of the training signal to form a defined group pattern, which will be used to train the structured dictionary in subsequent steps.
[0093] Step 13: Use each sub-band spectrum to train the mixed speech and noise structured dictionary for the corresponding modulation frequency. At a specific modulation frequency, first use the signals from different clustering frames of the sub-band spectrum to train the joint sub-dictionary, and then concatenate the sub-dictionaries into a structured dictionary.
[0094] At modulation frequency m, a joint sub-dictionary for different clusters is trained sequentially. For the i-th cluster, i∈{1,2,...,K}, a hybrid-speech matrix is constructed. and hybrid noise matrix Training a joint sub-dictionary using a sparse constraint learning algorithm and The learning process is as follows:
[0095]
[0096]
[0097] in, The sparse coefficients of the hybrid-speech matrix, C The kth column, The sparse coefficients of the mixture-noise matrix, for The k-th column, where q is a sparse constraint; This represents the Frobenius norm, and ||||1 represents the 1-norm; after the above learning process is carried out in K clusters, the sub-dictionaries are concatenated to obtain a structured dictionary:
[0098]
[0099]
[0100]
[0101]
[0102] in, and Two sets of hybrid structured dictionaries, A structured dictionary for speech. This step trains a structured dictionary for noise using subband spectra with predefined group patterns. Compared to the disordered distribution of atoms in unstructured dictionaries of traditional methods, the atoms in the structured dictionary exhibit obvious clustering characteristics.
[0103] Step 2, Testing Phase
[0104] Step 21: Calculate the subband spectrum of the mixed test signal, retain its acoustic phase and modulation phase, and use it for subsequent inverse transform;
[0105] Let the mixed test signal be x. test (t), calculate its subband spectrum m∈{1,2,...,M}, retain the acoustic phase in the two short-time Fourier transforms. and modulation phase This will be used in the subsequent inverse transform. This step extracts the components of the mixed test signal at different modulation frequencies and saves the noisy components and modulation phase for subsequent inverse transform.
[0106] Step 22: Calculate the projection coefficients of the subband spectrum of the hybrid test signal onto the hybrid structured dictionary using group sparse coding; calculate the subband spectrum of the hybrid test signal at the modulation frequency m. In hybrid structured dictionaries Projection coefficients on and because and Having group structure characteristics, the sparse coding process uses the following objective function:
[0107]
[0108]
[0109] in, for Zhongyu The corresponding submatrix composed of projection coefficients, for Zhongyu The corresponding submatrix composed of projection coefficients, for The kth column, for In the k-th column, the α1 and α2 regularization coefficients adjust the group sparsity, while the β1 and β2 regularization coefficients adjust the individual sparsity within the group. This step generates sparse coefficients optimized at both the group and individual levels through group sparse coding, which can reduce source confusion and thus more robustly recover speech from noise.
[0110] Step 23: Multiply the structured dictionary and projection coefficients to recover the modulation amplitude spectrum of speech and noise, and perform short-time Fourier inverse transform in combination with the modulation phase of the mixed test signal to obtain the acoustic amplitude spectrum of speech and noise.
[0111] At modulation frequency m, based on the mapping relationship between the mixed signal and speech and noise, the estimated speech and noise subband spectra are recovered by multiplying the structured dictionary and projection coefficients:
[0112]
[0113]
[0114] in, The speech subband spectrum estimated based on the hybrid-speech mapping. The noise sub-band spectrum is estimated based on the hybrid-noise mapping. After the above estimation process is performed at M modulation frequencies, the sub-band spectra of speech and noise are synthesized into a complete modulation amplitude spectrum, which is then combined with the noisy modulation phase. The estimated acoustic amplitude is obtained by inverse short-time Fourier transform. and Another set of complementary estimates was calculated based on the additive mixing model:
[0115]
[0116]
[0117] Among them, X test This step involves first recovering the different modulation frequency components of the speech and noise, and then performing an inverse transform to obtain the acoustic domain estimation result.
[0118] Step 24: Calculate the weights of the two sets of estimation results based on the Gini coefficient, further optimize the estimation using the Wiener filter, and obtain the enhanced speech in the time domain after short-time Fourier inverse transform.
[0119] Because the two sets of acoustic amplitude spectrum estimation results have different recovery accuracies, the weighted average result is calculated using the Gini coefficient:
[0120]
[0121]
[0122] Where λ is the Gini coefficient, which is related to the singular value distribution of the noise signal. The optimal estimate in the least mean square sense is obtained through the Wiener filter:
[0123]
[0124] in,() 2 This represents the operation of squaring matrix elements. S represents the matrix dot product operation. test This refers to the enhanced acoustic amplitude of the speech, combined with the noisy acoustic phase. The acoustic spectrum in complex form is obtained, and the enhanced speech in the time domain is obtained after inverse short-time Fourier transform. This step involves weighting the two sets of estimated acoustic domain results and optimizing them based on Wiener filtering, which can improve the estimation accuracy.
[0125] This invention processes speech features of different modulation frequencies separately, and at the same time uses the structured features of the modulation domain to train a structured dictionary. Based on a predetermined group pattern, it can more robustly recover speech signals from noise.
[0126] Figure 2 , Figure 3 and Figure 4 The results show the comparison of segSNR, PESQ, and SIR between this invention and the CJSR algorithm under the conditions of an average of 5 types of noise backgrounds, 20 speakers, and input signal-to-noise ratios of -5dB, 0dB, 5dB, and 10dB. Figure 2 , 3 As can be seen from points 1 and 4, the present invention can effectively improve the quality and intelligibility of speech under various conditions, especially in low signal-to-noise ratio environments.
[0127] Through the above description of the embodiments, those skilled in the art can clearly understand that the above embodiments can be implemented by software, or by using software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solutions of the above embodiments can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, mobile hard drive, etc.), including several instructions to cause a computer device (such as a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0128] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A modulation domain speech enhancement method based on group sparse representation, characterized in that, The method comprises the following steps: Step 1, training phase Step 11, calculating the modulation spectrum of the mixed training signal, the speech signal and the noise signal, discarding the acoustic phase and the modulation phase in the modulation transformation process to obtain the modulation amplitude spectrum, and splitting the modulation amplitude spectrum into sub-band spectra according to different modulation frequencies; Step 12, performing frame clustering analysis on each sub-band spectrum based on correlation distance to obtain signals of different clustering frames of the sub-band spectrum; Step 13, training the mixed, speech and noise structured dictionaries of the corresponding modulation frequencies by using each sub-band spectrum, and at a specific modulation frequency, first training a joint sub-dictionary by using the signals of the different clustering frames of the sub-band spectrum, and then splicing the sub-dictionary into a structured dictionary; Step 2, test phase Step 21, calculating the sub-band spectrum of the mixed test signal, retaining the acoustic phase and the modulation phase thereof for subsequent inverse transformation; Step 22, calculating the projection coefficient of the sub-band spectrum of the mixed test signal on the mixed structured dictionary by group sparse coding; Step 23, multiplying the structured dictionary and the projection coefficient to recover the speech and noise modulation amplitude spectrum, and performing inverse short-time Fourier transform on the modulation phase of the mixed test signal to obtain the acoustic amplitude spectrum of the speech and the noise; Step 24, calculating the weight of two groups of estimation results based on the Gini coefficient, the two groups of estimation results being estimated amplitudes, the two groups of estimation results being complementary, further optimizing the estimation by using a Wiener filter, and obtaining the enhanced speech in the time domain after inverse short-time Fourier transform.
2. The modulation domain speech enhancement method based on group sparse representation according to claim 1, characterized in that, In the step 11, the step of calculating the modulation spectrum of the mixed training signal, the speech signal and the noise signal, discarding the acoustic phase and the modulation phase in the modulation transformation process to obtain the modulation amplitude spectrum, and splitting the modulation amplitude spectrum into sub-band spectra according to different modulation frequencies comprises the following steps: Let the clean speech signal be , the noise signal be , and the mixed training signal be represented as: For discrete-time exponential, a Short-time Fourier transform gives the acoustic spectrum of the mixed training signal , For acoustic frequency, a is expressed in pole form: wherein is the acoustic magnitude spectrum, is the acoustic phase spectrum, denotes a modulo operation, the acoustic magnitude spectrum is written in matrix form , the acoustic phase spectrum is written as , and are the number of acoustic frequency bins and the number of acoustic time frames, respectively; likewise, the acoustic magnitude spectrum of the speech and the noise are obtained as and , the modulation transform requires a second short-time Fourier transform on . wherein is the modulation spectrum of the mixed training signal, is the modulation window function, is the modulation time exponent, is the modulation frequency exponent, is the modulation frame duration, is expressed in pole form: The modulation amplitude spectrum is denoted as The modulation phase spectrum is denoted as , , , , are the number of acoustic frequency points, the number of modulation time frames and the number of modulation frequencies, respectively; the subband spectrum at modulation frequency is denoted as , Similarly, the modulation amplitude spectrum of speech and noise and is obtained, as well as their subband spectra and . 3. The modulation domain speech enhancement method based on group sparse representation according to claim 1, characterized in that, In the step 12, the step of performing frame clustering analysis on each sub-band spectrum based on correlation distance to obtain signals of different clustering frames of the sub-band spectrum comprises the following steps: At the modulation frequency , the frames of speech subband spectrum and noise subband spectrum are divided into clusters according to the Pearson correlation distance using the K-means algorithm, obtaining the signals of different cluster frames of speech and noise subband spectrum: For the mixed training signal subband spectrum of clusters, make with and share frame labels, get the signal of different cluster frames of the mixed training signal subband spectrum: Where shared-Labels 1 indicates the use of frame labels of the form and shared-Labels 2 indicates the use of frame labels of the form 4. The modulation domain speech enhancement method based on group sparse representation according to claim 1, characterized in that, In the step 13, the step of training the structured dictionary of the corresponding modulation frequencies by using each sub-band spectrum, and at a specific modulation frequency, first training a joint sub-dictionary by using the signals of the different clustering frames of the sub-band spectrum, and then splicing the sub-dictionary into a structured dictionary comprises the following steps: At the modulation frequency , the joint sub-dictionaries of different clusters are trained in turn, for the first cluster, , a mixed-speech matrix and a mixed-noise matrix are constructed, the joint sub-dictionary and are trained using a sparse constraint learning algorithm, and the learning process is as follows: wherein, are sparse coefficients of the mixed-speech matrix, are sparse coefficients of the mixed-speech matrix, are the first columns of the mixed-speech matrix, are sparse coefficients of the mixed-noise matrix, are sparse coefficients of the mixed-noise matrix, are the first columns of the mixed-noise matrix, are sparse constraints; denotes the Frobenius norm, denotes the 1-norm; the above learning process is performed in clusters, the sub-dictionaries are concatenated to obtain the structured dictionary: wherein, and are two sets of mixed structured dictionaries, is a speech structured dictionary, is a noise structured dictionary.
5. The modulation domain speech enhancement method based on group sparse representation according to claim 1, characterized in that, In the step 21, the step of calculating the sub-band spectrum of the mixed test signal, retaining the acoustic phase and the modulation phase thereof for subsequent inverse transformation comprises the following steps: Let the mixed test signal be , compute its subband spectrum , , retain the acoustic phase and the modulation phase in both short-time Fourier transforms for later inverse transform usage.
6. The modulation domain speech enhancement method based on group sparse representation according to claim 4, characterized in that, In the step 22, the step of calculating the projection coefficient of the sub-band spectrum of the mixed test signal on the mixed structured dictionary by group sparse coding comprises the following steps: At modulation frequency At this point, calculate the subband spectrum of the mixed test signal. In hybrid structured dictionaries , Projection coefficients on and ,because and Having group structure characteristics, the sparse coding process uses the following objective function: wherein, is wherein a sub-matrix composed of projection coefficients corresponding to is wherein a sub-matrix composed of projection coefficients corresponding to is the first column of is the first column of , regular coefficients can adjust the group sparsity, , regular coefficients can adjust the individual sparsity within the group.
7. The modulation domain speech enhancement method based on group sparse representation according to claim 6, characterized in that, In the step 23, the step of multiplying the structured dictionary and the projection coefficient to recover the speech and noise modulation amplitude spectrum, and performing inverse short-time Fourier transform on the modulation phase of the mixed test signal to obtain the acoustic amplitude spectrum of the speech and the noise comprises the following steps: At the modulation frequency At the modulation frequency, the structured dictionary and the projection coefficients are multiplied to recover the estimated speech and noise subband spectra according to the mapping of the mixed signal to speech and noise: where is the speech subband spectrum estimated from the speech-noise map, is the noise subband spectrum estimated from the speech-noise map; the estimation process is described in After the estimation of the subband spectra of speech and noise in the The estimated acoustic amplitude is obtained by inverse short-time Fourier transform and A complementary set of estimates is computed from the additive mixture model: wherein, is the acoustic amplitude of the mixed test signal.
8. The modulation domain speech enhancement method based on group sparse representation according to claim 7, characterized in that, In the step 24, the step of calculating the weight of two groups of estimation results based on the Gini coefficient, further optimizing the estimation by using a Wiener filter, and obtaining the enhanced speech in the time domain after inverse short-time Fourier transform comprises the following steps: Because the recovery precisions of the two groups of acoustic amplitude spectrum estimation results are different, the Gini coefficient is used to calculate the weighted average result. wherein is the Gini coefficient, related to the singular value distribution of the noise signal, and the optimal estimation result in the sense of minimum mean square is obtained by the Wiener filter: wherein, denotes a matrix element square operation, denotes a matrix point multiplication operation, is the enhanced speech acoustic amplitude, combined with the noisy acoustic phase results in a complex acoustic spectrum, and an inverse short-time Fourier transform results in the time-domain enhanced speech.
Citation Information
Patent Citations
Speech enhancement method
CN109087664A
Adaptive sparse-tree structure noise reduction method of very noisy vibration signal of main reducer
CN108844617A
Single-channel speech enhancement method based on joint dictionary learning and sparse representation
CN111508518A