Speech Enhancement Algorithm Based on Improved Phase Spectrum Compensation and Fully Convolutional Neural Network

By improving the combination of phase spectrum compensation and fully convolutional neural networks, the phase compensation factor and simple loss function are optimized by frame signal-to-noise ratio, the problem of poor enhancement effect of traditional speech enhancement algorithms under non-stationary noise is solved, and the speech quality and intelligibility are improved.

CN114242099BActive Publication Date: 2025-07-22NANJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111534489.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-15
Publication Date
2025-07-22
Estimated Expiration
2041-12-15

AI Technical Summary

Technical Problem

When existing speech enhancement algorithms deal with non-stationary noise, traditional phase compensation factors cannot be dynamically adjusted, resulting in poor enhancement effect. Frequency domain methods ignore phase information, and time domain methods rely on complex loss functions to lead to high training difficulty and speech distortion.

Method used

The logarithmic power spectrum with noisy speech is used as the training feature, combined with the improved phase spectrum compensation algorithm and the fully convolutional neural network, the phase compensation factor is optimized through the frame signal-to-noise ratio, and a simple and effective loss function is designed, and the phase compensation and logarithmic power spectrum are fused as training goals to build a fully convolutional neural network model.

Benefits of technology

It improves the noise cancellation ability and intelligibility of speech enhancement, overcomes the shortcomings of traditional methods, and achieves better enhancement effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114242099B_ABST
    Figure CN114242099B_ABST
Patent Text Reader

Abstract

The present invention proposes a speech enhancement algorithm based on improved phase spectrum compensation and fully convolutional neural network. The feature data of noisy speech and clean speech are obtained by preprocessing the speech data in the training set. Combining the feature data of noisy speech and clean speech, the compensation factor calculation formula of the frame signal-to-noise ratio optimized phase compensation algorithm is introduced, and then the improved formula is used to calculate the phase compensation factor. A fully convolutional neural network model is built, and the phase compensation factor combined with the logarithmic power spectrum of clean speech is used as the training target of the fully convolutional neural network for network model training. The test speech is input into the trained model to obtain the estimated value of the logarithmic power spectrum and the phase compensation function. The amplitude spectrum and phase spectrum of the speech signal are reconstructed respectively using the estimated value of the logarithmic power spectrum and the phase compensation function to obtain the final enhanced speech. The present invention improves the noise cancellation ability of the algorithm while better ensuring the speech intelligibility, thus improving the overall effect of speech enhancement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a voice enhancement method, specifically to a voice enhancement algorithm based on improved phase spectrum compensation and fully convolutional neural network, belonging to the technical field of voice signal processing. Background Art

[0002] It is understood that voice is an important way of information exchange between people. However, during the process of people using voice for communication, it is always interfered by various noises. Noisy voice will not only increase people's auditory fatigue and reduce the quality of voice communication, but also degrade the performance of voice processing systems based on feature parameter extraction. Therefore, in order to reduce the impact of background noise on voice quality, voice enhancement is needed to suppress background noise.

[0003] Phase spectrum compensation is an enhancement algorithm that uses voice phase spectrum information to enhance voice signals. Its basic idea is: calculate the short-time amplitude spectrum of the noisy voice signal and the short-time amplitude spectrum of the estimated noise signal respectively, calculate the compensation factor using the phase spectrum compensation function, and then superimpose the compensation factor on the spectrum of the noisy voice to obtain the compensated voice spectrum. When restoring the enhanced voice signal, the compensated phase is obtained using the compensated voice spectrum, and then the amplitude spectrum of the noisy voice signal is inserted, and the inverse discrete Fourier transform is performed. The general form of the phase compensation function is:

[0004] is the estimated value of the noise amplitude spectrum, Ψ(k) is an antisymmetric function, λ is the compensation factor, and the traditional phase compensation factor λ is an empirical value, generally taken as 3.74.

[0005] The advantage of phase compensation is that the computational complexity is small, it is easy to implement, and the enhancement effect is also good. However, because most current voice enhancement algorithms deal with non-stationary noise and the noise energy is uncertain, a fixed compensation factor cannot be dynamically adjusted to a suitable value according to the change of noise. The fixed parameters cannot fully exert the effect of the phase compensation algorithm, nor can they combine DNN to correct the voice phase spectrum. Therefore, optimizing the compensation function with fixed parameters, introducing supervised parameter learning, and weighing the voice distortion and noise reduction effect after enhancement are the key points for improving the phase compensation algorithm to give full play to its own advantages.

[0006] The speech enhancement method based on deep neural network significantly improves the speech enhancement performance under non-stationary noise conditions compared with traditional methods, and has become a research hotspot in the field of speech enhancement in recent years. Existing work mainly focuses on the design of training features and training objectives as well as the improvement of network structure. According to the design methods of training features and objectives, the speech enhancement methods based on deep neural network can be divided into two categories: time domain and frequency domain. In frequency domain speech enhancement, the magnitude spectrum or log power spectrum of noisy speech is generally used as training features. In addition to the magnitude spectrum and log power spectrum, the magnitude spectrum mask can also be used as a training objective. In time domain speech enhancement, the time domain waveforms of noisy speech and clean speech are generally used as training features and training objectives respectively. However, there are still many problems in the existing solutions. Frequency domain speech enhancement often ignores the phase information of speech. Recent research has found that only enhancing the phase spectrum of speech and keeping the magnitude spectrum of noisy speech unchanged can effectively improve speech quality. While the time domain speech enhancement algorithm directly uses the waveforms of noisy speech and clean speech as training features and training objectives, its performance is very dependent on the design of the loss function. A complex loss function greatly increases the difficulty of training. On the contrary, using a simple time domain least mean square error function as the loss function requires a large amount of time for parameter tuning, which is prone to speech distortion problems, affecting the intelligibility of speech signals, damaging speech signals, and even reducing the signal-to-noise ratio. Summary of the Invention

[0007] The technical problem to be solved by the present invention is to overcome the deficiencies of the prior art and provide a speech enhancement algorithm based on improved phase spectrum compensation and fully convolutional neural network. The log power spectrum of noisy speech is used as training features, and the improved phase spectrum compensation algorithm is used to estimate the phase spectrum of the noisy signal. The obtained phase spectrum compensation factor is used as one of the training objectives of the network. Combined with a unique loss function design, the log power spectrum of the clean speech signal is used as the common training objective. Considering that the training features are correlated in both time and frequency, the present invention uses a convolutional neural network to obtain better training effects, so as to enhance speech signals.

[0008] The present invention provides a speech enhancement algorithm based on improved phase spectrum compensation and fully convolutional neural network, including the following steps:

[0009] Step 1, preprocess the speech data in the training set to obtain the feature data of noisy speech and clean speech;

[0010] Step 2, combine the obtained feature data of noisy speech and clean speech, introduce the compensation factor calculation formula of the frame signal-to-noise ratio optimized phase compensation algorithm, and then calculate the phase compensation factor using the improved formula;

[0011] Step 3: Build a fully convolutional neural network model. Use the obtained phase compensation factor and the logarithmic power spectrum of the clean speech as the training target of the fully convolutional neural network (FCNN), and train the fully convolutional neural network model;

[0012] Step 4: Input the test speech into the trained model to obtain the estimated value of the logarithmic power spectrum and the phase compensation function;

[0013] Step 5: Use the estimated value of the logarithmic power spectrum and the phase compensation function obtained in the previous step to reconstruct the amplitude spectrum and phase spectrum of the speech signal respectively to obtain the final enhanced speech. As a further technical solution of the present invention,

[0014] The further optimized technical solution of the present invention is as follows:

[0015] In the said Step 1, preprocess the noisy speech and the clean speech and extract the feature parameters of the noisy speech and the clean speech. The specific operations are as follows:

[0016] Step 1-1: Let y(n) represent the noisy speech signal and y(n)=d(n)+x(n), where d(n) is the noise signal and x(n) is the clean signal. Assume that x(n) and d(n) are statistically independent and have zero mean. Window M samples of y(n) with the window function w(n) and perform M-point FFT (Fast Fourier Transform) to transform the noisy speech into the frequency domain to obtain the spectrum Y(l,k) of the noisy speech, where l is the frame number label and k represents the frequency component and k = 0, 1, 2, …, M - 1; Similarly, the spectrum X(l,k) of the clean speech x(n) and the spectrum D(l,k) of the noise signal d(n) can be obtained;

[0017] Step 1-2: Represent the spectrum Y(l,k) of the noisy speech in polar coordinates, which can be divided into an amplitude spectrum and a phase spectrum, that is

[0018] Y(l,k)=|Y(l,k)|e j∠Y(l,k)

[0019] where |Y(l,k)| is the short-time amplitude spectrum of the noisy speech y(n), and ∠Y(l,k) is the phase spectrum; Similarly, the short-time amplitude spectrum |X(l,k)| of the clean speech x(n) and the short-time amplitude spectrum |D(l,k)| of the noise signal d(n) can be obtained;

[0020] Step 1-3: Use the following formula to calculate the logarithmic power spectrum S(n) of the noisy speech and the logarithmic power spectrum T(n) of the clean speech,

[0021] S(n)=[log e (|Y(n,1)| 2 ),log e(|Y(n,2)| 2 ),…, log e (|Y(n,k)| 2 ),…, log e (|Y(n,M - 1)| 2 )]

[0022] T(n)=[log e (|X(n,1)| 2 ), log e (|X(n,2)| 2 ),…, log e (|X(n,k)| 2 ),…, log e (|X(n,M - 1)| 2 )]

[0023] Where, |Y(n,1)| 2 is the power of the first frequency band of the n - th frame obtained by short - time Fourier transform of the noisy speech, |Y(n,2)| 2 is the power of the second frequency band of the n - th frame obtained by short - time Fourier transform of the noisy speech, |Y(n,k)| 2 is the power of the k - th frequency band of the n - th frame obtained by short - time Fourier transform of the noisy speech, |Y(n,M - 1)| 2 is the power of the (M - 1) - th frequency band of the n - th frame obtained by short - time Fourier transform of the noisy speech; |X(n,1)| 2 is the power of the first frequency band of the n - th frame obtained by short - time Fourier transform of the clean speech, |X(n,2)| 2 is the power of the second frequency band of the n - th frame obtained by short - time Fourier transform of the clean speech, |X(n,k)| 2 is the power of the k - th frequency band of the n - th frame obtained by short - time Fourier transform of the clean speech, |X(n,M - 1)| 2 is the power of the (M - 1) - th frequency band of the n - th frame obtained by short - time Fourier transform of the clean speech.

[0024] In step 2, optimize the phase compensation function, and the specific operation is as follows:

[0025] Step 2 - 1. The traditional phase compensation function is expressed as

[0026]

[0027] Where, Λ(l,k) is the phase compensation function, is the estimated value of the noise, generally replaced by the magnitude spectrum |Y(l,k)| of the noisy speech, λ is an empirical value, Ψ(k) is an antisymmetric function for correcting the phase, denoted as

[0028]

[0029] The function of Λ(l,k) is shown in the following formula

[0030] Y Λ (l,k) = Y(l,k) + Λ(l,k)

[0031] where Y Λ (l,k) is the compensated spectrum. By extracting the phase of the compensated spectrum, the phase spectrum ∠Y Λ (l,k) is obtained as follows

[0032] ∠Y Λ (l,k) = arg(Y Λ (l,k))

[0033] Combining ∠Y Λ (l,k) with the magnitude spectrum |Y(l,k)| of the noisy speech, the spectral expression of the enhanced speech is obtained as

[0034]

[0035] where the exponential form of a complex number is represented by Euler's formula re jα , r is the magnitude, α represents the argument (phase), and here j is the representation method in the complex number domain;

[0036] Step 2-2: Introduce the frame signal-to-noise ratio to optimize the phase compensation function, and obtain the optimized phase compensation function, that is

[0037]

[0038] where c is an empirical value, generally set to 2.7, Λ(l,k) decreases as SNR l increases. When the current frame is a speech frame, the influence of the phase compensation factor Λ(l,k) on the noisy speech decreases, and more speech details are retained. SNR l is the signal-to-noise ratio of the l-th frame.

[0039] Step 2 also includes: Step 2-3: To simplify the training objective of the subsequent neural network, the optimized phase compensation function is abbreviated as the following formula

[0040] Λ(l,k) = Ψ(k) × Q(l,k)

[0041] Among them, the above formula is the product of the new compensation factor and the estimated value of the noise amplitude spectrum, which can also be simply referred to as the new phase compensation factor and is also one of the training objectives of the subsequent neural network. Q(l,k) is the new phase compensation factor. After removing the antisymmetric function, Q(l,k) is used as one of the training objectives of the subsequent neural network. The calculation formula of Q(l,k) is expressed as,

[0042]

[0043] . Q(l,k) has a total of N frames. Q(l,k) is two-dimensional data of length M in each of the N frames (all symbols with (l,k) are of such dimensions). In actual operation, the data of length M in each of the N frames will be brought into the operation, which is Q(l,k). With it, Λ(l,k) can be obtained, and according to step 2-1, the spectrum S Λ (l,k) of the enhanced speech can be obtained, and then the time-domain waveform of the enhanced speech can be obtained by performing ISTFT.

[0044] In step 3, a training objective and a loss function design that fuse phase spectrum compensation are proposed. The specific method is as follows:

[0045] Step 3-1: For speech enhancement in the frequency domain, the logarithmic power spectrum of the noisy speech and the logarithmic power spectrum of the clean speech are used as the training features and training objectives of the neural network respectively, that is,

[0046]

[0047] Among them, is an approximation value T n of the logarithmic power spectrum of the clean speech obtained by training the logarithmic power spectrum of the noisy speech through the neural network, and S n is the logarithmic power spectrum of the noisy speech; then, combined with the phase α n of the nth frame of the noisy speech, the time-domain waveform of the nth frame of the enhanced speech is obtained, that is,

[0048]

[0049] Step 3-2: On the basis of speech enhancement in the frequency domain, a phase compensation algorithm is fused as a new training objective. The new training objective includes the logarithmic power spectrum T n of the clean speech and the phase compensation factor Q n in two parts, that is,

[0050]

[0051] Among them, Q n represents the data of the nth frame of Q(l,k) and Q n = Q(l,k) l=n=(Q(n, 0), Q(n, 1), …, Q(n, M - 1)), from the phase compensation factor the phase compensation function can be obtained, and then the enhanced phase of the nth frame can be obtained through the phase compensation function, which can replace the phase α of the above noisy speech n , achieving a better speech enhancement effect;

[0052] Step 3 - 3: For the training objective of the fusion phase spectrum compensation proposed in the previous step, define the loss function of the neural network as

[0053]

[0054] where MSE is the joint mean square error of the logarithmic power spectrum and phase spectrum compensation, and the logarithmic power spectrum and phase spectrum compensation jointly affect this parameter. M is the frame length, T(n, k) is the logarithmic power spectrum of the nth frame of the clean speech, is the approximation value of the logarithmic power spectrum of the clean speech obtained by training the logarithmic power spectrum of the nth frame of the noisy speech through the neural network, is the approximation value of the phase compensation factor obtained by training through the neural network; β is the modulation parameter, which is a constant and is used to balance the influence of the logarithmic power spectrum and phase compensation on the network;

[0055] Step 3 - 4: Build a fully convolutional neural network model, which is mainly divided into an input layer, an encoder layer, a decoder layer, and an output layer;

[0056] Step 3 - 5: Use the phase compensation factor combined with the logarithmic power spectrum of the clean speech as the training objective of the fully convolutional neural network, and train the fully convolutional neural network model to obtain a trained fully convolutional neural network model.

[0057] The present invention uses the conventional backpropagation algorithm to train the neural network, uses the joint mean square error in Step 3 - 3 as the loss function, and uses the Adam algorithm for optimization. The initial learning rate is set to 0.05, the learning rate decay period is set to 10, the decay parameter is set to 0.2, the mini - batch (Batchsize) is set to 128, and the number of iterations (epoch) is set to 100. After the above training, it is the trained model of the present invention.

[0058] In the said Step 3 - 4, the input layer is a 5×257 matrix, indicating that the context length of the input feature is 5 frames, and each frame has 257 data points. The input data is composed of the logarithmic power spectra of the noisy speech of adjacent 5 frames, increasing its correlation in both the time and frequency dimensions.

[0059] The encoder layer consists of two convolutional layers and two dropout layers. The convolutional layer has three hyperparameters, namely the convolutional kernel size, the stride, and the padding method. The parameters of convolutional layer 1 are set as follows: the convolutional kernel size is 5; the stride is 1; the padding method is the same mode; the number of convolutional kernels is specified as 32; the ReLU (rectified linear unit) function is selected as the activation function. Using the ReLU activation function can greatly reduce the computational amount. After passing through convolutional layer 1, it enters the dropout layer, and the dropout ratio is set to 0.2. The parameters of convolutional layer 2 are set as follows: the convolutional kernel size is 7; the padding method is the same mode; the stride is 1; the number of convolutional kernels is 16; the ReLU function is selected as the activation function; the parameter of the dropout layer is still set to 0.2. The role of the encoder is to initially extract the shallow features of the noisy speech through convolutional layer 1, and then further extract more abstract and deep speech features through convolutional layer 2 for encoding.

[0060] The decoder layer also consists of two convolutional layers and two dropout layers, and its network structure is symmetric to that of the encoder layer. The parameters of convolutional layer 3 are set as follows: the convolutional kernel size is 7; the padding method is set to the same mode; the stride is 1; the number of convolutional kernels is 16; the ReLU function is selected as the activation function; the parameter of the dropout layer is still set to 0.2. The parameters of convolutional layer 4 are set as follows: the convolutional kernel size is 5; the padding method is the same mode; the stride is 1; the number of convolutional kernels is 32; the ReLU function is selected as the activation function; the parameter of the dropout layer is still set to 0.2. The role of the decoder is the opposite of that of the encoder, which restores the features compressed by the encoder to generate pure speech features.

[0061] The output layer is a fully connected layer, which is directly connected to the decoder. It has 257 neurons, and the linear function is selected as the activation function. The output is one frame of speech feature data and one frame of phase spectrum compensation factor.

[0062] In step 4, the noisy speech to be tested is subjected to feature extraction according to the method of step 1, and the extracted feature parameters are input into the trained fully convolutional neural network model, and then the logarithmic power spectrum estimate value and the phase compensation factor of the speech signal can be obtained.

[0063] In the present invention, the logarithmic power spectrum S of the noisy speech is input into the trained neural network model for one forward propagation to obtain the estimated value of the logarithmic power spectrum of the speech signal and the phase compensation factor. The speech enhancement process based on the neural network can be divided into two parts: training and enhancement. When training the model (which can be considered as the laboratory environment), all the information of the noisy speech and the clean speech is available, and they are used to train the model, that is, S→(T,Q). This step obtains the model (arrow, or the previous function f). When enhancing the speech (which can be considered as the actual application), only the noisy speech information is available, and the information of the clean speech needs to be estimated through the trained model obtained in the training stage, that is What is obtained in this step is an approximation or an estimated value of the clean speech information and the compensation factor The clean speech (enhanced speech) is restored through the approximation value, and the estimated value is directly substituted into step 2-2 for calculation

[0064] In step 5, the magnitude spectrum and phase spectrum of the speech signal are reconstructed using the logarithmic power spectrum and the phase compensation factor (according to the method in step 2-1, and finally S Λ (l,k)) is obtained, and the final enhanced speech is obtained through the inverse short-time Fourier transform (ISTFT).

[0065] Compared with the prior art, the present invention adopts the above technical solutions and has the following technical effects:

[0066] (1) Aiming at the problem that the parameters of the phase spectrum compensation algorithm are fixed and cannot be dynamically adjusted, the present invention introduces the frame signal-to-noise ratio, improves the phase compensation function, overcomes the defect that the traditional phase compensation cannot change with the change of noise, can respond to the change of the noise spectrum faster, and is more suitable for the training and estimation of the neural network, maximizing the performance of the phase compensation algorithm and achieving a better enhancement effect;

[0067] (2) An improved phase compensation factor is introduced as one of the training objectives of the neural network, making up for the shortcoming that the traditional frequency-domain-based speech enhancement algorithm ignores the phase. Under the same number of data samples, a better enhancement effect is achieved through a simple and effective loss function design.

[0068] In summary, while improving the noise cancellation ability of the algorithm, the present invention better ensures the speech intelligibility, thereby improving the overall effect of speech enhancement. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 is a flow chart of the present invention.

[0070] Figure 2 is a structural diagram of the fully convolutional neural network in the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0071] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings: This embodiment is implemented on the premise of the technical solution of the present invention, and detailed implementation manners and specific operation processes are given, but the protection scope of the present invention is not limited to the following embodiments.

[0072] This embodiment proposes a speech enhancement algorithm based on improved phase spectrum compensation and fully convolutional neural network, as Figure 1 shown, including the following steps:

[0073] Step 1: Preprocess the speech data in the training set to obtain the feature data of the noisy speech and the clean speech.

[0074] In the experiment, the speech used for training comes from the training set of TIMIT. Fifteen kinds of noises from NOISEX-92 are selected as the noise set. 100 clean speeches in the TIMIT training set and the noise set are randomly mixed at signal-to-noise ratios of -5dB, 0dB, and 5dB. A total of 4500 noisy speeches are generated at each signal-to-noise ratio and divided into a training set and a test set in a ratio of 9:1 for training the network model.

[0075] Preprocess the noisy speech and the clean speech and extract the feature parameters of the noisy speech and the clean speech. The specific operations are as follows:

[0076] Step 1-1: Let y(n) represent the noisy speech signal and y(n) = d(n) + x(n), where d(n) is the noise signal and x(n) is the clean signal. Assume that x(n) and d(n) are statistically independent and have zero mean. By windowing M samples of y(n) with w(n) and performing an M-point FFT (Fast Fourier Transform), the noisy speech is transformed into the frequency domain to obtain the spectrum Y(l,k) of the noisy speech, where l is the frame number label and k represents the frequency component and k = 0, 1, 2,..., M-1.

[0077] Similarly, the spectrum X(l,k) of the clean speech x(n) and the spectrum D(l,k) of the noise signal d(n) can be obtained. That is:

[0078] By windowing M samples of x(n) with w(n) and performing an M-point FFT, the clean speech is transformed into the frequency domain to obtain the spectrum X(l,k) of the clean speech;

[0079] By windowing M samples of d(n) with w(n) and performing an M-point FFT, the noise signal is transformed into the frequency domain to obtain the spectrum D(l,k) of the noise signal.

[0080] Step 1-2: Represent the spectrum Y(l,k) of the noisy speech on the polar coordinate, which can be divided into the amplitude spectrum and the phase spectrum, that is

[0081] Y(l,k) = |Y(l,k)|e j∠Y(l,k)

[0082] where |Y(l,k)| is the short-time amplitude spectrum of the noisy speech y(n), and its phase spectrum is ∠Y(l,k). Similarly, the short-time amplitude spectrum |X(l,k)| of the clean speech x(n) and the short-time amplitude spectrum |D(l,k)| of the noise signal d(n) can be obtained. That is:

[0083] Representing the spectrum X(l,k) of the clean speech in polar coordinates, it can be divided into an amplitude spectrum and a phase spectrum, that is

[0084] X(l,k) = |X(l,k)|e j∠X(l,k)

[0085] |X(l,k)| is the short-time amplitude spectrum of the clean speech x(n), and its phase spectrum is ∠X(l,k);

[0086] Representing the spectrum D(l,k) of the noise signal in polar coordinates, it can be divided into an amplitude spectrum and a phase spectrum, that is

[0087] D(l,k) = |D(l,k)|e j∠D(l,k)

[0088] |D(l,k)| represents the short-time amplitude spectrum of the noise signal d(n), and the phase spectrum is ∠D(l,k).

[0089] Step 1 - 3: Calculate the log-power spectrum S(n) of the noisy speech and the log-power spectrum T(n) of the clean speech using the following formula

[0090] S(n) = [log e (|Y(n,1)| 2 ), log e (|Y(n,2)| 2 ), …, log e (|Y(n,k)| 2 ), …, log e (|Y(n,M - 1)| 2 )]

[0091] T(n) = [log e (|x(n,1)| 2 ), log e (|X(n,2)| 2 ), …, log e (|X(n,k)| 2 ), …, log e (|X(n,M - 1)| 2 )]

[0092] Among them, |Y(n,1)| 2 is the power of the first frequency band of the nth frame obtained by short-time Fourier transform of the noisy speech, and |Y(n,2)| 2 is the power of the second frequency band of the nth frame obtained by short-time Fourier transform of the noisy speech, and |Y(n,k)| 2 is the power of the kth frequency band of the nth frame obtained by short-time Fourier transform of the noisy speech, and |Y(n,M - 1)| 2 is the power of the (M - 1)th frequency band of the nth frame obtained by short-time Fourier transform of the noisy speech; |X(n,1)| 2 is the power of the first frequency band of the nth frame obtained by short-time Fourier transform of the clean speech, and |X(n,2)| 2 is the power of the second frequency band of the nth frame obtained by short-time Fourier transform of the clean speech,

[0093] |X(n,k)| 2 is the power of the kth frequency band of the nth frame obtained by short-time Fourier transform of the clean speech, and |X(n,M - 1)| 2 is the power of the (M - 1)th frequency band of the nth frame obtained by short-time Fourier transform of the clean speech.

[0094] Step 2: Combine the characteristic data of the noisy speech and the clean speech that have been obtained, introduce the compensation factor calculation formula of the frame signal-to-noise ratio optimized phase compensation algorithm, and then calculate the phase compensation factor using the improved formula.

[0095] Introduce the frame signal-to-noise ratio improved phase compensation function, and the specific operation is as follows:

[0096] Step 2 - 1: The traditional phase compensation function can be expressed as

[0097]

[0098] Among them, Λ(l,k) is the phase compensation function, λ is the phase compensation factor, is the estimated value of the noise, and generally the amplitude spectrum |Y(l,k)| of the noisy speech can be used instead; λ is an empirical value. λ was originally a constant and was later improved to When doing model training, is brought in, and the actual value Z of the amplitude of the noise is directly used instead of The new compensation factor is Ψ(k) is an antisymmetric function used to correct the phase, denoted as

[0099]

[0100] The function of Λ(l,k) is shown in the following formula,

[0101] Y Λ (l,k) = Y(l,k) + Λ(l,k)

[0102] Among them, Y Λ (l,k) is the compensated spectrum. By extracting the phase from the compensated spectrum, the phase spectrum ∠Y Λ (l,k) is obtained as follows

[0103] ∠Y Λ (l,k) = arg(Y Λ (l,k))

[0104] Combining ∠Y Λ (l,k) with the magnitude spectrum |Y(l,k)| of the noisy speech, the spectral expression of the enhanced speech is obtained as

[0105]

[0106] Step 2-2: Introduce the frame signal-to-noise ratio to optimize the phase compensation function, and obtain the optimized phase compensation function, that is

[0107]

[0108] Among them, c is an empirical value, generally set to 2.7. Λ(l,k) decreases as the SNR l increases. When the current frame is a speech frame, the influence of the phase compensation function Λ(l,k) on the noisy speech decreases, and more speech details are retained. SNR l is the signal-to-noise ratio of the l-th frame of the signal

[0109] Step 2-3: To simplify the training objective of the subsequent neural network, the optimized phase compensation function is abbreviated as the following formula

[0110] Λ(l,k) = Ψ(k) × Q(l,k)

[0111] Among them, Q(l,k) is the new compensation factor. After removing the antisymmetric function, Q(l,k) is used as one of the training objectives of the subsequent neural network. The calculation formula of Q(l,k) is expressed as

[0112]

[0113] Step 3: Build a fully convolutional neural network model. The jointly obtained phase compensation factor and the logarithmic power spectrum of the clean speech are used as the training objective of the fully convolutional neural network (FCNN), and the fully convolutional neural network model is trained

[0114] A training objective and loss function design that combines phase spectrum compensation are proposed. The specific method is as follows

[0115] Step 3-1: For speech enhancement in the frequency domain, the logarithmic power spectrum of the noisy speech and the logarithmic power spectrum of the clean speech are used as the training features and training target of the neural network respectively

[0116]

[0117] wherein is the approximation value T of the logarithmic power spectrum of the clean speech obtained by training the logarithmic power spectrum of the noisy speech through the neural network, and S n is the logarithmic power spectrum of the noisy speech. Then, combined with the phase α n of the nth frame of the noisy speech, the time-domain waveform of the nth frame of the enhanced speech is obtained n

[0118] Step 3-2: On the basis of speech enhancement in the frequency domain, a phase compensation algorithm is fused as the new training target, and the new training target includes the logarithmic power spectrum T n of the clean speech and the phase compensation factor Q n in two parts, that is

[0119]

[0120] wherein, from the phase compensation factor the phase compensation function can be obtained, and then the enhanced phase of the nth frame can be obtained through the phase compensation function, which can replace the phase α n of the above-mentioned noisy speech to achieve a better speech enhancement effect

[0121] Step 3-3: For the training target of fusing phase spectrum compensation proposed in the previous step, the loss function of the neural network is defined as

[0122]

[0123] wherein, MSE is the joint mean square error of the logarithmic power spectrum and phase spectrum compensation, and the logarithmic power spectrum and phase spectrum compensation jointly affect this parameter. M is the frame length, T(n,k) is the logarithmic power spectrum of the nth frame of the clean speech, is the approximation value of the logarithmic power spectrum of the clean speech obtained by training the logarithmic power spectrum of the nth frame of the noisy speech through the neural network, is the approximation value of the phase compensation factor obtained by training through the neural network. β is a modulation parameter, which is a constant and is used to balance the influence of the logarithmic power spectrum and phase compensation on the network

[0124] Step 3-4: Build a fully convolutional neural network model for training

[0125] ​The present invention uses a fully convolutional neural network (FCNN). Different from the convolutional neural network, the FCNN replaces the fully connected layer with a new convolutional layer, and uses the convolutional layer to extract deeper and more abstract speech features. The FCNN model mainly includes an input layer, an encoder layer, a decoder layer, and an output layer.

[0126] The input layer of the model is a 5×257 matrix, indicating that the context length of the input feature is 5 frames, and each frame has 257 data points. The input data is composed of the log power spectra of adjacent 5-frame noisy speech, increasing its correlation in both the time and frequency dimensions.

[0127] The encoder layer of the model consists of two convolutional layers and two dropout layers. The convolutional layer has three hyperparameters, namely the convolutional kernel size, the stride, and the padding method. The parameters of convolutional layer 1 are set as follows: the convolutional kernel size is 5; the stride is 1; the padding method is the same mode; the number of convolutional kernels is specified as 32; the activation function uses the ReLU (rectified linear unit, Re-LU) function. Using the ReLU activation function can greatly reduce the computational amount. After passing through convolutional layer 1, it enters the dropout layer, and the dropout ratio is set to 0.2. The parameters of convolutional layer 2 are set as follows: the convolutional kernel size is 7; the padding method is the same mode; the stride is 1; the number of convolutional kernels is 16; the activation function uses the ReLU function; the parameter of the dropout layer is still set to 0.2. The role of the encoder is to initially extract the shallow features of the noisy speech through convolutional layer 1, and then further extract more abstract and deeper speech features through convolutional layer 2 for encoding.

[0128] The decoder layer of the model also consists of two convolutional layers and two dropout layers, and its network structure is symmetric to the encoder layer. The parameters of convolutional layer 3 are set as follows: the convolutional kernel size is 7; the padding method is set to the same mode; the stride is 1; the number of convolutional kernels is 16; the activation function uses the ReLU function; the parameter of the dropout layer is still set to 0.2. The parameters of convolutional layer 4 are set as follows: the convolutional kernel size is 5; the padding method is the same mode; the stride is 1; the number of convolutional kernels is 32; the activation function uses the ReLU function; the parameter of the dropout layer is still set to 0.2. The role of the decoder is the opposite of that of the encoder, restoring the features compressed by the encoder to generate clean speech features.

[0129] The output layer of the model is a fully connected layer, directly connected to the decoder, with 257 neurons. The activation function uses the linear function, and the output is 1 frame of speech feature data and 1 frame of phase spectrum compensation factor.

[0130] Step 3-5: Using the phase compensation factor and the logarithmic power spectrum of the clean speech as the training objective of the fully convolutional neural network, train the fully convolutional neural network model to obtain a trained fully convolutional neural network model.

[0131] Step 4: Input the test speech into the trained model to obtain the estimated value of the logarithmic power spectrum and the phase compensation function.

[0132] Extract the features of the noisy speech to be tested according to the method in Step 1, and input the extracted feature parameters into the trained fully convolutional neural network model, then the estimated value of the logarithmic power spectrum of the speech signal and the phase compensation factor can be obtained.

[0133] Step 5: Use the estimated value of the logarithmic power spectrum and the phase compensation function obtained in the previous step to reconstruct the amplitude spectrum and phase spectrum of the speech signal respectively to obtain the final enhanced speech.

[0134] Use the logarithmic power spectrum and the phase compensation factor to reconstruct the amplitude spectrum and phase spectrum of the speech signal, and the final enhanced speech is obtained through ISTFT.

[0135] Embodiment

[0136] To verify the effectiveness of the method of the present invention, a comparative experiment was conducted on the traditional PSC speech enhancement algorithm, the FCNN algorithm and the method of the present invention, and the algorithm performance was tested under different types of interference noises and different input signal-to-noise ratios. The perceptual evaluation of speech quality (PESQ) and short-time objective intelligibility (STOI) were selected as the speech evaluation indicators. PESQ is obtained by comparing the enhanced speech signal with the clean speech signal, and the value range is [-0.5, 4.5]. The larger the value, the higher the speech quality. STOI is also obtained by comparing the enhanced speech signal and the clean speech signal, and the value range is [0, 1]. The larger the value, the higher the intelligibility of the speech. Specifically as follows:

[0137] Method 1: PSC

[0138] Method 2: FCNN algorithm

[0139] Method 3: The method of the present invention

[0140] Use these three methods respectively to enhance the noisy speech with signal-to-noise ratios of -5dB, 0dB, and 5dB, and the noise types are factory2 (fac), babble (bab), and buccaneer1 (buc). The results are shown in Tables 1 to 3.

[0141] Table 1 - 5dB Noise

[0142]

[0143] Table 2 0dB Noise

[0144]

[0145] Table 3 5dB Noise

[0146]

[0147] As can be seen from the results of the above table, compared with the traditional phase compensation algorithm (PSC), the method of the present invention can effectively suppress various noises in the speech signal and better ensure the intelligibility of the speech. And compared with the FCNN algorithm, the method of the present invention adds the estimation of the phase compensation factor, and the effect is improved under both fac and buc noise conditions, thus improving the overall effect of speech enhancement.

[0148] As described above, it is only the specific implementation manner in the present invention, but the protection scope of the present invention is not limited thereto. Any transformation or replacement that can be understood and conceived by those familiar with the technology within the technical scope disclosed by the present invention should be covered within the scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A speech enhancement algorithm based on improved phase spectrum compensation and fully convolutional neural network, characterized in that, It includes the following steps: Step 1: Preprocess the training set speech data to obtain the feature data of the noisy speech and the clean speech; Preprocess the noisy speech and the clean speech and extract the feature parameters of the noisy speech and the clean speech. The specific operations are as follows: Step 1-1: Use y(n) to represent the noisy speech signal and y(n) = d(n) + x(n), where d(n) is the noise signal and x(n) is the clean signal; by windowing M samples of y(n) with w(n) and performing M-point FFT, the noisy speech is transformed into the frequency domain to obtain the spectrum Y(l,k) of the noisy speech, where l is the frame number marker and k represents the frequency component and k = 0, 1, 2, …, M - 1; similarly, the spectrum X(l,k) of the clean speech x(n) and the spectrum D(l,k) of the noise signal d(n) can be obtained; Step 1-2: Represent the spectrum Y(l,k) of the noisy speech on the polar coordinate, which is divided into the amplitude spectrum and the phase spectrum, that is Y(l,k) = |Y(l,k)|e j∠Y(l,k) where |Y(l,k)| is the short-time amplitude spectrum of the noisy speech y(n), and ∠Y(l,k) is the phase spectrum; similarly, the short-time amplitude spectrum |X(l,k)| of the clean speech x(n) and the short-time amplitude spectrum |D(l,k)| of the noise signal d(n) can be obtained; Step 1-3: Use the following formula to calculate the logarithmic power spectrum S(n) of the noisy speech and the logarithmic power spectrum T(n) of the clean speech, S(n) = [log e (|Y(n, 1)| 2 ), log e (|Y(n, 2)| 2 ), …, log e (|Y(n, k)| 2 ), …, log e (|Y(n, M - 1)| 2 )] T(n) = [log e (|X(n, 1)| 2 ), log e (|X(n, 2)| 2 ), …, log e (|X(n, k)| 2 ), …, log e (|X(n, M - 1)| 2 )] Among them, |Y(n,1)| 2 is the power of the first frequency band of the n-th frame obtained by short-time Fourier transform of the noisy speech, and |Y(n,2)| 2 is the power of the second frequency band of the n-th frame obtained by short-time Fourier transform of the noisy speech, and |Y(n,k)| 2 is the power of the k-th frequency band of the n-th frame obtained by short-time Fourier transform of the noisy speech, and |Y(n,M - 1)| 2 is the power of the (M - 1)-th frequency band of the n-th frame obtained by short-time Fourier transform of the noisy speech; |X(n,1)| 2 is the power of the first frequency band of the n-th frame obtained by short-time Fourier transform of the clean speech, and |X(n,2)| 2 is the power of the second frequency band of the n-th frame obtained by short-time Fourier transform of the clean speech, and |X(n,k)| 2 is the power of the k-th frequency band of the n-th frame obtained by short-time Fourier transform of the clean speech, and |X(n,M - 1)| 2 is the power of the (M - 1)-th frequency band of the n-th frame obtained by short-time Fourier transform of the clean speech; Step 2: Combine the feature data of the noisy speech and the clean speech, introduce the compensation factor calculation formula of the frame signal-to-noise ratio optimized phase compensation algorithm, and then use the improved formula to calculate the phase compensation factor; Optimize the phase compensation function. The specific operations are as follows: Step 2-1: The traditional phase compensation function is expressed as where, Λ(l,k) is the phase compensation function, λ is the phase compensation factor, is the estimated value of the noise, which can generally be replaced by the magnitude spectrum |Y(l,k)| of the noisy speech. λ is an empirical value, and Ψ(k) is an antisymmetric function used for phase correction, denoted as The role of Λ(l,k) is shown in the following formula, Y Λ (l, k) = Y(l, k) + Λ(l, k) Among them, Y Λ (l,k) is the compensated spectrum. By extracting the phase of the compensated spectrum, the phase spectrum ∠Y Λ (l,k) is obtained. ∠Y Λ (l,k) = arg(Y Λ (l,k)) Rotate ∠Y Λ (l,k) is combined with the magnitude spectrum |Y(l,k)| of the noisy speech to obtain the spectrum of the enhanced speech, and its expression is Step 2-2: Introduce the frame signal-to-noise ratio to optimize the phase compensation function to obtain the optimized phase compensation function, that is where c is an empirical value, and SNR l is the signal-to-noise ratio of the l-th frame; Step 2-3: In order to simplify the training objective of the subsequent neural network, the optimized phase compensation function is abbreviated as the following formula, Λ(l,k) = Ψ(k) × Q(l,k) where the calculation formula of Q(l,k) is expressed as, Step 3: Build a fully convolutional neural network model. The phase compensation factor combined with the logarithmic power spectrum of the clean speech is used as the training objective of the fully convolutional neural network to train the fully convolutional neural network model; Step 4: Input the test speech into the trained model to obtain the estimated value of the logarithmic power spectrum and the phase compensation function; Step 5: Use the estimated value of the logarithmic power spectrum and the phase compensation function to reconstruct the amplitude spectrum and the phase spectrum of the speech signal respectively to obtain the final enhanced speech.

2. The speech enhancement algorithm based on improved phase spectrum compensation and fully convolutional neural network according to claim 1, wherein The design of the training objective and the loss function that fuses the phase spectrum compensation is proposed in Step 3. The specific method is as follows: Step 3-1: For speech enhancement in the frequency domain, use the logarithmic power spectrum of the noisy speech and the logarithmic power spectrum of the clean speech as the training feature and the training objective of the neural network respectively, that is Among them, is the approximation value of the clean speech logarithmic power spectrum T obtained by neural network training on the noisy speech logarithmic power spectrum, n S n is the logarithmic power spectrum of the noisy speech; then combined with the phase α n of the n-th frame of the noisy speech, the time-domain waveform of the n-th frame of the enhanced speech is obtained That is Step 3-2: On the basis of speech enhancement in the frequency domain, fuse the phase compensation algorithm as the new training objective, and the new training objective includes the logarithmic power spectrum T of the clean speech n and the phase compensation factor Q n in two parts, that is Among them, from the phase compensation factor the phase compensation function can be obtained, and then the enhanced phase of the nth frame is obtained through the phase compensation function to replace the phase α of the noisy speech n , achieving a better speech enhancement effect; Step 3-3: For the training objective that fuses the phase spectrum compensation proposed in the previous step, define the loss function of the neural network as where MSE is the joint mean square error of logarithmic power spectrum and phase spectrum compensation, and T(n,k) is the logarithmic power spectrum of the n-th frame of the clean speech. is the approximation of the logarithmic power spectrum of the clean speech obtained by training the neural network with the logarithmic power spectrum of the n-th frame of the noisy speech. is the approximation obtained by training the neural network with the phase compensation factor, and β is the modulation parameter. Step 3-4: Build a fully convolutional neural network model, which is divided into an input layer, an encoder layer, a decoder layer and an output layer; Step 3-5: Use the phase compensation factor combined with the logarithmic power spectrum of the clean speech as the training objective of the fully convolutional neural network, and train the fully convolutional neural network model to obtain a trained fully convolutional neural network model.

3. The speech enhancement algorithm based on improved phase spectrum compensation and fully convolutional neural network according to claim 2, characterized in that In the said step 3-4, the input layer is a 5×257 matrix, indicating that the context length of the input feature is 5 frames, and each frame has 257 data points. The input data is composed of the logarithmic power spectra of 5 adjacent frames of noisy speech, increasing its correlation in both the time and frequency dimensions. The encoder layer consists of two convolutional layers and two dropout layers. The convolutional layer has three hyperparameters, namely the convolutional kernel size, the moving step size, and the padding method. The decoder layer also consists of two convolutional layers and two dropout layers, and its network structure is symmetric to that of the encoder layer. The output layer is a fully connected layer, directly connected to the decoder, with 257 neurons. The activation function is selected as the linear function, and the output is 1 frame of speech feature data and 1 frame of phase spectrum compensation factor.

4. The speech enhancement algorithm based on improved phase spectrum compensation and fully convolutional neural network according to claim 3, characterized in that, In the said step 4, extract the features of the to-be-tested noisy speech according to the method of step 1, and input the extracted feature parameters into the trained fully convolutional neural network model, then the logarithmic power spectrum estimation value and the phase compensation factor of the speech signal can be obtained.

5. The speech enhancement algorithm based on improved phase spectrum compensation and fully convolutional neural network according to claim 4, characterized in that, In the said step 5, use the logarithmic power spectrum and the phase compensation factor to reconstruct the amplitude spectrum and phase spectrum of the speech signal, and obtain the final enhanced speech through the inverse short-time Fourier transform.

Citation Information

Patent Citations

  • Voice enhancement method and system based on phase compensation

    CN108735213A

  • Audio processing method and electronic equipment

    CN111050269A