Phase prediction method using parallel estimation architecture network trained with anti-convolution loss
Through a parallel estimation architecture network trained with anti-winding loss, the parallel linear convolutional layer and phase calculation unit are used to directly predict the winding phase spectrum of the speech signal, solving the problems of low efficiency and complexity in the prior art, and achieving efficient and accurate phase prediction.
Patent Information
- Application Number
- CN202211489291.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-25
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2042-11-25
AI Technical Summary
The existing speech phase prediction methods are inefficient and complex, and cannot output accurate phase spectrum directly through the neural network, which is limited by phase winding problems and difficulty in phase modeling.
The parallel estimation architecture network trained with anti-winding loss is adopted. Through parallel linear convolutional layers and phase calculation units, the real imaginary part of the short-time complex spectrum is simulated to calculate the phase spectrum, and the anti-winding function is used to activate instantaneous phase, group delay and instantaneous angular frequency loss, limiting the phase within the main value interval, and directly predicting the winding phase spectrum of the speech signal.
It improves the efficiency and accuracy of speech phase prediction, avoids the problem of amplification of errors caused by phase winding, and realizes that the neural network directly outputs an accurate phase spectrum.
Smart Images

Figure CN115862673B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech signal processing technology, and in particular to a method for predicting phase using a parallel estimation architecture network trained with anti-warping loss. Background Art
[0002] Speech phase prediction, also known as speech phase reconstruction, aims to predict the corresponding phase spectrum from the amplitude spectrum of a speech signal. Speech phase prediction is widely used in numerous speech generation tasks. However, due to phase wrapping and the difficulty of phase modeling, accurate speech phase prediction remains a major challenge.
[0003] Current phase prediction methods primarily fall into two categories: iterative algorithms and neural network-based methods. Iterative algorithms are susceptible to the effects of the initial phase during iteration, and the reconstructed speech contains noticeable, unnatural noise. However, existing neural network-based methods are unable to directly output an accurate phase spectrum through training due to the phase warping characteristics of speech signals. Therefore, most existing neural network-based methods require a two-step process: first, the speech signal is processed using a neural network, and then the neural network output is processed using a specific algorithm (such as the Griffin-Lim algorithm or the cyclic phase unwrapping algorithm) to obtain an accurate phase spectrum. Clearly, this two-step neural network-based method is inefficient and complex to operate. Summary of the Invention
[0004] In response to the shortcomings of the above-mentioned prior art, the present invention provides a method for predicting phase using a parallel estimation architecture network trained with anti-warping loss, so as to provide a solution for directly obtaining an accurate warping phase spectrum of a speech signal through a neural network, thereby improving the efficiency of the phase prediction method based on the neural network.
[0005] In a first aspect, the present application provides a method for predicting phase using a parallel estimation architecture network trained with anti-warping loss, the method comprising:
[0006] Network training process:
[0007] Determining a neural network to be trained; wherein the neural network to be trained includes a residual convolutional network, a first linear convolutional layer and a second linear convolutional layer in parallel, and a phase calculation unit;
[0008] Obtain the logarithmic magnitude spectrum and true warped phase spectrum of the sample speech signal;
[0009] Processing the logarithmic magnitude spectrum of the sample speech signal using the neural network to be trained to obtain a predicted warped phase spectrum of the sample speech signal; wherein the predicted warped phase spectrum is calculated by the phase calculation unit based on a pseudo-real part and a pseudo-imaginary part; the pseudo-real part and the pseudo-imaginary part are output by the first linear convolution layer and the second linear convolution layer, respectively; and the phase of the predicted warped phase spectrum is within a principal value interval;
[0010] Calculating an anti-warping loss of the predicted warping phase spectrum and the actual warping phase spectrum; wherein the anti-warping loss is a linear combination of an instantaneous phase loss, a group delay loss, and an instantaneous angular frequency loss between the predicted warping phase spectrum and the actual warping phase spectrum; the instantaneous phase loss, the group delay loss, and the instantaneous angular frequency loss are all activated by an anti-warping function;
[0011] If the anti-warping loss does not meet the preset convergence condition, updating the parameters of the neural network to be trained according to the anti-warping loss, and returning to the step of processing the logarithmic magnitude spectrum using the neural network to be trained to obtain a predicted warped phase spectrum of the sample speech signal;
[0012] If the anti-wrapping loss meets the convergence condition, determining the neural network to be trained as a phase prediction neural network;
[0013] Phase prediction process:
[0014] Obtaining the logarithmic magnitude spectrum of the speech signal to be predicted;
[0015] The phase prediction neural network is used to process the logarithmic magnitude spectrum of the speech signal to be predicted to obtain a warped phase spectrum of the speech signal to be predicted.
[0016] Optionally, the process of obtaining a true warped phase spectrum of the sample speech signal includes:
[0017] Performing a short-time Fourier transform on the sample speech signal to obtain a short-time complex spectrum of the sample speech signal;
[0018] Phase calculation is performed based on the real part and the imaginary part of the short-time complex spectrum of the sample speech signal to obtain a real warped phase spectrum of the sample speech signal.
[0019] Optionally, the calculating the anti-warping loss of the predicted warping phase spectrum and the actual warping phase spectrum includes:
[0020] Calculating an instantaneous phase loss according to the actual winding phase spectrum and the predicted winding phase spectrum;
[0021] Performing frequency differentiation on the real warping phase spectrum and the predicted warping phase spectrum respectively, and calculating group delay loss according to the frequency differentiation between the real warping phase spectrum and the predicted warping phase spectrum;
[0022] performing time differentiation on the real winding phase spectrum and the predicted winding phase spectrum respectively, and calculating instantaneous angular frequency loss according to the time differentiation between the real winding phase spectrum and the predicted winding phase spectrum;
[0023] The anti-winding loss is obtained by adding the instantaneous phase loss, the group delay loss and the instantaneous angular frequency loss.
[0024] Optionally, the convergence condition is that the number of training times is greater than or equal to a preset maximum number of training times, wherein the number of training times is defined as the number of times the step of processing the logarithmic amplitude spectrum of the sample speech signal using the neural network to be trained to obtain the predicted warped phase spectrum of the sample speech signal is executed.
[0025] Optionally, the residual convolutional network includes:
[0026] A linear convolution layer, multiple parallel residual convolution blocks connected to the linear convolution layer, an accumulation unit for calculating the mean of the outputs of the multiple residual convolution blocks, and a linear unit with leakage correction connected to the accumulation unit.
[0027] A second aspect of the present application provides a device for predicting phase using a parallel estimation architecture network trained with anti-warping loss, comprising:
[0028] A generating unit, configured to generate a neural network to be trained; wherein the neural network to be trained comprises a residual convolutional network, a first linear convolutional layer and a second linear convolutional layer in parallel, and a phase calculation unit;
[0029] An acquisition unit, configured to acquire a logarithmic amplitude spectrum and a real warped phase spectrum of a sample speech signal;
[0030] a processing unit, configured to process the logarithmic magnitude spectrum of the sample speech signal using the neural network to be trained to obtain a predicted warped phase spectrum of the sample speech signal; wherein the predicted warped phase spectrum is calculated by the phase calculation unit based on a pseudo-real part and a pseudo-imaginary part; the pseudo-real part and the pseudo-imaginary part are output by the first linear convolution layer and the second linear convolution layer, respectively; and the phase of the predicted warped phase spectrum is within a principal value interval;
[0031] a calculation unit, configured to calculate an anti-warping loss of the predicted warping phase spectrum and the actual warping phase spectrum; wherein the anti-warping loss is a linear combination of an instantaneous phase loss, a group delay loss, and an instantaneous angular frequency loss between the predicted warping phase spectrum and the actual warping phase spectrum; and the instantaneous phase loss, the group delay loss, and the instantaneous angular frequency loss are all activated by an anti-warping function;
[0032] an updating unit, configured to update the parameters of the neural network to be trained according to the anti-warping loss if the anti-warping loss does not meet a preset convergence condition, and return to the step of processing the logarithmic magnitude spectrum using the neural network to be trained to obtain a predicted warping phase spectrum of the sample speech signal;
[0033] a determining unit, configured to determine the neural network to be trained as a phase prediction neural network if the anti-wrapping loss meets the convergence condition;
[0034] The acquisition unit is used to obtain the logarithmic magnitude spectrum of the speech signal to be predicted;
[0035] The processing unit is used to process the logarithmic magnitude spectrum of the speech signal to be predicted using the phase prediction neural network to obtain the warped phase spectrum of the speech signal to be predicted.
[0036] Optionally, when the acquiring unit acquires the real warped phase spectrum of the sample speech signal, it is specifically used to:
[0037] Performing a short-time Fourier transform on the sample speech signal to obtain a short-time complex spectrum of the sample speech signal;
[0038] Phase calculation is performed based on the real part and the imaginary part of the short-time complex spectrum of the sample speech signal to obtain a real warped phase spectrum of the sample speech signal.
[0039] Optionally, when the calculation unit calculates the anti-winding loss of the predicted winding phase spectrum and the actual winding phase spectrum, it is specifically used to:
[0040] Calculating an instantaneous phase loss according to the actual winding phase spectrum and the predicted winding phase spectrum;
[0041] Performing frequency differentiation on the real warping phase spectrum and the predicted warping phase spectrum respectively, and calculating group delay loss according to the frequency differentiation between the real warping phase spectrum and the predicted warping phase spectrum;
[0042] performing time differentiation on the real winding phase spectrum and the predicted winding phase spectrum respectively, and calculating instantaneous angular frequency loss according to the time differentiation between the real winding phase spectrum and the predicted winding phase spectrum;
[0043] The anti-winding loss is obtained by adding the instantaneous phase loss, the group delay loss and the instantaneous angular frequency loss.
[0044] Optionally, the convergence condition is that the number of training times is greater than or equal to a preset maximum number of training times, wherein the number of training times is defined as the number of times the step of processing the logarithmic amplitude spectrum of the sample speech signal using the neural network to be trained to obtain the predicted warped phase spectrum of the sample speech signal is executed.
[0045] Optionally, the residual convolutional network includes:
[0046] A linear convolution layer, multiple parallel residual convolution blocks connected to the linear convolution layer, an accumulation unit for calculating the mean of the outputs of the multiple residual convolution blocks, and a linear unit with leakage correction connected to the accumulation unit.
[0047] The present application provides a method for predicting phase using a parallel estimation architecture network trained with anti-warping loss. The method includes, during the training process, simulating the process of calculating the phase spectrum from the real and imaginary parts of the short-time complex spectrum through two parallel linear convolution layers in the neural network to be trained, and a phase calculation unit, and limiting the predicted phase value to the main value interval to achieve the prediction of the warped phase spectrum, and the anti-warping loss used during training includes the instantaneous phase error, group delay error and instantaneous angular frequency error activated by the anti-warping function, thereby avoiding the error expansion problem caused by phase warping. After the training is completed, the trained phase prediction neural network is used to process the logarithmic amplitude spectrum of the speech signal to be predicted to obtain the warped phase spectrum. This solution directly predicts the warped phase spectrum of the speech signal through a neural network, and introduces an anti-warping function when calculating the loss to solve the error expansion problem caused by phase warping during training, with high efficiency and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0049] Figure 1 A flowchart of a method for predicting phase using a parallel estimation architecture network trained with anti-warping loss provided in an embodiment of the present application;
[0050] Figure 2 A schematic diagram of a network training process in a method for predicting phase using a parallel estimation architecture network trained with anti-warping loss provided in an embodiment of the present application;
[0051] Figure 3 A schematic diagram of error amplification caused by phase wrapping provided in an embodiment of the present application;
[0052] Figure 4 A schematic diagram of the structure of a device for predicting phase using a parallel estimation architecture network trained with anti-warping loss provided in an embodiment of the present application. DETAILED DESCRIPTION
[0053] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0054] In order to facilitate understanding of the technical solution of this application, some terms that may be involved in this application are first explained.
[0055] Phase wrapping. Assuming the main phase value interval is (-π, π], the phase will jump at the boundaries -π and π, showing a discontinuous phenomenon. This phenomenon is called phase wrapping of the speech signal.
[0056] Amplitude spectrum and phase spectrum. After Fourier transforming a speech signal, a short-time complex spectrum can be obtained. Then, the short-time complex spectrum is calculated according to the amplitude calculation formula to obtain the amplitude spectrum of the speech signal. The amplitude spectrum reflects the amplitudes of the sinusoidal signals of different frequencies that make up the speech signal. The phase spectrum of the speech signal predicted based on the amplitude spectrum reflects the phases of the sinusoidal signals of different frequencies in the speech signal.
[0057] In particular, after taking the natural logarithm of the amplitude spectrum, the logarithmic amplitude spectrum of the signal can be obtained.
[0058] The first aspect of the present application provides a method for predicting phase using a parallel estimation architecture network trained with anti-convolution loss, see Figure 1 , is a flowchart of the method, which may include the following steps.
[0059] S101, determining a neural network to be trained.
[0060] The structure of the neural network to be trained can be found in Figure 2 , it can be seen that the neural network to be trained includes a residual convolutional network, a parallel first linear convolutional layer and a second linear convolutional layer, and a phase calculation unit.
[0061] Among them, the residual convolution network is a general deep network, which can specifically include a linear convolution layer, multiple parallel residual convolution blocks connected to the linear convolution layer, an accumulation unit for calculating the mean of the outputs of multiple residual convolution blocks, and a leaky rectified linear unit (LReLU) connected to the accumulation unit.
[0062] The data input to the residual convolution network will pass through the linear convolution layer and multiple residual convolution blocks of the residual convolution network in sequence. Then the outputs of multiple residual convolution blocks are added and averaged in the accumulation unit (equivalent to jump connections between multiple residual convolution blocks). After that, the output of the accumulation unit is activated by a linear unit with leakage correction to obtain the output of the residual convolution network.
[0063] Furthermore, each residual convolution block is composed of multiple residual convolution sub-blocks in cascade. In each residual convolution sub-block, the input is first activated by LReLU, then passes through a linear dilation convolution layer, then activated by LReLU again and passes through a linear convolution layer, and finally added to the input of the residual convolution sub-block (i.e., residual connection) to obtain the output of the residual convolution sub-block.
[0064] The benefit of setting up a residual convolutional network is that it improves the network depth through residual connections and skip connections, and increases the network's receptive field through dilated convolution, thereby improving the network's modeling ability.
[0065] In step S101, the various parameters in the structure, that is, the parameters of the neural network to be trained, can be adjusted according to the structure of the neural network to be trained. Figure 2 The values of the parameters in the residual convolutional network, the first linear convolutional layer, and the second linear convolutional layer shown are randomly initialized, that is, the values of the parameters in the structure are randomly set within a certain value range. The structure after setting the values of all parameters in the structure is the neural network to be trained.
[0066] S102: Obtain a logarithmic magnitude spectrum and a true warped phase spectrum of the sample speech signal.
[0067] The sample speech signal may be a natural speech signal acquired using a speech acquisition device (e.g., a microphone). The logarithmic magnitude spectrum of the sample speech signal may be obtained by first performing a Fourier transform on the sample speech signal to obtain a short-time complex spectrum of the sample speech signal, calculating the short-time complex spectrum of the sample speech signal using an amplitude formula to obtain an magnitude spectrum of the sample speech signal, and then taking the natural logarithm of the magnitude spectrum to obtain the logarithmic magnitude spectrum of the sample speech signal.
[0068] The logarithmic magnitude spectrum of the sample speech signal can be expressed as Where F represents the total number of frames of the logarithmic magnitude spectrum, and N represents the number of frequency points of the logarithmic magnitude spectrum. It can be seen that the logarithmic magnitude spectrum of the sample speech signal can be regarded as a matrix with F rows and N columns, where each row corresponds to a signal frame in the sample speech signal, and each column corresponds to a specific frequency value f. The element in the i-th row and k-th column represents the frequency value f corresponding to the k-th column of the i-th signal frame in the sample speech signal. k The logarithm of the amplitude.
[0069] A signal frame can be considered as a segment of signal segmented from a sample speech signal with a duration equal to a preset segmentation duration. For example, if the preset segmentation duration is 1 second, then a segment of 1 second of signal segmented from a sample speech signal is equivalent to a signal frame.
[0070] See Figure 2 , the process of obtaining the true warped phase spectrum of the sample speech signal may include:
[0071] A1, performing short-time Fourier transform on the sample speech signal to obtain the short-time complex spectrum of the sample speech signal;
[0072] A2, performing phase calculation based on the real part and the imaginary part of the short-time complex spectrum of the sample speech signal to obtain the true warped phase spectrum of the sample speech signal.
[0073] Short-time Fourier transform is an existing Fourier transform algorithm, and its specific implementation will not be described in detail.
[0074] In step A2, the real part extracted from the short-time complex spectrum of the sample speech signal can be expressed as The imaginary part can be written as The true warped phase spectrum of the sample speech signal calculated based on the real part and the imaginary part is recorded as P=Φ(R, I), where Φ() represents the function used for phase calculation, and the expression of this function can be expressed by the following formula (1).
[0075]
[0076] In addition, define Φ(0,0)=0. Function Sgn * (x) is defined as follows:
[0077] When x≥0, Sgn * (x)=1; when x<0, Sgn * (x) = -1.
[0078] The phase calculation function strictly limits the phase of the true warping phase spectrum to the main value interval (-π, π], thus enabling the prediction of the warping phase spectrum.
[0079] It can be understood that the real and imaginary parts of the short-time complex spectrum of the above-mentioned sample speech signal, as well as the true warped phase spectrum of the sample speech signal, are matrices of F rows and N columns. Therefore, in formula (1), the two matrices of the real and imaginary parts of the input are calculated element by element to obtain the elements of the corresponding positions in the true warped phase spectrum. For example, the elements of the 1st row and 1st column of the real part R and the elements of the 1st row and 1st column of the imaginary part I are substituted into formula (1), and the calculated results are used as the elements of the 1st row and 1st column of the true warped phase spectrum. The elements of the 1st row and 2nd column of the real part R and the elements of the 1st row and 2nd column of the imaginary part I are substituted into formula (1), and the calculated results are used as the elements of the 1st row and 2nd column of the true warped phase spectrum. This can be deduced by analogy until all elements of the true warped phase spectrum are calculated.
[0080] S103 , using the neural network to be trained to process the logarithmic magnitude spectrum of the sample speech signal to obtain a predicted warped phase spectrum of the sample speech signal.
[0081] Among them, the predicted warped phase spectrum is calculated by the phase calculation unit according to the pseudo real part and the pseudo imaginary part; the pseudo real part and the pseudo imaginary part are output by the first linear convolution layer and the second linear convolution layer respectively; the phase of the predicted warped phase spectrum is within the main value interval.
[0082] The implementation of step S103 can be found in Figure 2 After the logarithmic amplitude spectrum logA of the sample speech signal is input into the neural network to be trained, it is first processed by the residual convolution network. Then the first linear convolution layer calculates the output of the residual convolution network to obtain the pseudo real part The second linear convolution layer calculates the output of the residual convolution network to obtain the pseudo-imaginary part After obtaining the pseudo real part and pseudo imaginary part, the phase calculation unit can calculate the pseudo real part and pseudo imaginary part using the phase calculation function shown in the above formula (1). The calculation result is the predicted warped phase spectrum of the sample speech signal, which is recorded as According to the above formula (1), it can be seen that the predicted warping phase calculated by the phase calculation function has the value of each element within the main value interval (-π, π], so the phase spectrum predicted by the phase prediction network trained by the method of this embodiment is a warped phase spectrum.
[0083] S104 , calculating the anti-warping loss of the predicted warping phase spectrum and the actual warping phase spectrum.
[0084] Among them, the anti-warping loss is a linear combination of the instantaneous phase loss, group delay loss and instantaneous angular frequency loss between the predicted warping phase spectrum and the true warping phase spectrum; the instantaneous phase loss, group delay loss and instantaneous angular frequency loss are all activated by the anti-warping function.
[0085] If the anti-winding loss does not meet the preset convergence condition, step S105 is executed; if the anti-winding loss meets the preset convergence condition, step S106 is executed.
[0086] The process of calculating the anti-winding loss in step S104 may specifically include the following steps:
[0087] B1, instantaneous phase loss is calculated based on the real winding phase spectrum and the predicted winding phase spectrum;
[0088] B2, perform frequency difference on the real winding phase spectrum and the predicted winding phase spectrum respectively, and calculate the group delay loss based on the frequency difference between the real winding phase spectrum and the predicted winding phase spectrum;
[0089] B3, performing time difference on the real winding phase spectrum and the predicted winding phase spectrum respectively, and calculating the instantaneous angular frequency loss based on the time difference between the real winding phase spectrum and the predicted winding phase spectrum;
[0090] B4, the anti-winding loss is obtained by adding the instantaneous phase loss, group delay loss and instantaneous angular frequency loss.
[0091] First, to facilitate understanding of the calculation method of instantaneous phase loss in this embodiment, the error amplification problem caused by phase wrapping during neural network training is explained. Figure 3 , is a schematic diagram of error amplification caused by phase wrapping provided in an embodiment of the present application.
[0092] Figure 3 The black dots in the middle represent the phase in the true warping phase spectrum, i.e., the true phase, and the striped dots represent the phase in the predicted warping phase spectrum, i.e., the predicted phase. Since both are calculated using the phase calculation function shown in formula (1), both are limited to the principal value interval (-π, π]. The goal of model training is to make the two as close as possible to each other, that is, the striped dots are as close as possible to the black dots, so the error between the two needs to be estimated during the training process.
[0093] However, due to the winding nature of the phase, there are two paths from the black dot to the striped dot, namely Figure 3 The direct path and the winding path in the equation are given by , and the actual error between the predicted phase and the natural phase is the minimum of the absolute error (i.e., the direct path length) and the winding error (i.e., the winding path length). For example, for the predicted phase at point A, With natural phase P A , the true error between the two is the absolute error, but for the predicted phase at point B With natural phase P B, the true error between the two becomes the warping error. Therefore, when evaluating the error between the true and predicted warped phase spectra, if either the true error or the warping error is used exclusively, the error between the two will increase with each training session. This phenomenon is the error amplification problem caused by phase warping.
[0094] In order to solve the error amplification problem caused by the phase wrapping, this embodiment defines the expression of the true error e as shown in the following formula (2):
[0095]
[0096] Where round() means rounding. It can be seen that the above formula (2) is equivalent to a formula about the error Therefore, we can define the function f as shown in the following formula (3): AW (x):
[0097]
[0098] When using error When x is substituted in formula (3), the above function f AW (x) is equal to the true error shown in formula (2), so the function f AW (x) is regarded as an anti-warping function, which can avoid the error amplification problem caused by phase warping.
[0099] Based on the anti-windup function defined above, the instantaneous phase loss L can be set as shown in the following formula (4): IP The calculation formula is:
[0100]
[0101] Among them, f AW (X) represents the element-by-element anti-warping function calculation of the matrix X, that is, in formula (4), The output result is a matrix. Each element in the matrix is converted into a matrix using the function shown in formula (3). The elements at the corresponding positions in are calculated, for example, the matrix Substitute the element in the first row and first column of the formula into x in formula (3), and the result is The element in the first row and first column of the output matrix, and so on for other elements.
[0102] avg() means calculating the average value of all elements in the matrix within the brackets.
[0103] In formula (4) In this embodiment, each sample speech signal can be averaged according to the above steps to obtain the corresponding real warping phase spectrum and predicted warping phase spectrum. Therefore, when there are multiple sample speech signals, multiple This means calculating multiple The average value of .
[0104] In summary, in step B1, the actual winding phase spectrum and the predicted winding phase spectrum can be substituted into the above formula (4) to calculate the instantaneous phase loss.
[0105] The group delay loss L in step B2 GD , the frequency difference of the real warping phase spectrum and the predicted warping phase spectrum can be substituted into the following formula (5) to calculate:
[0106]
[0107] In formula (5), represents the frequency difference of the predicted warping phase spectrum, Δ DF P represents the frequency difference of the true warping phase spectrum, Δ DF represents the difference along the frequency axis,
[0108] The instantaneous angular frequency loss L in step B3 IAF , the time difference between the predicted warping phase spectrum and the real warping phase spectrum can be substituted into the following formula (6) to calculate:
[0109]
[0110] In formula (6), represents the time difference of the predicted warping phase spectrum, Δ DT P represents the time difference of the true warped phase spectrum after time difference, Δ DT represents the difference along the time axis.
[0111] It can be seen from formulas (5) and (6) that both group delay loss and instantaneous angular frequency loss are activated by the anti-warping function of formula (3), thus avoiding the error expansion caused by phase warping when calculating these two losses.
[0112] In formulas (5) and (6), and The meaning of and formula (4) Consistent, no further elaboration.
[0113] In step B4, the instantaneous phase loss, group delay loss, and instantaneous angular frequency loss calculated in steps B1 to B3 can be substituted into the following formula (7) to calculate the final anti-winding loss L:
[0114] L=L IP +L GD +L IAF (7).
[0115] Optionally, the convergence condition may be set as: the number of training times is greater than or equal to a preset maximum number of training times, where the number of training times is defined as the number of times step S103 is performed.
[0116] In other optional embodiments, different convergence conditions may be set according to actual conditions, which are not limited here.
[0117] S105: Update the parameters of the neural network to be trained according to the anti-convolution loss.
[0118] After executing step S105, the process returns to executing step S103, that is, the process of processing the logarithmic magnitude spectrum using the neural network to be trained to obtain the predicted warped phase spectrum of the sample speech signal.
[0119] In step S105, a gradient back propagation algorithm can be used to calculate the update amount of each parameter in the neural network to be trained based on the anti-convolution loss, and then the value of each parameter in each neural network to be trained is updated according to the update amount.
[0120] S106: Determine the neural network to be trained as a phase prediction neural network.
[0121] The process from steps S101 to S106 can be regarded as the network training process in the method provided in this embodiment.
[0122] S107: Obtain the logarithmic magnitude spectrum of the speech signal to be predicted.
[0123] The method of obtaining the logarithmic magnitude spectrum of the speech signal to be predicted in step S107 is the same as the method of obtaining the logarithmic magnitude spectrum of the sample speech signal in step S102, and will not be repeated here.
[0124] S108 , using a phase prediction neural network to process the logarithmic magnitude spectrum of the speech signal to be predicted, to obtain a warped phase spectrum of the speech signal to be predicted.
[0125] The phase prediction neural network in step S108 is the neural network after the neural network to be trained has completed the training process from S101 to S106. The structures of the two and the functions of each structure are exactly the same. Therefore, the process of using the phase prediction neural network to process the logarithmic amplitude spectrum of the speech signal to be predicted in step S108 to obtain the corresponding warped phase spectrum is the same as the process of using the neural network to be trained to process the sample speech signal in step S103 to obtain the predicted warped phase spectrum of the sample speech signal, and will not be repeated here.
[0126] The processes of steps S107 and S108 can be regarded as the phase prediction process in the method of this embodiment.
[0127] Optionally, after obtaining the warped phase spectrum of the speech signal to be predicted through S108, the speech waveform can be reconstructed by combining the logarithmic amplitude spectrum and the warped phase spectrum of the speech signal to be predicted to obtain the reconstructed speech waveform corresponding to the speech signal to be predicted. Specifically, the logarithmic amplitude spectrum and the warped phase spectrum of the speech signal to be predicted can be combined into a short-time complex spectrum, and then the short-time complex spectrum is inversely short-time Fourier transformed to obtain the corresponding reconstructed speech waveform. The above process can be expressed as follows:
[0128]
[0129] In formula (8), ISTFT() represents the inverse short-time Fourier transform, P0 represents the warped phase spectrum of the speech signal to be predicted output by the phase prediction network in step S108, A represents the amplitude spectrum of the speech signal to be predicted, and i is a complex unit, that is, i is equal to
[0130] The present application provides a method for predicting phase using a parallel estimation architecture network trained with anti-warping loss. The method includes, during the training process, simulating the process of calculating the phase spectrum from the real and imaginary parts of the short-time complex spectrum through two parallel linear convolution layers in the neural network to be trained, and a phase calculation unit, and limiting the predicted phase value to the main value interval to achieve the prediction of the warped phase spectrum, and the anti-warping loss used during training includes the instantaneous phase error, group delay error and instantaneous angular frequency error activated by the anti-warping function, thereby avoiding the error expansion problem caused by phase warping. After the training is completed, the trained phase prediction neural network is used to process the logarithmic amplitude spectrum of the speech signal to be predicted to obtain the warped phase spectrum. This solution directly predicts the warped phase spectrum of the speech signal through a neural network, and introduces an anti-warping function when calculating the loss to solve the error expansion problem caused by phase warping during training, with high efficiency and accuracy.
[0131] According to the method for predicting phase using a parallel estimation architecture network trained with anti-warping loss provided in the embodiment of the present application, the embodiment of the present application also provides a device for predicting phase using a parallel estimation architecture network trained with anti-warping loss, see Figure 4 , is a structural diagram of the device, which may include the following units.
[0132] A generating unit 401 is configured to generate a neural network to be trained; wherein the neural network to be trained includes a residual convolutional network, a first linear convolutional layer and a second linear convolutional layer in parallel, and a phase calculation unit;
[0133] An acquisition unit 402 is configured to acquire a logarithmic magnitude spectrum and a true warped phase spectrum of a sample speech signal;
[0134] a processing unit 403 configured to process the logarithmic magnitude spectrum of the sample speech signal using the neural network to be trained to obtain a predicted warped phase spectrum of the sample speech signal; wherein the predicted warped phase spectrum is calculated by the phase calculation unit based on the pseudo-real part and the pseudo-imaginary part; the pseudo-real part and the pseudo-imaginary part are output by the first linear convolution layer and the second linear convolution layer, respectively; and the phase of the predicted warped phase spectrum is within a principal value interval;
[0135] a calculation unit 404 for calculating an anti-warping loss of the predicted warping phase spectrum and the actual warping phase spectrum; wherein the anti-warping loss is a linear combination of an instantaneous phase loss, a group delay loss, and an instantaneous angular frequency loss between the predicted warping phase spectrum and the actual warping phase spectrum; the instantaneous phase loss, the group delay loss, and the instantaneous angular frequency loss are all activated by an anti-warping function;
[0136] An updating unit 405 is configured to update the parameters of the neural network to be trained according to the anti-warping loss if the anti-warping loss does not meet the preset convergence condition, and return to the step of processing the logarithmic magnitude spectrum using the neural network to be trained to obtain a predicted warping phase spectrum of the sample speech signal;
[0137] a determining unit 406 configured to determine the neural network to be trained as a phase prediction neural network if the anti-wrapping loss meets the convergence condition;
[0138] The acquisition unit 402 is used to obtain the logarithmic magnitude spectrum of the speech signal to be predicted;
[0139] The processing unit 403 is configured to process the logarithmic magnitude spectrum of the speech signal to be predicted using a phase prediction neural network to obtain a warped phase spectrum of the speech signal to be predicted.
[0140] Optionally, when the acquiring unit 402 acquires the real warped phase spectrum of the sample speech signal, it is specifically configured to:
[0141] Performing short-time Fourier transform on the sample speech signal to obtain a short-time complex spectrum of the sample speech signal;
[0142] The phase is calculated based on the real part and the imaginary part of the short-time complex spectrum of the sample speech signal to obtain the real warped phase spectrum of the sample speech signal.
[0143] Optionally, when the calculation unit 404 calculates the anti-winding loss of the predicted winding phase spectrum and the actual winding phase spectrum, it is specifically used to:
[0144] The instantaneous phase loss is calculated based on the real winding phase spectrum and the predicted winding phase spectrum;
[0145] Performing frequency differences on the real winding phase spectrum and the predicted winding phase spectrum respectively, and calculating the group delay loss based on the frequency differences between the real winding phase spectrum and the predicted winding phase spectrum;
[0146] Performing time difference on the real winding phase spectrum and the predicted winding phase spectrum respectively, and calculating the instantaneous angular frequency loss based on the real winding phase spectrum and the predicted winding phase spectrum after time difference;
[0147] The anti-winding loss is obtained by adding the instantaneous phase loss, group delay loss and instantaneous angular frequency loss.
[0148] Optionally, the convergence condition is that the number of training times is greater than or equal to a preset maximum number of training times, wherein the number of training times is defined as the number of times the step of processing the logarithmic amplitude spectrum of the sample speech signal using the neural network to be trained to obtain the predicted warped phase spectrum of the sample speech signal is executed.
[0149] Optionally, the residual convolutional network includes:
[0150] A linear convolution layer, multiple parallel residual convolution blocks connected to the linear convolution layer, an accumulation unit for calculating the mean of the outputs of multiple residual convolution blocks, and a linear unit with leakage correction connected to the accumulation unit.
[0151] The specific working principle and beneficial effects of the device for predicting phase using a parallel estimation architecture network trained with anti-warping loss provided in this embodiment can be found in the relevant steps and beneficial effects of the method for predicting phase using a parallel estimation architecture network trained with anti-warping loss provided in the embodiment of the present application, and will not be repeated here.
[0152] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0153] It should be noted that the concepts of "first" and "second" mentioned in the present invention are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0154] The present application is capable of being implemented or used by those skilled in the art. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to be embodied in the widest possible manner consistent with the principles and novel features disclosed herein.
Claims
1. A method for predicting phase using a parallel estimation architecture network trained with anti-convolution loss, characterized in that: The method comprises: Network training process: Determining a neural network to be trained; wherein the neural network to be trained includes a residual convolutional network, a first linear convolutional layer and a second linear convolutional layer in parallel, and a phase calculation unit; Obtain the logarithmic magnitude spectrum and true warped phase spectrum of the sample speech signal; Processing the logarithmic magnitude spectrum of the sample speech signal using the neural network to be trained to obtain a predicted warped phase spectrum of the sample speech signal; wherein the predicted warped phase spectrum is calculated by the phase calculation unit based on a pseudo-real part and a pseudo-imaginary part; the pseudo-real part and the pseudo-imaginary part are output by the first linear convolution layer and the second linear convolution layer, respectively; and the phase of the predicted warped phase spectrum is within a principal value interval; Calculating an anti-warping loss of the predicted warping phase spectrum and the actual warping phase spectrum; wherein the anti-warping loss is a linear combination of an instantaneous phase loss, a group delay loss, and an instantaneous angular frequency loss between the predicted warping phase spectrum and the actual warping phase spectrum; the instantaneous phase loss, the group delay loss, and the instantaneous angular frequency loss are all activated by an anti-warping function; If the anti-warping loss does not meet the preset convergence condition, updating the parameters of the neural network to be trained according to the anti-warping loss, and returning to the step of processing the logarithmic magnitude spectrum using the neural network to be trained to obtain a predicted warped phase spectrum of the sample speech signal; If the anti-wrapping loss meets the convergence condition, determining the neural network to be trained as a phase prediction neural network; Phase prediction process: Obtaining the logarithmic magnitude spectrum of the speech signal to be predicted; The phase prediction neural network is used to process the logarithmic magnitude spectrum of the speech signal to be predicted to obtain a warped phase spectrum of the speech signal to be predicted.
2. The method according to claim 1, characterized in that The process of obtaining the true warped phase spectrum of the sample speech signal includes: Performing a short-time Fourier transform on the sample speech signal to obtain a short-time complex spectrum of the sample speech signal; Phase calculation is performed based on the real part and the imaginary part of the short-time complex spectrum of the sample speech signal to obtain a real warped phase spectrum of the sample speech signal.
3. The method according to claim 1, characterized in that The calculating the anti-winding loss of the predicted winding phase spectrum and the actual winding phase spectrum includes: Calculating an instantaneous phase loss according to the actual winding phase spectrum and the predicted winding phase spectrum; Performing frequency differentiation on the real warping phase spectrum and the predicted warping phase spectrum respectively, and calculating group delay loss according to the frequency differentiation between the real warping phase spectrum and the predicted warping phase spectrum; performing time differentiation on the real winding phase spectrum and the predicted winding phase spectrum respectively, and calculating instantaneous angular frequency loss according to the time differentiation between the real winding phase spectrum and the predicted winding phase spectrum; The anti-winding loss is obtained by adding the instantaneous phase loss, the group delay loss and the instantaneous angular frequency loss.
4. The method according to claim 1, wherein The convergence condition is that the number of training times is greater than or equal to a preset maximum number of training times, wherein the number of training times is defined as the number of times the step of processing the logarithmic amplitude spectrum of the sample speech signal using the neural network to be trained to obtain the predicted warped phase spectrum of the sample speech signal is executed.
5. The method according to claim 1, wherein The residual convolutional network includes: A linear convolution layer, multiple parallel residual convolution blocks connected to the linear convolution layer, an accumulation unit for calculating the mean of the outputs of the multiple residual convolution blocks, and a linear unit with leakage correction connected to the accumulation unit.
Citation Information
Patent Citations
Passive positioning method based on amplitude and phase information of CSI
CN112147573A
Speech enhancement method based on mask mapping and hybrid cavity convolutional network
CN113936681A