A high-quality vocoder model based on generative adversarial neural network

By adopting a vocoder model based on a generative adversarial neural network in the vocoder technology, combining multi-field fusion and Unet-type hourglass-shaped convolutional neural network, the shortcomings of existing vocoder technologies in sound quality and controllability are solved, and high-quality audio decoding and speech synthesis are achieved.

CN115035904BActive Publication Date: 2025-05-06NANJING UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210391848.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-14
Publication Date
2025-05-06
Estimated Expiration
2042-04-14

AI Technical Summary

Technical Problem

The existing vocoder technology has shortcomings in sound quality and controllability, and the training overhead of vocoder based on neural networks is large, ignoring the frequency domain phase and time domain self-similar information of audio, resulting in slow training convergence and defects in the synthetic waveform details.

Method used

A high-quality vocoder model based on generative adversarial neural network is adopted, including generators, acoustic feature extractors, multi-scale discriminators, multi-period discriminators and multi-phase discriminators. Waveform generation is generated through multi-field fusion and a convolutional neural network with Unet-type hourglass structure, and multiple discriminators and acoustic features are optimized.

Benefits of technology

It realizes high-quality audio decoding, reduces the difficulty of learning neural networks, saves training time and computing resources, improves voice quality, and can naturally and smoothly synthesize long audio sequences of any length.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115035904B_ABST
    Figure CN115035904B_ABST
Patent Text Reader

Abstract

The present invention discloses a high-quality vocoder model based on a generative adversarial neural network. The model first uses a generator module to convert the Mel spectrum of the audio into a waveform form. The model is constructed by a Unet-type hourglass-shaped convolutional neural network with a multi-view fusion block; an acoustic feature extractor and multiple discriminator modules are used to optimize the generated waveform from multiple angles; wherein the acoustic feature extractor is constructed using a traditional signal processing method, and the discriminator module is composed of three parts: a multi-scale discriminator, a multi-cycle discriminator, and a multi-phase discriminator, and is constructed based on a convolutional neural network. The present invention greatly reduces the learning difficulty of the neural network, saves training time and computing resource overhead; uses phase information and self-similar features in the time domain to optimize the generated waveform to obtain a waveform with higher sound quality; uses a localized training strategy, and can synthesize long audio sequences of any length more naturally and smoothly.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a vocoder model, in particular to a high-quality vocoder model based on a generative adversarial neural network. Background Art

[0002] Vocoder or sound synthesizer technology is a digital signal processing technology for encoding and decoding audio waveform data. Vocoder technology has been widely used, including signal data compression, voice and voiceprint recognition, voice and singing synthesis, audio editing and effects, etc.

[0003] In a neural network speech synthesis system, the output of the upstream model is usually the encoding of the target audio data in a latent space of the model, or some more general frequency domain audio encoding designed by humans, such as Mel spectrum, MFCC (Mel-Frequency Ceptral Coefficients) features, etc. However, these encodings cannot directly generate sound waves that can be heard by the human ear through acoustic output devices, but need to use a vocoder to decode these encoded data into time-domain audio waveforms before they can be played through devices such as speakers. Therefore, the vocoder is an indispensable component in such sound processing systems.

[0004] Currently, traditional vocoders based on digital signal processing methods have poor sound quality and low controllability, while vocoders based on neural networks have high training costs and ignore the effective use of information such as frequency domain phase and time domain self-similarity of audio, resulting in slow training convergence and detailed flaws in the synthesized waveform. Summary of the invention

[0005] Purpose of the invention: The technical problem to be solved by the present invention is to provide a high-quality vocoder model based on a generative adversarial neural network in view of the shortcomings of the prior art.

[0006] In order to solve the above technical problems, the present invention discloses a high-quality vocoder model based on a generative adversarial neural network, comprising the following steps:

[0007] Step 1, construct a high-quality vocoder model based on a generative adversarial neural network, which includes: a generator, an acoustic feature extractor, a multi-scale discriminator, a multi-cycle discriminator, and a multi-phase discriminator;

[0008] Step 2, obtain PCM-encoded audio data from the data set to obtain the real waveform;

[0009] Step 3, preprocessing the real waveform obtained in step 2, dividing the training set and the validation set, slicing the training set, and obtaining the Mel spectrum and rough waveform;

[0010] Step 4, sending the Mel spectrum and rough waveform obtained in step 3 to a generator to obtain a generated waveform;

[0011] Step 5, the real waveform in step 2 and the corresponding generated waveform in step 4 are sent to the acoustic feature extractor and three discriminators, namely, the multi-scale discriminator, the multi-cycle discriminator and the multi-phase discriminator, to obtain the acoustic features, the scores of the three discriminators and the feature maps of the three discriminators, and then substituted into the discriminator loss function to calculate the three discriminator loss values ​​and optimize the discriminator parameters;

[0012] Step 6, substituting the acoustic features, the discriminator scores and the feature maps described in step 5 into the generator loss function to calculate the generator loss and optimize the generator parameters; repeating the training process of steps 5 and 6 until the vocoder model converges;

[0013] Step 7: Use the validation set data obtained in step 3 to evaluate the model performance and complete the construction and training of a high-quality vocoder model based on a generative adversarial neural network.

[0014] In step 2 of the present invention, the data set does not restrict the audio data content to be music, human voice or noise, and the audio data is a group of audio files encoded by PCM.

[0015] The preprocessing in step 3 of the present invention includes: extraction of linear amplitude spectrum, phase spectrum, Mel spectrum, rough waveform and level envelope features, the method is as follows:

[0016] First, all audio data are resampled at a uniform sampling rate, and then audio features are extracted, including: extracting linear amplitude spectrum and phase spectrum through short-time Fourier transform; extracting Mel spectrum through Mel filter bank, and then obtaining rough waveform through Griffin-Lim algorithm; extracting level envelope through MaxPooling pooling layer.

[0017] The division of the training set and the validation set in step 3 of the present invention includes: dividing the data into a non-overlapping training set and a test set.

[0018] The slicing of the training set in step 3 of the present invention includes: slicing the data of the training set into overlapping, fixed-length slices to implement a localized training strategy.

[0019] The generator in step 4 of the present invention is a convolutional neural network with multi-view fusion and Unet-type hourglass structure; the network uses a given Mel spectrum as a reference, and obtains a generated waveform by multi-step transformation of the rough waveform through encoder shortening and decoder stretching; the network includes:

[0020] The encoder consists of Conv1D downsampling layers, which transforms the rough waveform from the temporal space to the spectral space;

[0021] The decoder, which consists of ConvTransposed1D upsampling layers, restores the hidden layer encoding in the spectral space to the time domain space;

[0022] The encoder and decoder contain multiple multi-view fusion blocks with residuals, ResBlock, which serve as the backbone network for feature mapping;

[0023] The multiple Conv1D concatenation layers contained in the decoder are used to fuse the hidden layer encoding information from the peer layers in the encoder to obtain the generated waveform.

[0024] In step 5 of the present invention, the acoustic feature extractor is a short-time Fourier transform process for extracting a phase spectrum;

[0025] The three discriminators are: a multi-scale discriminator, a multi-cycle discriminator and a multi-phase discriminator;

[0026] Among them, the multi-scale discriminator uses Conv1D network to identify the authenticity of the generated waveform at three different waveform scales, including the original waveform, the two-fold downsampled waveform and the four-fold downsampled waveform; the multi-period discriminator uses Conv2D network to identify the authenticity of the generated waveform after grouping under five conditions of grouping period of 2, 3, 5, 7 and 11; the multi-phase discriminator uses Conv2D network to identify the authenticity of the phase spectrum obtained by the acoustic feature extractor of the generated waveform under three settings of FFT points of 512, 1024 and 2048 respectively;

[0027] The discriminator loss is the sum of three discriminator adversarial losses, and the optimizer used is Adam.

[0028] In step 5 of the present invention, the method for calculating the loss values ​​of the three discriminators includes:

[0029] Each discriminator has two outputs: the discriminator score D x , the feature map of the discriminator where the subscript x is taken as s, f, and p to refer to the multi-scale discriminator, multi-period discriminator, and multi-phase discriminator, respectively;

[0030] The discriminator loss includes: the discriminator adversarial loss composed of the scores from the three discriminators, the method is:

[0031] loss d =d s +d f +d p

[0032] Among them, the three discriminators confront the loss d s d f and d p They are:

[0033]

[0034]

[0035]

[0036] Among them, the score of the multi-scale discriminator is D s , the score of the multi-cycle discriminator is D f , the score of the multi-phase discriminator is D p , the generator is G, the target true waveform is y, the Mel spectrum to be decoded is mel, and the rough waveform of the Mel spectrum to be decoded is wav; the short horizontal lines on the three discriminator scores represent the mean.

[0037] In step 6 of the present invention, the acoustic feature extractor comprises: a short-time Fourier transform process for extracting a phase spectrum, a MaxPooling layer for extracting an actual level envelope of a waveform;

[0038] The discriminator and optimizer are consistent with those in step 5;

[0039] Generator loss g Including: three generator adversarial losses, three generator feature map losses, multiple spectrum amplitude losses, level envelope losses, waveform self-similarity losses. The specific calculation methods include:

[0040] loss g =(g s +g f +g p )+α*(fm s +fm f +fm p )+β*mstft+γ*dyn+δ*sm

[0041] Three generators against the loss g s , g f and g p for:

[0042]

[0043]

[0044]

[0045] Three feature map matching losses fm s 、fm f and FM p for:

[0046]

[0047]

[0048]

[0049] The multiple spectrum amplitude loss mstft is:

[0050]

[0051] The level envelope loss dyn is:

[0052] dyn=|MaxPooling(y)-MaxPooling(G(mel,wav)))|+|MaxPooling(-y)-MaxPooling(-G(mel,wav)))|

[0053] The waveform self-similarity loss sm is:

[0054] sm=|y even -y odd |

[0055] Among them, the feature map of the multi-scale discriminator is The feature map of the multi-cycle discriminator is The feature map of the multi-phase discriminator is The logarithmic-scale amplitude spectrum obtained by the i-th set of short-time Fourier transform is stft i ,y even and odd are the sampling points on the even and odd bits of the original signal respectively; α, β, γ and δ are balance factor constants; and the double vertical lines represent absolute values.

[0056] The performance evaluation described in step 7 of the present invention is obtained based on the validation set data in step 3, including objective model loss, generalization evaluation and subjective sound quality auditory evaluation.

[0057] Beneficial effects:

[0058] The present invention utilizes a Unet-type network structure to fuse the Mel spectrum and the rough waveform for waveform generation, and simultaneously uses a variety of discriminators and acoustic features to optimize the generated waveform, providing a high-quality neural network vocoder model based on a generative adversarial network architecture. Compared with the existing vocoder model, the biggest feature of the present invention is that it utilizes the Unet structure to fuse the information of the rough waveform to greatly reduce the learning difficulty of the neural network, while effectively utilizing the phase information and the level envelope and self-similar features in the time domain to optimize the generated waveform, ultimately achieving the dual goals of reducing training resources and improving speech quality.

[0059] High-quality decoding of speech data based on Mel spectrum is achieved through this neural network vocoder model. Since a rough version of the target real waveform is used as a reference input, the learning difficulty of the neural network is greatly reduced, thereby saving training time and computing resource overhead. Since the generated waveform is optimized by utilizing phase information and self-similar features in the time domain, a waveform with higher sound quality can be obtained. Due to the use of a localized training strategy, the vocoder can synthesize long audio sequences of arbitrary length more naturally and smoothly. The present invention applies the generative adversarial network architecture to the construction of a neural network vocoder, and designs the above-mentioned module to achieve high-quality audio decoding, with low training overhead, good sound quality, and high controllability, and can be used as a core basic component of various audio processing systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments, and the above and / or other advantages of the present invention will become more clear.

[0061] Figure 1 It is a schematic diagram of the architecture of the present invention.

[0062] Figure 2 This is the data flow diagram during training of the present invention.

[0063] Figure 3 This is the data flow diagram when the present invention is inferred.

[0064] Figure 4 The figure is a schematic diagram for comparing the waveform generated by the present invention with the original waveform.

[0065] Figure 5 The figure is a schematic diagram for comparing the frequency spectrum corresponding to the waveform generated by the present invention and the frequency spectrum corresponding to the original waveform. DETAILED DESCRIPTION

[0066] like Figure 1 As shown, a high-quality vocoder model based on a generative adversarial neural network includes the following steps:

[0067] Step 1, construct a high-quality vocoder model based on a generative adversarial neural network, which includes: a generator, an acoustic feature extractor, a multi-scale discriminator, a multi-cycle discriminator, and a multi-phase discriminator;

[0068] Step 2, obtaining PCM (Pulse Code Modulation) encoded audio data from the data set to obtain a real waveform; wherein the data set does not restrict the audio data content to be music, human voice or noise, and the audio data is a set of PCM encoded audio files.

[0069] Step 3, preprocessing the real waveform obtained in step 2, dividing the training set and the validation set, slicing the training set, and obtaining the Mel spectrum and rough waveform;

[0070] The preprocessing includes: extraction of linear amplitude spectrum, phase spectrum, Mel spectrum, rough waveform and level envelope features, the method is as follows:

[0071] First, all audio data are resampled at a uniform sampling rate, and then audio features are extracted, including: extracting linear amplitude spectrum and phase spectrum through short-time Fourier transform; extracting Mel spectrum through Mel filter bank, and then obtaining rough waveform through Griffin-Lim algorithm (reference: Griffin, Daniel, and Jae Lim. "Signal estimation from modified short-time Fourier transform." IEEE Transactions on Acoustics, Speech, and Signal Processing 32.2 (1984): 236-243.); extracting level envelope through MaxPooling pooling layer.

[0072] The training set and validation set division includes: dividing the data into a non-overlapping training set and a test set.

[0073] The slicing of the training set includes: slicing the data of the training set into overlapping, fixed-length slices to implement a localized training strategy.

[0074] Step 4, sending the Mel spectrum and the rough waveform obtained in step 3 to a generator to obtain a generated waveform; wherein the generator is a convolutional neural network (CNN) with a multi-view fusion and Unet (reference: U-Net model, Ronneberger, O., P. Fischer, and T. Brox. "U-Net: Convolutional Networks for Biomedical Image Segmentation." Springer International Publishing (2015).) hourglass-shaped structure; the network uses a given Mel spectrum as a reference, and transforms the rough waveform through a multi-step transformation of an encoder shortening and a decoder stretching to obtain a generated waveform; the network includes:

[0075] The encoder composed of Conv1D downsampling layers (can be multiple, usually three or four) converts the rough waveform from the time domain space to the spectral space;

[0076] The decoder composed of ConvTransposed1D upsampling layers (can be multiple, as long as the number of encoders mentioned above is consistent) restores the hidden layer encoding in the spectral space to the time domain space;

[0077] The encoder and decoder contain multiple multi-view fusion blocks with residuals, ResBlock, which serve as the backbone network for feature mapping;

[0078] The multiple Conv1D concatenation layers contained in the decoder are used to fuse the hidden layer encoding information from the peer layers in the encoder to obtain the generated waveform.

[0079] Step 5, the real waveform in step 2 and the corresponding generated waveform in step 4 are sent to the acoustic feature extractor and three discriminators, namely, the multi-scale discriminator, the multi-cycle discriminator and the multi-phase discriminator, to obtain the acoustic features, the scores of the three discriminators and the feature maps of the three discriminators, and then substituted into the discriminator loss function to calculate the three discriminator loss values ​​and optimize the discriminator parameters;

[0080] Wherein, the acoustic feature extractor is a short-time Fourier transform process for extracting a phase spectrum;

[0081] The three discriminators are: a multi-scale discriminator, a multi-cycle discriminator and a multi-phase discriminator;

[0082] Among them, the multi-scale discriminator uses Conv1D network to identify the authenticity of the generated waveform at three different waveform scales, including the original waveform, the two-fold downsampled waveform and the four-fold downsampled waveform; the multi-period discriminator uses Conv2D network to identify the authenticity of the generated waveform after grouping under five conditions of grouping period of 2, 3, 5, 7 and 11; the multi-phase discriminator uses Conv2D network to identify the authenticity of the phase spectrum obtained by the acoustic feature extractor of the generated waveform under three settings of FFT points of 512, 1024 and 2048 respectively;

[0083] The discriminator loss is the sum of three discriminator adversarial losses, and the optimizer used is Adam (reference: Kingma, D. and J. Ba "Adam: A Method for Stochastic Optimization." Computer Science (2014).).

[0084] The methods for calculating the three discriminator loss values ​​include:

[0085] Each discriminator has two outputs: the discriminator score D x , the feature map of the discriminator where the subscript x is taken as s, f, and p to refer to the multi-scale discriminator, multi-period discriminator, and multi-phase discriminator, respectively;

[0086] The discriminator loss includes: the discriminator adversarial loss composed of the scores from the three discriminators, the method is:

[0087] loss d =d s +d f +d p

[0088] Among them, the three discriminators confront the loss d s d f and d p They are:

[0089]

[0090]

[0091]

[0092] Among them, the score of the multi-scale discriminator is D s , the score of the multi-cycle discriminator is D f , the score of the multi-phase discriminator is D p , the generator is G, the target true waveform is y, the Mel spectrum to be decoded is mel, and the rough waveform of the Mel spectrum to be decoded is wav; the short horizontal lines on the three discriminator scores represent the mean.

[0093] Step 6, substituting the acoustic features, the discriminator scores and the feature maps described in step 5 into the generator loss function to calculate the generator loss and optimize the generator parameters; repeating the training process of steps 5 and 6 until the vocoder model converges;

[0094] The acoustic feature extractor comprises: a short-time Fourier transform process for extracting a phase spectrum, a MaxPooling layer for extracting an actual level envelope of a waveform;

[0095] The discriminator and optimizer are consistent with those in step 5;

[0096] Generator loss g Including: three generator adversarial losses, three generator feature map losses, multiple spectrum amplitude losses, level envelope losses, waveform self-similarity losses. The specific calculation methods include:

[0097] loss g =(g s +g f +g p )+α*(fm s+fm f +fm p )+β*mstft+γ*dyn+δ*sm

[0098] Three generators against the loss g s , g f and g p for:

[0099]

[0100]

[0101]

[0102] Three feature map matching losses fm s 、fm f and FM p for:

[0103]

[0104]

[0105]

[0106] The multiple spectrum amplitude loss mstft is:

[0107]

[0108] The level envelope loss dyn is:

[0109] dyn=|MaxPooling(y)-MaxPooling(G(mel,wav)))|+|MaxPooling(-y)-MaxPooling(-G(mel,wav)))|

[0110] The waveform self-similarity loss sm is:

[0111] sm=|y even -y odd |

[0112] Among them, the feature map of the multi-scale discriminator is The feature map of the multi-cycle discriminator is The feature map of the multi-phase discriminator is The logarithmic-scale amplitude spectrum obtained by the i-th set of short-time Fourier transform is stft i ,y even and odd are the sampling points on the even and odd bits of the original signal respectively; α, β, γ and δ are balance factor constants; and the double vertical lines represent absolute values.

[0113] Step 7, use the validation set data obtained in step 3 to evaluate the model performance, and complete the construction and training of a high-quality vocoder model based on a generative adversarial neural network; wherein the performance evaluation is based on the validation set data in step 3, including objective model loss, generalization evaluation, and subjective sound quality auditory evaluation.

[0114] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0115] Example

[0116] This embodiment provides a neural network vocoder model construction and training method based on a generative adversarial neural network, and its model structure is as follows: Figure 1 As shown in Figure 2 The detailed process is as follows:

[0117] 1. Build a high-quality vocoder model based on generative adversarial neural network

[0118] In this embodiment, Python language and PyTorch framework are used to implement the vocoder model. It consists of the following parts: 1. A generator module built by multi-view fusion blocks and Unet-style hourglass-shaped structure CNN (Convolutional Neural Network); 2. An acoustic feature extractor composed of multiple traditional signal processing methods; 3. Three discriminator modules built by CNN, namely multi-scale discriminator, multi-cycle discriminator, and multi-phase discriminator; 4. Relative adversarial and feature map matching losses, three acoustic feature losses. The details of each module will be given in the relevant places in the following steps.

[0119] 2. Get audio data from the dataset

[0120] The data set used in this embodiment is the Chinese female voice data set "Chinese Mandarin Female DB-1" open sourced and published by DataBaker Technology Co., Ltd., but any other audio data set may also be used, for example, a set of PCM-encoded wav format files collected arbitrarily.

[0121] 3. Data preprocessing

[0122] First, resample all audio to a uniform sampling rate, such as 16kHz. Then, for each audio file, go through the short-time Fourier transform, Mel filter bank, numerical truncation, logarithm, etc., to obtain the 80-segment Mel spectrum in the 125-7600Hz frequency band under the logarithmic scale. Based on the obtained Mel spectrum, go through the inverse Mel filter bank, the Griffin-Lim algorithm with a small number of iterations, etc., to obtain a rough waveform with background noise. Then pair the Mel spectrum with its corresponding rough waveform into a binary data = (mel, wav).

[0123] The entire dataset is divided into two disjoint subsets, namely the training set and the validation set, in an appropriate ratio. For the training set, each tuple is sliced ​​overlappingly with a fixed segment size segment_size and aligned in length so that segment_size = length (wav i )=length(mel i )×hop_length, where hop_length is the frame hopping length used by the short-time Fourier transform, such as 256; thus obtaining a set of fixed-length data pairs data i =(mel i ,wav i ), which is used as input for model training. The training technique of this localized training strategy helps improve the quality of long sequence synthesis.

[0124] 4. Obtain the generated waveform through the generator

[0125] Use the preprocessed data to train a CNN generator module with multi-view fusion and Unet-style hourglass structure. The module includes: 1. An encoder composed of multiple Conv1D downsampling layers to convert the time domain space to the spectral space; 2. A decoder composed of multiple ConvTransposed1D upsampling layers to restore the spectral space to the time domain space; 3. The encoder and decoder contain multiple multi-view fusion blocks with residuals, which are used as the backbone network for feature mapping; 4. The decoder contains multiple Conv1D splicing layers to fuse the hidden layer coding information from the peer layer in the encoder. The overall construction forms a Unet-style hourglass-shaped symmetrical structure, so that parameter reuse can be achieved in the peer layer, saving nearly half of the parameters.

[0126] The specific data flow is as follows. First, the rough waveform wav is converted into the code E in the spectrum space through the encoder wav , a concatenation layer is used in the spectrum space to fuse the binary formed by the code and the Mel spectrum (E wav,mel) information. Then the information fused code is restored to the time domain space through the decoder to obtain the generated waveform. In particular, a splicing layer is used in each layer of the decoder to fuse the intermediate output information from the encoder peer layer.

[0127] 5. Calculate the discriminator loss and optimize the discriminator parameters

[0128] The generated waveform and the target true waveform are sent to the acoustic feature extractor and three discriminators to obtain the discriminator loss, and then the Adam optimizer is used to optimize the parameters of the discriminator.

[0129] The acoustic feature extractor includes: a short-time Fourier transform process to extract the phase spectrum.

[0130] The three discriminators are: multi-scale discriminator, multi-period discriminator, and multi-phase discriminator. Among them, the multi-scale discriminator uses Conv1D network to identify the authenticity of the generated waveform at three scales: the original waveform, the two-fold down-sampled version, and the four-fold down-sampled version of the generated waveform; the multi-period discriminator uses Conv2D network to identify the authenticity of the generated waveform after grouping under five conditions of grouping period of 2, 3, 5, 7, and 11; the multi-phase discriminator uses Conv2D network to identify the authenticity of the phase spectrum obtained by the acoustic feature extractor of the generated waveform under three settings of FFT points of 512, 1024, and 2048. Note that each discriminator has two outputs: the score D of the discriminator and the score D of the acoustic feature extractor. x , the feature map of the discriminator The subscript x is s, f, and p to refer to the multi-scale discriminator, multi-period discriminator, and multi-phase discriminator, respectively.

[0131] The discriminator loss includes: the discriminator adversarial loss composed of the scores from the three discriminators, the formula is:

[0132] loss d =d s +d f +d p

[0133] Discriminator adversarial loss:

[0134]

[0135]

[0136]

[0137] The symbols are: The score D of the multi-scale discriminator s , the score D of the multi-cycle discriminator f , the score D of the multi-phase discriminatorp , generator G, target true waveform y, Mel spectrum to be decoded mel, rough waveform wav of Mel spectrum to be decoded; the short horizontal line represents the mean.

[0138] 6. Calculate the generator loss and optimize the generator parameters

[0139] The generated waveform and the target true waveform are sent to the acoustic feature extractor and three discriminators to obtain the generator loss, and then the Adam optimizer is used to optimize the parameters of the generator.

[0140] The acoustic feature extractor includes: a short-time Fourier transform process to extract the phase spectrum, and a MaxPooling layer to extract the actual level envelope of the waveform.

[0141] The three discriminators are consistent with the description in step 4.

[0142] The generator loss includes: the generator adversarial loss composed of the scores from the three discriminators, the feature map matching loss composed of the feature maps from the three discriminators, the multi-spectral amplitude loss, level envelope loss, and waveform self-similarity loss from the acoustic feature extractor, and the formula is:

[0143] loss g =(g s +g f +g p )+α*(fm s +fm f +fm p )+β*mstft+γ*dyn+δ*sm

[0144] Generator adversarial loss:

[0145]

[0146]

[0147]

[0148] Feature map matching loss:

[0149]

[0150]

[0151]

[0152] Multiplet amplitude loss:

[0153]

[0154] Level envelope loss:

[0155] dyn=|MaxPooling(y)-MaxPooling(G(mel,wav)))|+|MaxPooling(-y)-MaxPooling(-G(mel,wav)))|

[0156] Waveform self-similarity loss:

[0157] sm=|y even -y odd |

[0158] The symbols are: Feature map of multi-scale discriminator Feature map of multi-cycle discriminator Feature map of multi-phase discriminator The logarithmic-scale amplitude spectrum stft obtained by the i-th set of short-time Fourier transform i ,y even and odd are the sampling points on the even and odd bits of the original signal, α, β, γ, δ are the balance factor constants, and the double vertical lines represent the absolute values. The remaining symbols are the same as those described in step 4.

[0159] It is particularly noted that among the acoustic feature losses, the multi-spectral amplitude loss focuses on coarse-grained constraints on the acoustic features of the generated waveform in the frequency domain, while the level envelope loss and waveform self-similarity loss focus on more fine-grained constraints on the waveform in the time domain; multi-faceted loss feedback can prompt the generator to produce better waveforms.

[0160] 7. Model performance evaluation

[0161] Use validation set data to objectively evaluate the generator loss to evaluate the generalization ability of the model and detect overfitting risks. Figure 4 , Figure 5 As shown in the figure, the waveforms are similar and the spectrum loss is small, which shows the fidelity of the audio quality. At the same time, waveform generation and subjective evaluation of human hearing experiments are carried out to evaluate the sound quality of the model output.

[0162] This embodiment provides a neural network vocoder model inference (use) method based on a generative adversarial neural network. The inference process is as follows: Figure 3 The detailed process is as follows:

[0163] 1. Obtain the trained vocoder model

[0164] This embodiment uses the vocoder model trained in the above-mentioned embodiment.

[0165] 2. Get the Mel spectrum to be decoded

[0166] The Mel spectrum used in this embodiment is an 80-segment Mel spectrum on a logarithmic scale in the frequency band of 125 to 7600 Hz, which is predicted by an upstream model of an audio processing system, or obtained by sequentially performing short-time Fourier transform, Mel filter bank, numerical truncation, logarithm, and the like on the original audio waveform.

[0167] 3. Calculate the rough waveform

[0168] For a given Mel spectrum, it is first converted into a linear spectrum through an inverse Mel filter bank, and then the Griffin-Lim algorithm with fewer iterations is used to obtain a rough waveform containing background noise.

[0169] 4. Obtain the generated waveform through the generator

[0170] The given Mel spectrum and its corresponding rough waveform are sent to the generator module of the model described in step 1, and a high-quality generated waveform is obtained by calculation.

[0171] The present invention provides a concept and method for a high-quality vocoder model based on a generative adversarial neural network. There are many methods and approaches to implement the technical solution. The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention. All components not specified in this embodiment can be implemented by existing technologies.

Claims

1. A high-quality vocoder model based on a generative adversarial neural network, characterized in that: The following steps are involved: Step 1, construct a high-quality vocoder model based on a generative adversarial neural network, which includes: a generator, an acoustic feature extractor, a multi-scale discriminator, a multi-cycle discriminator, and a multi-phase discriminator; Step 2, obtain the PCM-encoded audio data from the data set to obtain the real waveform; Step 3, preprocessing the real waveform obtained in step 2, dividing the training set and the validation set, slicing the training set, and obtaining the Mel spectrum and rough waveform; Step 4, sending the Mel spectrum and rough waveform obtained in step 3 to a generator to obtain a generated waveform; Step 5, the real waveform in step 2 and the corresponding generated waveform in step 4 are sent to the acoustic feature extractor and three discriminators, namely, the multi-scale discriminator, the multi-cycle discriminator and the multi-phase discriminator, to obtain the acoustic features, the scores of the three discriminators and the feature maps of the three discriminators, and then substituted into the discriminator loss function to calculate the three discriminator loss values ​​and optimize the discriminator parameters; Step 6, substituting the acoustic features, the discriminator scores and the feature maps described in step 5 into the generator loss function to calculate the generator loss and optimize the generator parameters; repeating the training process of steps 5 and 6 until the vocoder model converges; Step 7: Use the validation set data obtained in step 3 to evaluate the model performance and complete the construction and training of a high-quality vocoder model based on a generative adversarial neural network.

2. A high-quality vocoder model based on a generative adversarial neural network according to claim 1, characterized in that: In step 2, the data set does not restrict the audio data content to be music, human voice or noise, and the audio data is a set of audio files encoded in PCM.

3. A high-quality vocoder model based on a generative adversarial neural network according to claim 2, characterized in that: The preprocessing in step 3 includes: extraction of linear amplitude spectrum, phase spectrum, Mel spectrum, rough waveform and level envelope features, the method is as follows: First, all audio data are resampled at a uniform sampling rate, and then audio features are extracted, including: extracting linear amplitude spectrum and phase spectrum through short-time Fourier transform; extracting Mel spectrum through Mel filter bank, and then obtaining rough waveform through Griffin-Lim algorithm; extracting level envelope through MaxPooling pooling layer.

4. A high-quality vocoder model based on a generative adversarial neural network according to claim 3, characterized in that: The division of the training set and the validation set in step 3 includes: dividing the data into non-overlapping training sets and test sets.

5. A high-quality vocoder model based on a generative adversarial neural network according to claim 4, characterized in that: The slicing of the training set in step 3 includes: slicing the data of the training set into overlapping, fixed-length slices to implement a localized training strategy.

6. A high-quality vocoder model based on a generative adversarial neural network according to claim 5, characterized in that: The generator in step 4 is a convolutional neural network with multi-view fusion and Unet-type hourglass structure; the network uses a given Mel spectrum as a reference, and transforms the rough waveform through a multi-step transformation of the encoder shortening and the decoder stretching to obtain a generated waveform; the network includes: The encoder consists of a one-dimensional convolutional Conv1D downsampling layer, which transforms the rough waveform from the time domain space to the spectral space; The decoder, which consists of a one-dimensional transposed convolution ConvTransposed1D upsampling layer, restores the hidden layer encoding in the spectral space to the time domain space; The encoder and decoder contain multiple multi-view fusion blocks with residuals, ResBlock, which serve as the backbone network for feature mapping; The multiple Conv1D concatenation layers contained in the decoder are used to fuse the hidden layer encoding information from the peer layers in the encoder to obtain the generated waveform.

7. A high-quality vocoder model based on a generative adversarial neural network according to claim 6, characterized in that: In step 5, the acoustic feature extractor is a short-time Fourier transform process for extracting a phase spectrum; The three discriminators are: a multi-scale discriminator, a multi-cycle discriminator and a multi-phase discriminator; Among them, the multi-scale discriminator uses Conv1D network to identify the authenticity of the generated waveform at three different waveform scales, including the original waveform, the two-fold downsampled waveform and the four-fold downsampled waveform; the multi-period discriminator uses the two-dimensional convolution Conv2D network to identify the authenticity of the generated waveform after grouping under the five conditions of grouping period of 2, 3, 5, 7 and 11; the multi-phase discriminator uses Conv2D network to identify the authenticity of the phase spectrum obtained by the acoustic feature extractor of the generated waveform under the three settings of fast Fourier transform point number FFT of 512, 1024 and 2048 respectively; The discriminator loss is the sum of three discriminator adversarial losses, and the optimizer used is the Adam optimizer.

8. A high-quality vocoder model based on a generative adversarial neural network according to claim 7, characterized in that: In step 5, the method for calculating the loss values ​​of the three discriminators includes: Each discriminator has two outputs: the discriminator score D x , the feature map of the discriminator where the subscript x is taken as s, f, and p to refer to the multi-scale discriminator, multi-period discriminator, and multi-phase discriminator, respectively; The discriminator loss includes: the discriminator adversarial loss composed of the scores from the three discriminators, the method is: loss d =d s +d f +d p Among them, the three discriminators confront the loss d s d f and d p They are: Among them, the score of the multi-scale discriminator is D s , the score of the multi-cycle discriminator is D f , the score of the multi-phase discriminator is D p , the generator is G, the target true waveform is y, the Mel spectrum to be decoded is mel, and the rough waveform of the Mel spectrum to be decoded is wav; the short horizontal lines on the three discriminator scores represent the mean.

9. A high-quality vocoder model based on a generative adversarial neural network according to claim 8, characterized in that: In step 6, the acoustic feature extractor includes: a short-time Fourier transform process for extracting a phase spectrum, and a MaxPooling layer for extracting an actual level envelope of a waveform; The discriminator and optimizer are consistent with those in step 5; Generator loss g Including: three generator adversarial losses, three generator feature map losses, multiple spectrum amplitude losses, level envelope losses, waveform self-similarity losses. The specific calculation methods include: loss g =(g s +g f +g p )+α*(fm s +fm f +fm p )+β*mstft+γ*dyn+δ*sm Three generators against the loss g s , g f and g p for: Three feature map matching losses fm s 、fm f and FM p for: The multiple spectrum amplitude loss mstft is: The level envelope loss dyn is: dyn=|MaxPooling(y)-MaxPooling(G(mel,wav))|+|MaxPooling(-y)-MaxPooling(-G(mel,wav))| The waveform self-similarity loss sm is: sm=|y even -y odd | Among them, the feature map of the multi-scale discriminator is The feature map of the multi-cycle discriminator is The feature map of the multi-phase discriminator is The logarithmic-scale amplitude spectrum obtained by the i-th set of short-time Fourier transform is stft i ,y even and odd are the sampling points on the even and odd bits of the original signal respectively; α, β, γ and δ are balance factor constants; and the double vertical lines represent absolute values.

10. A high-quality vocoder model based on a generative adversarial neural network according to claim 9, characterized in that: The performance evaluation described in step 7 is based on the validation set data in step 3, including objective model loss, generalization evaluation, and subjective sound quality auditory evaluation.

Citation Information

Patent Citations

  • Multi-scale StarGAN voice conversion method based on shared training

    CN111462768A

  • Speech synthesis method based on generative adversarial network

    CN113066475A