Working method of a zero-shot general vocoder based on contrastive learning and generative adversarial network
By extracting the speaker characteristics in the Mel score and conducting adversarial training in the generator based on the method of comparative learning and generative adversarial network, the problem of synthesis quality reduction in general vocoder in multi-speaker, multi-language or multi-style speech synthesis is solved, and high-quality audio generation and timbre restoration are achieved.
Patent Information
- Application Number
- CN202211192592.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-28
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-09-28
AI Technical Summary
When generating multi-speaker, multi-language or multi-style voice, existing universal vocoders have problems such as degraded synthesis quality, mechanical sound intensity, and poor audio high-frequency details recovery, especially when sound synthesis outside the dataset is not performed well.
Using a method based on contrast learning and generative adversarial network, the speaker features in the Mel score are extracted through the speaker encoder, and a fusion module and a multi-scale discriminator are introduced into the generator for adversarial training to generate high-quality audio.
The audio synthesis quality of the model for the speaker who has never seen before is improved, the model's generalization ability is enhanced, and the synthetic audio can highly reflect the speaker's personal characteristics and tone.
Smart Images

Figure CN115662451B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech synthesis, and in particular to a working method of a zero-shot general vocoder based on contrastive learning and generative adversarial network. Background Art
[0002] With the development of artificial intelligence and the popularization of smart cities and smart homes, speech synthesis appears more and more in people's lives. Therefore, the generation and modeling of waveforms are also tasks that are very much in need but challenging recently.
[0003] At present, a large number of studies have proved that vocoders have excellent performance in terms of generation speed and audio fidelity when trained with single-speaker utterances. However, some models still face difficulties in generating natural-sounding voices in multiple domains, such as multi-speaker, multi-language, or multi-style speech. The capabilities of these models can be evaluated by the sound quality when the model is trained on data of multiple speakers and the sound quality of audio that does not exist in the training set. A vocoder that can generate high-fidelity audio in various domains and can handle situations whether the input is encountered during training or from outside the training set is usually called a general vocoder. With the maturity of neural network vocoder technology, from autoregressive vocoders such as the causal-convolution-based vocoder WaveNet to vocoders based on generative adversarial networks, through continuous optimization of the model structure and adjustment of the loss function, etc., recent vocoders have gradually achieved a balance between synthesis quality and synthesis speed. The SOTA vocoder has improved the generality for voices outside the dataset to a certain extent, but due to the large differences in timbre characteristics between different individuals and populations, when the vocoder synthesizes voices outside the dataset, there are still problems such as a decrease in quality, a strong mechanical sound, and poor recovery of high-frequency details of the audio. Therefore, the general vocoder is still a task worthy of research. Summary of the Invention
[0004] The present invention provides a working method of a zero-shot general vocoder based on contrastive learning and generative adversarial network, including the following steps:
[0005] Step 1, input the target synthetic Mel spectrogram into the model and perform a logarithmic transformation on the values.
[0006] Step 2, input the input Mel spectrogram into the speaker encoder to obtain the speaker encoding representation.
[0007] Step 3, input the Mel spectrogram input in Step 1 and the speaker encoding representation obtained in Step 2 into the generator, and after multiple upsamplings and convolutions in the adversarial-trained generation module, finally the generation module outputs the synthesized audible waveform.
[0008] As a further improvement of the present invention, in step 1, the input Mel-spectrum values need to be normalized, and then the normalized Mel-spectrum is input into the model.
[0009] As a further improvement of the present invention, in step 2, the speaker encoder encodes the speaker feature information hidden in the Mel-spectrum through an unsupervised method, and uses a residual network trained by a pre-trained contrastive learning method to learn and encode the Mel-spectrum.
[0010] As a further improvement of the present invention, in step 2, the speaker encoder includes a 34-layer residual network for extracting speaker features, a pooling layer for integrating speaker features, a linear layer, an activation layer, and an average pooling layer along the time direction for aggregating features of each frame. Specifically, it further includes the following steps:
[0011] Step 211: Extract features from the input Mel-spectrum through the residual network. The residual network contains 16 residual blocks for feature extraction and downsampling to obtain two-dimensional speaker features.
[0012] Step 212: Perform average pooling on the obtained two-dimensional speaker features to obtain global features.
[0013] Step 213: Map the global features obtained in step 212 through the linear layer to obtain the mapped speaker features.
[0014] Step 214: Perform average pooling on the speaker features obtained in step 213 along the time axis direction to obtain the final one-dimensional speaker encoding representation.
[0015] As a further improvement of the present invention, in step 2, a contrastive learning method is introduced to pre-train the speaker encoder. The specific steps are as follows:
[0016] Step 221: In the training stage, the speaker encoder uses a set of random audio to form a training batch as input. In this training batch, two fixed-length and non-overlapping sub-audio segments are randomly selected from each audio, and the Mel-spectrum corresponding to each audio segment is the input of the Mel-spectrum encoder. Among them, the Mel-spectrums derived from the same source audio are positive examples of each other, while the Mel-spectrums from different source audio are negative examples of each other.
[0017] Step 222: After inputting the training set of the same batch in step 211 into the speaker encoder, a set of corresponding feature representations of the Mel-spectrum segments are obtained.
[0018] Step 223: According to the contrastive learning method, use the contrastive loss to calculate the distance between each feature representation vector in the output representation matrix and calculate the loss.
[0019] As a further improvement of the present invention, in step 3, the input Mel spectrogram passes through a fusion module in the generator to fuse the speaker encoding representation with the generated intermediate speaker representation, and through multiple upsampling operations, the feature dimension is upsampled to the dimension of the waveform, and feature processing is performed through convolution. The features are then processed by the MRF block to process the waveform features at different scales. The specific steps are as follows:
[0020] Step 31: The input Mel spectrogram first undergoes preliminary feature processing through the input layer. The input layer includes a one-dimensional convolutional layer to obtain a preliminary waveform feature representation.
[0021] Step 32: The waveform feature representation obtained in step 31 and the speaker encoding representation obtained in step 2 are input into the fusion module for fusion.
[0022] Step 33: The representation obtained after fusing the waveform feature representation and the speaker encoding representation through the fusion module is upsampled through a transposed convolutional layer. The waveform feature representation can reach the dimension of the final waveform through multiple upsamplings.
[0023] Step 34: The upsampled waveform representation is input into the MRF module to process the waveform features at different scales. The specific form of the MRF module is that the waveform representation passes through residual blocks with different dilation rates in parallel to learn the feature paradigms of different waveforms at different scales, so that the synthesized waveform can better recover the features in different frequency bands. The dilation rates used are 1, 3, and 5 respectively.
[0024] As a further improvement of the present invention, steps 32, 33, and 34 are one waveform representation upsampling process. Such an upsampling process needs to be performed 4 times in total during the waveform generation process, and the upsampling rates are 8, 8, 2, and 2 respectively, converting the Mel spectrogram to an audible waveform with a sampling rate of 22050 kHz.
[0025] As a further improvement of the present invention, in step 32, the fusion module uses an attention mechanism to perform feature fusion on the two input representations. The intermediate result of the input waveform feature representation is used as a query along the time-axis sample points, the speaker encoding representation in step 2 is used as the key and value, and fusion is performed using the fusion module. The specific steps are as follows:
[0026] Step 311: The intermediate result of the input waveform feature representation is used as a query, and the speaker encoding representation in step 2 is used as the key and value.
[0027] Step 312: The query, key, and value obtained in step 311 are respectively multiplied by the corresponding transformation matrices to obtain the transformed query, key, and value.
[0028] Step 313: Query the cross product with the key to obtain the correlation as the weight, and obtain the fused feature by performing weighted summation of the normalized weight and the corresponding value.
[0029] Step 314: Add the input waveform representation in Step 311 and the representation in Step 313 to obtain the waveform representation after fusing the speaker features.
[0030] As a further improvement of the present invention, in the said Step 3, it includes using a multi-period discriminator and a generator for adversarial training, and the specific steps are as follows:
[0031] Step 321: Input the generated waveform into the multi-period discriminator, and the multi-period discriminator outputs the output distribution for determining whether it is a real waveform.
[0032] Step 322: Input the corresponding real waveform into the multi-period discriminator, and the multi-period discriminator outputs the output distribution for determining whether it is a real waveform.
[0033] Step 323: The multi-period discriminator includes 5 sub-discriminators. The input layer of each sub-discriminator is transformed according to different given periods. The input waveform is sampled at a fixed period p to obtain p equally long sub-waveforms, which are concatenated to obtain a two-dimensional input matrix and input into the sub-discriminator for discrimination, where the 5 sub-discriminators are sampled and transformed at fixed periods 2, 3, 5, 7, and 11 respectively.
[0034] Step 324: The sub-discriminator of the multi-period discriminator includes five two-dimensional convolutional layers. The input waveform representation obtains a feature map as the output after passing through each convolutional layer for calculating the loss, and a convolutional layer is added at the end to map the output to one dimension to obtain the discrimination result.
[0035] Step 323: Minimize the KL divergence between the distribution obtained in Step 321 and the distribution in Step 322.
[0036] Step 324: Minimize the L1 distance between the generated waveform converted to the Mel spectrogram and the real Mel spectrogram;
[0037] Step 325: Minimize the distance between a series of feature maps obtained during the process of the generated waveform and the real waveform passing through the discriminator.
[0038] As a further improvement of the present invention, in the said Step 3, it also includes using a multi-scale discriminator and a generator for adversarial training, and the specific steps are as follows:
[0039] Step 331: Input the generated waveform into the multi-scale discriminator, and output the output distribution for determining whether it is a real waveform.
[0040] Step 332: Input the corresponding real waveform into the multi-scale discriminator, and output the output distribution for determining whether it is a real waveform.
[0041] Step 333: The multi-scale discriminator contains 4 sub-discriminators. The input of each sub-discriminator is the original input waveform. Each sub-discriminator convolves the waveform through a convolutional layer with a different convolutional kernel size. The different convolutional kernels are used to capture the discrimination of the waveform at different scales. Each sub-discriminator includes 8 convolutional layers. The output feature maps after each convolutional layer are used to calculate the loss. The output of the last convolutional layer is mapped to a one-dimensional discrimination result.
[0042] Step 334: Minimize the KL divergence between the distribution obtained in Step 331 and the distribution in Step 332.
[0043] Step 335: Minimize the L1 distance between the generated waveform converted to a Mel spectrogram and the real Mel spectrogram.
[0044] Step 336: Minimize the distance between a series of feature maps obtained during the process of the generated waveform and the real waveform passing through the discriminator.
[0045] The beneficial effects of the present invention are as follows: (1) In the existing research on general vocoders, there has been no work on adding speaker encoding based on the generative adversarial network. The working method of the vocoder based on the generative adversarial network is currently the optimal solution for balancing synthesis quality and synthesis speed. The work of integrating speaker representation on the vocoder based on the generative adversarial network in the present invention is a supplement to the current work of general vocoders, providing a working method for a zero-shot general vocoder based on contrastive learning and speaker encoding; (2) The present invention introduces speaker representation based on contrastive learning into the general vocoder. Compared with using the speaker ID as an additional input to the model in the past, the speaker encoder based on contrastive learning can better extract the features of the speaker, better improve the generalization of the model, and improve the synthesis quality of the speaker audio that the model has not seen; (3) The present invention introduces a fusion module into the generator. Compared with the traditional method of using splicing for feature fusion, in the upsampling process of the present invention, the attention mechanism is used multiple times to fuse the features of the speaker into the waveform representation, ensuring that the synthesized audio can highly reflect the personal characteristics of the speaker and restoring the timbre of the synthesized audio. Description of the Drawings
[0046] Figure 1 is the flowchart of the working method of the zero-shot general vocoder of the present invention;
[0047] Figure 2 is the schematic diagram of the process of the speaker encoder obtaining the speaker encoding method of the present invention;
[0048] Figure 3 is the schematic diagram of the process of the speaker encoder training to extract speaker features of the present invention;
[0049] Figure 4 It is the flowchart of the method for synthesizing audible waveforms by the generator of the present invention;
[0050] Figure 5 It is the flowchart of the working method of the zero-shot general vocoder. Detailed implementation manners
[0051] The following is a further detailed description of the present invention. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. In addition, it should be noted that for the convenience of description, only parts related to the present invention rather than all structures are shown in the drawings.
[0052] Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the steps as sequential processes, many of the steps can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the steps can be rearranged. The process can be terminated when its operation is completed, but it can also have additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, and so on.
[0053] The present invention discloses a working method of a zero-shot general vocoder based on contrast learning and generative adversarial networks, and its purpose is to achieve a general vocoder that can have generalization and synthesize high-quality audio under the training with a small amount of mel-spectrum audio.
[0054] The present invention discloses a working method of a zero-shot general vocoder based on contrast learning and generative adversarial networks. The input of the model is the mel-spectrum representation extracted from the audio, and in actual use, it is the synthesized mel-spectrum. After the mel-spectrum representation is input, the model first extracts the speaker features contained in the input mel-spectrum through a pre-trained speaker representation extraction module, and then the mel-spectrum is upsampled by the generator to synthesize speech. At the same time, through the embedding module, the extracted speaker features are embedded into the speech representation in multiple layers of the speech synthesis module during the audio synthesis process of the model, so as to guide the model to synthesize high-quality audio with a timbre matching the target timbre.
[0055] As Figure 1 shown, the present invention discloses a working method of a zero-shot general vocoder based on contrast learning and generative adversarial networks, including the following steps:
[0056] Step 1, input the target synthesized mel-spectrum into the model and perform a logarithmic transformation on the values.
[0057] In the preferred implementation manner, the input mel-spectrum is scaled, specifically including: normalizing the values of the input mel-spectrum; inputting the normalized mel-spectrum into the model.
[0058] Step 2: Input the normalized Mel spectrogram into the speaker encoder to obtain the speaker encoding representation. The speaker encoder encodes the speaker feature information hidden in the Mel spectrogram through an unsupervised method, and uses a residual network trained by a pre-trained contrastive learning method to learn and encode the Mel spectrogram.
[0059] To ensure that the residual network can successfully extract the speaker features in the Mel spectrogram during the unsupervised training process, the contrastive learning method is used during training. By minimizing the distance between the representations of non-overlapping Mel spectrogram segments from the same training audio and maximizing the distance between the representations of non-overlapping Mel spectrogram segments from different training audio, the network extracts the information in the Mel spectrogram that is different from other Mel spectrograms, that is, the information of the speaker in the audio.
[0060] In the preferred embodiment, as Figure 2 shown, the speaker encoder includes a 34-layer residual network for extracting speaker features, and a pooling layer, a linear layer, an activation layer, and an average pooling layer along the time direction for integrating features. The speaker encoder is mainly composed of a convolutional layer and 16 residual blocks.
[0061] The input Mel spectrogram undergoes a feature dimension transformation through a convolutional layer with a kernel size of 3. The speaker features are extracted through 34 residual blocks to obtain two-dimensional speaker features. The obtained two-dimensional speaker features are subjected to average pooling to obtain global features. The obtained global features are mapped through a linear layer to obtain the mapped speaker features. The obtained features are averaged along the time axis to obtain the final 1*256-dimensional speaker representation.
[0062] The speaker encoder utilizes the easy optimization of the residual network and avoids the problems of gradient disappearance and gradient explosion. It can also extract high-level features of speech by stacking deep residual blocks. Specifically:
[0063] When the speaker representation features enter the residual block, they are first activated through a two-dimensional convolutional layer, the corresponding batch normalization, and the activation layer for feature processing. The output features are then activated again through a two-dimensional convolutional layer, batch normalization, and the activation layer. Among the 4th, 8th, and 14th residual blocks in the stacked residual blocks, they will also pass through a convolutional layer with the number of channels doubled for downsampling and avoiding information loss, and then sequentially through batch normalization and the activation layer; in the last layer, the output is added to the input of the residual block and activated through an activation layer to obtain the feature map output by the residual block.
[0064] In the specific implementation process, the steps executed by the speaker encoder are as Figure 2 shown. The specific steps are:
[0065] Step 211: The input Mel spectrogram is subjected to feature extraction through a residual network, where the residual network contains 16 residual blocks for feature extraction and downsampling to obtain two-dimensional speaker features.
[0066] Step 212: Average pooling is performed on the obtained two-dimensional speaker features to obtain global features.
[0067] Step 213: The global features obtained in Step 212 are mapped through a linear layer to obtain the mapped speaker features.
[0068] Step 214: Average pooling is performed on the features obtained in Step 213 along the time axis direction to obtain the final one-dimensional speaker encoding representation.
[0069] Since the speaker encoder is unsupervised feature extraction, in order to ensure that the speaker encoder extracts the target features, the present invention introduces contrastive learning to pre-train the speaker encoder. During the pre-training process, the specific implementation method is as Figure 3 shown, and the specific steps are as follows:
[0070] Step 221: In the training stage, the speaker encoder uses a set of random audio to form a training batch as input. In this training batch, two fixed-length and non-overlapping sub-audio segments are randomly selected from each audio, and the Mel spectrogram corresponding to each audio segment is the input of the Mel spectrogram encoder. Among them, the Mel spectrograms derived from the same source audio are positive examples of each other, while the Mel spectrograms from different source audio are negative examples of each other.
[0071] Step 222: After inputting the training set of the same batch in Step 211 into the speaker encoder, a set of corresponding feature representations of the Mel spectrogram segments are obtained.
[0072] Step 223: According to the method of contrastive learning, it is hoped that the representations obtained in Step 212 that are positive examples of each other are as close as possible, and those that are negative examples of each other are as far away as possible. Therefore, a contrastive loss is used for the output representation matrix to calculate the distance between each feature representation vector and calculate the loss.
[0073] Step 3: The normalized Mel spectrogram and the speaker encoding representation obtained in Step 2 are input into the generator. After multiple upsamplings and convolutions in the adversarial training generation module, finally, the generation module outputs a synthesized audible waveform for the human ear.
[0074] In the preferred embodiment, as Figure 4 shown, the input Mel spectrogram passes through a fusion module in the generator to fuse the speaker encoding and the generated intermediate speaker representation, and through multiple upsampling operations, the feature dimension is upsampled to the dimension of the waveform, and feature processing is performed through convolution. The features are then processed by the MRF block to process the waveform features at different scales. MRF is a multi-receptive field block. The specific steps are as follows:
[0075] Step 31: The input Mel spectrogram first undergoes preliminary feature processing through the input layer. The input layer includes a one-dimensional convolutional layer to obtain a preliminary feature representation of the waveform.
[0076] Step 32: Input the waveform feature representation obtained in Step 31 and the speaker feature representation obtained in Step 2 into the fusion module for fusion.
[0077] In the preferred embodiment, the present invention uses an attention mechanism to perform feature fusion on the two input representations. The waveform feature representation input along the time-axis sample points is used as the query, and the speaker encoding representation in Step 2 is used as the key and value, and is fused through the fusion module. The specific steps are as follows:
[0078] Step 311: Use the input waveform representation, such as the Mel spectrogram or the intermediate result in the generator generation process, as the query, and the speaker encoding representation in Step 2 as the key and value.
[0079] Step 312: Multiply the query, key, and value obtained in Step 311 by their corresponding transformation matrices to obtain the transformed query, key, and value.
[0080] Step 313: The cross product of the query and the key is used as the weight, and the fused feature is obtained by weighted summing the normalized weight and the corresponding value.
[0081] Step 314: Add the input waveform representation in Step 311 to the representation in Step 313 to obtain the waveform representation after fusing the speaker features.
[0082] Step 33: Upsample the representation obtained by fusing the waveform feature representation and the speaker encoding representation through the fusion module using a transposed convolutional layer. The waveform feature representation can reach the dimension of the final waveform through multiple upsamplings.
[0083] Step 34: Input the upsampled waveform representation into the MRF module to process the waveform features at different scales. The specific form of the MRF module is that the waveform representation passes through residual blocks with different dilation rates in parallel to learn the feature patterns of different waveforms at different scales, so that the synthesized waveform can better recover the features in different frequency bands. The dilation rates used are 1, 3, and 5 respectively;
[0084] Steps 32, 33, and 34 are one upsampling process of the waveform representation. Such an upsampling process needs to be performed 4 times in the waveform generation process, and the upsampling rates are 8, 8, 2, and 2 respectively, converting the Mel spectrogram to an audible waveform with a sampling rate of 22050 kHz for the human ear.
[0085] In the preferred embodiment, in order to ensure that the output waveform synthesized by the generator is as close as possible to the real corresponding waveform, an adversarial training method is used to enable the generation module to synthesize the target audio corresponding to the input Mel spectrogram, including using a multi-scale discriminator and a multi-period discriminator. The purpose of the multi-period discriminator is to discriminate the features of the input at different periodic sample points, such as Figure 5 as shown, specifically:
[0086] Step 321: Input the generated waveform in Step 3 into the multi-period discriminator, and output the output distribution for discriminating whether it is a real waveform.
[0087] Step 322: Input the corresponding real waveform into the multi-period discriminator, and output the output distribution for discriminating whether it is a real waveform.
[0088] Step 323: The multi-period discriminator contains 5 sub-discriminators. The input layer of each sub-discriminator is transformed according to different given periods. The input waveform is sampled at a fixed period p to obtain p equally long sub-waveforms, which are concatenated to obtain a two-dimensional input matrix and input into the sub-discriminator for discrimination. Among them, the 5 sub-discriminators are sampled and transformed at fixed periods of 2, 3, 5, 7, and 11 respectively.
[0089] Step 324: The sub-discriminator of the multi-period discriminator includes five two-dimensional convolutional layers. The input waveform representation obtains a feature map as the output after passing through each convolutional layer for calculating the loss, and a convolutional layer is added at the end to map the output to one dimension to obtain the discrimination result.
[0090] Step 323: Minimize the KL divergence between the distribution obtained in Step 321 and the distribution in Step 322. The KL divergence, that is, relative entropy, is used to measure the degree of difference between distributions.
[0091] Step 324: Minimize the L1 distance, that is, the Manhattan distance, between the generated waveform converted into a Mel spectrogram and the real Mel spectrogram.
[0092] Step 325: Minimize the distance between a series of feature maps obtained during the process of the generated waveform and the real waveform passing through the discriminator.
[0093] The purpose of the multi-scale discriminator is to supplement the defect that the multi-period discriminator does not perform discrimination on continuous waveforms and to discriminate the feature patterns existing in the input at different scales, such as Figure 5 as shown, specifically:
[0094] Step 331: Input the generated waveform in Step 3 into the multi-scale discriminator, and output the output distribution for discriminating whether it is a real waveform.
[0095] Step 332: Input the corresponding real waveform into the multi-scale discriminator, and output the output distribution for discriminating whether it is a real waveform.
[0096] Step 333: The multi-scale discriminator includes 4 sub-discriminators. The input of each sub-discriminator is the original input waveform. Each sub-discriminator convolves the waveform through a convolutional layer with a different convolutional kernel size. The different convolutional kernels are used to capture the discrimination of the waveform at different scales. Each sub-discriminator includes 8 convolutional layers. The output feature map after each convolutional layer is used to calculate the loss. The output of the last convolutional layer is mapped to a one-dimensional discrimination result.
[0097] Step 334: Minimize the KL divergence between the distribution obtained in Step 331 and the distribution in Step 332.
[0098] Step 335: Minimize the L1 distance between the generated waveform converted to a Mel spectrogram and the real Mel spectrogram.
[0099] Step 336: Minimize the distance between a series of feature maps obtained during the process of the discriminator for the generated waveform and the real waveform.
[0100] The beneficial effects of the present invention are as follows: (1) In the existing research on general vocoders, there has been no work on adding speaker encoding based on the generative adversarial network. The vocoder method based on the generative adversarial network is currently the optimal solution for balancing synthesis quality and synthesis speed. The work of the present invention on fusing speaker representation on the vocoder based on the generative adversarial network is a supplement to the current work on general vocoders, providing a working method for a zero-shot general vocoder based on contrastive learning and speaker encoding; (2) The present invention introduces a speaker representation based on contrastive learning into the general vocoder. Compared with using the speaker ID as an additional input to the model in the past, the speaker encoder based on contrastive learning can better extract the features of the speaker, better improve the generalization of the model, and improve the synthesis quality of the speaker audio that the model has not seen; (3) The present invention introduces a fusion module into the generator. Compared with the traditional method of using splicing for feature fusion, in the upsampling process of the present invention, the attention mechanism is used multiple times to fuse the features of the speaker into the waveform representation, ensuring that the synthesized audio can highly reflect the personal characteristics of the speaker and restoring the timbre of the synthesized audio.
[0101] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention belongs, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should all be regarded as belonging to the protection scope of the present invention.
Claims
1. A working method of a zero-shot general vocoder based on contrastive learning and generative adversarial network, characterized in that, Including the following steps: Step 1, input the Mel spectrogram to be synthesized into the model and perform a logarithmic transformation on the values; Step 2, input the input Mel spectrogram into the speaker encoder to obtain the speaker encoding representation; Step 3, input the Mel spectrogram input in Step 1 and the speaker encoding representation obtained in Step 2 into the generator. After multiple upsamplings and convolutions in the adversarially trained generation module, finally the generation module outputs the synthesized audible waveform for the human ear; In Step 2, the speaker encoder encodes the speaker feature information hidden in the Mel spectrogram through an unsupervised method, and uses a residual network trained by a pre-trained contrastive learning method to learn and encode the Mel spectrogram; In Step 2, a contrastive learning method is introduced to pre-train the speaker encoder. The specific steps are as follows: Step 221: In the training stage, the speaker encoder uses a set of random audio to form a training batch as the input. In this training batch, two non-overlapping sub-audio segments of fixed length are randomly selected for each audio, and the Mel spectrogram corresponding to each audio segment is the input of the Mel spectrogram encoder. Among them, the Mel spectrograms derived from the same source audio are positive examples of each other, while the Mel spectrograms from different source audio are negative examples of each other; Step 222: After inputting the training set of the same batch in Step 211 into the speaker encoder, a set of corresponding feature representations of the Mel spectrogram segments are obtained; Step 223: According to the contrastive learning method, use the contrastive loss to calculate the distance between each feature representation vector of the output representation matrix and calculate the loss.
2. The working method of the zero-shot general vocoder according to claim 1, characterized in that, In Step 1, the input Mel spectrogram values need to be normalized, and then the normalized Mel spectrogram is input into the model.
3. The working method of the zero-shot general vocoder according to claim 1, characterized in that In Step 2, the speaker encoder includes a 34-layer residual network for extracting speaker features, a pooling layer for integrating speaker features, a linear layer, an activation layer, and an average pooling layer along the time direction for aggregating the features of each frame. Specifically, it also includes the following steps: Step 211: The input Mel spectrogram undergoes feature extraction through the residual network. The residual network contains 16 residual blocks for feature extraction and downsampling to obtain two-dimensional speaker features; Step 212: Perform average pooling on the obtained two-dimensional speaker features to obtain global features; Step 213: Map the global features obtained in Step 212 through the linear layer to obtain the mapped speaker features; Step 214: Perform average pooling on the speaker features obtained in Step 213 along the time axis direction to obtain the final one-dimensional speaker encoding representation.
4. The working method of the zero-shot universal vocoder according to claim 1, characterized in that, In Step 3, the input Mel spectrogram in the generator fuses the speaker encoding representation with the generated intermediate speaker representation through the fusion module, and undergoes multiple upsampling operations to upsample the feature dimension to the dimension of the waveform, and performs feature processing through convolution. The features are then processed by the MRF block to process the waveform features at different scales. The specific steps are as follows: Step 31: The input Mel spectrogram first undergoes preliminary feature processing through the input layer. The input layer includes a one-dimensional convolutional layer to obtain a preliminary feature representation of the waveform; Step 32: Input the waveform feature representation obtained in Step 31 and the speaker encoding representation obtained in Step 2 into the fusion module for fusion; Step 33: Upsample the representation obtained by fusing the waveform feature representation and the speaker encoding representation through the fusion module using a transposed convolutional layer. The waveform feature representation can reach the dimension of the final waveform through multiple upsamplings; Step 34: Input the upsampled waveform representation into the MRF module to process waveform features at different scales. The specific form of the MRF module is that the waveform representation passes through residual blocks with different dilation rates in parallel to learn the feature patterns of different waveforms at different scales, so that the synthesized waveform can better recover the features in different frequency bands. The dilation rates used are 1, 3, and 5 respectively.
5. The working method of the zero-shot general vocoder according to claim 4, characterized in that, The above Steps 32, 33, and 34 are one upsampling process of the waveform representation. Such upsampling processes are required to be performed 4 times in total during the waveform generation process, and the upsampling rates are 8, 8, 2, and 2 respectively, to convert the Mel spectrogram into an audible waveform with a sampling rate of 22050 kHz.
6. The working method of the zero-shot universal vocoder according to claim 4, characterized in that, In Step 32, the fusion module uses an attention mechanism to fuse the features of the two input representations. The sample points of the input waveform feature representation along the time axis are used as queries, and the speaker encoding representation in Step 2 is used as keys and values, and fused using the fusion module. The specific steps are as follows: Step 311: Use the intermediate result of the input waveform feature representation as the query, and the speaker encoding representation in Step 2 as the key and value; Step 312: Multiply the query, key, and value obtained in Step 311 by their corresponding transformation matrices to obtain the transformed query, key, and value; Step 313: Multiply the query and the key to obtain the correlation as the weight, and obtain the fused feature by weighted summing the normalized weight and the corresponding value; Step 314: Add the input waveform representation in Step 311 and the representation in Step 313 to obtain the waveform representation after fusing the speaker features.
7. The working method of the zero-shot general vocoder according to claim 4, characterized in that, In Step 3, it includes adversarial training using a multi-period discriminator and a generator. The specific steps are as follows: Step 321: Input the generated waveform into the multi-period discriminator, and the multi-period discriminator outputs the output distribution for determining whether it is a real waveform; Step 322: Input the corresponding real waveform into the multi-period discriminator, and the multi-period discriminator outputs the output distribution for determining whether it is a real waveform; Step 323: The multi-period discriminator contains 5 sub-discriminators. The input layer of each sub-discriminator is transformed according to different given periods. The input waveform is sampled at a fixed period p to obtain p equally long sub-waveforms, which are concatenated to obtain a two-dimensional input matrix and input into the sub-discriminator for discrimination. Among them, the 5 sub-discriminators are sampled and transformed at fixed periods of 2, 3, 5, 7, and 11 respectively; Step 324: The sub-discriminator of the multi-period discriminator includes five two-dimensional convolutional layers. The input waveform representation obtains a feature map as the output after passing through each convolutional layer for calculating the loss, and a convolutional layer is added at the end to map the output to one dimension to obtain the discrimination result; Step 323: Minimize the KL divergence between the distribution obtained in Step 321 and the distribution in Step 322; Step 324: Minimize the L1 distance between the generated waveform after being converted into a Mel spectrogram and the real Mel spectrogram; Step 325: Minimize the distance between a series of feature maps obtained during the discriminator process for the generated waveform and the real waveform.
8. The working method of the zero-shot general vocoder according to claim 7, characterized in that, In step 3, it also includes adversarial training of the multi-scale discriminator and the generator, and the specific steps are as follows: Step 331: Input the generated waveform into the multi-scale discriminator to output the output distribution for discriminating whether it is a real waveform; Step 332: Input the corresponding real waveform into the multi-scale discriminator to output the output distribution for discriminating whether it is a real waveform; Step 333: The multi-scale discriminator contains 4 sub-discriminators. The input of each sub-discriminator is the original input waveform. Each sub-discriminator convolves the waveform through a convolutional layer with a different convolutional kernel size. The different convolutional kernels are used to capture the discrimination of the waveform at different scales; each sub-discriminator includes 8 convolutional layers. The feature maps output after each convolutional layer are used to calculate the loss, and the output of the last convolutional layer is mapped to a one-dimensional discrimination result; Step 334: Minimize the KL divergence between the distribution obtained in step 331 and the distribution in step 332; Step 335: Minimize the L1 distance between the generated waveform after being converted into a Mel spectrogram and the real Mel spectrogram; Step 336: Minimize the distance between a series of feature maps obtained during the discriminator process for the generated waveform and the real waveform.
Citation Information
Patent Citations
Cross-language timbre conversion system and method based on zero-order learning
CN112767958A
Generating speech signals using both neural network-based vocoding and generative adversarial training
US20210366461A1