Zero-Shot Voice Cloning Method and Device Based on Audio Decoupling and Fusion
By using audio decoupling and fusion technology in zero-sample voice cloning, the timbre information of the target speaker is extracted and fused, the problem of insufficient timbre details of the speaker in speech synthesis is solved, and the similarity of the speaker in speech is improved.
Patent Information
- Application Number
- CN202211012716.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-23
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-08-23
AI Technical Summary
Zero-sample voice clones have biases in obtaining the details of the target speaker's tone, resulting in bias in the speaker's similarity.
Using an audio decoupling and fusion method, the audio tone information and content information are extracted respectively through the audio tone encoder and the audio content encoder, and the decoupling capability is improved through mutual information constraints. At the same time, multiple reference audios are used to fusion of tone information to generate speaker embeddings containing the target speaker's timbre information.
It improves the accuracy and detail of the speaker's tone in speech synthesis, and reduces the deviation of synthetic speech in speaker similarity.
Smart Images

Figure CN115497449B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of speech synthesis technology, and in particular to a zero-sample speech cloning method and device based on audio decoupling and fusion. Background Art
[0002] With the development of human-computer voice interaction technology, the application scope of speech synthesis is becoming wider and wider. For example, common voice assistants, smart speakers, map navigation, etc. in life, as well as audio books, AI anchors, singing synthesis and other applications that have gradually developed in recent years have gradually penetrated into people's lives. Speech synthesis aims to synthesize high-quality speech for a given text. Among them, the research goal of small-sample speech synthesis is to learn the characteristics of the speaker's voice and perform speech synthesis with only a small amount of speech data. Zero-sample speech cloning is a type of small-sample speech synthesis. This direction aims to explore small-sample speech synthesis in extreme cases. That is, after obtaining a few audios of the target speaker, the voice is cloned quickly and accurately without training the model. However, due to the rich expressiveness of human natural speech and the great changes in the speaker's timbre and rhythm, modeling is very difficult. Therefore, zero-sample speech cloning is a very challenging task. Summary of the invention
[0003] In view of the above problems, the present invention provides a zero-sample speech cloning method and device based on audio decoupling and fusion. Zero-sample speech synthesis only needs to provide a few audios spoken by the target speaker (called reference audio). The speaker encoding part can extract the speaker embedding related to the speaker's timbre from the audio, add the extracted speaker embedding to the speech synthesis module, and synthesize the speech of the target speaker's timbre. In order to better obtain the speaker embedding from the reference audio, the present invention respectively designs an audio timbre encoder and an audio content encoder to extract the content information and timbre information in the reference audio. At the same time, in order to improve the decoupling capability, the mutual information constraint is introduced to make the coupling degree between the extracted content embedding and the timbre embedding as low as possible. In zero-sample speech synthesis, a single audio of the target speaker is usually used to obtain the speaker representation. However, a single audio cannot cover all the details of the target speaker's timbre, and the timbre information contained therein is often limited, which makes the synthesized speech have deviations in speaker similarity. In order to obtain sufficient timbre information, the present invention uses multiple reference audios of the target speaker as input, extracts timbre information from each reference audio, and then fuses the multiple timbre information through a timbre information fusion module, so that the final speaker embedding contains the timbre information of the target speaker as much as possible.
[0004] The first aspect of the present invention is a zero-sample voice cloning method based on audio decoupling and fusion, comprising the following steps:
[0005] Extract the acoustic features of the audio of the target speaker to obtain the acoustic feature Mel spectrogram;
[0006] Use the voice and content decoupling module to separate the voice information and content information of the acoustic feature Mel spectrogram. The voice and content decoupling module includes an audio content encoder and an audio voice encoder. The audio content encoder is used to extract the content information related to the language content in the acoustic feature Mel spectrogram, and the audio voice encoder is used to extract the voice information related to the speaker identity in the acoustic feature Mel spectrogram;
[0007] Fuse the voice information set to obtain the speaker voice embedding representation;
[0008] Input the speaker voice embedding representation and the text into the zero-shot voice cloning model to synthesize the Mel spectrogram corresponding to the text and with the voice of the target speaker. The zero-shot voice cloning model includes an encoder, a variational adaptor, and a decoder. The encoder is used to extract the hidden state related to the text, the variational adaptor is used to predict the fundamental frequency and the duration of the input text features for the input features, and the decoder is used to synthesize the Mel spectrogram from the input features;
[0009] Input the Mel spectrogram with the voice of the target speaker into the vocoder to convert the Mel spectrogram with the voice of the target speaker into an audible waveform signal for the human ear.
[0010] Further, extract the acoustic features of the audio of the target speaker to obtain the audio feature Mel spectrogram, specifically including:
[0011] Perform pre-emphasis, framing, and windowing processing on the input audio signal in sequence to obtain the preprocessed audio signal;
[0012] Perform short-time Fourier transform and Mel spectrogram transform on the preprocessed audio signal in sequence to obtain the acoustic feature Mel spectrogram.
[0013] Further, the audio content encoder includes an encoding module, a vector quantization layer, and a contrast prediction module. The encoding module includes two consecutive convolutional layers, an activation layer, and an instance regularization layer. The encoding module is used to generate a continuous feature vector from the input acoustic feature Mel spectrogram; the vector quantization layer consists of K different codes to form a trainable codebook E, and is used to map each continuous feature vector to one of the K codebooks through nearest neighbor search to form a discrete feature vector; the contrast prediction module is an autoregressive structure, and is used to generate a context representation using the discrete feature vectors at the previous t time moments, and predict the discrete feature vectors at the subsequent n time moments using the context representation.
[0014] Furthermore, the audio timbre encoder includes a convolutional layer, multiple convolutional banks, and an average pooling layer. The input acoustic feature Mel spectrogram undergoes a convolutional layer with a convolutional kernel size of 8 for feature dimension conversion to obtain a dimension-converted feature vector. The dimension-converted feature vector passes through 6 convolutional banks to expand the receptive field. The dimension-converted feature vector after expanding the receptive field undergoes an average pooling operation to obtain a global feature vector. The global feature vector passes through 2 convolutional banks for feature conversion to obtain the speaker's timbre information.
[0015] Furthermore, the decoupling degree of the timbre information and the content information is improved by reducing the mutual information between the timbre information and the content information separated by the timbre and content decoupling module, specifically including: taking the timbre information and the content information separated by the timbre and content decoupling module as two constraint variables to be constrained, and adding an unbiased estimation loss of the upper bound of the mutual information during the training process of the entire method model.
[0016] Furthermore, the information fusion module is used to fuse the timbre information set to obtain the speaker timbre embedding representation. The information fusion module includes a multi-head attention mechanism module and a timbre library. The timbre library uses the basis vectors of the speaker timbre as the values and keys of the attention mechanism. The multi-head attention mechanism module uses the timbre information as the query, the speaker timbre library as the values and keys, and takes the average of the similarity scores between each timbre vector and the timbre library to obtain the speaker timbre embedding representation.
[0017] Furthermore, the speaker timbre embedding representation and the text are input into the zero-shot voice cloning model to synthesize the Mel spectrogram corresponding to the text and with the target speaker's timbre, specifically including:
[0018] Input the text into the encoder to extract the text-related hidden state;
[0019] After expanding the speaker timbre embedding representation to the same length as the text, add it to the text hidden state to obtain a fused feature;
[0020] Input the fused feature into the variational adaptor to predict the pitch, energy, and duration information of the text features of the fused feature;
[0021] Expand the fused feature to the frame-level length according to the duration information;
[0022] Input the expanded fused feature into the decoder to synthesize the Mel spectrogram with the target speaker's timbre.
[0023] In the second aspect of the present invention, a zero-shot voice cloning device based on audio decoupling and fusion is provided, including:
[0024] An acoustic feature Mel spectrogram acquisition module, configured to perform acoustic feature extraction on the audio of the target speaker to obtain an acoustic feature Mel spectrogram;
[0025] A timbre and content decoupling module, which is used to separate the timbre information and content information of the acoustic feature Mel spectrogram. The timbre and content decoupling module includes an audio content encoder and an audio timbre encoder. The audio content encoder is used to extract the content information related to the language content in the acoustic feature Mel spectrogram, and the audio timbre encoder is used to extract the timbre information related to the speaker identity in the acoustic feature Mel spectrogram;
[0026] A timbre information fusion module, which is used to perform feature fusion on the timbre information set to obtain a speaker timbre embedding representation;
[0027] A timbre Mel spectrogram acquisition module, which is used to input the speaker timbre embedding representation and text into a zero-shot voice cloning model to synthesize a Mel spectrogram corresponding to the text and with the target speaker's timbre. The zero-shot voice cloning model includes an encoder, a variational adaptor, and a decoder. The encoder is used to extract the hidden state related to the text, the variational adaptor is used to predict the fundamental frequency and the duration of the input text features for the input features, and the decoder is used to synthesize the Mel spectrogram from the input features;
[0028] A Mel spectrogram conversion module, which is used to input the Mel spectrogram with the target speaker's timbre into a vocoder to convert the Mel spectrogram with the target speaker's timbre into an audible waveform signal for the human ear.
[0029] In a third aspect of the present invention, there is provided a zero-shot voice cloning device based on audio decoupling and fusion, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above-mentioned zero-shot voice cloning method based on audio decoupling and fusion.
[0030] In a fourth aspect of the present invention, there is provided a computer-readable storage medium, on which instructions are stored, and when the instructions are executed by a processor, the processor is made to execute the above-mentioned zero-shot voice cloning method based on audio decoupling and fusion.
[0031] A zero-shot voice cloning method and device based on audio decoupling and fusion provided by the present invention have the following specific beneficial effects:
[0032] (1) Existing zero-shot voice synthesis usually does not constrain the extracted speaker embeddings and cannot obtain pure speaker information. The zero-shot voice cloning method provided by the present invention is improved on an open-source voice synthesis basic model, and provides a zero-shot voice cloning method based on mutual information audio decoupling and multi-reference audio fusion;
[0033] (2) The key to zero-shot voice cloning lies in how to quickly and accurately extract the voiceprint embedding from the audio provided by the target speaker, and use this voiceprint embedding to guide the speech synthesis module to generate speech with the corresponding voiceprint. The present invention improves the decoupling degree of the voiceprint information and the content information by reducing the mutual information between the voiceprint information and the content information separated by the voiceprint and content decoupling module, avoiding the leakage of the content information in the audio into the voiceprint representation, which may lead to a decline in the quality of the synthesized speech;
[0034] (3) The present invention introduces a voiceprint information fusion module. Compared with obtaining the speaker information from only a single audio, the present invention can input several audios, and the voiceprint information fusion module can fuse the voiceprint information to ensure that the voiceprint embedding contains as much speaker voiceprint information as possible. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 is the flowchart of the zero-shot voice cloning method based on audio decoupling and fusion in the embodiment of the present invention;
[0036] Figure 2 is the schematic flowchart of the method for obtaining the acoustic feature mel spectrogram in the embodiment of the present invention;
[0037] Figure 3 is the schematic flowchart of the method of the voiceprint and content decoupling module in the embodiment of the present invention;
[0038] Figure 4 is the schematic flowchart of the zero-shot voice cloning model in the embodiment of the present invention;
[0039] Figure 5 is the schematic structural diagram of the zero-shot voice cloning device based on audio decoupling and fusion in the embodiment of the present invention;
[0040] Figure 6 is the architecture diagram of the computer device in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] The present invention will be further described in detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. In addition, it should be noted that only the parts related to the present invention are shown in the drawings for the convenience of description, rather than all the structures.
[0042] Before discussing the exemplary embodiments in more detail, it should be noted that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts depict the steps as sequential processes, many of the steps can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the steps can be rearranged. The process can be terminated when its operations are completed, but it can also have additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, and so on.
[0043] Embodiments of the present invention are directed to a zero-shot voice cloning method and apparatus based on audio decoupling and fusion, and provide the following embodiments:
[0044] Based on Embodiment 1 of the present invention
[0045] As Figure 1 shown, it is a flowchart of the zero-shot voice cloning method based on audio decoupling and fusion according to Embodiment 1 of the present invention, and the specific steps are as follows:
[0046] Step 1: Extract acoustic features from the audio of the target speaker to obtain an acoustic feature Mel spectrogram.
[0047] In a preferred embodiment, extracting acoustic features from the audio of the target speaker to obtain an audio feature Mel spectrogram specifically includes: sequentially performing pre-emphasis, framing, and windowing on the input audio signal to obtain a preprocessed audio signal; sequentially performing short-time Fourier transform and Mel spectrogram transform on the preprocessed audio signal to obtain an acoustic feature Mel spectrogram.
[0048] Specifically, as Figure 2 shown, converting several audio (≤5) into Mel spectrograms by first performing feature conversion, and the specific steps include:
[0049] Step 11: Perform pre-emphasis on the input audio to compensate for the high-frequency part of the speech signal.
[0050] Step 12: Perform framing on the speech signal.
[0051] Step 13: Window the framed speech signal to avoid spectral leakage.
[0052] Step 14: Perform short-time Fourier transform and Mel spectrogram transform on the audio signal obtained in Step 13 to convert the waveform signal into an acoustic feature Mel spectrogram.
[0053] Step 2: Use the timbre and content decoupling module to separate the timbre information and content information of the acoustic feature Mel spectrogram. The timbre and content decoupling module includes an audio content encoder and an audio timbre encoder. The audio content encoder is used to extract the content information related to the language content in the acoustic feature Mel spectrogram, and the audio timbre encoder is used to extract the timbre information related to the speaker identity in the acoustic feature Mel spectrogram;
[0054] In the preferred embodiment, as Figure 3 shown, the audio content encoder includes an encoding module, a vector quantization layer, and a contrast prediction module. The encoding module includes two consecutive convolutional layers, an activation layer, and an instance regularization layer. The encoding module is used to generate a continuous feature vector from the input acoustic feature Mel spectrogram. The vector quantization layer consists of K different codes to form a trainable codebook E, and is used to map each continuous feature vector to one of the K codebooks through nearest neighbor search to form a discrete feature vector. The contrast prediction module is an autoregressive structure, and is used to generate a context representation using the discrete feature vectors at the previous t time steps, and predict the discrete feature vectors at the subsequent n time steps using the context representation. The audio timbre encoder includes a convolutional layer, a multi-layer convolutional bank, and an average pooling layer. The input acoustic feature Mel spectrogram is subjected to a convolutional layer with a convolutional kernel size of 8 for feature dimension conversion to obtain a dimension conversion feature vector. The dimension conversion feature vector is passed through a 6-layer convolutional bank to expand the receptive field. The dimension conversion feature vector after expanding the receptive field is subjected to an average pooling operation to obtain a global feature vector. The global feature vector is passed through a 2-layer convolutional bank for feature conversion to obtain the timbre information of the speaker.
[0055] In the specific implementation process, the steps executed by the audio content encoder (Content Encoder) are as Figure 4 shown, which are:
[0056] Step 211: The input reference Mel spectrogram first passes through two (Convolutional Layer + Rectified Linear Unit (RELU) + Instance Normalization) modules to obtain a continuous feature vector.
[0057] Step 212: Input the continuous feature vector extracted in Step 211 into the vector quantization layer (VectorQuantization). The vector quantization layer consists of a trainable codebook E composed of K different codes. During the training process, the vector quantization layer maps each continuous feature vector to one of the K codebooks through nearest neighbor search;
[0058] Step 213: To ensure that the discrete feature vectors extracted in Step 212 contain content information, a contrast prediction module is added after the vector quantization layer. This contrast prediction module has an autoregressive structure that uses the discrete feature vectors at the previous t time steps to generate a context representation c, and uses this context representation to predict the discrete feature vectors at the subsequent n time steps.
[0059] In the specific implementation process, the audio timbre encoder is an extensible and accurate module for speaker feature extraction, and the execution steps are as follows:
[0060] Step 221: The input reference Mel spectrogram passes through a convolutional layer with a convolutional kernel size of 8 for feature dimension conversion;
[0061] Step 222: The features in Step 221 pass through a convolutional bank with 6 layers to expand the receptive field;
[0062] Step 223: Perform average pooling operation on the features in Step 222 to extract global features.
[0063] Step 224: Pass the global features extracted in Step 223 through a convolutional bank with two layers for further feature conversion to obtain speaker information.
[0064] Preferably, the decoupling degree of the timbre information and the content information is improved by reducing the mutual information between the timbre information and the content information separated by the timbre and content decoupling module, specifically including: taking the timbre information and the content information separated by the timbre and content decoupling module as two constraint variables to be constrained, and adding an unbiased estimation loss of the upper bound of the mutual information during the training process of the entire method model.
[0065] In the specific implementation process, the audio content encoder and the audio timbre encoder are unsupervised feature extractions. To ensure that the encoded vectors extract corresponding features, the present invention introduces mutual information to measure the dependence relationship between the two. And during training, the decoupling degree of the content information and the timbre information is improved by reducing the mutual information between the two. Since it is very difficult to directly estimate the mutual information between two random variables in actual situations, it is only possible to approximately obtain the mutual information amount by estimating its upper bound or lower bound. The present invention uses the method of estimating the upper bound of vCLUB to estimate the upper bound of the mutual information, specifically as follows:
[0066] Step 231: Take the content information and the timbre information as two constraint variables to be constrained, and add an unbiased estimation loss of the upper bound of the mutual information during the model training process. This unbiased estimation loss estimates the mutual information between the content information and the timbre information. Denote the content embedding as z, the timbre embedding as s, N represents the number of input reference audios, T represents the number of frames of each reference audio, and q is the mutual information estimation network. The unbiased estimation of vCLUB between the content embedding z and the timbre embedding s can be written as:
[0067]
[0068] Step 3: Perform feature fusion on the timbre information set to obtain the speaker timbre embedding representation;
[0069] Preferably, an information fusion module is used to perform feature fusion on the timbre information set to obtain the speaker timbre embedding representation. The information fusion module includes a multi-head attention mechanism module and a timbre library. The timbre library uses the base vectors of the speaker timbres as the values and keys of the attention mechanism. The multi-head attention mechanism module uses the timbre information as the query, and the speaker timbre library as the values and keys, and takes the average of the similarity scores between each timbre vector and the timbre library to obtain the speaker timbre embedding representation.
[0070] In the specific implementation process, the number of several timbre information extracted in step 2 is determined by the input. The multi-timbre information fusion module is composed of a multi-head attention mechanism and a timbre library. Specifically:
[0071] Step 31: The present invention maintains a speaker timbre library P, which learns the base vectors of the timbres. This timbre library serves as the values and keys of the attention mechanism.
[0072] Step 32: In order to learn different aspect attributes of the timbre, a multi-head attention mechanism is used, which can process different attributes on different attention heads. Use the several timbre information S in step 2 as the query, the speaker timbre library P as the values V and keys K, W q 、W k 、W v are all transformation matrices, and d m is the feature dimension. The calculation process of the multi-head attention mechanism is as follows:
[0073] Q = SW q , K = PW k , V = PW v
[0074]
[0075] Step 4: Input the speaker timbre embedding representation and the text into the zero-shot voice cloning model to synthesize the Mel spectrogram corresponding to the text with the target speaker's timbre. The zero-shot voice cloning model includes an encoder, a variational adapter, and a decoder. The encoder is used to extract the hidden states related to the text. The variational adapter is used to predict the fundamental frequency and the duration of the input text features for the input features. The decoder is used to synthesize the Mel spectrogram from the input features;
[0076] Preferably, the speaker voice embedding representation and the text are input into a zero-shot voice cloning model to synthesize a Mel spectrogram corresponding to the text with the target speaker voice. Specifically, it includes: inputting the text into an encoder to extract text-related hidden states; extending the speaker voice embedding representation to the same length as the text and adding it to the text hidden states to obtain a fused feature; inputting the fused feature into a variational adaptor to predict the pitch, energy, and duration information of the text features of the fused feature; extending the fused feature to the frame-level length according to the duration information; inputting the extended fused feature into a decoder to synthesize a Mel spectrogram with the target speaker voice.
[0077] In the specific implementation process, since the final speaker embedding representation is extracted in step 3, this speaker embedding representation is input into the zero-shot voice cloning model as a condition to guide the zero-shot voice cloning model to synthesize speech similar to the target speaker. The input of the zero-shot voice cloning model is the text and the speaker embedding representation in the dataset, and the output is the intermediate acoustic feature: Mel spectrogram. As Figure 4 shown, the zero-shot voice cloning model mainly consists of an encoder (Encoder), a variational adaptor (variance Adaptor), and a decoder (Decoder). Specifically:
[0078] Step 41: Input the text into the encoder to extract text-related hidden states; as Figure 4 shown, the encoder contains four encoding blocks. However, no speaker embedding information is added at the Encoder end
[0079] Step 42: Extend the speaker embedding representation to the same length as the text and add it to the text hidden states to obtain the fused feature;
[0080] Step 43: Input the fused feature obtained in step 42 into the variational adaptor, which is used to predict the fundamental frequency and the duration of each text feature. After obtaining the duration information, extend the fused feature to the frame-level length.
[0081] Step 44: Input the feature extended in step 43 into the decoder to synthesize the corresponding Mel spectrogram. The decoder consists of 6 layers of encoding blocks. In this encoding block, the speaker embedding is added to the regularization module.
[0082] Step 5: Input the Mel spectrogram with the target speaker voice into a vocoder to convert the Mel spectrogram with the target speaker voice into a waveform signal audible to the human ear.
[0083] Specifically, the Mel spectrogram obtained in step 4 cannot be directly perceived by the human ear. Therefore, a vocoder is required to convert the Mel spectrogram into a waveform signal audible to the human ear. The vocoder used in the embodiment is the open-source model HiFiGAN.
[0084] Based on Embodiment 2 of the present invention
[0085] A zero-shot voice cloning device 500 based on audio decoupling and fusion provided by Embodiment 2 of the present invention can execute the zero-shot voice cloning method based on audio decoupling and fusion provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method. The device can be implemented in the form of software and / or hardware (integrated circuit), and is generally integrated in a server or a terminal device.
[0086] Figure 5 It is a schematic structural diagram of a zero-shot voice cloning device 500 based on audio decoupling and fusion in Embodiment 2 of the present invention. Refer to Figure 5 , the zero-shot voice cloning device 500 based on audio decoupling and fusion in the embodiment of the present invention may specifically include:
[0087] An acoustic feature Mel spectrum acquisition module 510, configured to extract acoustic features from the audio of the target speaker to obtain an acoustic feature Mel spectrum;
[0088] A timbre and content decoupling module 520, configured to separate the timbre information and content information of the acoustic feature Mel spectrum, where the timbre and content decoupling module includes an audio content encoder and an audio timbre encoder. The audio content encoder is configured to extract content information related to the language content in the acoustic feature Mel spectrum, and the audio timbre encoder is configured to extract timbre information related to the speaker identity in the acoustic feature Mel spectrum;
[0089] A timbre information fusion module 530, configured to perform feature fusion on the timbre information set to obtain a speaker timbre embedding representation;
[0090] A timbre Mel spectrum acquisition module 540, configured to input the speaker timbre embedding representation and text into a zero-shot voice cloning model to synthesize a Mel spectrum corresponding to the text and with the timbre of the target speaker, where the zero-shot voice cloning model includes an encoder, a variational adaptor, and a decoder. The encoder is configured to extract text-related hidden states, the variational adaptor is configured to predict the fundamental frequency and the duration of the input text features for the input features, and the decoder is configured to synthesize a Mel spectrum from the input features;
[0091] A Mel spectrum conversion module 550, configured to input the Mel spectrum with the timbre of the target speaker into a vocoder to convert the Mel spectrum with the timbre of the target speaker into a waveform signal audible to the human ear.
[0092] In addition to the above 5 modules, the device 500 may further include other components. However, since these components are not related to the content of the embodiments of the present disclosure, their illustrations and descriptions are omitted here.
[0093] The specific working process of the zero-shot voice cloning device 500 based on audio decoupling and fusion may refer to the description of Embodiment 1 of the zero-shot voice cloning method based on audio decoupling and fusion above, and will not be elaborated here.
[0094] Based on Embodiment 3 of the present invention
[0095] The system according to the embodiments of the present invention can also be implemented with the aid of Figure 6 the architecture of the computing device shown. Figure 6 The architecture of the computing device is shown. As Figure 6 shown, a computer system 610, a system bus 630, one or more CPUs 640, input / output components 620, a memory 650, etc. The memory 650 can store various data or files used for computer processing and / or communication, as well as program instructions executed by the CPU, including the method of Embodiment 1. Figure 6 The architecture shown is only exemplary. When implementing different devices, one or more components in Figure 6 are adjusted according to actual needs.
[0096] Based on Embodiment 4 of the present invention
[0097] The embodiments of the present invention can also be implemented as a computer-readable storage medium. Computer-readable instructions are stored on the computer-readable storage medium according to Embodiment 4. When the computer-readable instructions are run by a processor, the zero-shot voice cloning method based on audio decoupling and fusion according to Embodiment 1 of the present invention described with reference to the above drawings can be executed.
[0098] Based on the zero-shot voice cloning method and device for audio decoupling and fusion provided by the above embodiments, zero-shot voice synthesis only needs to provide several sentences of audio spoken by the target speaker (referred to as reference audio). The speaker encoding part can extract the speaker embedding related to the speaker's timbre from this audio, and add the extracted speaker embedding to the voice synthesis module to synthesize the voice with the timbre of the target speaker. To better obtain the speaker embedding from the reference audio, the present invention separately designs an audio timbre encoder and an audio content encoder to extract the content information and timbre information in the reference audio respectively. At the same time, in order to improve the decoupling ability, mutual information constraints are introduced to make the coupling degree between the extracted content embedding and timbre embedding as low as possible. In zero-shot voice synthesis, usually a single audio of the target speaker is used to obtain the speaker representation. However, a single audio cannot cover all the details of the target speaker's timbre, and the timbre information contained therein is often limited, which makes the synthesized voice deviate in terms of speaker similarity. To be able to obtain sufficient timbre information, the present invention will use multiple reference audios of the target speaker as inputs, extract the timbre information for each reference audio respectively, and then fuse several timbre information through a timbre information fusion module, so that the final speaker embedding can contain as much timbre information of the target speaker as possible.
[0099] Note that the above is only the preferred embodiment of the present invention and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described here, and various obvious changes, re-adjustments and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments only. Without departing from the concept of the present invention, more other equivalent embodiments can be included, and the scope of the present invention is determined by the scope of the appended claims.
Claims
1. A zero-shot voice cloning method based on audio decoupling and fusion, characterized in that, It includes the following steps: Extract the acoustic features of the audio of the target speaker to obtain the acoustic feature Mel spectrogram; Use the voice and content decoupling module to separate the voice information and content information of the acoustic feature Mel spectrogram. The voice and content decoupling module includes an audio content encoder and an audio voice encoder. The audio content encoder is used to extract the content information related to the language content in the acoustic feature Mel spectrogram, and the audio voice encoder is used to extract the voice information related to the speaker identity in the acoustic feature Mel spectrogram; Fuse the voice information set to obtain the speaker voice embedding representation; Input the speaker voice embedding representation and the text into the zero-shot voice cloning model to synthesize the Mel spectrogram corresponding to the text and with the target speaker's voice. The zero-shot voice cloning model includes an encoder, a variational adaptor and a decoder. The encoder is used to extract the hidden state related to the text, the variational adaptor is used to predict the fundamental frequency and the duration of the input text features for the input features, and the decoder is used to synthesize the Mel spectrogram from the input features; Input the Mel spectrogram with the target speaker's voice into the vocoder to convert the Mel spectrogram with the target speaker's voice into a waveform signal audible to the human ear; Input the speaker voice embedding representation and the text into the zero-shot voice cloning model to synthesize the Mel spectrogram corresponding to the text and with the target speaker's voice, specifically including: Input the text into the encoder to extract the hidden state related to the text; Extend the speaker voice embedding representation to the same length as the text and add it to the text hidden state to obtain the fused feature; Input the fused feature into the variational adaptor to predict the pitch, energy and duration information of the text features of the fused feature; Extend the fused feature to the frame-level length according to the duration information; Input the extended fused feature into the decoder to synthesize the Mel spectrogram with the target speaker's voice.
2. The zero-shot voice cloning method based on audio decoupling and fusion according to claim 1, characterized in that, Extract the acoustic features of the audio of the target speaker to obtain the audio feature Mel spectrogram, specifically including: Perform pre-emphasis, framing and windowing processing on the input audio signal in sequence to obtain the preprocessed audio signal; Perform short-time Fourier transform and Mel spectrogram transform on the preprocessed audio signal in sequence to obtain the acoustic feature Mel spectrogram.
3. The zero-shot voice cloning method based on audio decoupling and fusion according to claim 1, characterized in that, The audio content encoder includes an encoding module, a vector quantization layer and a contrast prediction module. The encoding module includes two consecutive convolutional layers, an activation layer and an instance regularization layer. The encoding module is used to generate a continuous feature vector from the input acoustic feature Mel spectrogram; the vector quantization layer consists of K different codes to form a trainable codebook E, which is used to map each continuous feature vector to one of the K codebooks through nearest neighbor search to form a discrete feature vector; the contrast prediction module is an autoregressive structure, which is used to generate a context representation using the discrete feature vectors at the previous t moments and predict the discrete feature vectors at the subsequent n moments using the context representation.
4. The zero-shot voice cloning method based on audio decoupling and fusion according to claim 1, characterized in that, The audio timbre encoder includes a convolutional layer, multiple convolutional banks, and an average pooling layer. The input acoustic feature Mel spectrogram undergoes a convolutional layer with a convolutional kernel size of 8 for feature dimension transformation to obtain a dimension-transformed feature vector. The dimension-transformed feature vector passes through 6 convolutional banks to expand the receptive field. The dimension-transformed feature vector with the expanded receptive field undergoes an average pooling operation to obtain a global feature vector. The global feature vector passes through 2 convolutional banks for feature transformation to obtain the speaker's timbre information.
5. The zero-shot voice cloning method based on audio decoupling and fusion according to claim 1, characterized in that, By reducing the mutual information between the timbre information and content information separated by the timbre and content decoupling module, the decoupling degree of the timbre information and content information is improved. Specifically, it includes: taking the timbre information and content information separated by the timbre and content decoupling module as two constraint variables to be constrained, and adding an unbiased estimation loss of the upper bound of mutual information during the training process of the entire method model.
6. The zero-shot voice cloning method based on audio decoupling and fusion according to claim 1, characterized in that, The information fusion module is used to perform feature fusion on the timbre information set to obtain the speaker timbre embedding representation. The information fusion module includes a multi-head attention mechanism module and a timbre library. The timbre library uses the basis vectors of the speaker timbre as the values and keys of the attention mechanism. The multi-head attention mechanism module takes the timbre information as the query, and the speaker timbre library as the values and keys, and averages the similarity scores between each timbre vector and the timbre library to obtain the speaker timbre embedding representation.
7. A zero-shot voice cloning device based on audio decoupling and fusion, characterized in that, The device includes: An acoustic feature Mel spectrogram acquisition module for extracting the acoustic feature Mel spectrogram of the target speaker's audio. A timbre and content decoupling module for separating the timbre information and content information of the acoustic feature Mel spectrogram. The timbre and content decoupling module includes an audio content encoder and an audio timbre encoder. The audio content encoder is used to extract the content information related to the language content in the acoustic feature Mel spectrogram, and the audio timbre encoder is used to extract the timbre information related to the speaker identity in the acoustic feature Mel spectrogram. A timbre information fusion module for performing feature fusion on the timbre information set to obtain the speaker timbre embedding representation. A timbre Mel spectrogram acquisition module for inputting the speaker timbre embedding representation and text into a zero-shot voice cloning model to synthesize a Mel spectrogram corresponding to the text and with the target speaker's timbre. The zero-shot voice cloning model includes an encoder, a variational adaptor, and a decoder. The encoder is used to extract the hidden state related to the text, the variational adaptor is used to predict the fundamental frequency and the duration of the input text feature for the input feature, and the decoder is used to synthesize the Mel spectrogram from the input feature. A Mel spectrogram conversion module for inputting the Mel spectrogram with the target speaker's timbre into a vocoder to convert the Mel spectrogram with the target speaker's timbre into an audible waveform signal for the human ear. Inputting the speaker timbre embedding representation and text into a zero-shot voice cloning model to synthesize a Mel spectrogram corresponding to the text and with the target speaker's timbre, specifically including: Inputting the text into the encoder to extract the hidden state related to the text. Expanding the speaker timbre embedding representation to the same length as the text and adding it to the text hidden state to obtain a fused feature. Input the fused features into the variational adaptor to predict the pitch, energy, and duration information of the text features of the fused features; Expand the fused features to the frame-level length according to the duration information; Input the expanded fused features into the decoder to synthesize the Mel spectrogram with the target speaker's timbre.
8. A zero-shot voice cloning device based on audio decoupling and fusion, characterized in that, Comprising: At least one processor; And, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the zero-shot voice cloning method based on audio decoupling and fusion according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for causing a computer to execute the zero-shot voice cloning method based on audio decoupling and fusion according to any one of claims 1 to 6.