Voice generation method and apparatus, storage medium, and electronic device
The speech generation method addresses the challenge of fusing speech synthesis and voice timbre conversion by processing speech and text feature vectors through a Sequence-to-Sequence model and a vocoder, achieving improved performance and efficiency even with limited data.
Patent Information
- Application Number
- JP2024568723
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-05-27
- Filing Date
- 2022-09-22
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-09-22
AI Technical Summary
Existing technologies face challenges in achieving effective fusion of speech synthesis and voice timbre conversion, particularly when a large amount of recorded voice data is not available, leading to decreased performance and complex training methods.
A speech generation method that involves obtaining speech and text feature vectors, processing them through a Sequence-to-Sequence model and a vocoder, and using a voice generation model to improve the fusion of voice synthesis and voice timbre conversion tasks, even with limited data.
This method enhances the performance of voice synthesis and voice timbre conversion tasks, improves timbre cloning with limited data, reduces training difficulty and time, and supports multiple application scenarios.
Smart Images

Figure 2025517777000001_ABST
Abstract
Description
Related Applications
[0001] This disclosure claims the priority of a Chinese patent application with the application number 202210593870.9, filed on May 27, 2022, and the title "Voice Generation Method, Apparatus, Storage Medium, and Electronic Device", and all the contents of the Chinese patent application are incorporated herein by reference.
Technical Field
[0002] This disclosure relates to the field of voice processing technology, and in particular, to a voice generation method, a voice generation apparatus, a computer-readable storage medium, and an electronic device.
Background Art
[0003] In recent years, with the rapid development of the field of deep learning, voice synthesis technology (Text to Speech, TTS) has made remarkable progress. At the same time, due to the development of various deep learning technologies, voice color conversion (voice conversion, VC) has also been rapidly developing. However, both the TTS model and the VC model require a large amount of recorded voice data (more than a dozen hours) to achieve the expected effect. However, voice recording is very expensive and complicated. Therefore, how to guarantee the voice synthesis effect and the voice color conversion effect when only a small amount of speaker data can be obtained has become a research hot spot. This research is called speaker adaptation or speaker cloning.
[0004] Currently, some research has been conducted on combining speech synthesis technology and voice timbre conversion. For example, different input source contents can be encoded using different encoders and then decoded using the same decoder to process the two tasks simultaneously. However, effectively speaking, the effect of TTS always decreases, or different input source contents are encoded using different encoders, but the model training is complex and many loss functions and hyperparameters are also required. Therefore, the research on combining speech synthesis technology and voice timbre conversion not only has a complex training method but also cannot effectively process speech synthesis and voice timbre conversion, so the performance of the two tasks cannot be improved simultaneously.
[0005] In view of this, in this field, it is necessary to develop a new speech generation method and device.
[0006] Note that the information partially disclosed in the above background technology is only used to strengthen the understanding of the background of the present disclosure, so it may include information that does not constitute prior art known to those skilled in the art.
Summary of the Invention
[0007] An object of the present disclosure is to provide a speech generation method, a speech generation device, a computer-readable storage medium, and an electronic device that can at least overcome to some extent the technical problem that the fusion effect of speech synthesis technology and voice timbre conversion is low due to the limitations of related technologies, and it is difficult to improve the performance of the two tasks simultaneously.
[0008] Other features and advantages of the present disclosure will become apparent from the following detailed description or will be partially learned from the implementation of the present disclosure.
[0009] According to a first aspect of an embodiment of the present invention, a speech generation method is provided, and the speech generation method includes: Obtaining a speech feature vector of a speech to be processed, and inputting the speech feature vector into a speech generation model to obtain a language unit vector; Obtain a text feature vector, and determine a feature vector to be processed based on the text feature vector and the language unit vector; Input the feature vector to be processed into a Sequence-to-Sequence model to obtain an acoustic feature vector, and input the acoustic feature vector into a vocoder to obtain a target voice corresponding to the voice to be processed or the text feature vector.
[0010] In an exemplary embodiment of the present invention, obtaining a language unit vector by inputting the acoustic feature vector into a voice generation model includes: Inputting the acoustic feature vector into a voice generation model such that the voice generation model outputs an acoustic coding vector and a self-reconstructed voice; Calculating a loss for the voice to be processed and the self-reconstructed voice to obtain a first loss value, and determining the acoustic coding vector as a language unit vector based on the first loss value.
[0011] In an exemplary embodiment of the present invention, inputting the acoustic feature vector into a voice generation model such that the voice generation model outputs an acoustic coding vector and a self-reconstructed voice includes: Inputting the acoustic feature vector into a voice generation model, and non-linearly transforming the acoustic feature vector by an encoder module of the voice generation model to obtain an acoustic coding vector; Quantizing the acoustic coding vector by a vector quantization module of the voice generation model to obtain an acoustic quantization sequence, and obtaining a speaker vector corresponding to the voice to be processed; Non-linearly transforming the acoustic quantization sequence and the speaker vector by a decoder module of the voice generation model to obtain a self-reconstructed voice.
[0012] In an exemplary embodiment of the present invention, obtaining a speaker vector corresponding to the voice to be processed includes: Obtaining a speaker identifier corresponding to the speech to be processed, and determining a correspondence relationship between the speaker identifier and a speaker vector, wherein the correspondence relationship is determined based on the speech generation model, querying the speaker vector corresponding to the speaker identifier based on the correspondence relationship.
[0013] In an exemplary embodiment of the present invention, obtaining the speech quantization sequence by quantizing the speech coding vector by a vector quantization module of the speech generation model includes quantizing the speech coding vector by a nearest neighbor search algorithm based on a codebook in the vector quantization module of the speech generation model to obtain a speech quantization sequence.
[0014] In an exemplary embodiment of the present invention, obtaining the speech quantization sequence by quantizing the speech coding vector by the nearest neighbor search algorithm includes obtaining an updated codebook by updating the codebook, and quantizing the speech coding vector by a nearest neighbor search algorithm based on the updated codebook to obtain a speech quantization sequence.
[0015] In an exemplary embodiment of the present invention, obtaining an updated codebook by updating the codebook includes obtaining a codebook identifier of the codebook for each frame, comparing the codebook identifiers to obtain a comparison result, and obtaining an updated codebook by merging the codebooks according to the comparison result.
[0016] In an exemplary embodiment of the present invention, obtaining an acoustic feature vector by inputting the feature vector to be processed into a sequence-to-sequence model Obtaining a processing target acoustic vector of the processing target feature vector, and inputting the processing target feature vector and the processing target acoustic vector into a sequence-to-sequence model, so that the sequence-to-sequence model outputs a processed acoustic vector, and Performing loss calculation on the processing target acoustic vector and the processed acoustic vector to obtain a second loss value, and determining the processed acoustic vector as an acoustic feature vector based on the second loss value.
[0017] In an exemplary embodiment of the present invention, the step of inputting the processing target feature vector and the processing target acoustic vector into a sequence-to-sequence model so that the sequence-to-sequence model outputs a processed acoustic vector includes: Inputting the processing target feature vector and the processing target acoustic vector into a sequence-to-sequence model, and non-linearly mapping the processing target feature vector and the processing target acoustic vector by an encoder module of the sequence-to-sequence model to obtain a spatial encoding vector; Adding the spatial encoding vector and the speaker vector to obtain an alignment target vector, and obtaining an audio feature sequence; Aligning the alignment target vector and the audio feature sequence by an attention mechanism of the sequence-to-sequence model to obtain a context representation vector, and non-linearly mapping the context representation vector by a decoder of the sequence-to-sequence model to obtain a processed acoustic vector.
[0018] In an exemplary embodiment of the present invention, the step of determining a processing target feature vector based on the text feature vector and the language unit vector includes: Determining the text feature vector or the speech unit vector as the processing target feature vector, or Adding the text feature vector and the language unit vector to obtain a processing target feature vector.
[0019] In an exemplary embodiment of the present invention, inputting the acoustic feature vector into a vocoder to obtain a target voice corresponding to the voice to be processed or the text feature vector is extracting the acoustic features of the acoustic feature vector via a post-processing network, and inputting the acoustic features into the vocoder, so that the vocoder outputs a pending voice corresponding to the voice to be processed or the text feature vector, and performing loss calculation on the pending voice and the voice to be processed to obtain a third loss value, and determining the pending voice as a target voice based on the third loss value.
[0020] In an exemplary embodiment of the present invention, performing loss calculation on the pending voice and the voice to be processed to obtain a third loss value is when the vocoder is an adversarial generation network, performing loss calculation on the pending voice and the voice to be processed to obtain an adversarial network loss value of the adversarial generation network, and performing loss calculation on the pending voice and the voice to be processed to obtain a voice feature loss value, and obtaining a third loss value by weighted addition of the adversarial network loss value and the voice feature loss value.
[0021] According to a second aspect of an embodiment of the present invention, there is provided a voice generation device, the voice generation device including a data acquisition module configured to acquire an acoustic feature vector of a voice to be processed, and input the acoustic feature vector into a voice generation model to acquire a language unit vector, and a vector determination module configured to acquire a text feature vector, and determine a feature vector to be processed based on the text feature vector and the language unit vector, and The voice generation module is configured to input the processing target feature vector into a sequence-to-sequence model to obtain an acoustic feature vector, and input the acoustic feature vector into a vocoder to obtain a target voice corresponding to the processing target voice or the text feature vector.
[0022] According to a third aspect of an embodiment of the present invention, an electronic device including a processor and a memory is provided, and computer-readable instructions are stored in the memory. When the computer-readable instructions are executed by the processor, the voice generation method in any of the above exemplary embodiments is realized.
[0023] According to a fourth aspect of an embodiment of the present invention, a computer-readable storage medium storing a computer program is provided. When the computer program is executed by a processor, the voice generation method in any of the above exemplary embodiments is realized.
[0024] As can be seen from the above technical solutions, the voice generation method, voice generation device, computer storage medium, and electronic device in the exemplary embodiments of the present disclosure have at least the following advantages and excellent effects.
[0025] In the method and device according to the exemplary embodiments of the present disclosure, by obtaining an acoustic feature vector and a text feature vector, voice and text can be received as inputs, fusing a voice synthesis task and a voice timbre conversion task, performing multimodal modeling, and improving the performance of the voice synthesis task and the voice timbre conversion task. Furthermore, in the case of a small amount of data, an acoustic feature vector and a text feature vector are obtained, a strategy for multiple types of timbre cloning is provided, the effect of timbre cloning with a small amount of data is improved, the training difficulty and training time of multiple models are reduced, and a timbre cloning method in multiple application scenarios is supported.
[0026] It should be noted that the above general description and subsequent detailed description are merely exemplary and explanatory, and do not limit the present disclosure.
Brief Description of the Drawings
[0027]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Embodiments for Carrying Out the Invention
[0028] Hereinafter, the present embodiment will be described in more detail with reference to the drawings. However, the exemplary embodiments can be implemented in various forms and are not limited to the embodiments described herein. These embodiments are provided to make the present disclosure more comprehensive and complete, and to comprehensively convey the concept of the exemplary embodiments to those skilled in the art. The features, structures, or characteristics described may be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to fully understand the embodiments of the present disclosure. However, those skilled in the art should recognize that one or more of the specific details may be omitted when practicing the technical solution of the present disclosure, or other methods, components, devices, steps, etc. may be adopted. In other cases, well-known technical solutions are not shown or described in detail to avoid obscuring various aspects of the present disclosure.
[0029] In this specification, the terms "one", "a", "the", and "said" are used to indicate the presence of one or more elements / components / etc., the terms "comprising" and "having" mean an open inclusion, and other elements / components / etc. may exist in addition to the listed elements / components / etc., and the terms "first", "second", etc. do not limit the number of their objects and are only used as mere marks.
[0030] Note that the drawings are schematic of the present disclosure and are not necessarily drawn to scale. The same or corresponding parts in the drawings are denoted by the same reference numerals, and duplicate descriptions are omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily need to correspond to physically or logically independent entities.
[0031] In recent years, with the rapid development of the field of deep learning, voice synthesis technology has made remarkable progress. At the same time, due to the development of various deep learning technologies, voice timbre conversion has also been rapidly developing. However, whether it is a TTS model or a VC model, a large amount of recorded voice data (more than a dozen hours) is required to achieve the expected effect. However, voice recording is very expensive and cumbersome. Therefore, how to guarantee the voice synthesis effect and timbre conversion effect when only a small amount of speaker data can be obtained has become a research hot spot. This research is called speaker adaptation or speaker cloning.
[0032] Here, speaker adaptation is a technology in which a deep learning model is quickly and automatically adapted to a target speaker, so that the deep learning model brings a significant improvement in the performance of the speaker.
[0033] From the perspective of speaker cloning, voice synthesis technology and voice timbre conversion should be regarded as systems that generate the voice of a target speaker based on different inputs.
[0034] In the field of TTS, speaker adaptation can be divided into supervised speaker adaptation and unsupervised speaker adaptation.
[0035] Here, supervised speaker adaptation refers to data that requires a <text, voice> pair when adapting, and unsupervised speaker adaptation refers to requiring only voice data and not requiring the corresponding text when adapting.
[0036] As is clear from previous research, in teacher-assisted speaker adaptation, high-quality effects can be achieved by fine-tuning a multi-speaker base model using a small amount of <text, audio> pair data of the target speaker.
[0037] However, in speaker adaptation without a teacher, the model cannot be fine-tuned.
[0038] A common method for speaker adaptation without a teacher is to extract one speaker vector from an audio segment using a speaker identification model, and then use that speaker vector to synthesize the audio of that speaker.
[0039] However, as the amount of data increases, the performance of this method does not improve further.
[0040] Currently, some research is being conducted on combining speech synthesis technology and voice color conversion.
[0041] For example, a sequence-to-sequence TTS model can be used to extract speaker-independent representations to model a VC model, a TTS pre-training model can be used to improve the VC model effect, different encoders can be used to encode different input source contents, and then the same decoder can be used for decoding, and the two tasks can be processed simultaneously. However, effectively speaking, the effect of TTS always decreases.
[0042] Alternatively, different encoders are used to encode different input source contents, but the training of the model is complex and many loss functions and hyperparameters are also required. Therefore, the research on combining speech synthesis technology and voice color conversion is not only complex in the training method, but also cannot handle speech synthesis and voice color conversion well, and cannot improve the performance of the two tasks simultaneously.
[0043] In response to the problems existing in the related art, the present disclosure provides a voice generation method. FIG. 1 shows a flowchart of the voice generation method. As shown in FIG. 1, the voice generation method includes at least the following steps.
[0044] In step S110, an acoustic feature vector of the voice to be processed is obtained, and the acoustic feature vector is input into a voice generation model to obtain a language unit vector.
[0045] In step S120, a text feature vector is obtained, and a feature vector to be processed is determined based on the text feature vector and the language unit vector.
[0046] In step S130, the feature vector to be processed is input into a sequence-to-sequence model to obtain an acoustic feature vector, and the acoustic feature vector is input into a vocoder to obtain a target voice corresponding to the voice to be processed or the text feature vector.
[0047] In an exemplary embodiment of the present disclosure, by obtaining the acoustic feature vector and the text feature vector, voice and text can be received as inputs, the voice synthesis task and the voice timbre conversion task are fused, multimodal modeling is performed, and the performance of the voice synthesis task and the voice timbre conversion task is improved. Further, in the case of a small amount of data, the acoustic feature vector and the text feature vector are obtained, a strategy for multiple types of timbre cloning is provided, the effect of timbre cloning with a small amount of data is improved, the training difficulty and training time of multiple types of models are reduced, and the timbre cloning method in multiple application scenarios is supported.
[0048] Hereinafter, each step of the voice generation method will be described in detail.
[0049] In step S110, an acoustic feature vector of the voice to be processed is obtained, and the acoustic feature vector is input into a voice generation model to obtain a language unit vector.
[0050] In an exemplary embodiment of the present disclosure, the voice to be processed may be a voice to be converted for voice timbre conversion.
[0051] Here, voice timbre conversion is a system that automatically converts the voice of speaker A into the voice of speaker B while keeping the content of the conversation unchanged.
[0052] Therefore, the voice to be processed can be understood as the voice of speaker A.
[0053] Correspondingly, the voice feature vector of the voice to be processed may be a Mel Frequency Spectrum feature vector extracted based on the voice to be processed.
[0054] The Mel Frequency Spectrum feature can simulate the processing characteristics of the human ear for voice to a certain extent, so it can better reflect the human auditory characteristics, thereby improving the user's auditory experience.
[0055] After obtaining the voice feature vector, the voice feature vector can be input into the voice generation model so that the voice generation model outputs the corresponding language unit vector.
[0056] In an alternative embodiment, FIG. 2 is a flowchart of a method for the voice generation model to output a language unit vector. As shown in FIG. 2, the method at least includes the following steps. In step S210, the voice feature vector is input into the voice generation model so that the voice generation model outputs a voice encoding vector and a self-recovering voice.
[0057] In an alternative embodiment, FIG. 3 is a flowchart of a processing method of the voice generation model. As shown in FIG. 3, the method at least includes the following steps. In step S310, the voice feature vector is input into the voice generation model, and the voice feature vector is non-linearly transformed by the encoder module of the voice generation model to obtain a voice encoding vector.
[0058] Here, the voice generation model may be a VQ-VAE (Vector Quatization-Variational AutoEncoder) model, or may be other models, and the exemplary embodiments do not particularly limit this.
[0059] The VQ-VAE model is an autoencoder that is clearly characterized in that the encoded encoding vector is discrete. VQ-VAE includes an encoding layer and a decoding layer. The voice feature vector can be encoded into a discrete encoding vector through the encoding layer, and the discrete encoding vector can be decoded into a vector through the decoding layer.
[0060] Specifically, when the voice generation model is a VQ-VAE model, the voice feature vector is input into the voice generation model, and the encoder module in the VQ-VAE model maps the voice feature vector to a high-dimensional voice encoding vector through non-linear transformation. This voice encoding vector can be represented by z 1:N and can be represented as such.
[0061] The encoder can extract more abstract and higher-dimensional features from the input features through non-linear transformation of the neural network.
[0062] Here, the encoder module in the VQ-VAE model may be composed of a CNN (Convolutional Neural Networks) and an LSTM (Long short-term memory).
[0063] Specifically, the CNN model may include an input layer, a convolution layer, a pooling layer, a fully connected layer (FC), and an output layer.
[0064] Here, the activation function of the convolutional layer uses ReLU (Rectified Linear Unit), and the pooling layer has no activation function. The combination of the convolutional layer and the pooling layer can appear multiple times in the hidden layer, and actually, this number of times is determined according to the needs of the model.
[0065] Of course, combinations of convolutional layers with each other, or combinations of convolutional layers with convolutional layers and pooling layers, may also be utilized, and there are no restrictions when constructing the model. However, the most common CNNs are combinations of several convolutional layers and pooling layers.
[0066] After several convolutional layers and pooling layers, there is a fully connected layer, which is actually a DNN (Deep Neural Networks) structure, and the output layer only performs tasks such as classification using the Softmax activation function.
[0067] LSTM is a special RNN (Recurrent Neural Network), mainly to solve the problems of vanishing gradients and exploding gradients in the long sequence training process. Briefly speaking, compared with general RNNs, LSTM can have better performance in longer sequences.
[0068] In step S320, the vector quantization module of the voice generation model quantizes the voice encoding vector to obtain a voice quantization sequence, and obtains the speaker vector corresponding to the voice to be processed.
[0069] After the encoder module of the VQ-VAE model obtains the voice encoding vector z 1:N it uses the vector quantization TIFF2025517777000002.tif7170 module in the VQ-VAE model to quantize the high-dimensional voice encoding vector into a voice quantization sequence. This voice quantization sequence It may be shown at TIFF2025517777000003.tif7170.
[0070] In an alternative embodiment, based on the codebook in the vector quantization module of the voice generation model, a voice quantization sequence is obtained by quantizing voice coding vectors using a nearest neighbor search algorithm.
[0071] Specifically, based on the codebook in the VQ-VAE model, consecutive voice coding vectors are quantized into a discrete voice quantization sequence by a nearest neighbor search algorithm.
[0072] In an alternative embodiment, FIG. 4 is a flowchart of a method for quantizing a voice quantization sequence. As shown in FIG. 4, the method at least includes the following steps. In step S410, the codebook is updated to obtain an updated codebook.
[0073] In an alternative embodiment, FIG. 5 is a flowchart of a method for updating a codebook. As shown in FIG. 5, the method at least includes the following steps. In step S510, the codebook identifiers of the codebooks for each frame are obtained, and the codebook identifiers are compared to obtain a comparison result.
[0074] Since the voice feature vectors of each frame correspond to their respective codebooks, for example, the codebooks for the voice feature vectors of 5 frames may be book1, book1, book1, book2, and book2.
[0075] Considering that the text feature vector may be processed by a sequence-to-sequence model later, in order to obtain a sequence representation adapted by the text feature vector, the codebook identifiers can be compared to obtain a comparison result.
[0076] Here, the comparison result can reflect whether two or more consecutive codebook identifiers are the same.
[0077] In step S520, based on the comparison result, the codebook is merged to obtain an updated codebook.
[0078] When the comparison result reflects that two or more consecutive codebook identifiers are the same, the same codebook identifiers can be merged into one for querying by the nearest neighbor search algorithm.
[0079] For example, when the codebooks of the voice feature vectors of 5 frames are book1, book1, book1, book2, and book2, the codebook identifiers can be merged into book1 and book2 to obtain an updated codebook.
[0080] In this exemplary embodiment, by updating the codebook through the post-processing network, a sequence representation in which the voice quantization sequence is more compatible with the text feature vector can be obtained, providing data support for the multi-modal voice color conversion task and being convenient for improving the performance of voice color conversion.
[0081] In step S420, based on the updated codebook, the voice coding vector is quantized by the nearest neighbor search algorithm to obtain a voice quantization sequence.
[0082] Here, each updated codebook is a K*D-dimensional codebook maintained in the VQ-VAE model.
[0083] For example, each codebook is composed of K D-dimensional coding vectors e 1 , e 2 , ···, e Kmay be included. Using the encoding layer of the VQ-VAE model to encode the audio feature vector of H’*W’*D dimensions, and further for each D-dimensional vector in the audio feature vector of H’*W’*D dimensions, the encoding vector e i closest to the D-dimensional vector can be found in the codebook respectively. Here, the encoding vector e i is one vector in the codebook, and by representing the D-dimensional vector with the index of the encoding vector e i , a discrete vector of H’*W’ dimensions is obtained. Here, K, D, H’ and W’ represent dimensions respectively.
[0084] Furthermore, based on a preset discrete encoding method, the discrete vector of H’*W’ dimensions is converted into an audio quantization sequence.
[0085] Here, the preset discrete encoding method may be one-hot encoding or other types of encoding methods, and the exemplary embodiments are not particularly limited thereto.
[0086] Specifically, using the codebook of the one-hot encoding method, by the look-up table method, the discrete vector of H′*W′ dimensions is converted into another discrete encoding vector of H′*W′ dimensions encoded by the codebook of the one-hot encoding method, and further an audio quantization sequence is obtained based on the converted discrete encoding vector of H′*W′ dimensions.
[0087] For example, after converting a 3*3 discrete vector into a 3*3 discrete encoding vector encoded by another codebook of the one-hot encoding method, a 1*9 audio quantization sequence can be obtained based on each element in the converted 3*3 discrete encoding vector.
[0088] In the exemplary embodiment, a discrete encoding process is performed on the speech encoding vector by the vector quantization module of the speech generation model to obtain a corresponding speech quantization sequence, and a database and theoretical support are provided to output the speech encoding vector and the self-referential speech to the speech generation model.
[0089] In addition, a speaker vector of the speaker who spoke or pronounced the speech to be processed may be obtained.
[0090] In an alternative embodiment, FIG. 6 is a flowchart of a method for obtaining a speaker vector. As shown in FIG. 6, the method at least includes the following steps. In step S610, a speaker identifier corresponding to the speech to be processed is obtained, and a correspondence between the speaker identifier and the speaker vector is determined. Here, the correspondence is determined based on the speech generation model.
[0091] Here, the speaker identifier can uniquely represent the identifier information of the speaker who spoke or pronounced the speech to be processed.
[0092] In addition, a table for storing the correspondence between the speaker identifier and the speaker vector can be simultaneously maintained by the first loss value calculated between the self-referential speech output by the decoder in the speech generation model and the speech to be processed.
[0093] In step S620, based on the correspondence, a speaker vector corresponding to the speaker identifier is queried.
[0094] In the table storing the correspondence between the speaker identifier and the speaker vector, a speaker vector corresponding to the speaker identifier can be queried based on the speaker identifier.
[0095] In the exemplary embodiment, by obtaining the speaker vector maintained by the voice generation model, data support is provided to the decoder module of the voice generation model, and the support provided by the training of the decoder module and the encoder module via the speaker vector and the voice quantization sequence can be useful for the generation of the encoder of the voice generation model and the determination of the speaker vector.
[0096] In step S330, the decoder module of the voice generation model non-linearly transforms the voice quantization sequence and the speaker vector to obtain a self-reconstructed voice.
[0097] The decoder module of the voice generation model can receive the quantized voice quantization sequence, obtain the speaker vector, then add the voice quantization sequence and the speaker vector, and then reduce it by non-linear transformation to obtain a self-reconstructed voice.
[0098] This decoder transforms high-dimensional and abstract hidden features into more explicit features through the non-linear transformation of the neural network.
[0099] In addition, when the voice generation model is a VQ-VAE model, the decoder module in the VQ-VAE model may be composed of a CNN and an LSTM.
[0100] In the exemplary embodiment, through the corresponding processing of the encoder module, vector quantization module, and decoder module in the voice generation model, a self-reconstructed voice can be output to support the training of the voice generation model.
[0101] In step S220, loss calculation is performed on the voice to be processed and the self-reconstructed voice to obtain a first loss value, and based on the first loss value, the voice encoding vector is determined as the language unit vector.
[0102] After the voice generation model outputs the self-returned voice, the first loss value of the voice generation model can be calculated for the voice to be processed and the self-returned voice.
[0103] Specifically, the first loss value may be calculated by an L2 norm loss function.
[0104] The L2 norm loss function is shown in Equation (1). TIFF2025517777000004.tif11170
[0105] The L2 norm loss function is also called the least square error (LSE). The L2 norm loss function minimizes the sum of the squares of the differences between the target value y i and the estimated value f(x i ).
[0106] In a general regression problem, this loss is used, and outliers have a large impact on this loss.
[0107] When the first loss value calculated according to Equation (1) reaches a stable value and does not decrease any further, it indicates that the training of the voice generation model has been completed.
[0108] In this case, the voice encoding vector output from the encoder module in the voice generation model can be determined as the teacherless language unit vector. The language unit vector can be represented by k i .
[0109] In this exemplary embodiment, a language unit vector can be obtained by the voice generation model, providing data input of the voice modality to the sequence-to-sequence model, and supporting a multi-module voice generation method.
[0110] In step S120, a text feature vector is obtained, and a feature vector to be processed is determined based on the text feature vector and the language unit vector.
[0111] In an exemplary embodiment of the present disclosure, speech synthesis is a system that automatically converts natural text into speech.
[0112] Therefore, after obtaining natural text, the phoneme sequences of the natural text can be extracted as text feature vectors. The method of extracting the phoneme sequences of the natural text may be realized using an LSTM model, and the exemplary embodiment does not particularly limit this.
[0113] After obtaining the text feature vector, a processing target feature vector can be determined based on the text feature vector and the language unit vector.
[0114] In an alternative embodiment, FIG. 7 shows a flowchart of a method for determining a processing target feature vector based on a text feature vector and a language unit vector. As shown in FIG. 7, the method at least includes the following steps. In step S710, a text feature vector or an audio unit vector is determined as the processing target feature vector.
[0115] In the field of TTS, speaker adaptation can be divided into supervised speaker adaptation and unsupervised speaker adaptation. Here, supervised speaker adaptation refers to data that requires a <text, speech> pair when adapting, and unsupervised speaker adaptation refers to only requiring speech data when adapting and not requiring the corresponding text.
[0116] As is clear from previous research, in supervised speaker adaptation, high-quality effects can be achieved by fine-tuning a multi-speaker basic model using a small amount of <text, speech> pair data of the target speaker.
[0117] However, in unsupervised speaker adaptation, the model cannot be fine-tuned.
[0118] Here, a general speaker adaptation method without a teacher is to extract one speaker vector from an audio segment using a speaker identification model, and then synthesize the audio of the speaker using the speaker vector.
[0119] Obviously, speaker adaptation without a teacher cannot be achieved by a trained model. Also, as the amount of data increases, the performance of this method will not improve further. Therefore, in speaker adaptation without a teacher, a text feature vector can be determined as a processing target feature vector, and a text synthesis effect can be realized by a subsequent sequence-to-sequence model and a vocoder.
[0120] Voice timbre conversion can determine a voice unit vector as a processing target feature vector, and the achievement of a voice timbre conversion task can be realized by a subsequent sequence-to-sequence model and a vocoder.
[0121] In step S720, a text feature vector and a language unit vector are added to obtain a processing target feature vector.
[0122] In order to improve the effect of the voice timbre conversion task, a text feature vector and a language unit vector can be added to obtain a processing target feature vector.
[0123] Note that in the process of obtaining the language unit vector, since the codebook in the vector quantization module of the voice generation model has already been updated, the language unit vector and the text feature vector are very compatible, so the text feature vector and the language unit vector can be directly added.
[0124] When obtaining a processing target feature vector by adding a text feature vector and a language unit vector, since the processing target feature vector corresponds to an expression in which the text modality is added to the language modality, the processing target feature vector at this time is an enhanced data representation. Based on this, the performance of the voice timbre conversion task realized by the processing target feature vector is better.
[0125] Here, Modality refers to the source or form of information.
[0126] For example, information may be expressed in various forms such as voice, video, text, image, etc., and each form of expression of the information may be called a modality of the information. Based on this, multimodality is fused in ways such as text, voice, vision, motion, and environment. Multimodal Machine Learning (MMML) refers to the ability to process and understand multi-source modality information by machine learning methods. For example, currently, a popular research direction is multimodal learning between images, videos, audio, and semantics.
[0127] In this exemplary embodiment, based on the differences between the voice synthesis task and the voice timbre conversion task, different processing target feature vectors can be determined as the basis for subsequent model processing, and data support is provided to improve the performance of the voice synthesis task and the voice timbre conversion task.
[0128] In step S130, the processing target feature vector is input into a sequence-to-sequence model to obtain an acoustic feature vector, and the acoustic feature vector is input into a vocoder to obtain a target voice corresponding to the processing target voice or text feature vector.
[0129] In the exemplary embodiment of the present disclosure, after determining the processing target feature vector, the corresponding acoustic feature vector can be obtained by inputting the processing target feature vector into a sequence-to-sequence model.
[0130] In an optional embodiment, FIG. 8 shows a flowchart of a method for outputting an acoustic feature vector by a sequence-to-sequence model. As shown in FIG. 8, the method at least includes the following steps. In step S810, an acoustic vector to be processed of a feature vector to be processed is obtained, and by inputting the feature vector to be processed and the acoustic vector to be processed into the sequence-to-sequence model, the sequence-to-sequence model outputs a processed acoustic vector.
[0131] The acoustic vector to be processed may be a mel-frequency spectrum feature vector.
[0132] In an optional embodiment, FIG. 9 shows a flowchart of a processing method of a sequence-to-sequence model. As shown in FIG. 9, the method at least includes the following steps. In step S910, a feature vector to be processed and an acoustic vector to be processed are input into the sequence-to-sequence model, and the encoder module of the sequence-to-sequence model performs a non-linear mapping on the feature vector to be processed and the acoustic vector to be processed to obtain a spatial encoding vector.
[0133] The sequence-to-sequence model may be a sequence-to-sequence (Seq2seq) model with an attention mechanism, or may be other models. The exemplary embodiments herein do not particularly limit this.
[0134] When the sequence-to-sequence model is a Seq2seq model based on an attention mechanism, the sequence-to-sequence model may include an encoder module, an attention mechanism, and a decoder module.
[0135] Here, the encoder module of the sequence-to-sequence model may be configured to obtain a sequence of representations corresponding to the feature vector to be processed and the acoustic vector to be processed. The attention mechanism may be configured to generate a fixed-length semantic representation according to the sequence of representations. The decoder module may obtain an acoustic vector according to the semantic representation.
[0136] Specifically, the encoder module of the sequence-to-sequence model may include a Feature Embedding layer, a Convolutional Pre-net, a Dense Pre-net, a CBHG (Convolution Bank + Highway network + bidirectional Gated Recurrent Unit, that is, a convolutional layer + a highway network + a bidirectional recurrent neural network, that is, CBHG consists of a convolutional layer, a highway network, and a bidirectional recurrent neural network) sub-model, and a Down-sampling Convolution layer.
[0137] First, use the Feature Embedding layer to encode the feature vector to be processed and then input it into the Convolutional Pre-net. By performing non-linear transformation on the encoded feature vector to be processed and the acoustic vector to be processed, the convergence and generalization ability of the sequence-to-sequence model based on the attention mechanism can be improved. At the same time, by inputting the number of audio frames corresponding to the acoustic vector to be processed into the Dense Pre-net, the corresponding depth features can be obtained. Then, the outputs of the Convolutional Pre-net and the Dense Pre-net are collectively input into the CBHG sub-model to extract the corresponding context features. After that, this is input into the Down-sampling Convolution to reduce the computational complexity and receptive field, and finally, the corresponding spatial encoding vector is obtained.
[0138] Therefore, the encoder module of the sequence-to-sequence model performs a non-linear transformation on the feature vector to be processed and the acoustic vector to be processed, and maps them to a high-dimensional spatial encoding vector. The spatial encoding vector can be represented by h t and can be represented as such.
[0139] In step S920, the spatial encoding vector and the speaker vector are added to obtain an alignment target vector, and an acoustic feature sequence is acquired.
[0140] To perform multi-speaker modeling, the attention mechanism of the sequence-to-sequence model can also receive the speaker vector as an input.
[0141] To input the speaker vector, the spatial encoding vector and the speaker vector can be added to obtain an alignment target vector.
[0142] Furthermore, since the attention mechanism is an autoregressive model, an acoustic feature sequence can also be obtained. The acoustic feature sequence can be represented by m t — 1 and can be represented as such. When t = 1, the acoustic feature sequence is initialized to a sequence of all zeros, and at t = 2 and subsequent times, the acoustic feature sequence is the feedback sequence for the previous time of the decoder module.
[0143] In step S930, the attention mechanism of the sequence-to-sequence model aligns the alignment target vector and the acoustic feature sequence to obtain a context representation vector, and the decoder of the sequence-to-sequence model performs a non-linear mapping on the context representation vector to obtain a processed acoustic vector.
[0144] Normally, since the acoustic feature vector is longer than the alignment target vector, the alignment target vector and the acoustic feature sequence can be aligned to obtain a context representation vector.
[0145] Specifically, the method of aligning the alignment target vector and the speech feature sequence may perform dot product calculation on the alignment target vector and the speech feature sequence.
[0146] In addition, the context representation vector obtained by aligning the alignment target vector and the speech feature sequence reflects the context relationship of the context and guarantees the effect of speech generation.
[0147] Furthermore, the decoder module of the sequence-to-sequence model mainly obtains the processed acoustic vector by returning the context representation vector obtained by performing alignment according to the alignment target vector and the speech feature sequence to the original speech acoustic feature space by non-linear mapping. Therefore, the processed acoustic vector may be a mel-frequency spectrum, and the processed acoustic vector may be represented by m.
[0148] In this exemplary embodiment, through the corresponding processing of the encoder module, the attention mechanism, and the decoder module in the sequence-to-sequence model for the processing target feature vector and the processing target acoustic vector, a fusion method can be provided for the speech synthesis task and the speech timbre conversion task, improving the timbre cloning effect with a small amount of data. Also, since multiple types of input data can be received, timbre cloning in various scenarios is supported.
[0149] In step S820, loss calculation is performed on the processing target acoustic vector and the processed acoustic vector to obtain a second loss value, and based on the second loss value, the processed acoustic vector is determined as the acoustic feature vector.
[0150] After the sequence-to-sequence model outputs the processed acoustic vector, the second loss value between the processing target acoustic vector and the processed acoustic vector can be calculated according to formula (1).
[0151] When the second loss value calculated according to formula (1) reaches a stable value and does not decrease any further, it indicates that the training of the sequence-to-sequence model has already been completed.
[0152] In this case, the processed acoustic vector output by the sequence-to-sequence model converged through training can be determined as the acoustic feature vector.
[0153] After the sequence-to-sequence model outputs the acoustic feature vector, further, the acoustic feature vector can be input into a vocoder to obtain the target voice for the text synthesis task or voice timbre conversion.
[0154] In an alternative embodiment, FIG. 10 is a flowchart of a method for outputting a target voice by a generator. As shown in FIG. 10, the method at least includes the following steps. In step S1010, the acoustic features of the acoustic feature vector are extracted through a post-processing network, and by inputting the acoustic features into a vocoder, the vocoder outputs a pending voice corresponding to the voice to be processed or the text feature vector.
[0155] Among them, the post-processing network is mainly set to generate acoustic characteristics with higher accuracy. The acoustic features may be represented by TIFF2025517777000005.tif7170.
[0156] The post-processing network may be a CNN network, or may be an LSTM network or the like. The exemplary embodiments herein are not particularly limited thereto.
[0157] A vocoder is a system that converts acoustic features, such as mel-frequency spectra, into voice audio.
[0158] Here, the vocoder may be a WaveNet model, a Griffin-Lim algorithm, a GAN network (Generative Adversarial Network), etc., and the exemplary embodiments are not particularly limited thereto.
[0159] Specifically, the WaveNet model is a sequence generation model and can be used for voice generation modeling. In the modeling of the acoustic model of voice synthesis, since WaveNet can directly learn the mapping to the sample value sequence, it has a good synthesis effect. Currently, WaveNet has great potential in the field of voice synthesis in the modeling of the acoustic model of voice synthesis.
[0160] Since the WaveNet model can predict the result of the t-th point according to the previous t-1 points of a sequence, it can be used to predict the numerical values of the sampling points in voice.
[0161] Griffin-Lim is an algorithm for reconstructing voice under the condition that only the amplitude spectrum is known and the phase spectrum is unknown.
[0162] The implementation of the Griffin-Lim algorithm is simple. The Griffin-Lim algorithm is an iterative algorithm. The iterative process first randomly initializes a phase spectrum, and then uses the phase spectrum and the known amplitude spectrum to synthesize new voice by ISTFT (Inverse Short-Time Fourier Transform). Furthermore, STFT (Short-Time Fourier Transform) is performed on the synthesized voice to obtain a new amplitude spectrum and phase spectrum. Finally, the new amplitude spectrum is discarded, and voice is synthesized using the phase spectrum and the known amplitude spectrum, and this is repeated.
[0163] The GAN network is a machine learning method proposed by Ian J. Goodfello et al. in the 2014 paper "Generative Adversarial Nets". Here, in the GAN network, there are two models, namely the generative model (generative model G) and the discriminative model (discriminative model D).
[0164] Taking the generation of images as an example, G is a network that generates images. It receives random noise z, and then generates an image based on this noise, and the generated data is denoted as G(z).
[0165] D is a discriminative network that discriminates whether an image is "real" (i.e., whether it is forged). Its input parameter is x, where x represents an image, and the output D(x) represents the probability that x is a real image. If it is 1, it means it is a real image, and if the output is 0, it means there is no possibility that it is a real image.
[0166] In the training process, the goal of the generative network G is to generate fake images to deceive the discriminative network D, and the goal of the discriminative network D is to discriminate whether an image is generated by G. Thus, a process of one game is formed. At the same time, the capabilities of G and D are also gradually improved in the training process. In the most ideal case, D(G(z)) = 0.5.
[0167] The vocoder can process the audio acoustic features extracted by the post - processing network and output the pending audio.
[0168] When performing the speech synthesis task, the pending audio may be the audio synthesized based on the text feature vector. When performing the voice color conversion task, the pending audio may be the audio converted based on the audio to be processed.
[0169] In step S1020, loss calculation is performed on the pending voice and the voice to be processed to obtain a third loss value, and based on the third loss value, the pending voice is determined as the target voice.
[0170] In an alternative embodiment, FIG. 11 shows a flowchart of a method for performing loss calculation to obtain a third loss value. As shown in FIG. 11, the method at least includes the following steps. In step S1110, if the vocoder is an adversarial generation network, loss calculation is performed on the pending voice and the voice to be processed to obtain the adversarial network loss value of the adversarial generation network.
[0171] When the vocoder adopts an adversarial generation network, the loss function of the adversarial generation network is shown in Equation (2). TIFF2025517777000006.tif16170
[0172] Here, D(x) identifies the true sample. Since the closer the identification result is to 1 here, the better, the loss function is log(D(x)). z is a random input, and G(z) represents the generated sample. For the generated sample, it is desirable that the identification result D(G(z)) of the discriminator is closer to 0, that is, the total numerical value is maximized. Therefore, the overall expression form is as shown in Equation (2).
[0173] Therefore, according to Equation (2), loss calculation can be performed on the pending voice and the voice to be processed to obtain the adversarial network loss value of the adversarial generation network.
[0174] In step S1120, loss calculation is performed on the pending voice and the voice to be processed to obtain a voice feature loss value, and the adversarial network loss value and the voice feature loss value are weighted and added to obtain a third loss value.
[0175] Furthermore, according to Equation (1), loss calculation can also be performed on the pending voice and the voice to be processed to obtain a voice feature loss value.
[0176] Furthermore, based on the experience value, weights corresponding to the adversarial network loss value and the voice feature loss value can be set, and the adversarial network loss value and the voice feature loss value can be weighted and added to obtain a third loss value.
[0177] Note that when the vocoder adopts another network or model, the corresponding loss value can be calculated as the third loss value only according to Equation (1).
[0178] In this exemplary embodiment, corresponding loss value calculation methods are set based on different vocoder contents, with stronger correspondence, ensuring the accuracy of the training results of different types of vocoders and further ensuring the reliability of target voice generation.
[0179] When the third loss value reaches a stable value and does not decrease further, it indicates that the vocoder has been trained to converge and the vocoder can be put into the application stage. Therefore, it can be determined that the pending voice is the target voice.
[0180] Hereinafter, the voice generation method in the embodiments of the present disclosure will be described in detail with reference to the application scenario.
[0181] FIG. 12 shows a framework schematic diagram of a voice generation model in an application scenario. As shown in FIG. 12, the VQ-VAE model includes an encoder module 1210, a vector quantization module 1220, and a decoder module 1230.
[0182] First, the voice feature vector of the voice is input into the encoder module 1210 of the VQ-VAE model.
[0183] Here, the voice to be processed may be the voice to be converted for voice timbre conversion. Correspondingly, the voice feature vector of the voice to be processed may be a mel-frequency spectrum feature vector extracted based on the voice to be processed.
[0184] Furthermore, the encoder module of the voice generation model non-linearly transforms the voice feature vector to obtain a high-dimensional voice encoding vector. This voice encoding vector is represented by z 1:N and can be represented as such.
[0185] Here, the encoder module in the VQ-VAE model may be composed of a CNN and an LSTM.
[0186] Then, the vector quantization module 1220 of the voice generation model quantizes the voice encoding vector to obtain a voice quantization sequence.
[0187] After the encoder module of the VQ-VAE model obtains the voice encoding vector z 1:N the vector quantization in the VQ-VAE model quantizes the high-dimensional voice encoding vector into a voice quantization sequence by the TIFF2025517777000007.tif9170 module. The voice quantization sequence may be represented by TIFF2025517777000008.tif8170.
[0188] Based on the codebook in the vector quantization module of the voice generation model, the voice encoding vector is quantized by the nearest neighbor search algorithm to obtain a voice quantization sequence.
[0189] Specifically, based on the codebook updated in the VQ-VAE model, the continuous voice encoding vectors are quantized into a discrete voice quantization sequence by the nearest neighbor search algorithm.
[0190] When updating the codebook, the codebook identifier of the codebook for each frame can be obtained, and the comparison result can be obtained by comparing the codebook identifiers.
[0191] Since the speech feature vectors of each frame correspond to their respective codebooks, for example, the codebooks for the speech feature vectors of 5 frames may be book1, book1, book1, book2, and book2.
[0192] Considering that the text feature vector may be processed by a sequence-to-sequence model later, in order to obtain a sequence representation adapted to the text feature vector, the codebook identifiers can be compared to obtain a comparison result.
[0193] Here, the comparison result can reflect whether two or more consecutive codebook identifiers are the same.
[0194] Based on the comparison result, the codebooks are merged to obtain an updated codebook.
[0195] When the comparison result reflects that two or more consecutive codebook identifiers are the same, the same codebook identifiers can be merged into one by querying with the nearest neighbor search algorithm.
[0196] For example, when the codebooks for the speech feature vectors of 5 frames are book1, book1, book1, book2, and book2, the codebook identifiers can be merged into book1 and book2 to obtain an updated codebook.
[0197] Based on the updated codebook, the speech coding vector is quantized by the nearest neighbor search algorithm to obtain a speech quantization sequence.
[0198] Here, each updated codebook is a K*D-dimensional codebook maintained in the VQ-VAE model.
[0199] For example, each codebook contains K D-dimensional coding vectors e 1 、e 2 、···、eK may include. Using the encoding layer of the VQ-VAE model to encode an audio feature vector of H’*W’*D dimensions, and further for each D-dimensional vector in the audio feature vector of H’*W’*D dimensions, the encoding vector e closest to the D-dimensional vector i can be found in the codebook respectively. Here, the encoding vector e i is one vector in the codebook, and by representing the D-dimensional vector with the index of the encoding vector e i a discrete vector of H’*W’ dimensions is obtained. Here, K, D, H’ and W’ represent dimensions respectively.
[0200] Furthermore, based on a preset discrete encoding method, the discrete vector of H’*W’ dimensions is converted into an audio quantization sequence.
[0201] Here, the preset discrete encoding method may be one-hot encoding or other types of encoding methods, and the exemplary embodiments are not particularly limited thereto.
[0202] Specifically, using the codebook of the one-hot encoding method, by the look-up table method, the discrete vector of H′*W′ dimensions is converted into another discrete encoding vector of H′*W′ dimensions encoded by the codebook of the one-hot encoding method, and further an audio quantization sequence is obtained based on the converted discrete encoding vector of H′*W′ dimensions.
[0203] For example, after converting a 3*3 discrete vector into a 3*3 discrete encoding vector encoded by another codebook of the one-hot encoding method, a 1*9 audio quantization sequence can be obtained based on each element in the converted 3*3 discrete encoding vector.
[0204] In addition, a speaker vector of the speaker who spoke or pronounced the audio to be processed may be obtained.
[0205] Obtain a speaker identifier corresponding to the voice to be processed, and determine the correspondence relationship between the speaker identifier and the speaker vector, where the correspondence relationship is determined based on the voice generation model.
[0206] Here, the speaker identifier can uniquely represent the identifier information of the speaker who uttered or pronounced the voice to be processed.
[0207] Note that a table for storing the correspondence relationship between the speaker identifier and the speaker vector can be simultaneously maintained by the first loss value calculated between the self-recovery voice output by the decoder in the voice generation model and the voice to be processed.
[0208] Query the speaker vector corresponding to the speaker identifier based on the correspondence relationship.
[0209] In the table for storing the correspondence relationship between the speaker identifier and the speaker vector, the speaker vector corresponding to the speaker identifier can be queried based on the speaker identifier.
[0210] Finally, the voice quantization sequence and the speaker vector are non-linearly transformed by the decoder module of the voice generation model to obtain the self-recovery voice.
[0211] The decoder module of the voice generation model can receive the quantized voice quantization sequence, obtain the speaker vector, then add the voice quantization sequence and the speaker vector, and finally reduce them by non-linear transformation to obtain the self-recovery voice.
[0212] Perform loss calculation on the voice to be processed and the self-recovery voice to obtain the first loss value, and determine the voice encoding vector as the language unit vector based on the first loss value.
[0213] After the voice generation model outputs the self-recovery voice, the first loss value of the voice generation model can be calculated for the voice to be processed and the self-recovery voice.
[0214] Specifically, the first loss value may be calculated by an L2 norm loss function. The L2 norm loss function is shown in Equation (1).
[0215] When the first loss value calculated according to Equation (1) reaches a stable value and does not decrease any further, it indicates that the training of the voice generation model has been completed.
[0216] In this case, the voice encoding vector output from the encoder module in the voice generation model can be determined as an unsupervised language unit vector. The language unit vector can be represented by k i and can be expressed as such.
[0217] Unsupervised learning can discover or extract useful information representations from its own data. The unsupervised algorithm of VQ-VAE in this application scenario can extract discrete information representations from data in different formats.
[0218] Since such discrete representation units are very close to the phonemes in language texts, it is very suitable to use such unsupervised, discrete language units as the input of the end-to-end language synthesis model.
[0219] Also, this also well matches the problems to be solved.
[0220] In order to combine the voice synthesis task and the voice timbre conversion task into one system, the common input of this system, that is, the phonemes extracted from the text and the unsupervised language units extracted from the VQ-VAE model, can be found.
[0221] In FIG. 13, the sequence-to-sequence model may include an encoder module 1240, an attention mechanism 1250, a decoder module 1260, and a post-processing network 1270.
[0222] In order to fuse the speech synthesis task and the voice color conversion task, it is also possible to obtain a text feature vector.
[0223] After obtaining the natural text, the phoneme sequence of the natural text can be extracted as a text feature vector.
[0224] After obtaining the text feature vector, a processing target feature vector can be determined based on the text feature vector and the language unit vector.
[0225] In speaker adaptation without a teacher, the text feature vector is determined as the processing target feature vector, and the text synthesis effect can be realized by the subsequent sequence-to-sequence model and vocoder.
[0226] Voice color conversion determines the voice unit vector as the processing target feature vector, and the achievement of the voice color conversion task can be realized by the subsequent sequence-to-sequence model and vocoder.
[0227] In order to improve the effect of the voice color conversion task, the text feature vector and the language unit vector can be added to obtain the processing target feature vector.
[0228] In addition, in the process of obtaining the language unit vector, since the codebook in the vector quantization module of the voice generation model has already been updated, the language unit vector and the text feature vector are very compatible, so the text feature vector and the language unit vector can be directly added.
[0229] When adding the text feature vector and the language unit vector to obtain the processing target feature vector, since the processing target feature vector corresponds to the expression obtained by adding the text modality to the language modality, the processing target feature vector at this time is an enhanced data representation. Based on this, the performance of the voice color conversion task realized by the processing target feature vector is better.
[0230] Furthermore, an acoustic vector to be processed of the feature vector to be processed is obtained, and the acoustic vector to be processed may be a mel-frequency spectrum feature vector.
[0231] The feature vector to be processed and the acoustic vector to be processed are input into a sequence-to-sequence model, and a non-linear mapping is performed on the feature vector to be processed and the acoustic vector to be processed by an encoder module 1240 of the sequence-to-sequence model to obtain a spatial encoding vector.
[0232] The sequence-to-sequence model may be a sequence-to-sequence model based on an attention mechanism.
[0233] Specifically, the encoder module of the sequence-to-sequence model may include a feature embedding layer, a convolutional pre-processing network, a dense pre-processing network, a CBHG sub-model, and a downsampling convolutional layer.
[0234] First, the feature vector to be processed is encoded using a FeatureEmbedding layer and then input into a Convolutional Pre-net. By performing non-linear transformation on the encoded feature vector to be processed and the acoustic vector to be processed, the convergence and generalization capabilities of the sequence-to-sequence model based on the attention mechanism are improved. At the same time, by inputting the number of audio frames corresponding to the acoustic vector to be processed into a Dense Pre-net, the corresponding depth features are obtained. Then, the outputs of the Convolutional Pre-net and the Dense Pre-net are collectively input into a CBHG sub-model to extract the corresponding context features, and then this is input into a Down-sampling Convolution to reduce the computational amount and receptive field, and finally the corresponding spatial encoding vector is obtained.
[0235] Therefore, the encoder module of the sequence-to-sequence model performs a non-linear transformation on the feature vector to be processed and the acoustic vector to be processed, and maps them to a high-dimensional spatial encoding vector. The spatial encoding vector can be represented by h t and can be represented as such.
[0236] The spatial encoding vector and the speaker vector are added to obtain an alignment target vector, and an audio feature sequence is obtained.
[0237] To model a multi-speaker model, the attention mechanism 1250 of the sequence-to-sequence model can also receive the speaker vector as an input.
[0238] To input the speaker vector, the spatial encoding vector and the speaker vector can be added to obtain an alignment target vector.
[0239] Furthermore, since the attention mechanism 1250 is an autoregressive model, it can also obtain an audio feature sequence. The audio feature sequence can be represented by m t — 1 and can be represented as such. When t = 1, the audio feature sequence is initialized to a sequence of all zeros. At t = 2 and subsequent times, the audio feature sequence is the feedback sequence for the previous time of the decoder module 1260.
[0240] The attention mechanism 1250 of the sequence-to-sequence model aligns the alignment target vector and the audio feature sequence to obtain a context representation vector, and the decoder 1260 of the sequence-to-sequence model performs a non-linear mapping on the context representation vector to obtain a processed acoustic vector.
[0241] Normally, since the audio feature vector is longer than the alignment target vector, the alignment target vector and the audio feature sequence can be aligned to obtain a context representation vector.
[0242] Specifically, the method of aligning the alignment target vector and the speech feature sequence may perform dot product calculation on the alignment target vector and the speech feature sequence.
[0243] In addition, the context representation vector obtained by aligning the alignment target vector and the speech feature sequence reflects the context relationship of the context and ensures the effect of speech generation.
[0244] Furthermore, the decoder module 1260 of the sequence-to-sequence model mainly obtains a processed acoustic vector by returning the context representation vector obtained by alignment according to the alignment target vector and the speech feature sequence to the original speech acoustic feature space by non-linear mapping. Therefore, the processed acoustic vector may be a mel-frequency spectrum, and the processed acoustic vector may be represented by m.
[0245] After the sequence-to-sequence model outputs the processed acoustic vector, the second loss value between the acoustic vector to be processed and the processed acoustic vector can be calculated according to formula (1).
[0246] When the second loss value calculated according to formula (1) reaches a stable value and does not decrease any further, it indicates that the training of the sequence-to-sequence model has been completed.
[0247] In this case, the processed acoustic vector output by the sequence-to-sequence model converged by training can be determined as the acoustic feature vector.
[0248] The speech acoustic feature of the acoustic feature vector is extracted through the post-processing network 1270, and by inputting the speech acoustic feature into the vocoder 1280, the vocoder 1280 outputs the pending speech corresponding to the acoustic vector to be processed or the text feature vector.
[0249] Among them, the post-processing network 1270 is mainly set to generate higher-precision voice acoustic characteristics. The voice acoustic features are It may be represented by TIFF2025517777000009.tif7170.
[0250] Among them, the vocoder may be a WaveNet model, a Griffin-Lim algorithm, a GAN network, etc., and the exemplary embodiments are not particularly limited thereto.
[0251] The voice acoustic features extracted by the post-processing network 1270 can be processed by the vocoder 1280 to output pending voice.
[0252] When performing a voice synthesis task, the pending voice may be a voice synthesized based on a text feature vector. When performing a voice timbre conversion task, the pending voice may be a voice converted based on the voice to be processed.
[0253] If the vocoder 1280 is an adversarial generation network, loss calculation is performed on the pending voice and the voice to be processed to obtain the adversarial network loss value of the adversarial generation network.
[0254] When the vocoder 1280 adopts an adversarial generation network, the loss function of the adversarial generation network is shown in Equation (2). Therefore, according to Equation (2), loss calculation can be performed on the pending voice and the voice to be processed to obtain the adversarial network loss value of the adversarial generation network.
[0255] Loss calculation is performed on the pending voice and the voice to be processed to obtain a voice feature loss value, and the adversarial network loss value and the voice feature loss value are weighted and added to obtain a third loss value.
[0256] Furthermore, according to formula (1), loss calculation can also be performed on the pending voice and the voice to be processed to obtain a voice feature loss value.
[0257] Furthermore, based on empirical values, weights corresponding to the adversarial network loss value and the voice feature loss value can be set, and the adversarial network loss value and the voice feature loss value can be weighted and added to obtain a third loss value.
[0258] Note that when the vocoder 1280 adopts another network or model, the corresponding loss value can be calculated as the third loss value only according to formula (1).
[0259] When the third loss value reaches a stable value and does not decrease any further, it indicates that the vocoder 1280 has already been trained to convergence, and since the vocoder 1280 can be put into the application stage, it can be determined that the pending voice is the target voice.
[0260] According to the voice generation method in the exemplary embodiments of the present disclosure, by obtaining a voice feature vector and a text feature vector, voice and text can be received as inputs, the voice synthesis task and the voice timbre conversion task can be fused, multimodal modeling can be performed, and the performance of the voice synthesis task and the voice timbre conversion task can be improved. Furthermore, in the case of a small amount of data, a voice feature vector and a text feature vector are obtained, a strategy for multiple types of timbre clones is provided, the effect of timbre cloning with a small amount of data is improved, the training difficulty and training time of multiple types of models are reduced, and the timbre cloning method in multiple application scenarios is supported.
[0261] In addition, the voice generation method in this application scenario can obtain good effects through fine-tuning in teacher-based speaker adaptation, and the performance can also be improved as the amount of data increases in speaker-independent speaker adaptation.
[0262] In the teacher-assisted voice clone, such performance improvement is attributed to the use of the teacherless language unit as a means of data augmentation, thereby enabling the support of the model's performance with fewer data samples.
[0263] Also, in both teacher-assisted and teacherless voice clones, the teacherless language unit helps the attention mechanism learn more robust alignment results, thereby improving the model's performance with fewer samples.
[0264] In the voice color conversion task, the performance of this method is superior to that of a single-task VC model and has a larger improvement margin.
[0265] Therefore, overall, the multimodal voice clone system provided in this application scenario is superior to single-task TTS or VC models in various scenarios and has strong practical value.
[0266] Also, in an exemplary embodiment of the present disclosure, a voice generation device is further provided. FIG. 13 is a schematic diagram showing the configuration of the voice generation device. As shown in FIG. 13, the voice generation device 1300 may include a data acquisition module 1310, a vector determination module 1320, and a voice generation module 1330.
[0267] The data acquisition module 1310 is configured to acquire the voice feature vector of the voice to be processed and input the voice feature vector into a voice generation model to obtain a language unit vector.
[0268] The vector determination module 1320 is configured to acquire a text feature vector and determine a feature vector to be processed based on the text feature vector and the language unit vector.
[0269] The voice generation module 1330 is configured to input the feature vector to be processed into a sequence-to-sequence model to obtain an acoustic feature vector, and input the acoustic feature vector into a vocoder to obtain a target voice corresponding to the voice to be processed or the text feature vector.
[0270] In an exemplary embodiment of the present invention, obtaining a language unit vector by inputting the acoustic feature vector into a voice generation model includes: inputting the acoustic feature vector into a voice generation model so that the voice generation model outputs a voice encoding vector and a self-reconstructed voice; and calculating a loss for the voice to be processed and the self-reconstructed voice to obtain a first loss value, and determining the voice encoding vector as a language unit vector based on the first loss value.
[0271] In an exemplary embodiment of the present invention, inputting the acoustic feature vector into a voice generation model so that the voice generation model outputs a voice encoding vector and a self-reconstructed voice includes: inputting the acoustic feature vector into a voice generation model, and non-linearly transforming the acoustic feature vector by an encoder module of the voice generation model to obtain a voice encoding vector; quantizing the voice encoding vector by a vector quantization module of the voice generation model to obtain a voice quantization sequence, and obtaining a speaker vector corresponding to the voice to be processed; and non-linearly transforming the voice quantization sequence and the speaker vector by a decoder module of the voice generation model to obtain a self-reconstructed voice.
[0272] In an exemplary embodiment of the present invention, obtaining a speaker vector corresponding to the voice to be processed includes: obtaining a speaker identifier corresponding to the voice to be processed, and determining a correspondence between the speaker identifier and the speaker vector, where the correspondence is determined based on the voice generation model. querying, based on the correspondence relationship, the speaker vector corresponding to the speaker identifier;
[0273] In an exemplary embodiment of the present invention, obtaining the voice quantization sequence by quantizing the voice encoding vector by the vector quantization module of the voice generation model includes quantizing the voice encoding vector by a nearest neighbor search algorithm based on a codebook in the vector quantization module of the voice generation model to obtain a voice quantization sequence.
[0274] In an exemplary embodiment of the present invention, obtaining the voice quantization sequence by quantizing the voice encoding vector by the nearest neighbor search algorithm includes obtaining an updated codebook by updating the codebook; and obtaining the voice quantization sequence by quantizing the voice encoding vector by a nearest neighbor search algorithm based on the updated codebook.
[0275] In an exemplary embodiment of the present invention, obtaining an updated codebook by updating the codebook includes obtaining a codebook identifier of the codebook for each frame, comparing the codebook identifiers to obtain a comparison result; and obtaining an updated codebook by merging the codebooks according to the comparison result.
[0276] In an exemplary embodiment of the present invention, obtaining an acoustic feature vector by inputting the feature vector to be processed into a sequence-to-sequence model includes obtaining a voice vector to be processed of the feature vector to be processed, and inputting the feature vector to be processed and the voice vector to be processed into the sequence-to-sequence model, so that the sequence-to-sequence model outputs a processed voice vector; Performing loss calculation on the acoustic vector to be processed and the processed acoustic vector to obtain a second loss value, and determining the processed acoustic vector as an acoustic feature vector based on the second loss value.
[0277] In an exemplary embodiment of the present invention, inputting the feature vector to be processed and the acoustic vector to be processed into a sequence-to-sequence model, and the sequence-to-sequence model outputting a processed acoustic vector Inputting the feature vector to be processed and the acoustic vector to be processed into a sequence-to-sequence model, and obtaining a spatial encoding vector by non-linearly mapping the feature vector to be processed and the acoustic vector to be processed by an encoder module of the sequence-to-sequence model Adding the spatial encoding vector and the speaker vector to obtain an alignment target vector, and obtaining an audio feature sequence Aligning the alignment target vector and the audio feature sequence by an attention mechanism of the sequence-to-sequence model to obtain a context representation vector, and obtaining a processed acoustic vector by non-linearly mapping the context representation vector by a decoder of the sequence-to-sequence model.
[0278] In an exemplary embodiment of the present invention, determining a feature vector to be processed based on the text feature vector and the language unit vector Determining the text feature vector or the speech unit vector as the feature vector to be processed, or Including obtaining a feature vector to be processed by adding the text feature vector and the language unit vector.
[0279] In an exemplary embodiment of the present invention, inputting the acoustic feature vector into a vocoder to obtain a target speech corresponding to the speech to be processed or the text feature vector Extract the acoustic features of the acoustic feature vectors through a post-processing network, and input the acoustic features into a vocoder, so that the vocoder outputs a pending voice corresponding to the voice to be processed or the text feature vector. Performing loss calculation on the pending voice and the voice to be processed to obtain a third loss value, and determining the pending voice as a target voice based on the third loss value.
[0280] In an exemplary embodiment of the present invention, obtaining a third loss value by performing loss calculation on the pending voice and the voice to be processed includes: When the vocoder is an adversarial generation network, performing loss calculation on the pending voice and the voice to be processed to obtain an adversarial network loss value of the adversarial generation network. Performing loss calculation on the pending voice and the voice to be processed to obtain a voice feature loss value, and obtaining a third loss value by weighted addition of the adversarial network loss value and the voice feature loss value.
[0281] Specific details of the above voice generation device 1300 are described in detail in the corresponding voice generation method, so the description is omitted here.
[0282] In addition, in the above, some modules or units of the voice generation device 1300 are mentioned, but such division is not essential. In fact, according to the embodiments of the present disclosure, the features and functions of the above two or more modules or units may be embodied in one module or unit. Conversely, the features and functions of the above one module or unit may be embodied by a plurality of modules or units.
[0283] Also, in an exemplary embodiment of the present disclosure, an electronic device capable of implementing the above method is further provided.
[0284] Next, with reference to FIG. 14, an electronic device 1400 according to an embodiment of the present invention will be described. The electronic device 1400 shown in FIG. 14 is merely an example and does not impose any limitations on the functions and usage scope of the embodiments of the present invention.
[0285] As shown in FIG. 14, the electronic device 1400 is represented as a general-purpose computing device. The components of the electronic device 1400 may include, but are not limited to, the at least one processing unit 1410, the at least one storage unit 1420, a bus 1430 that connects different system components (including the storage unit 1420 and the processing unit 1410), and a display unit 1440.
[0286] Here, program code is stored in the storage unit, and the program code may be executed by the processing unit 1410, whereby the processing unit 1410 executes the steps of various exemplary embodiments of the present invention described in the "Exemplary Method" section of this specification.
[0287] The method includes: acquiring a voice feature vector of a voice to be processed, and inputting the voice feature vector into a voice generation model to obtain a language unit vector; acquiring a text feature vector, and determining a feature vector to be processed based on the text feature vector and the language unit vector; inputting the feature vector to be processed into a sequence-to-sequence model to obtain an acoustic feature vector, and inputting the acoustic feature vector into a vocoder to obtain a target voice corresponding to the voice to be processed or the text feature vector.
[0288] Optionally, acquiring a language unit vector by inputting the voice feature vector into a voice generation model includes: inputting the voice feature vector into the voice generation model such that the voice generation model outputs a voice coding vector and a self-recovering voice. Performing loss calculation on the voice to be processed and the self-loopback voice to obtain a first loss value, and determining the voice encoding vector as a language unit vector based on the first loss value.
[0289] Inputting the voice feature vector into a voice generation model so that the voice generation model outputs a voice encoding vector and a self-loopback voice is selectable, Inputting the voice feature vector into a voice generation model, and non-linearly transforming the voice feature vector by an encoder module of the voice generation model to obtain a voice encoding vector. Quantizing the voice encoding vector by a vector quantization module of the voice generation model to obtain a voice quantization sequence, and obtaining a speaker vector corresponding to the voice to be processed. Non-linearly transforming the voice quantization sequence and the speaker vector by a decoder module of the voice generation model to obtain a self-loopback voice.
[0290] Obtaining a speaker vector corresponding to the voice to be processed is selectable, Obtaining a speaker identifier corresponding to the voice to be processed, and determining a correspondence relationship between the speaker identifier and the speaker vector, where the correspondence relationship is determined based on the voice generation model. Querying the speaker vector corresponding to the speaker identifier based on the correspondence relationship.
[0291] Quantizing the voice encoding vector by a vector quantization module of the voice generation model to obtain a voice quantization sequence is selectable, Including quantizing the voice encoding vector by a nearest neighbor search algorithm based on a codebook in the vector quantization module of the voice generation model to obtain a voice quantization sequence.
[0292] Optionally, quantizing the voice encoding vector by the nearest neighbor search algorithm to obtain a voice quantization sequence includes: updating the codebook to obtain an updated codebook; quantizing the voice encoding vector by the nearest neighbor search algorithm based on the updated codebook to obtain a voice quantization sequence.
[0293] Optionally, updating the codebook to obtain an updated codebook includes: obtaining a codebook identifier of the codebook for each frame, comparing the codebook identifiers to obtain a comparison result; merging the codebooks according to the comparison result to obtain an updated codebook.
[0294] Optionally, inputting the feature vector to be processed into a sequence-to-sequence model to obtain an acoustic feature vector includes: obtaining a voice vector to be processed of the feature vector to be processed, inputting the feature vector to be processed and the voice vector to be processed into the sequence-to-sequence model, so that the sequence-to-sequence model outputs a processed voice vector; calculating a loss for the voice vector to be processed and the processed voice vector to obtain a second loss value, and determining the processed voice vector as an acoustic feature vector based on the second loss value.
[0295] Optionally, inputting the feature vector to be processed and the voice vector to be processed into the sequence-to-sequence model, so that the sequence-to-sequence model outputs a processed voice vector includes: Input the processing target feature vector and the processing target acoustic vector into a sequence-to-sequence model, and non-linearly map the processing target feature vector and the processing target acoustic vector by an encoder module of the sequence-to-sequence model to obtain a spatial encoding vector, and add the spatial encoding vector and the speaker vector to obtain an alignment target vector, and obtain an audio feature sequence, and align the alignment target vector and the audio feature sequence by an attention mechanism of the sequence-to-sequence model to obtain a context representation vector, and non-linearly map the context representation vector by a decoder of the sequence-to-sequence model to obtain a processed acoustic vector, including.
[0296] Selectable, determining a processing target feature vector based on the text feature vector and the language unit vector, determine the text feature vector or the speech unit vector as the processing target feature vector, or add the text feature vector and the language unit vector to obtain a processing target feature vector.
[0297] Selectable, inputting the acoustic feature vector into a vocoder to obtain a target voice corresponding to the processing target voice or the text feature vector, extract the voice acoustic features of the acoustic feature vector through a post-processing network, and input the voice acoustic features into the vocoder, so that the vocoder outputs a pending voice corresponding to the processing target voice or the text feature vector, and perform loss calculation on the pending voice and the processing target voice to obtain a third loss value, and determine the pending voice as the target voice based on the third loss value, including.
[0298] Selectable, calculating a loss for the pending voice and the voice to be processed to obtain a third loss value, when the vocoder is an adversarial generation network, calculating a loss for the pending voice and the voice to be processed to obtain an adversarial network loss value of the adversarial generation network, calculating a loss for the pending voice and the voice to be processed to obtain a voice feature loss value, and weighted adding the adversarial network loss value and the voice feature loss value to obtain a third loss value, including.
[0299] By the above method, by obtaining a voice feature vector and a text feature vector, voice and text can be received as inputs, fusing a voice synthesis task and a voice timbre conversion task, performing multimodal modeling, and improving the performance of the voice synthesis task and the voice timbre conversion task. Furthermore, in the case of a small amount of data, a voice feature vector and a text feature vector are obtained, a strategy for multiple types of timbre clones is provided, the effect of timbre cloning with a small amount of data is improved, the training difficulty and training time of multiple models are reduced, and a timbre cloning method in multiple application scenarios is supported.
[0300] The storage unit 1420 may include a readable medium in the form of a volatile storage unit such as a random access memory unit (RAM) 1421 and / or a cache storage unit 1422, and may further include a read-only memory unit (ROM) 1423.
[0301] The storage unit 1420 may further include a program / utilities 1424 having a set (at least one) of program modules 1425, such program modules 1425 including, but not limited to, an operating system, one or more application programs, other program modules, and program data, and the realization of a network environment may be included in each or a certain combination of these examples.
[0302] Bus 1430 represents any one or more of several types of bus structures, such as a memory unit bus or memory unit controller, a peripheral device bus, an accelerated graphics port, and a processor or local bus using any of various bus architectures.
[0303] Electronic device 1400 may communicate with one or more external devices 1600 (such as a keyboard, a pointing device, a Bluetooth device, etc.), may communicate with one or more devices through which a user can interact with the electronic device 1400, and / or may communicate with any device (such as a router, a modem, etc.) through which the electronic device 1400 can communicate with one or more other computing devices. Such communication can be performed via an input / output interface (I / O) 1450. Also, electronic device 1400 can communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN) and / or a public network, such as the Internet) via a network adapter 1460. As shown in the figure, network adapter 1460 communicates with other modules of electronic device 1400 via bus 1430. Although not shown, other hardware and / or software modules may be used in combination with electronic device 1400, including, but not limited to, microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0304] From the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein may be implemented by software or by combining the necessary hardware with software. Therefore, the technical solution according to the embodiments of the present disclosure may be embodied in the form of a software product, and the software product may be stored on a non-volatile storage medium (which may be a CD-ROM, a USB disk, a portable hard disk, etc.) or on a network, and may include several instructions for a computing device (which may be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0305] In an exemplary embodiment of the present disclosure, there is further provided a computer-readable storage medium storing a program product capable of implementing the above method in this specification. In some possible embodiments, each aspect of the present invention includes program code, and when the program product is executed on a terminal device, the program code may be implemented in the form of a program product used to cause the terminal device to execute the steps according to various exemplary embodiments of the present invention described in the "exemplary method" section of this specification.
[0306] The method includes: acquiring a voice feature vector of the voice to be processed, and inputting the voice feature vector into a voice generation model to obtain a language unit vector; acquiring a text feature vector, and determining a feature vector to be processed based on the text feature vector and the language unit vector; inputting the feature vector to be processed into a sequence-to-sequence model to obtain an acoustic feature vector, and inputting the acoustic feature vector into a vocoder to obtain a target voice corresponding to the voice to be processed or the text feature vector.
[0307] Optionally, acquiring the voice feature vector and inputting the voice feature vector into a voice generation model to obtain a language unit vector includes: Inputting the voice feature vector into a voice generation model such that the voice generation model outputs a voice encoding vector and a self-recovering voice; Performing loss calculation on the voice to be processed and the self-recovering voice to obtain a first loss value, and determining the voice encoding vector as a language unit vector based on the first loss value.
[0308] Optionally, inputting the voice feature vector into a voice generation model such that the voice generation model outputs a voice encoding vector and a self-recovering voice includes: Inputting the voice feature vector into a voice generation model, non-linearly transforming the voice feature vector by an encoder module of the voice generation model to obtain a voice encoding vector; Quantizing the voice encoding vector by a vector quantization module of the voice generation model to obtain a voice quantization sequence, and obtaining a speaker vector corresponding to the voice to be processed; Non-linearly transforming the voice quantization sequence and the speaker vector by a decoder module of the voice generation model to obtain a self-recovering voice.
[0309] Optionally, obtaining a speaker vector corresponding to the voice to be processed includes: Obtaining a speaker identifier corresponding to the voice to be processed, and determining a correspondence between the speaker identifier and the speaker vector, where the correspondence is determined based on the voice generation model; Querying the speaker vector corresponding to the speaker identifier based on the correspondence.
[0310] Optionally, quantizing the voice encoding vector by a vector quantization module of the voice generation model to obtain a voice quantization sequence includes: Quantizing the speech encoding vector by a nearest neighbor search algorithm based on a codebook in the vector quantization module of the speech generation model to obtain a speech quantization sequence.
[0311] Optionally, quantizing the speech encoding vector by the nearest neighbor search algorithm to obtain a speech quantization sequence includes updating the codebook to obtain an updated codebook, and quantizing the speech encoding vector by a nearest neighbor search algorithm based on the updated codebook to obtain a speech quantization sequence.
[0312] Optionally, updating the codebook to obtain an updated codebook includes obtaining a codebook identifier of the codebook for each frame, comparing the codebook identifiers to obtain a comparison result, and merging the codebooks according to the comparison result to obtain an updated codebook.
[0313] Optionally, inputting the feature vector to be processed into a sequence-to-sequence model to obtain an acoustic feature vector includes obtaining a processing target acoustic vector of the feature vector to be processed, inputting the feature vector to be processed and the processing target acoustic vector into the sequence-to-sequence model, so that the sequence-to-sequence model outputs a processed acoustic vector, and performing loss calculation on the processing target acoustic vector and the processed acoustic vector to obtain a second loss value, and determining the processed acoustic vector as an acoustic feature vector based on the second loss value.
[0314] Optionally, inputting the feature vector to be processed and the processing target acoustic vector into the sequence-to-sequence model, so that the sequence-to-sequence model outputs a processed acoustic vector includes Input the processing target feature vector and the processing target acoustic vector into a sequence-to-sequence model, and non-linearly map the processing target feature vector and the processing target acoustic vector by the encoder module of the sequence-to-sequence model to obtain a spatial encoding vector; Add the spatial encoding vector and the speaker vector to obtain an alignment target vector, and obtain an audio feature sequence; Align the alignment target vector and the audio feature sequence by the attention mechanism of the sequence-to-sequence model to obtain a context representation vector, and non-linearly map the context representation vector by the decoder of the sequence-to-sequence model to obtain a processed acoustic vector, including.
[0315] Being selectable, determining a processing target feature vector based on the text feature vector and the language unit vector includes: Determining the text feature vector or the speech unit vector as the processing target feature vector, or Adding the text feature vector and the language unit vector to obtain a processing target feature vector.
[0316] Being selectable, inputting the acoustic feature vector into a vocoder to obtain a target voice corresponding to the processing target voice or the text feature vector includes: Extracting the voice acoustic features of the acoustic feature vector through a post-processing network, and inputting the voice acoustic features into the vocoder, so that the vocoder outputs a pending voice corresponding to the processing target voice or the text feature vector; Performing loss calculation on the pending voice and the processing target voice to obtain a third loss value, and determining the pending voice as the target voice based on the third loss value.
[0317] It is possible to perform loss calculation on the pending voice and the voice to be processed to obtain a third loss value, when the vocoder is an adversarial generation network, performing loss calculation on the pending voice and the voice to be processed to obtain an adversarial network loss value of the adversarial generation network, performing loss calculation on the pending voice and the voice to be processed to obtain a voice feature loss value, and weighted adding the adversarial network loss value and the voice feature loss value to obtain a third loss value, including.
[0318] By the above method, by obtaining a voice feature vector and a text feature vector, voice and text can be received as inputs, fusing a voice synthesis task and a voice timbre conversion task, performing multimodal modeling, and improving the performance of the voice synthesis task and the voice timbre conversion task. Furthermore, in the case of a small amount of data, a voice feature vector and a text feature vector are obtained, a strategy for multiple types of timbre clones is provided, the effect of timbre cloning with a small amount of data is improved, the training difficulty and training time of multiple models are reduced, and a timbre cloning method in multiple application scenarios is supported.
[0319] As shown in FIG. 15, a program product 1500 for realizing the above method according to an embodiment of the present invention employs a portable compact disc read only memory (CD-ROM) and includes program code, which can be executed on a terminal device, for example, a personal computer. However, the program product of the present invention is not limited thereto. In this specification, a readable storage medium may be any tangible medium including a program, and the program may be instructed to be used in or in combination with the use of a system, apparatus, or device.
[0320] The program product can use any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared or semiconductor system, apparatus or device. More specific examples (non-exhaustive list) of the readable storage medium include the following. An electrical connection having one or more conductors, a portable computer diskette, a hard disk, a random access memory (RAM), a read only memory (ROM), an erasable programmable read only memory (EPROM or flash memory), an optical fiber, a portable compact disk read only memory (CD-ROM), an optical storage device, a magnetic storage device or any combination of the above.
[0321] The computer readable signal medium can include a propagated data signal that includes readable program code, either in baseband or as part of a carrier wave. Such a propagated data signal can take various forms including, but not limited to, electromagnetic signals, optical signals or any suitable combination of the above. The readable signal medium can be any readable medium other than the readable storage medium, which can transmit, propagate or transfer a program used in conjunction with an instruction execution system, apparatus or device.
[0322] The program code included in the readable media can be transmitted using any appropriate medium including, but not limited to, wireless, wireline, optical fiber cable, RF, etc. or any suitable combination of the above.
[0323] The program code for executing the operation of the present invention can be described using one or a combination of multiple programming languages, including object-oriented programming languages such as Java and C++, and traditional procedural programming languages such as the C programming language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's computing device, as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, entirely on a remote computing device or server. In the scenario of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can also be made to an external computing device (for example, through the Internet using an Internet service provider).
[0324] After considering the specification and implementing the invention disclosed herein, those skilled in the art can easily conceive of other embodiments of the present disclosure. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, and these variations, uses, or adaptations follow the general principles of the present disclosure and include well-known common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The present specification and examples are merely illustrative, and the true scope and spirit of the present disclosure are indicated by the claims.
Claims
1. Obtaining a voice feature vector of a voice to be processed, inputting the voice feature vector into a voice generation model to obtain a language unit vector, obtaining a text feature vector, and determining a feature vector to be processed based on the text feature vector and the language unit vector, inputting the feature vector to be processed into a sequence-to-sequence model to obtain an acoustic feature vector, and inputting the acoustic feature vector into a vocoder to obtain a target voice corresponding to the voice to be processed or the text feature vector, A voice generation method characterized by the above.
2. The step of obtaining a language unit vector by inputting the voice feature vector into a voice generation model includes inputting the voice feature vector into the voice generation model so that the voice generation model outputs a voice encoding vector and a self-recovering voice, performing loss calculation on the voice to be processed and the self-recovering voice to obtain a first loss value, and determining the voice encoding vector as a language unit vector based on the first loss value. The voice generation method according to claim 1, characterized by the above.
3. The step of inputting the voice feature vector into the voice generation model so that the voice generation model outputs a voice encoding vector and a self-recovering voice includes inputting the voice feature vector into the voice generation model, non-linearly transforming the voice feature vector by an encoder module of the voice generation model to obtain a voice encoding vector, quantizing the voice encoding vector by a vector quantization module of the voice generation model to obtain a voice quantization sequence, and obtaining a speaker vector corresponding to the voice to be processed, and non-linearly transforming the voice quantization sequence and the speaker vector by a decoder module of the voice generation model to obtain a self-recovering voice. The voice generation method according to claim 2, characterized by the above.
4. The step of obtaining a speaker vector corresponding to the voice to be processed includes obtaining a speaker identifier corresponding to the voice to be processed, and determining a correspondence relationship between the speaker identifier and the speaker vector, where the correspondence relationship is determined based on the voice generation model, and querying the speaker vector corresponding to the speaker identifier based on the correspondence relationship. The voice generation method according to claim 3, characterized in that...
5. Obtaining a voice quantization sequence by quantizing the voice encoding vector by the vector quantization module of the voice generation model... includes obtaining a voice quantization sequence by quantizing the voice encoding vector by a nearest neighbor (KNN, K-Nearest Neighbor) search algorithm based on the codebook in the vector quantization module of the voice generation model. The voice generation method according to claim 3, characterized in that...
6. Obtaining a voice quantization sequence by quantizing the voice encoding vector by the nearest neighbor search algorithm... includes obtaining an updated codebook by updating the codebook. and obtaining a voice quantization sequence by quantizing the voice encoding vector by a nearest neighbor search algorithm based on the updated codebook. The voice generation method according to claim 5, characterized in that...
7. Obtaining an updated codebook by updating the codebook... includes obtaining a codebook identifier of the codebook for each frame, comparing the codebook identifiers to obtain a comparison result. and obtaining an updated codebook by merging the codebooks according to the comparison result. The voice generation method according to claim 6, characterized in that...
8. Obtaining an acoustic feature vector by inputting the processing target feature vector into a sequence-to-sequence model... includes obtaining a processing target acoustic vector of the processing target feature vector, inputting the processing target feature vector and the processing target acoustic vector into a sequence-to-sequence model, so that the sequence-to-sequence model outputs a processed acoustic vector. performing loss calculation on the processing target acoustic vector and the processed acoustic vector to obtain a second loss value, and determining the processed acoustic vector as an acoustic feature vector based on the second loss value. The voice generation method according to claim 3, characterized in that...
9. By inputting the processing target feature vector and the processing target acoustic vector into a sequence-to-sequence model, the sequence-to-sequence model outputs a processed acoustic vector... Input the processing target feature vector and the processing target acoustic vector into a sequence-to-sequence model, and non-linearly map the processing target feature vector and the processing target acoustic vector by an encoder module of the sequence-to-sequence model to obtain a spatial encoding vector. Add the spatial encoding vector and the speaker vector to obtain an alignment target vector, and obtain an audio feature sequence. Align the alignment target vector and the audio feature sequence by an attention mechanism of the sequence-to-sequence model to obtain a context representation vector, and non-linearly map the context representation vector by a decoder of the sequence-to-sequence model to obtain a processed acoustic vector, including: The method for generating speech according to claim 8, characterized in that.
10. Determining a processing target feature vector based on the text feature vector and the language unit vector includes: Determining the text feature vector or the speech unit vector as the processing target feature vector, or Adding the text feature vector and the language unit vector to obtain a processing target feature vector. The method for generating speech according to claim 1, characterized in that.
11. Inputting the acoustic feature vector into a vocoder to obtain a target speech corresponding to the processing target speech or the text feature vector includes: Extracting the speech acoustic features of the acoustic feature vector through a post-processing network, and inputting the speech acoustic features into the vocoder, so that the vocoder outputs a pending speech corresponding to the processing target speech or the text feature vector. Performing loss calculation on the pending speech and the processing target speech to obtain a third loss value, and determining the pending speech as the target speech based on the third loss value, including: The method for generating speech according to claim 1, characterized in that.
12. Performing loss calculation on the pending speech and the processing target speech to obtain a third loss value includes: When the vocoder is an adversarial generation network, performing loss calculation on the pending speech and the processing target speech to obtain an adversarial network loss value of the adversarial generation network. Performing loss calculation on the pending voice and the voice to be processed to obtain a voice feature loss value, and obtaining a third loss value by weighted addition of the adversarial network loss value and the voice feature loss value, The voice generation method according to claim 11, characterized in that.
13. A data acquisition module configured to acquire a voice feature vector of a voice to be processed and input the voice feature vector into a voice generation model to obtain a language unit vector, A vector determination module configured to acquire a text feature vector and determine a feature vector to be processed based on the text feature vector and the language unit vector, A voice generation module configured to input the feature vector to be processed into a sequence-to-sequence model to obtain an acoustic feature vector, and input the acoustic feature vector into a vocoder to obtain a target voice corresponding to the voice to be processed or the text feature vector, A voice generation device, characterized in that.
14. A computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the voice generation method according to any one of claims 1-12 is realized, A computer-readable storage medium, characterized in that.
15. An electronic device comprising a processor and a memory, The memory is configured to store executable instructions of the processor, The processor is configured to execute the executable instructions to execute the voice generation method according to any one of claims 1-12, An electronic device, characterized in that.
Citation Information
Patent Citations
Speech encoding method with plural code books
JP1993265496A
Voice conversion device, voice conversion method and program
JP2019109306A
Acoustic model learning device, voice synthesizer and program
JP2020060633A