Pitch-based speech conversion model training method and speech conversion system
The pitch-based voice conversion model addresses the challenge of pitch similarity in limited data scenarios by aligning and refining voice features, enhancing the quality of synthesized voices.
Patent Information
- Application Number
- JP2024228437
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-28
- Filing Date
- 2024-12-25
- Publication Date
- 2025-07-10
- Estimated Expiration
- 2044-12-25
AI Technical Summary
Existing voice conversion models struggle to maintain pitch similarity when trained with a small amount of sample data, resulting in significant differences between synthesized and real human voices.
A method for training a pitch-based voice conversion model using a pre-encoder, post-encoder, timing alignment module, decoder, and pitch extraction module, which involves extracting audio and pitch features, aligning them, and iteratively refining the model until convergence to enhance pitch similarity.
Improves the pitch similarity of synthesized voices by aligning and refining voice features, ensuring higher accuracy and quality even with limited sample data.
Smart Images

Figure 2025105553000001_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of voice conversion, and more particularly, to a method for training a voice conversion model based on pitch and a voice conversion system.
Background Art
[0002] Voice conversion technology is a technology that converts the voice of one person into the voice of another person using a voice conversion model. However, in order to obtain high-quality human voices, it is necessary to train the voice conversion model using a large amount of sample data, and by doing so, more realistic human voices can be obtained through voice conversion technology.
[0003] In the actual training process, it is often difficult to obtain a large amount of sample data, so the voice conversion model can only be trained using a small amount of sample data. However, since the expressive power of human natural voices is rich, the changes in human voices in terms of timbre and rhythm are large. On the other hand, the voices of people generated by the voice conversion model trained with a small amount of sample data have a certain difference from the voices of actual people.
[0004] In order to reduce the difference in voices, when training a voice conversion model using a small amount of sample data, a method of pre-training and fine-tuning the model can be adopted, that is, first pre-train the voice conversion model on a dataset of a large amount of audio data, and then fine-tune the voice conversion model using a small amount of sample data. However, in the case of a small amount of sample data, the problem that the pitch similarity between the voice of the person after conversion and the voice of the real person is low still exists.
Summary of the Invention
Problems to be Solved by the Invention
[0005] This is to reduce the problem that the pitch similarity between the voice of the person after conversion and the voice of the real person is low in the case of a small amount of sample data.
Means for Solving the Problem
[0006] In a first aspect, some embodiments of the present application provide a method for training a pitch-based voice conversion model applied to training a voice conversion model. The voice conversion model includes a pre-encoder, a post-encoder, a timing alignment module, a decoder, and a pitch extraction module. The method includes: Inputting a reference voice into the pre-encoder and the pitch extraction module, extracting an audio feature code by the pre-encoder, and extracting a pitch feature by the pitch extraction module; Performing feature stitching on the audio feature code and the pitch feature to obtain a voice stitching feature; Inputting a linear spectrum corresponding to the reference voice into the post-encoder to obtain a hidden variable of the audio; Aligning the timing sequences of the voice stitching feature and the hidden variable of the audio by the timing alignment module to obtain a converted voice code; Decoding the converted voice code by the decoder to obtain a converted voice; Calculating a training loss of the converted voice, and if the training loss is less than or equal to a training loss threshold, outputting a voice conversion model based on the current parameters of the training target model; if the training loss is greater than the training loss threshold, performing iterative training on the voice conversion model, where the training target model is a voice conversion model that has not been trained until convergence.
[0007] In some embodiments, the pitch extraction module includes an encoder layer, a filter layer, an intermediate layer, and a decoder layer. The encoder layer, the intermediate layer, and the decoder layer form a first encoding branch of the pitch extraction module, and the encoder layer, the filter layer, and the decoder layer form a second encoding branch of the pitch extraction module.
[0008] In some embodiments, the encoder layer includes an average pooling layer and a convolutional network, and the steps of extracting pitch features by the pitch extraction module are as follows: extracting a pitch feature vector of the reference audio by the convolutional network; performing downsampling on the pitch feature vector by the average pooling layer to obtain a pitch feature code; and decoding the pitch feature code by the decoder layer to obtain the pitch features.
[0009] In some embodiments, the convolutional network includes a convolutional block, the convolutional block includes a 2D convolutional layer, a batch normalization layer, and a relu function, and the steps of extracting a pitch feature vector of the reference audio by the convolutional network are as follows: extracting a deep audio vector by the 2D convolutional layer; performing fast convergence processing on the deep audio vector by the batch normalization layer to extract a convergent pitch feature from the deep audio vector; and adding a non-linear relationship to the convergent pitch feature by the relu function to obtain the pitch feature vector.
[0010] In some embodiments, a shortcut convolutional layer is installed between the input end and the output end of the convolutional network. Before the step of adding a non-linear relationship to the convergent pitch feature by the relu function, extracting a shortcut pitch feature by the shortcut convolutional layer; Step of stitching the shortcut pitch feature and the convergence pitch feature to obtain a pitch stitching feature; Further including the step of adding a non-linear relationship between the pitch stitching features by the relu function to obtain the pitch feature vector.
[0011] In some embodiments, the decoder layer includes a transposed convolutional layer and the convolutional network, and the step of performing a decoding operation on the pitch feature code by the decoder layer is: Performing a transposed convolution calculation on the pitch feature code by the transposed convolutional layer to obtain a transposed convolution feature vector; Performing decoding on the transposed convolution feature vector by the convolutional network to obtain the pitch feature.
[0012] In some embodiments, the step of aligning the voice stitching feature and the timing sequence of the hidden variable of the audio by the timing alignment module is: Obtaining a template voice sequence of the timing alignment module; Aligning the voice stitching feature and the timing sequence of the hidden variable of the audio based on the template voice sequence; Encoding the aligned voice stitching feature and the hidden variable of the audio to obtain a converted voice code.
[0013] In some embodiments, the voice conversion model further includes a style encoder. After the step of performing feature stitching on the audio feature code and the pitch feature, the method includes: Extracting the style feature of the reference voice by the style encoder; mapping the style feature to the voice stitching feature to update the voice stitching feature;
[0014] In some embodiments, the training loss includes a spectral loss, and the step of calculating the training loss of the converted voice includes: obtaining the spectral accuracy of the reference voice and obtaining the spectral accuracy of the converted voice; calculating the spectral loss based on the spectral accuracy of the reference voice and the spectral accuracy of the converted voice by the following formula:
Equation
[0015] In a second aspect, the present application provides a voice conversion system, the voice conversion system includes a voice conversion model, the voice conversion model is obtained by training based on the training method of the pitch-based voice conversion model described in the first aspect, the voice conversion model includes a pre-encoder, a post-encoder, a timing alignment module, a decoder and a pitch extraction module, where the pre-encoder is configured to extract the audio feature code of the reference voice, the pitch extraction module is configured to extract the pitch feature of the reference voice, the post-encoder is configured to generate a hidden variable of the audio based on the linear spectrum corresponding to the reference voice, The timing alignment module is configured to align the voice stitching features and the timing sequence of the hidden variables of the audio to obtain a converted voice code. The voice stitching features are obtained by combining the audio feature code and the pitch feature. The decoder is configured to perform decoding on the converted voice code to obtain the converted voice.
Advantages of the Invention
[0016] As can be seen from the above technical solution, the present application provides a method for training a pitch-based voice conversion model and a voice conversion system. The method is used for training a voice conversion model, and the voice conversion model includes a pre-encoder, a post-encoder, a timing alignment module, a decoder, and a pitch extraction module. The method includes inputting a reference voice into the pre-encoder and the pitch extraction module, so that the pre-encoder outputs an audio feature code and the pitch extraction module extracts a pitch feature. Subsequently, input the linear spectrum corresponding to the reference voice into the post-encoder to obtain the hidden variable of the audio. Input the voice stitching features obtained by stitching the audio feature code and the pitch feature and the hidden variable of the audio into the timing alignment module to obtain a converted voice code, and decode the converted voice code by the decoder to obtain the converted voice. Furthermore, calculate the training loss of the converted voice to judge the convergence degree of the voice conversion model. The present application extracts the pitch feature of the reference voice by the pitch extraction module, stitches it with the audio feature code, performs alignment, and makes the pitch feature of the converted voice match that of a real person's voice, so as to improve the pitch similarity of the converted voice when the voice sample is insufficient.
Brief Description of the Drawings
[0017] To more clearly explain the technical solution of the present application, the drawings necessary for use in the following embodiments will be briefly introduced below. Obviously, those skilled in the art can obtain other drawings based on these drawings without creative efforts.
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Modes for Carrying Out the Invention
[0018] Hereinafter, the present application will be described in detail in combination with embodiments with reference to the drawings. Unless there is no contradiction, the embodiments of the present application and the features in the embodiments can be combined with each other.
[0019] It should be noted that terms such as "first" and "second" in the specification, claims and the above drawings of the present application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0020] Voice conversion technology is a technology that converts the voice of one person into the voice of another person using a voice conversion model. For example, when a user drives a digital human generated according to their own image, if the digital human speaks in a voice other than the user's own voice to improve the interest of the interaction with the digital human, the user can convert their own voice into the voice of another person through voice conversion technology, thereby achieving the purpose of voice conversion.
[0021] In order to improve the accuracy of voice conversion and obtain high-quality converted voices of people, it is necessary to train a voice conversion model using a large amount of sample data, and by doing so, obtain a more realistic voice of a person through voice conversion technology.
[0022] However, in the actual training process, it is often difficult to obtain a large amount of sample data, so the voice conversion model can only be trained using a small amount of sample data. However, since the expressive power of human natural voices is rich, while the changes in the voices of people in terms of timbre and rhythm are large, there is a certain difference between the converted voices generated by the voice conversion model trained with a small amount of sample data and the voices of actual people.
[0023] When the content of the language is fixed, the timbre and pitch of the speaker are factors that affect the speaker's personality. When the amount of speaker data is large, voice conversion technology can express good conversion results. However, when the amount of speaker data is small, voice conversion technology has a low conversion effect in terms of the timbre and pitch of the speaker.
[0024] In some embodiments, when training a voice conversion model using a small amount of sample data, the difference between the synthesized voice of a person and the actual voice of a person can be reduced by adopting the methods of pre-training and fine-tuning the model. That is, first, pre-train the voice conversion model on a large dataset of audio data, and then fine-tune the voice conversion model using a small amount of sample data. However, in the case of a small amount of sample data, the problem that the pitch similarity between the synthesized voice of a person and the real voice of a person is low still exists.
[0025] In order to reduce the problem that the pitch similarity between the voice of a person after conversion and the real voice of a person is low in the case of a small amount of sample data, some embodiments of the present application provide a training method for a pitch-based voice conversion model. The training method is used to train a voice conversion model. Here, the usage stage of the voice conversion model includes a pre-training stage, a learning and training stage, and an application stage.
[0026] In the pre-training stage of the voice conversion model, the voice conversion model can be pre-trained by training the voice, so that the voice conversion model is equipped with basic voice conversion capabilities. In the process of pre-training, the training voice used can be obtained in various ways. For example, the corresponding audio file can be extracted from the video media resource as the training voice. In this example, if the duration of the audio file is long, the audio file can be split to obtain a plurality of training voice segments. Or, some embodiments of the present application can also perform voice synthesis on the preset text to obtain the training voice. The present application does not overly limit the acquisition method of the training voice.
[0027] Note that since the training voice is used to execute the pre-training process for the voice conversion model, the training voice should include the voice of at least one person, so that voice conversion training is performed on the voice of the person in the pre-training process of the voice conversion model.
[0028] The training method provided in the embodiments of the present application is used to train the learning and training stage of the voice conversion model. In the learning and training stage, the voice conversion model includes a pre-encoder, a post-encoder, a timing alignment module, a decoder, and a pitch extraction module. Here, the pre-encoder includes a normalizing flow layer, a projection layer, and a content encoder.
[0029] Note that in the embodiments of the present application, based on the voice synthesis algorithm, an audio feature code and a pitch feature can be added and used as the core algorithm of the voice conversion model. For example, the VITS (Versatile and Interpretable Text-to-Speech) algorithm, which synthesizes widely used and interpretable text into voice, can be used as the basic synthesis algorithm of the voice conversion model. Based on this algorithm, an audio encoder and an RMPVE pitch extraction algorithm can be added and used as the core algorithm of the voice conversion model.
[0030] The VITS algorithm is a high-expressive voice synthesis algorithm that combines variational inference, normalizing flow, and adversarial training. Instead of using spectra, latent variables are used to connect the acoustic model and the vocoder in voice synthesis. Random modeling is performed on the latent variables, and a random length predictor is used to improve the diversity of the synthesized voice. The above is only an exemplary description of the embodiments of the present application. In actual applications, other voice synthesis algorithms can be further used to be combined with the voice conversion model of the embodiments of the present application.
[0031] Figure 1 is a flowchart of a method for training a pitch-based voice conversion model provided in an embodiment of the present application. Referring to Figure 1, the training method includes the following steps.
[0032] In S100, the reference voice is input into the pre-encoder and the pitch extraction module, the audio feature code is extracted by the pre-encoder, and the pitch feature is extracted by the pitch extraction module.
[0033] The reference voice is the voice to be learned by the voice conversion model. For example, the user inputs the voice of person A into the voice conversion model, and the voice conversion model outputs the voice of person B based on the same text content as the voice of person A. Here, the voice of person B is the reference voice. After the voice conversion model completes the learning training for the reference voice, when the target that needs to be converted is input into the voice conversion model, the voice conversion model can perform voice conversion on the target voice with the voice features of the reference voice and obtain the converted voice.
[0034] It should be noted that the reference voice and the training voice are voice samples applied in two different stages of the voice conversion model. Here, the training voice is used to perform training on the voice conversion model in the pre-training stage of the voice conversion model, so that the voice conversion model is equipped with basic voice conversion capabilities. The reference voice is used to perform training on the voice conversion model in the learning training stage of the voice conversion model, so that the voice conversion model can convert the voice features of the input voice into the voice features of the reference voice.
[0035] In this embodiment, during the training process of the voice conversion model, it is necessary to input the reference voice into the audio encoder in the pre-encoder, and by doing so, the audio encoder extracts the audio feature code of the reference voice. Also, in order to make the pitch of the converted voice output by the voice conversion model more consistent with the pitch of a real person's voice, the reference voice can be further input into the pitch extraction module synchronously, so that the pitch extraction module extracts the pitch feature of the reference voice and improves the learning accuracy of the voice conversion model with respect to the reference voice.
[0036] In some embodiments, the pre-encoder can further recognize the text content of the reference voice and extract the features of the text information based on the text content, thereby improving the learning accuracy of the voice conversion model with respect to the reference voice.
[0037] Note that in order to extract audio features better, preprocessing can be performed on the reference voice before inputting it into the audio encoder. For example, the reference voice is processed into a Mel spectrum, and the Mel spectrum is input into the audio encoder, so that the audio encoder extracts the audio features of the Mel spectrum and the text features corresponding to the reference voice, compresses and encodes the audio features and text features, outputs the audio feature code, and improves the accuracy of the voice obtained by subsequent conversion.
[0038] In some embodiments, as shown in FIG. 2, the pitch extraction module can include an encoder layer (Encoder Layers), a filter layer (Skip Hidden Feature Filters), an intermediate layer (Intermediate layers), and a decoder layer (Decoder layers). Here, the pitch extraction module can include two encoding branches. The encoder layer, the intermediate layer, and the decoder layer form the first encoding branch of the pitch extraction module, and the encoder layer, the filter layer, and the decoder layer form the second encoding branch of the pitch extraction module.
[0039] Here, the encoder layer is used to encode the reference audio. In the first encoding branch, the intermediate layer is used to extract the high-level pitch features of the reference audio by a plurality of convolutional layers. In the second encoding branch, the filter layer skips the hidden features in the reference audio and is used to extract the low-level pitch features in the reference audio by a convolutional layer. The decoder layer is used to perform a transposed convolution operation on the high-level and low-level pitch features extracted by the first and second encoding branches, thereby completing the decoding and obtaining the pitch features.
[0040] In some embodiments, the encoder layer can include a preset number of REB network structures. Here, as shown in FIG. 3, the REB network structure includes a plurality of RCB convolutional networks and one 2×2 average pooling layer (Avgpool). Therefore, in the process of the pitch extraction module extracting the pitch features of the reference audio, the RCB convolutional network can extract the pitch feature vector of the reference audio.
[0041] The REB network structure includes a plurality of RCB convolutional networks. The 2×2 average pooling layer is installed after the last layer of the RCB convolutional network. The average pooling layer can perform downsampling on the pitch feature vector output by the RCB convolutional network to obtain the pitch feature code, and then the decoder layer further performs decoding on the pitch feature code to obtain the pitch features.
[0042] Since the encoder layer includes a plurality of layers of REB network structures, these REB network structures can sequentially increase the dimension for extracting pitch features according to the input order, and can extract pitch features from the shallow layer to the deep layer. For example, the extraction dimension of the REB network structure of the first layer is (1, 16), the extraction dimension of the REB network structure of the second layer is (16, 32), the extraction dimension of the REB network structure of the third layer is (32, 64), and so on. Thereby, pitch features are extracted step by step and stably, and the extraction accuracy of pitch features is improved.
[0043] In some embodiments, the RCB convolutional network can include a plurality of convolutional blocks. Here, as shown in FIG. 4, the convolutional block can include a 2D convolutional layer, a batch normalization layer, and a relu function, and the batch normalization layer and the relu function are installed after the 2D convolutional layer of each layer. In the process of extracting the pitch feature vector by the RCB convolutional network, the 2D convolutional layer can extract the deep audio vector. In this embodiment, the accuracy of the audio vector can be determined by setting the number of layers of the 2D convolutional layer. It should be noted that the more 2D convolutional layers are installed, the higher the accuracy of the audio vector should be. When the accuracy of the audio vector reaches the maximum, the number of layers of the 2D convolutional layer is the maximum number of layers. Based on the maximum number of layers, increasing the 2D convolutional layer cannot improve the accuracy of the audio vector.
[0044] After the 2D convolutional layer extracts the audio vector, it is necessary to input the audio vector into the batch normalization layer of the corresponding convolutional block. The batch normalization layer performs fast convergence processing on the deep audio vector to extract the convergent pitch feature from the deep audio vector, improve the convergence speed of the RCB convolutional network, and thereby improve the efficiency of voice conversion.
[0045] After obtaining the converged pitch features, a non-linear relationship is added to the converged pitch features by the relu activation function, introducing non-linear characteristics into the convolutional block to obtain a pitch feature vector. Thus, the pitch feature vector output by the current convolutional block can be used as the input for the next convolutional block, enabling the extraction of the pitch features of the reference speech by multiple layers of 2D convolutional layers and obtaining the pitch feature vector finally output by the convolutional block. Note that, in order to improve the efficiency of extracting the pitch features of the reference speech, in the embodiments of the present application, the optimal convolutional core of the 2D convolutional layer of the convolutional block may be 3×3. In actual applications, based on the efficiency required for the application, the convolutional core can be adjusted to 2×2 or 4×4, and the present application does not specifically limit the size of the convolutional core.
[0046] In some embodiments, since a shortcut convolutional layer is further installed between the input end and the output end of the RCB convolutional network, before the relu activation function adds a non-linear relationship to the converged pitch features, the shortcut convolutional layer can extract shortcut pitch features. Here, the shortcut convolutional layer includes only one layer of 2D convolutional layer. Note that, since the speed at which the shortcut convolutional layer extracts the shortcut pitch features is greater than the speed at which the convolutional block extracts the pitch feature vector, the convolutional core of the shortcut convolutional layer should be smaller than the convolutional core of the 2D convolutional layer of the convolutional block. For example, the convolutional core of the shortcut convolutional layer may be 1×1.
[0047] After the shortcut convolutional layer extracts the shortcut pitch features, stitching can be performed on the shortcut pitch features and the converged pitch features output by the batch normalization layer, thereby combining the pitch features extracted based on different dimensions to obtain pitch stitching features. Subsequently, a non-linear relationship between the pitch stitching features is added by the relu activation function, using the voice stitching features as the input for the next convolutional block, and thereby obtaining the pitch feature vector output by the next convolutional block.
[0048] In some embodiments, the intermediate layer may include a plurality of ICB network structures, where the ICB network structure is similar to the REB network structure, and the only difference is that there is no average pooling layer in the ICB network structure. The intermediate layer can extract deeper pitch features in the pitch feature code by the ICB network structure and output them in the format of the pitch feature code.
[0049] In some embodiments, the filter layer includes a plurality of layers of RCB convolutional networks. The RCB convolutional network in the filter layer may be the same as the dimensional arrangement of the REB network structure according to the input order, sequentially increasing the dimension of pitch feature extraction, and extracting pitch features from the shallow layer to the deep layer. The embodiments of the present application can easily improve the pitch similarity between the converted voice output by the subsequent voice conversion model and the voice of a real person by obtaining the deep pitch features of the reference voice through the U-Net structure composed of the 2D convolutional layer and the transposed convolutional layer according to the above network structure.
[0050] In some embodiments, the decoder layer includes a plurality of layers of RDB network structures. As shown in FIG. 5, the RDB network structure includes a transposed convolutional layer and an RCB convolutional network. A batch normalization layer and a relu activation function can be connected between the transposed convolutional layer and the RCB convolutional network. The functions of the batch normalization layer and the relu activation function in the decoder layer can refer to the functions in the above encoder layer, and the description is omitted in the embodiments of the present application.
[0051] In this embodiment, since the decoder layer needs to decode the pitch feature code finally output by the intermediate layer and the filter layer, first, a transposed convolution calculation is performed on the pitch feature code by the transposed convolution layer to obtain a transposed convolution feature vector. Subsequently, the batch normalization layer is used to improve the convergence speed of the decoder layer, and the relu activation function is used to introduce non-linear characteristics into the RCB convolutional network in the decoder layer.
[0052] After the batch normalization layer outputs the transposed convolution feature vector, the RCB convolutional network hierarchically decodes the transposed convolution feature vector, thereby completing the extraction of the pitch feature of the reference speech.
[0053] In S200, feature stitching is performed on the audio feature code and the pitch feature to obtain a voice stitching feature.
[0054] In order to easily combine the pitch feature and the audio feature code, a stitching process is performed on the audio feature code and the pitch feature to obtain the voice stitching feature of the reference speech. In the voice stitching feature, the pitch feature can be mapped to the audio feature code, thereby better combining the pitch feature with the audio feature code.
[0055] In S300, the linear spectrum corresponding to the reference speech is input into the post-encoder to obtain the audio latent variable.
[0056] In the training process of the voice conversion model, a corresponding linear spectrum is generated based on the reference speech, and then the linear spectrum is input into the post-encoder, and the audio latent variable corresponding to the reference speech can be output based on the post-encoder.
[0057] Note that the hidden variable of the audio is generated by the post-encoder only in the learning and training process, and in the application process, it is generated by the normalization flow in the pre-encoder of the voice conversion model. Here, the post-encoder can adopt the non-causal WaveNet residual module in WaveGlow and Glow-TTS. The non-causal WaveNet residual module in WaveGlow and Glow-TTS is applied only in the learning and training process and does not participate in the application process of the voice conversion model.
[0058] In S400, the timing alignment module aligns the voice stitching feature and the timing sequence of the hidden variable of the audio to obtain a converted voice code.
[0059] Before performing voice conversion, it is necessary to ensure that the voice stitching feature and the hidden variable of the audio can correspond to the timing, thereby alleviating the problem that the video and voice between the converted voice and the reference voice are not synchronized. Therefore, the voice stitching feature and the hidden variable of the audio can be input into the timing alignment module, and thereby the timing alignment module aligns the timing sequences of the voice stitching feature and the hidden variable of the audio based on the timing alignment module, and outputs a converted voice code.
[0060] In some embodiments, the reference voice is further input into the projection layer of the pre-encoder, so that the voice stitching feature of the reference voice is projected by the projection layer into the timing alignment module, so as to better combine the voice stitching feature and the hidden variable of the audio, and improve the accuracy of the voice stitching feature being fused into the converted voice code.
[0061] The timing alignment module can adopt the MAS alignment estimation algorithm (Monotonic Alignment Search). The MAS alignment estimation algorithm is an algorithm used for audio signal processing. It is used to compare one audio sequence with one template, thereby performing an alignment operation.
[0062] Therefore, in some embodiments, in the process of inputting the voice stitching feature and the hidden variable of the audio into the timing alignment module, it is necessary to perform timing alignment on the voice stitching feature and the hidden variable of the audio based on a specific template voice sequence. In this embodiment, the timing alignment module can obtain the template voice sequence, and the template voice sequence is a voice sequence that speaks the text content specified by a preset speech rate, intonation, and pitch within a preset time. The template voice sequence plays a role of reference in the process of aligning the voice stitching feature and the hidden variable of the audio. The timing alignment module can align the voice stitching feature and the hidden variable of the audio according to the template voice sequence based on the MAS alignment estimation algorithm, and can perform encoding on the aligned voice stitching feature and the hidden variable of the audio to obtain a converted voice code.
[0063] In some embodiments, before inputting the voice stitching feature and the hidden variable of the audio into the timing alignment module, first input the voice stitching feature and the hidden variable of the audio into the normalization flow. The normalization flow improves the complexity of the prior distribution of the voice stitching feature and the hidden variable of the audio, thereby improving the complexity of the pitch stitching feature, enhancing the efficiency of the voice conversion model in learning the pitch feature, and improving the pitch similarity between the converted voice output by the voice conversion model that has completed subsequent training and the voice of a real person.
[0064] In S500, the decoder decodes the converted voice code to obtain the converted voice.
[0065] In this embodiment, in order to better combine the voice stitching features of the reference voice, it is necessary to perform combination processing and timing alignment processing using the voice stitching features in the code state and the hidden variables of the audio. However, since the voice conversion model cannot output the converted voice in the code state, it is necessary to input the converted voice code into the decoder of the voice conversion model, and thereby the decoder decodes the converted voice code to obtain the converted voice.
[0066] In this embodiment, the decoder may use the generator of the Vocoder HiFi-GAN V1, or may use the generator of other Vocoders. The embodiments of the present application do not overly limit the type of Vocoder used in the decoder.
[0067] In S600, calculate the training loss of the converted voice. If the training loss is less than or equal to the training loss threshold, output the voice conversion model based on the current parameters of the training target model. If the training loss is greater than the training loss threshold, perform iterative training on the voice conversion model.
[0068] After obtaining the converted voice, it indicates that the voice conversion model has completed one learning training process. At this time, calculate the training loss of the current converted voice, and thereby the learning training progress of the voice conversion model can be judged. As shown in FIG. 6, if the training loss is less than or equal to the training loss threshold, it indicates that the voice conversion model has been trained until it converges. At this time, the voice conversion model can be output based on the current parameters of the training target model.
[0069] Note that after the voice conversion model completes the learning training, it reaches a convergent state, and thus can enter the application stage of the voice conversion model. In order to make it easier to distinguish voice conversion models with different training levels, in the embodiments of the present application, a voice conversion model that has not been trained until convergence is defined as the model to be trained.
[0070] If the training loss is greater than the training loss threshold, it indicates that the voice conversion model has not been trained until convergence. At this time, it is necessary to continuously train the model to be trained with the reference voice, thereby performing iterative training on the model to be trained. After each iterative training, the training loss is calculated until the training loss is less than or equal to the training loss threshold and the voice conversion model is trained until convergence. In this way, in the application process, the voice conversion model can output a converted voice with a high pitch similarity to the real person's voice and a high voice accuracy based on the target voice.
[0071] In some embodiments, the voice conversion model can further include a style encoder and a rhythm encoder. Here, after performing feature stitching on the audio feature code and the pitch feature, the reference voice or the mel spectrum corresponding to the reference voice is further input into the style encoder to extract the style feature of the reference voice by the style encoder, and the obtained style feature is mapped to the voice stitching feature, thereby updating the voice stitching feature. The voice stitching feature simultaneously retains the voice stitching feature and the style feature of the reference voice.
[0072] In some embodiments, the style encoder can select the MelStyleEncoder module, and the MelStyleEncoder module includes three sub-modules: a spectral processing layer, a temporal processing layer, and a multi-head attention layer.
[0073] Specifically, the configuration and operation method of the style encoder are as follows.
[0074] The spectral processing layer is composed of a single fully connected layer, and is used to obtain the mel spectrum of the input training speech and convert it into a feature sequence.
[0075] The timing processing layer includes a single gated convolution layer and a single residual layer, and is used to obtain the timing information in the feature sequence.
[0076] Based on the timing information, the attention layer is used to extract the style features corresponding to the feature sequence corresponding to the timing information within a first preset time, and this operation is repeated. The first preset time is a short time at the frame level, and the above operation is to extract the style features corresponding to the feature sequence within a plurality of short times respectively. Based on this, at a second preset time, the plurality of style features corresponding to the plurality of first preset times are averaged to obtain a style vector. Usually, the second preset time is a long time, and the second preset time includes the first preset time.
[0077] After inputting the mel-spectrogram corresponding to the reference audio into the style encoder, the Spectral processing sub-module in the style encoder can convert the input mel-spectrogram into a sequence of frame-level hidden states by means of a fully-connected layer. The Temporal processing sub-module can capture the timing information in the reference audio by means of a Gated CNN and residual connections. The Multi-head attention sub-module can encode global information by means of a multi-head self-attention mechanism and residual connections. Here, the multi-head self-attention is used at the frame level to better extract style features from short reference audio and outputs style vectors Style Embeddings obtained by averaging over time.
[0078] In some embodiments, the style encoder can further include a style adaptive unit. The style adaptive unit includes one layer of normalization layer and one layer of fully-connected layer, predicts corresponding feature biases and feature gains based on the style vector, and is used to output as the style information of the audio. The output style information of the audio is used in subsequent voice conversion. In the conventional voice conversion process, subsequent operations are performed only based on the style features obtained based on the audio. In order to achieve high effectiveness, a large number of training samples are required to achieve more accurate style extraction. With the above improvement, the style information changes adaptively based on the change of the style vector, the reproduction of the style becomes more accurate, and the required amount of training audio also becomes smaller.
[0079] In some embodiments, the pre-encoder may further include a text encoder, which can cooperate with the rhythm encoder to extract the rhythm features of the reference audio. In this embodiment, the reference text can be input into the text encoder, where the reference text is the text corresponding to the audio content in the reference audio, and the text encoder can extract text features based on the reference text. Also, the text features can be input into the rhythm encoder together with the reference audio, and the rhythm encoder can output rhythm features based on the text features and the reference audio.
[0080] In this embodiment, the rhythm encoder can use the ProsodyEncoder module, and the ProsodyEncoder module can extract rhythm features from the reference audio through word-level vector quantization bottleneck.
[0081] In some embodiments, as shown in FIG. 7, the rhythm encoder includes a rhythm convolutional layer and a pooling layer. Here, the rhythm convolutional layer includes a relu activation function and a normalization layer. The relu activation function can remove the linearization of the rhythm encoder, thereby enabling the rhythm convolutional layer to have a non-linear representation ability, fitting deeper rhythm features, and improving the accuracy of extracting rhythm features. The normalization layer is used to perform normalization processing on each sample of the reference audio, improving the convergence speed of the rhythm encoder, reducing the phenomenon of overfitting, and improving the efficiency of extracting rhythm features.
[0082] The structure of the rhythm encoder may also be one layer of rhythm convolutional layer, pooling layer, and another layer of rhythm convolutional layer. For the convenience of description, in the embodiments of the present application, the rhythm convolutional layer before the pooling layer is defined as the first rhythm convolutional layer, and the rhythm convolutional layer after the pooling layer is defined as the second rhythm convolutional layer.
[0083] In the process of inputting text features and reference audio into the rhythm encoder, the text features are directly input into the rhythm convolution layer, and thereby the rhythm convolution layer compresses the text features to obtain word-level hidden features from the text features. The reference audio is first preprocessed into a mel spectrum, and based on the structure of the rhythm encoder, the mel spectrum is sequentially input into the first rhythm convolution layer, the pooling layer, and the second rhythm convolution layer. After the mel spectrum of the reference audio is input into the first rhythm convolution layer, the first rhythm convolution layer can perform word-level rhythm quantization on the reference audio based on the compressed word-level hidden features to obtain rhythm features.
[0084] After the first rhythm convolution layer outputs the rhythm features, the pooling layer performs feature dimension reduction on the word-level hidden features and the rhythm attribute features to obtain a rhythm code, which can reduce the computational complexity of the voice conversion model, reduce the problem of overfitting, and improve the extraction efficiency of features. After the pooling layer outputs the rhythm code, the rhythm code can be input into the second rhythm convolution layer to extract deep rhythm features, and finally the deep rhythm features can be compressed by the vector quantization layer, thereby outputting the final rhythm features, improving the feature extraction accuracy, and making the rhythm code obtained in the training process of the voice conversion model match the rhythm of a real person's speech.
[0085] In some embodiments, the voice stitching features, the style features, and the rhythm features can be simultaneously input into the normalization flow of the pre-encoder, thereby improving the complexity of the voice stitching features, the style features, and the rhythm features. Also, based on the timing alignment module, timing alignment is performed on the voice stitching features, the style features, and the rhythm features with improved complexity and the hidden variables of the audio, and the obtained audio code is combined with the voice stitching features, the style features, and the rhythm features at the same time to reduce the difference in audio between the converted voice and the voice of a real person.
[0086] In some embodiments, the audio encoder can be constructed based on a pre-trained Hubert model. Through pre-training the audio encoder, the audio encoder can better extract the audio characteristics of the reference audio. The following describes the above pre-training process.
[0087] First, a clustering model constructed based on a k-means network is preset. The clustering model can include a feature extraction layer, a clustering processing layer, and a category labeling layer. Here, the feature extraction layer is used to perform feature extraction on the training data in the pre-training process, using a self-supervised model, such as the above Hubert model. Also, the feature extraction layer is also used to constitute the feature extraction part of the audio encoder. The clustering processing layer is constructed based on the K-mens model and is used to perform clustering processing on the extracted audio features, that is, cluster the audio features of the same category. The category labeling layer is used to assign a category code corresponding to the audio features of a certain category after clustering.
[0088] After completing the construction of the clustering model, pre-training can be performed on the clustering model using general training data, and the general training data can be based on LibriSpeech-960 and AISHELL-3 data. For example, voice sample data of 200 speakers is obtained, and the number of clusters is 200. In the clustering model, the feature extraction layer performs audio feature processing on the training data, and the clustering processing layer performs clustering on the corresponding audio features, so that the clustering model can perform clustering processing on voice samples of different speaking categories. Note that the clustering process is to automatically complete the clustering of voices with similar voice styles by means of unsupervised training and similarity calculation. Voices in the same category may not be the same speaker, and it is only necessary for the similarity of the voice styles to reach a preset similarity threshold.
[0089] The category labeling layer can perform category encoding on the categories after clustering for the subsequent training of the audio encoder. For example, after clustering the voice sample data by the clustering model, different categories are obtained, and category codes such as ID1.1, ID1.2, …… ID1.9 can be assigned to different categories respectively. The category code gives a unique label for distinguishing each audio feature category after being clustered by the clustering model, thereby facilitating the mapping and encoding of categories in the subsequent learning and training process of the voice conversion model.
[0090] Note that in the above clustering model, except for the part of the feature extraction layer, the remaining parts are not involved in the construction of the audio encoder, and only provide the audio feature category code to the audio encoder in the training stage of the audio encoder. After the voice conversion model completes training, the clustering model is similarly not involved in the inference operation in the application process of the voice conversion model.
[0091] After the audio encoder completes the pre-training, the audio encoder can be constructed based on the model and each layer structure constructed after the pre-training. In some embodiments, the audio encoder includes the following parts.
[0092] The feature code unit is constructed by the Hubert model that has completed the pre-training of clustering in the above clustering model and is used to extract and encode the audio features of the reference audio.
[0093] The category mapping unit is composed of one layer of mapping layer and is used for mapping the audio feature category code, that is, it is used to map the corresponding category code to the audio features extracted by the feature code unit.
[0094] The category code unit is composed of one layer of embedding layer and is used to assign the category code defined in the clustering model to the audio features extracted by the feature code unit during the training process of the audio encoder.
[0095] During the training process of the audio encoder, first, initialization processing is performed on the feature code unit and the category mapping unit, that is, some parameters of the Hubert model and the mapping layer are randomly initialized. After the initialization is completed, the audio encoder can be trained with general training data.
[0096] During the actual training process of the audio encoder, in addition to normally performing model training and parameter updating, the audio encoder can be trained based on the predicted category code of the category code corresponding to the general training data and the real category code obtained by the clustering model.
[0097] The category mapping unit and the category code unit map one category code to the audio features extracted by the feature extraction unit, and this category code is the predicted category code for the general training data. Then, minimization processing is performed on the average cross-entropy between the predicted category code of the Hubert model for the general training sample and the real category code given in the training process of the clustering model. Based on this, the loss function of the audio encoder is updated, and at the same time, the related parameters of the audio encoder are updated to complete the training of the audio encoder.
[0098] In some embodiments, various training losses can be included in the process of training the voice conversion model. Therefore, in the embodiments of the present application, the training loss can include spectral loss, variance loss, decoder loss, random time length prediction loss, and feature matching loss when training the generator. In this embodiment, the training loss can be represented by the following formula. [Number] Here, L total is the training loss, L recon is the spectral loss, L kl is the variance loss, L dur is the random time length prediction loss, L adv is the loss in the training process of the decoder, and L fm (G) is the feature matching loss when training the generator.
[0099] In some embodiments, the spectral loss is the training loss between the training audio and the training synthesized audio. To calculate the spectral loss, obtain the spectral accuracy of the training audio and the spectral accuracy of the training synthesized audio, and calculate the spectral loss based on the spectral accuracy of the training audio and the spectral accuracy of the training synthesized audio according to the following formula. [Number] Here, L recon is the spectral loss, x mel is the spectral accuracy of the training audio, and x ^ mel is the spectral accuracy of the training synthesized audio.
[0100] In some embodiments, L kl is the KL divergence loss, which is the loss between the posterior distribution estimation of the final hidden variable obtained by combining the variable obtained by processing the linear spectrum of the audio by the post-encoder with the style vector output by the style encoder and the rhythm code output by the rhythm encoder, and the prior distribution estimation of the hidden variable between the given conditional text and its information.
[0101] In this embodiment, calculate the posterior distribution result based on the hidden variable of the audio, obtain the alignment information for synthesizing the audio code, then calculate the prior distribution result based on the alignment information and the hidden variable of the audio, and finally calculate the divergence loss based on the posterior distribution result and the prior distribution result according to the following formula. [Number] Here, L kl is the divergence loss, z is the hidden variable of the audio, and logq φ (z│x lin ) is the posterior distribution result, and logp θ (z│c text, A) is the prior distribution result, c text is the preset text, and A is the alignment information.
[0102] In some embodiments, the decoder further includes a discriminator, and the discriminator can identify the synthesized speech obtained after being decoded in the application process of the voice conversion model. If the discriminator cannot distinguish between the synthesized speech and the voice of a real person, it indicates that the accuracy of the synthesized speech has reached the accuracy of the voice of a real person, and thereby outputs the synthesized speech.
[0103] In the training process of the voice conversion model, the training synthesized speech can be input into the discriminator to obtain the discriminant features output by the discriminator, and the feature matching loss when training the generator can be calculated. Here, the feature matching loss can be regarded as the reconstruction loss and is used to constrain the output of the intermediate layer of the discriminator. The discriminator is used for adversarial training with the decoder in the training voice conversion model.
[0104] In some embodiments, the feature matching loss when training the generator can be calculated by the following formula.
Number
[0105] In some embodiments, L adv ×L fm (G) is the loss of the decoder module. Here, L advis the least squares loss function for adversarial training, and the loss of the discriminator can be calculated by the following formula. [Number] Here, L adv (D) is the loss of the discriminator, Ε is the expression form of the expected value, x is the true waveform of the training spectrum, D(x) indicates the discrimination result of the true waveform, G(z) indicates the feature representation generated after inputting the hidden variable z, and D(G(z)) indicates the discrimination result for the feature G(z) generated by the generator.
[0106] The loss of the generator can be calculated by the following formula. [Number] L adv (G) is the loss of the generator, D(G(z)) indicates the discrimination result for the feature G(z) generated by the generator, and Ε is the expression form of the expected value.
[0107] Therefore, the least squares loss function of adversarial training can be obtained by calculating based on the loss of the discriminator and the loss of the generator, and thereby the loss of the decoder can be obtained by calculating based on the least squares loss function of adversarial training and the feature matching loss when training the generator.
[0108] In some embodiments, L dur is the random time length prediction loss. After the timing alignment module obtains the text code of the training text based on the MAS algorithm, the predicted mean variance and the hidden variable Z are processed by the normalization flow to obtain the optimal alignment matrix of the normal distribution, and the random time length prediction loss can be calculated.
[0109] As can be seen from the above technical solutions, the present application provides a method for training a pitch-based voice conversion model and a voice conversion system. The method is used to train a voice conversion model, which includes a pre-encoder, a post-encoder, a timing alignment module, a decoder, and a pitch extraction module. The method inputs a reference voice into the pre-encoder and the pitch extraction module, outputs an audio feature code by the pre-encoder, and extracts a pitch feature by the pitch extraction module. Subsequently, the linear spectrum corresponding to the reference voice is input into the post-encoder to obtain a hidden variable of the audio. The voice stitching feature obtained by stitching the audio feature code and the pitch feature and the hidden variable of the audio are input into the timing alignment module to obtain a converted voice code, and the converted voice code is decoded by the decoder to obtain the converted voice. Further, the training loss of the converted voice is calculated to determine the convergence degree of the voice conversion model. The present application extracts the pitch feature of the reference voice by the pitch extraction module, stitches it with the audio feature code, performs alignment, and makes the pitch feature of the converted voice match that of a real person's voice, so as to improve the pitch similarity of the converted voice when the voice sample is insufficient.
[0110] As used throughout this specification, phrases such as "a plurality of embodiments", "some embodiments", "one embodiment", or "an embodiment" mean that the specific features, components, or characteristics described with reference to this embodiment are included in at least one embodiment. Therefore, the conjunctions such as "in a plurality of embodiments", "in some embodiments", "in at least another embodiment", or "in an embodiment" referred to throughout this specification do not necessarily refer to the same embodiment. Also, in one or more embodiments, the specific features, components, or characteristics can be combined in any suitable manner. Therefore, unless otherwise limited, the specific features, components, or characteristics shown or described with reference to one embodiment can be combined in whole or in part with the features, components, or characteristics of one or more other embodiments. Such modifications and variations are within the scope of the present application.
[0111] Similar parts among the embodiments provided in this application may be referred to each other. The specific embodiments provided above are only some examples in all concepts of this application and do not limit the protection scope of this application. Any other embodiments extended based on the solution means of this application without creative efforts by those skilled in the art shall fall within the protection scope of this application.
[0112] What is described above is only a preferred embodiment of this application. It should be noted that those skilled in the art can make some improvements and modifications on the premise of not departing from the principle of this application, and these improvements and modifications should also be regarded as falling within the protection scope of this application.
Claims
1. It is used for training a voice conversion model including a pre-encoder, a post-encoder, a timing alignment module, a decoder, and a pitch extraction module. Inputting a reference voice into the pre-encoder and the pitch extraction module, extracting an audio feature code by the pre-encoder, and extracting a pitch feature by the pitch extraction module. Performing feature stitching on the audio feature code and the pitch feature to obtain a voice stitching feature. Inputting a linear spectrum corresponding to the reference voice into the post-encoder to obtain a hidden variable of the audio. Aligning the timing sequences of the voice stitching feature and the hidden variable of the audio by the timing alignment module to obtain a converted voice code. Decoding the converted voice code by the decoder to obtain a converted voice. Calculating the training loss of the converted voice. If the training loss is less than or equal to a training loss threshold, outputting a voice conversion model based on the current parameters of the training target model. If the training loss is greater than the training loss threshold, performing iterative training on the voice conversion model. The training target model is a voice conversion model that has not been trained until convergence. This is a method for training a pitch-based voice conversion model.
2. The pitch extraction module includes an encoder layer, a filter layer, an intermediate layer, and a decoder layer. The encoder layer, the intermediate layer, and the decoder layer form a first encoding branch of the pitch extraction module. The encoder layer, the filter layer, and the decoder layer form a second encoding branch of the pitch extraction module. This is the method for training a pitch-based voice conversion model according to Claim 1.
3. The encoder layer includes an average pooling layer and a convolutional network. The step of extracting a pitch feature by the pitch extraction module is as follows: Extracting a pitch feature vector of the reference voice by the convolutional network. Executing downsampling on the pitch feature vector by the average pooling layer to obtain a pitch feature code; Executing decoding on the pitch feature code by the decoder layer to obtain the pitch feature, including the steps of, characterized in that the method for training a voice conversion model based on pitch according to claim 2.
4. The convolutional network includes a convolutional block, the convolutional block includes a 2D convolutional layer, a batch normalization layer, and a relu function, and the step of extracting the pitch feature vector of the reference voice by the convolutional network is: Extracting a deep audio vector by the 2D convolutional layer; Executing fast convergence processing on the deep audio vector by the batch normalization layer to extract a convergent pitch feature from the deep audio vector; Adding a non-linear relationship to the convergent pitch feature by the relu function to obtain the pitch feature vector, including the steps of, characterized in that the method for training a voice conversion model based on pitch according to claim 3.
5. A shortcut convolutional layer is installed between the input end and the output end of the convolutional network. Before the step of adding a non-linear relationship to the convergent pitch feature by the relu function, Extracting a shortcut pitch feature by the shortcut convolutional layer; Stitching the shortcut pitch feature and the convergent pitch feature to obtain a pitch stitching feature; Adding a non-linear relationship between the pitch stitching features by the relu function to obtain the pitch feature vector, further including the steps of, characterized in that the method for training a voice conversion model based on pitch according to claim 4.
6. The decoder layer includes a transposed convolutional layer and the convolutional network. The step of executing decoding on the pitch feature code by the decoder layer is: Executing a transposed convolution calculation on the pitch feature code by the transposed convolutional layer to obtain a transposed convolution feature vector; Executing decoding on the transposed convolution feature vector by the convolutional network to obtain the pitch feature, including the steps of, characterized in that the method for training a voice conversion model based on pitch according to claim 4.
7. The step of aligning the voice stitching feature and the timing sequence of the hidden variable of the audio by the timing alignment module is as follows: Obtaining the template voice sequence of the timing alignment module; Aligning the voice stitching feature and the timing sequence of the hidden variable of the audio based on the template voice sequence; Executing encoding on the aligned voice stitching feature and the hidden variable of the audio to obtain the converted voice code. The method for training a pitch-based voice conversion model according to claim 1 is characterized by comprising the above steps.
8. The voice conversion model further includes a style encoder. After the step of performing feature stitching on the audio feature code and the pitch feature, Extracting the style feature of the reference voice by the style encoder; Mapping the style feature to the voice stitching feature to update the voice stitching feature. The method for training a pitch-based voice conversion model according to claim 1 is characterized by further comprising the above steps.
9. The training loss includes a spectral loss. The step of calculating the training loss of the converted voice is as follows: Obtaining the spectral accuracy of the reference voice and obtaining the spectral accuracy of the converted voice; Calculating the spectral loss based on the spectral accuracy of the reference voice and the spectral accuracy of the converted voice according to the following formula: 【Number 1】 Here, L recon is the spectral loss, and x mel is the spectral accuracy of the reference speech, and x ^ mel is the spectral accuracy of the converted speech, and a step of: A method for training a pitch-based voice conversion model according to claim 1, characterized by comprising.
10. It is obtained by training based on the method for training a pitch-based voice conversion model according to any one of claims 1 to 9, and includes a voice conversion model including a pre-encoder, a post-encoder, a timing alignment module, a decoder, and a pitch extraction module. The pre-encoder is configured to extract the audio feature code of the reference voice. The pitch extraction module is configured to extract the pitch feature of the reference voice. The post-encoder is configured to generate a hidden variable of the audio based on the linear spectrum corresponding to the reference voice. The timing alignment module is configured to align the voice stitching feature and the timing sequence of the hidden variable of the audio to obtain a converted voice code, and the voice stitching feature is obtained by combining the audio feature code and the pitch feature. The decoder is configured to perform decoding on the converted voice code to obtain the converted voice. A voice conversion system characterized by the above is provided.