Joint training method and device for speech synthesis model and tone conversion model

By setting the encoder independently in the speech synthesis and tone conversion models and sharing only the decoder, the robustness and training efficiency of the model are improved, and the problem of poor joint model effect in the existing technology is solved, and higher task prediction accuracy is achieved.

CN120544535APending Publication Date: 2025-08-26GUANGZHOU SHIYUAN ELECTRONICS CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410207915.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-26
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

In the prior art, the combined model of speech synthesis and tone conversion is poor, which affects the robustness of the model.

Method used

The joint modeling architecture is adopted, including text encoder, diffusion decoder, vocoder, voice content recognition network and voice feature encoding network. These modules are trained in a specific order, and the encoder for tone conversion is independently set up, and only the decoder is shared to improve the accuracy of content and voiceprint extraction.

Benefits of technology

It improves the robustness and training efficiency of the combined model of speech synthesis and tone conversion, and enhances the accuracy of the model's task prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544535A_ABST
    Figure CN120544535A_ABST
Patent Text Reader

Abstract

The invention provides a joint training method and device for a speech synthesis model and a tone conversion model, and the method and device are applied to the field of speech processing, and the method comprises the steps: training a text encoder, a diffusion decoder and a vocoder based on sample speech data and sample text data, and obtaining the trained text encoder, diffusion decoder and vocoder; based on the trained diffusion decoder and the sample voice data, respectively training a voice content recognition network and a voice feature coding network to obtain the trained voice content recognition network and the trained voice feature coding network; based on the trained voice feature coding network and the target personalized voice, training the trained diffusion decoder to obtain a personalized diffusion decoder; and generating a speech synthesis model and a tone conversion model based on the trained text encoder, the personalized diffusion decoder, the trained vocoder and the trained speech content recognition network. The method can improve the joint modeling effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of speech processing, and in particular to a joint training method and device for a speech synthesis model and a timbre conversion model. Background Art

[0002] Personalized speech synthesis is a technology that generates customized speech synthesis results based on the user's individual characteristics and needs; personalized voice conversion is a technology that generates customized voice conversion results based on the user's individual characteristics and needs. The main difference between the two is that personalized speech synthesis takes text as input and produces a personalized voice, while personalized voice conversion takes the original speaker's voice as input and produces a personalized voice. However, since both processes process the input content and produce a personalized voice, personalized speech synthesis and personalized voice conversion can be jointly modeled, allowing the model to simultaneously achieve personalized speech synthesis and personalized voice conversion.

[0003] Therefore, how to improve the performance of the joint model of speech synthesis and timbre conversion is an urgent problem to be solved. Summary of the Invention

[0004] This application provides a joint training method and device for a speech synthesis model and a timbre conversion model, which can improve the performance of the joint speech synthesis and timbre conversion model. The technical solution is as follows:

[0005] In a first aspect, a joint training method for a speech synthesis model and a timbre conversion model is provided, the method comprising: training a text encoder, a diffusion decoder, and a vocoder based on sample speech data and sample text data to obtain trained text encoder, diffusion decoder, and vocoder; training a speech content recognition network and a speech feature coding network based on the trained diffusion decoder and sample speech data, respectively, to obtain trained speech content recognition network and speech feature coding network, wherein the speech content recognition network is used to extract content features from the sample speech data, and the speech feature coding network is used to extract content features and voiceprint features from the sample speech data; training a trained diffusion decoder based on the trained speech feature coding network and target personalized speech to obtain a personalized diffusion decoder; generating a speech synthesis model and a timbre conversion model based on the trained text encoder, personalized diffusion decoder, trained vocoder, and trained speech content recognition network, wherein the speech synthesis model is used to process input text into personalized speech with a target timbre, and the timbre conversion model is used to process input speech into personalized speech with a target timbre, where the target timbre is the timbre of the target personalized speech.

[0006] In the above technical solution, an embodiment of the present application provides a joint modeling method for a speech synthesis model and a timbre conversion model. The joint modeling architecture mainly includes a text encoder, a diffusion decoder, a vocoder, a speech content recognition network, and a speech feature coding network. By training each module in the joint modeling architecture separately in a specific order, the joint modeling purpose of personalized speech synthesis and personalized timbre conversion can be achieved, thereby reducing the time for individual modeling. Moreover, the joint modeling architecture has an independent encoder-speech content recognition network for timbre conversion, and only a decoder (diffusion decoder) can be shared during the joint modeling process. Compared with the shared encoder and decoder in the related art, the accuracy of content extraction and voiceprint extraction can be enhanced, thereby improving the model robustness of the joint modeling.

[0007] In combination with the first aspect, in some possible implementations, a speech synthesis model and a timbre conversion model are generated based on a trained text encoder, a personalized diffusion decoder, a trained vocoder, and a trained speech content recognition network, including: determining the trained text encoder, the personalized diffusion decoder, and the trained vocoder as the speech synthesis model; and determining the trained speech content recognition network, the personalized diffusion decoder, and the trained vocoder as the timbre conversion model.

[0008] In combination with the first aspect and the above-mentioned implementation methods, in some possible implementation methods, the method further includes: inputting the target text data into a trained text encoder to obtain target text features; inputting the target text features into a personalized diffusion decoder to obtain a first target Mel spectrum; inputting the first target Mel spectrum into a trained vocoder to obtain personalized synthesized speech, and the timbre of the personalized synthesized speech matches the timbre of the target personalized speech.

[0009] In combination with the first aspect and the above-mentioned implementation methods, in some possible implementation methods, the method further includes: inputting the target speech data into a trained speech content recognition network to obtain target content features; inputting the target content features into a personalized diffusion decoder to obtain a second target Mel spectrum; inputting the second target Mel spectrum into a trained vocoder to obtain personalized converted speech, and the timbre of the personalized converted speech matches the timbre of the target personalized speech.

[0010] In the above scheme, since the speech feature encoding network is only used to fine-tune the diffusion decoder during training, it is not required during model application. Based on the functions of the trained modules and the purposes of speech synthesis and voice conversion, during model application, the trained speech content recognition network, vocoder, and personalized diffusion decoder can be combined to generate a voice conversion model for personalized voice conversion; the trained text encoder, personalized diffusion decoder, and vocoder can be combined to generate a speech synthesis model for personalized speech synthesis. This achieves the goal of training two functional models simultaneously, improving model training efficiency.

[0011] In combination with the first aspect and the above-mentioned implementation methods, in some possible implementation methods, the speech content recognition network includes a speech feature extraction network and a content feature encoding network, and the method also includes: training the speech feature extraction network based on sample speech data and sample text data to obtain a trained speech feature extraction network, and the speech feature extraction network is used to extract content features and voiceprint features in the sample speech data; based on the trained diffusion decoder and sample speech data, respectively training the speech content recognition network and the speech feature encoding network to obtain a trained speech content recognition network and a speech feature encoding network, including: based on the trained diffusion decoder, the trained speech feature extraction network and the sample speech data, training the content feature encoding network in the speech content recognition network to obtain a trained content feature encoding network; training the speech feature encoding network based on the trained diffusion decoder and the sample speech data to obtain a trained speech feature encoding network.

[0012] In combination with the first aspect, in some possible implementations, the content feature coding network in the speech content recognition network is trained based on the trained diffusion decoder and sample speech data to obtain a trained content feature coding network, including: inputting the sample speech data into the trained speech feature extraction network to obtain a first sample speech feature, the first sample speech feature including a sample text feature and a sample voiceprint feature; inputting the sample text feature into the content feature coding network to obtain a second sample speech feature; inputting the second sample speech feature into the trained diffusion decoder to obtain a first sample Mel spectrum; determining a first model loss based on the first sample Mel spectrum and the standard Mel spectrum corresponding to the sample speech data; and updating the model parameters of the content feature coding network based on the first model loss to obtain a trained content feature coding network.

[0013] In combination with the first aspect, in some possible implementations, a speech feature coding network is trained based on a trained diffusion decoder and sample speech data to obtain a trained speech feature coding network, including: inputting the sample speech data into the speech feature coding network to obtain a third sample speech feature; inputting the third sample speech feature into the trained diffusion decoder to obtain a second sample Mel spectrum; determining a second model loss based on the second sample Mel spectrum and the standard Mel spectrum corresponding to the sample speech data; and updating the model parameters of the speech feature coding network based on the second model loss to obtain a trained speech feature coding network.

[0014] In combination with the first aspect, in some possible implementations, a trained diffusion decoder is trained based on a trained speech feature coding network and a target personalized speech to obtain a personalized diffusion decoder, including: inputting the target personalized speech into the trained speech feature coding network to obtain a fourth sample speech feature; inputting the fourth sample speech feature into the trained diffusion decoder to obtain a third sample Mel spectrum; determining a third model loss based on the third sample Mel spectrum and the standard Mel spectrum corresponding to the target personalized speech; and updating the model parameters of the trained diffusion decoder based on the third model loss to obtain a personalized diffusion decoder.

[0015] In this approach, fine-tuning the diffusion decoder through the speech feature encoding network allows inputting only the target personalized speech, without the need for text data. This reduces the time required for manual text annotation and further improves model training efficiency. Furthermore, during model training, a corresponding model loss is applied, and model parameters are updated based on this loss to align the model with the standard prediction results, thereby improving training accuracy.

[0016] In a second aspect, a joint training device for a speech synthesis model and a timbre conversion model is provided, which includes: a first training module for training a text encoder, a diffusion decoder, and a vocoder based on sample speech data and sample text data to obtain trained text encoders, diffusion decoders, and vocoders; a second training module for training a speech content recognition network and a speech feature coding network based on the trained diffusion decoder and sample speech data, respectively, to obtain trained speech content recognition network and speech feature coding network, wherein the speech content recognition network is used to extract content features from the sample speech data, and the speech feature coding network is used to extract content features and voiceprint features in the sample speech data; a third training module, used to train the trained diffusion decoder based on the trained speech feature encoding network and the target personalized speech to obtain a personalized diffusion decoder; a model generation module, used to generate a speech synthesis model and a timbre conversion model based on the trained text encoder, the personalized diffusion decoder, the trained vocoder, and the trained speech content recognition network, the speech synthesis model is used to process the input text into personalized speech with a target timbre, and the timbre conversion model is used to process the input speech into personalized speech with a target timbre, and the target timbre is the timbre of the target personalized speech.

[0017] In combination with the second aspect, in some possible implementations, the model generation module is also used to: determine the trained text encoder, personalized diffusion decoder, and trained vocoder as a speech synthesis model; and determine the trained speech content recognition network, personalized diffusion decoder, and trained vocoder as a timbre conversion model.

[0018] In combination with the second aspect and the above-mentioned implementation methods, in some possible implementation methods, the device also includes: a first acquisition module, used to input the target text data into a trained text encoder to obtain target text features; a second acquisition module, used to input the target text features into a personalized diffusion decoder to obtain a first target Mel spectrum; and a third acquisition module, used to input the first target Mel spectrum into a trained vocoder to obtain personalized synthesized speech, wherein the timbre of the personalized synthesized speech matches the timbre of the target personalized speech.

[0019] In combination with the second aspect and the above-mentioned implementation methods, in some possible implementation methods, the device also includes: a fourth acquisition module, used to input the target speech data into a trained speech content recognition network to obtain target content features; a fifth acquisition module, used to input the target content features into a personalized diffusion decoder to obtain a second target Mel spectrum; and a sixth acquisition module, used to input the second target Mel spectrum into a trained vocoder to obtain personalized converted speech, wherein the timbre of the personalized converted speech matches the timbre of the target personalized speech.

[0020] In combination with the second aspect and the above-mentioned implementation methods, in some possible implementation methods, the speech content recognition network includes a speech feature extraction network and a content feature encoding network; the device also includes: a fourth training module, which is used to train the speech feature extraction network based on sample speech data and sample text data to obtain a trained speech feature extraction network, and the speech feature extraction network is used to extract content features and voiceprint features in the sample speech data; the second training module is also used to: train the content feature encoding network in the speech content recognition network based on the trained diffusion decoder, the trained speech feature extraction network and the sample speech data to obtain a trained content feature encoding network; train the speech feature encoding network based on the trained diffusion decoder and the sample speech data to obtain a trained speech feature encoding network.

[0021] In combination with the second aspect and the above-mentioned implementation methods, in some possible implementation methods, the second training module is also used to: input the sample speech data into the trained speech feature extraction network to obtain a first sample speech feature, the first sample speech feature includes a sample text feature and a sample voiceprint feature; input the sample text feature into the content feature encoding network to obtain a second sample speech feature; input the second sample speech feature into the trained diffusion decoder to obtain a first sample Mel spectrum; determine the first model loss based on the first sample Mel spectrum and the standard Mel spectrum corresponding to the sample speech data; based on the first model loss, update the model parameters of the content feature encoding network to obtain a trained content feature encoding network.

[0022] In combination with the second aspect and the above-mentioned implementation methods, in some possible implementation methods, the second training module is further used to: input the sample speech data into the speech feature coding network to obtain a third sample speech feature; input the third sample speech feature into the trained diffusion decoder to obtain a second sample Mel spectrum; determine the second model loss based on the second sample Mel spectrum and the standard Mel spectrum corresponding to the sample speech data; based on the second model loss, update the model parameters of the speech feature coding network to obtain a trained speech feature coding network.

[0023] In combination with the second aspect and the above-mentioned implementation methods, in some possible implementation methods, the third training module is further used to: input the target personalized speech into the trained speech feature encoding network to obtain a fourth sample speech feature; input the fourth sample speech feature into the trained diffusion decoder to obtain a third sample Mel spectrum; determine the third model loss based on the third sample Mel spectrum and the standard Mel spectrum corresponding to the target personalized speech; based on the third model loss, update the model parameters of the trained diffusion decoder to obtain a personalized diffusion decoder.

[0024] In a third aspect, a computer device is provided, comprising: a processor and a memory, wherein the memory stores a computer program, and the computer program is loaded and executed by the processor to implement the above-mentioned joint training method of the speech synthesis model and the timbre conversion model.

[0025] In a fourth aspect, a computer program product is provided, comprising: a computer program code, which, when executed on a computer, enables the computer to execute the method in the first aspect or any possible implementation of the first aspect.

[0026] In a fifth aspect, a computer-readable storage medium is provided, which stores a computer program code. When the computer program code runs on a computer, the computer executes the method in the above-mentioned first aspect or any possible implementation of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0028] Figure 1 This is a diagram of the training architecture of the speech synthesis model and timbre conversion model provided by related technologies;

[0029] Figure 2 Schematic diagram of the training architecture of the speech synthesis model and timbre conversion model provided in the embodiment of the present application;

[0030] Figure 3 This is a flowchart of a joint training method for a speech synthesis model and a timbre conversion model provided in an embodiment of the present application;

[0031] Figure 4 Schematic diagram of the training process of the text encoder and diffusion decoder provided in the embodiment of the present application;

[0032] Figure 5 Schematic diagram of the training process of the vocoder provided in an embodiment of the present application;

[0033] Figure 6 This is a flowchart of another joint training method for a speech synthesis model and a timbre conversion model provided in an embodiment of the present application;

[0034] Figure 7 This is a schematic diagram of the application process of the speech synthesis model provided in the embodiment of the present application;

[0035] Figure 8 Schematic diagram of the application process of the timbre conversion model provided in the embodiment of the present application;

[0036] Figure 9 This is a flowchart of another joint training method for a speech synthesis model and a timbre conversion model provided in an embodiment of the present application;

[0037] Figure 10 This is a schematic diagram of the process of speech feature extraction network training provided in an embodiment of the present application;

[0038] Figure 11 This is a schematic diagram of the process of content feature coding network training provided by an embodiment of the present application;

[0039] Figure 12 This is a schematic diagram of the process of speech feature coding network training provided in an embodiment of the present application;

[0040] Figure 13 Schematic diagram of the process of personalized diffusion decoder training provided by an embodiment of the present application;

[0041] Figure 14 This is a structural block diagram of a joint training device for a speech synthesis model and a timbre conversion model provided in an embodiment of the present application;

[0042] Figure 15 It is a structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0043] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0044] In the related art, during the joint modeling process of the speech synthesis model and the timbre conversion model, the speech synthesis and the timbre synthesis share a decoder and an encoder. Figure 1 This is a diagram of the training architecture of the speech synthesis model and timbre conversion model provided by related technologies. Figure 1 As shown, the existing model training process is to train the voice conversion encoder 103, decoder 105 and vocoder 106 based on sample speech data 101 and sample text data 102; train the text encoder 104 based on the sample text data 102; if the decoder 105 is to be personalized trained, the personalized speech is input into the voice conversion encoder again to train the decoder 105.

[0045] As can be seen, during the joint modeling process, the speech synthesis model and the voice conversion model use the voice conversion encoder to fine-tune the decoders for personalized speech synthesis and voice conversion during the personalized decoder training, thus achieving unified modeling. In other words, during the joint modeling process, speech synthesis and voice conversion share the same decoder and encoder. However, sharing the encoder and decoder can affect the robustness of the model, making the joint model's performance inferior to that of individual modeling.

[0046] In order to improve the problem of poor joint modeling effect in the prior art, the embodiment of the present application provides a joint modeling method for speech synthesis and timbre conversion, which can effectively improve the joint modeling effect of the speech synthesis model and the timbre conversion model. Figure 2 As shown, it is a schematic diagram of the training architecture of the speech synthesis model and the timbre conversion model provided in the embodiment of the present application. The training architecture mainly includes a text encoder 204, a diffusion decoder 205, a vocoder 206, a speech content recognition network and a speech feature encoding network 208. The speech content recognition network includes a speech feature extraction network 203 and a content feature encoding network 207.

[0047] The joint modeling process mainly includes two training stages: the average model training stage and the personalized fine-tuning stage. The average model training stage includes: (1) using the large-scale multi-speaker speech recognition corpus data 201 to train the speech feature extraction network 203 in the speech content recognition network, and obtain the trained speech feature extraction network 203; (2) using the large-scale multi-speaker speech synthesis corpus data 202 to train the text encoder 204 and the diffusion decoder 205, and obtain the trained text encoder 204 and the diffusion decoder 205; (3) using the same large-scale multi-speaker speech synthesis corpus data 202 to train the vocoder 206, and obtain the trained vocoder 206; (4) fixing the model parameters of the diffusion decoder 205, training the speech feature encoding network 208, and obtain the trained speech feature encoding network 208; (5) fixing the model parameters of the diffusion decoder 205, and training the content feature encoding network 207 in the speech content recognition network. The personalized fine-tuning stage includes: (1) collecting personalized voice data through the constructed platform; (2) using the collected personalized voice data, fixing the model parameters of the voice feature coding network 208, and performing personalized fine-tuning on the trained diffusion decoder 205 to obtain a trained personalized decoder.

[0048] During the model application process, the speech text to be synthesized is passed through the trained text encoder 204, personalized diffusion decoder 205 and vocoder 206 to obtain synthesized personalized speech; or the original speaker's speech is passed through the trained speech content recognition network, personalized diffusion decoder 205 and vocoder 206 to obtain personalized speech after timbre conversion.

[0049] It can be seen that the joint modeling method provided in the embodiment of the present application performs timbre conversion encoding separately (that is, a speech content recognition network is set up separately), only shares a personalized decoder, and can set different encoders for different tasks. While realizing the joint modeling of two tasks, it improves the encoding accuracy of different tasks, thereby improving the modeling effect of the joint modeling and improving the task prediction accuracy of subsequent different task models.

[0050] Figure 3 This is a flowchart of a joint training method for a speech synthesis model and a timbre conversion model provided in an embodiment of the present application. This method is described using an example of a computer device. The method includes:

[0051] Step 301 : training a text encoder, a diffusion decoder, and a vocoder based on sample speech data and sample text data to obtain trained text encoder, diffusion decoder, and vocoder.

[0052] The speech synthesis model is used to convert text into speech, that is, the input is text and the output is speech; the timbre conversion model is used to convert the timbre of the speech into the timbre of the target speaker while retaining the text content in the speech, that is, the input is the first speech and the output is the second speech with a different timbre.

[0053] Depend on Figure 1 As can be seen from the joint modeling architecture shown, in the process of jointly establishing the speech synthesis model and the timbre conversion model, the text encoder, diffusion decoder and vocoder are first trained using sample speech data and sample text data to obtain the trained text encoder, diffusion decoder and vocoder.

[0054] The text encoder is used to extract contextual information from the input text, the diffusion decoder is used to convert the input text into a mel-spectrogram, and the vocoder is used to convert the mel-spectrogram into speech. The text encoder and diffusion decoder can use the same network architecture as Grad-TTS. For example, the text encoder can be a six-layer transformer feedforward network; the diffusion decoder can be composed of a U-Net network structure; and the vocoder can use the BigVGAN model structure, which consists of a generator and multiple discriminators. The generator incorporates periodic nonlinearity and anti-aliasing markers, which bring the desired inductive bias to speech waveform synthesis to improve the output speech quality. Optionally, the diffusion decoder can also be called a diffusion model-based decoder.

[0055] Optionally, before training the model, a large amount of sample speech data and sample text data corresponding to the sample speech data are first collected, and the sample speech data and sample text data are annotated, and the emotional features, phoneme features, and rhythm features of the sample speech data are annotated, and the phoneme features, semantic features, word and sentence division features of the sample text data and the duration features of the sample speech data corresponding to different words are annotated, etc., for subsequent determination of the model training loss.

[0056] Optionally, when training the text encoder, diffusion decoder, and vocoder, you can train them simultaneously. Alternatively, you can first train the text encoder and diffusion decoder, then fix the model parameters of the text encoder and diffusion decoder before training the vocoder. Training them separately can improve model training efficiency.

[0057] When training a text encoder, a diffusion decoder, and a vocoder simultaneously, the total loss of the model may include two parts, the output loss of the diffusion decoder and the output loss of the vocoder. The output loss of the diffusion decoder may include at least one of mel spectrum loss, diffusion loss, and duration loss, and the output loss of the vocoder may include at least one of duration loss and speech waveform loss.

[0058] Take the example of training the text encoder + diffusion decoder and vocoder separately, Figure 4 Schematic diagram of the training process of the text encoder and diffusion decoder provided in the embodiment of the present application. Figure 4 As shown, the training process of the text encoder and diffusion decoder is as follows: the sample text data 401 is input into the text encoder 403, and a high-level representation with contextual information (for example, the emotional information implied by the text, etc.) - context representation is output. The context representation is then input into the diffusion decoder 404 to obtain a predicted Mel spectrum 405. Then, based on the predicted Mel spectrum 405 and the standard Mel spectrum 406 corresponding to the sample speech data 402 corresponding to the sample text data 401, the Mel spectrum loss is determined, so as to train the text encoder 403 and the diffusion decoder 404 based on the Mel spectrum loss. Optionally, the diffusion loss of the diffusion decoder 404 and the duration loss can also be calculated to jointly train the text encoder 403 and the diffusion decoder 404. The duration loss can be: calculating the loss between the spectrum duration of the predicted Mel spectrum 405 corresponding to each word in the sample text data 401 and the spectrum duration of the standard Mel spectrum 406 corresponding to each word in the sample text data 401.

[0059] Figure 5 Schematic diagram of the training process of the vocoder provided in the embodiment of the present application. Figure 5As shown, the training process of the vocoder includes: inputting sample text data 501 into the trained text encoder 503 to obtain a context representation, then inputting the context representation into the trained diffusion decoder 504 to output a mel spectrum 505, inputting the mel spectrum 505 into the vocoder 506, and outputting a predicted speech 507; determining a duration loss based on the duration of each word in the sample text data 501 in the predicted speech 507 and the duration of each word in the sample text data 501 in the sample speech data 502, and determining a speech waveform loss based on the speech waveform of the predicted speech 507 and the speech waveform of the sample speech data 502 corresponding to the sample text data 501, and then training the vocoder 506 based on at least one of the duration loss and the speech waveform loss to obtain the trained vocoder 506. It should be noted that during the training of the vocoder 506, it is not necessary to update the model parameters of the diffusion decoder 504 and the text encoder 503.

[0060] Optionally, in the process of training the text encoder, diffusion decoder and vocoder, the same sample speech data and sample text data may be selected, or different sample speech data and sample text data may be selected.

[0061] Step 302: Based on the trained diffusion decoder and sample speech data, the speech content recognition network and the speech feature coding network are trained respectively to obtain trained speech content recognition network and speech feature coding network. The speech content recognition network is used to extract content features from the sample speech data, and the speech feature coding network is used to extract content features and voiceprint features from the sample speech data.

[0062] To avoid interference between the voiceprint feature extraction process and the content feature extraction process, the present application performs feature encoding for timbre conversion and speech synthesis separately during the training process, that is, the encoder is not shared. Correspondingly, in one possible implementation, the sample speech data is used as input, and the speech content recognition network and speech feature encoding network are trained separately based on the trained diffusion decoder, resulting in trained speech content recognition network and speech feature encoding network. The speech content recognition network is used to extract content features from the sample speech data, and the speech feature encoder network is used to separate and extract content features and voiceprint features from the sample speech data.

[0063] Optionally, when training the speech content recognition network, the mel spectrum loss, temporal loss, and diffusion loss can be calculated as model losses to train the speech content recognition network. Similarly, when training the speech feature coding network, the mel spectrum loss, temporal loss, and diffusion loss can also be calculated as model losses to train the speech feature coding network. Furthermore, when updating network parameters based on the loss function, the model parameters of the trained diffusion decoder are not updated; in other words, the model parameters of the diffusion decoder are fixed, and the speech content recognition network and the speech feature coding network are trained separately.

[0064] Optionally, the speech content recognition network may include a speech feature extraction network and a content feature encoding network. The speech feature extraction network is used to extract content features and voiceprint features from sample speech data. The content feature encoding network is used to further separate voiceprint features from the features extracted by the speech feature extraction network and extract higher-dimensional contextual information to obtain content features corresponding to the sample speech data. The speech feature extraction network can also be called a pre-trained speech recognition model, and the content feature encoding network can also be called a pre-trained speech recognition feature encoder. The speech feature encoding network may include a speech self-supervised model and a speech self-supervised feature encoder. The speech self-supervised model is used to extract content features and voiceprint features from sample speech data. The speech self-supervised feature encoder is used to obtain high-level representations of content features and voiceprint features with contextual information.

[0065] Step 303: Based on the trained speech feature coding network and the target personalized speech, the trained diffusion decoder is trained to obtain a personalized diffusion decoder.

[0066] In order for this joint model to convert input text or speech into speech output with a specific timbre—that is, to achieve personalized speech synthesis and personalized timbre conversion—the trained diffuse decoder needs to be fine-tuned for personalized timbre. In one possible implementation, the trained diffuse decoder is retrained based on the trained speech feature encoding network, using the target personalized speech as input, to obtain a personalized diffuse decoder.

[0067] Optionally, users can record their own personalized voice. The collection and processing of the target personalized voice includes: Upon detecting that a user has logged into the personalized voice collection platform, the system determines whether the ambient noise passes the signal-to-noise ratio test. If the ambient noise exceeds a preset threshold, the system prompts the user to change the environment and retest. If the ambient noise passes the signal-to-noise ratio test, the system begins recording valid voice for a preset duration. The valid voice is segmented using an endpoint detection algorithm to produce 2- to 15-second voice segments for subsequent training of the personalized diffusion model.

[0068] Step 304: Based on the trained text encoder, personalized diffusion decoder, trained vocoder, and trained speech content recognition network, a speech synthesis model and a timbre conversion model are generated. The speech synthesis model is used to process the input text into personalized speech with a target timbre, and the timbre conversion model is used to process the input speech into personalized speech with a target timbre, where the target timbre is the timbre of the target personalized speech.

[0069] After the above training steps, the joint model has the functions of personalized speech synthesis and personalized timbre conversion. Correspondingly, a speech synthesis model and a timbre conversion model can be generated based on the trained text encoder, personalized diffusion decoder, trained vocoder, and trained speech content recognition network. Among them, the speech synthesis model is used to process the input text into personalized speech with a target timbre; the timbre conversion model is used to process the input speech into personalized speech with a target timbre; the target timbre is the timbre corresponding to the target personalized speech.

[0070] In summary, an embodiment of the present application provides a joint modeling method for a speech synthesis model and a timbre conversion model: the joint modeling architecture mainly includes a text encoder, a diffusion decoder, a vocoder, a speech content recognition network, and a speech feature coding network. By training each module in the joint modeling architecture separately in a specific order, the joint modeling purpose of personalized speech synthesis and personalized timbre conversion can be achieved, thereby reducing the time for individual modeling. Moreover, the joint modeling architecture has an independent encoder-speech content recognition network for timbre conversion, and only a decoder (diffusion decoder) can be shared during the joint modeling process. Compared with the shared encoder and decoder in the related art, the accuracy of content extraction and voiceprint extraction can be enhanced, thereby improving the model robustness of the joint modeling.

[0071] During the model application process, based on the purpose of speech synthesis and timbre conversion, as well as the role of each module in the joint model, the modules required for speech synthesis and timbre conversion are organized into different links. Speech synthesis and timbre conversion functions can be implemented separately through different links.

[0072] Figure 6 This is a flowchart of another joint training method for a speech synthesis model and a timbre conversion model provided in an embodiment of the present application. This method is described using a computer device as an example. The method includes:

[0073] Step 601: training a text encoder, a diffusion decoder, and a vocoder based on sample speech data and sample text data to obtain trained text encoder, diffusion decoder, and vocoder.

[0074] Step 602: Based on the trained diffusion decoder and sample speech data, the speech content recognition network and the speech feature coding network are trained respectively to obtain trained speech content recognition network and speech feature coding network.

[0075] Among them, the speech content recognition network is used to extract content features in sample speech data, and the speech feature encoding network is used to extract content features and voiceprint features in sample speech data.

[0076] Step 603: Based on the trained speech feature coding network and the target personalized speech, the trained diffusion decoder is trained to obtain a personalized diffusion decoder.

[0077] The implementation of steps 601 to 603 can refer to the above embodiment, and will not be described in detail in this embodiment.

[0078] Step 604: The trained text encoder, personalized diffusion decoder, and trained vocoder are determined as a speech synthesis model.

[0079] The speech synthesis model is used to extract text features of the target text, combine the text features with the timbre of the target personalized speech, and output personalized synthesized speech. The corresponding speech synthesis model may include a text encoder, a personalized diffusion decoder, and a vocoder, so that the trained text encoder, personalized diffusion decoder, and trained vocoder can be determined as a speech synthesis model, which is a personalized speech synthesis model.

[0080] Step 605: Input the target text data into the trained text encoder to obtain target text features.

[0081] Target text features include phoneme features, semantic features, word and sentence segmentation features, and duration features of sample speech data corresponding to different words. During speech synthesis, the target text data can be input into a trained text encoder, which then performs feature encoding on the target text data to extract contextual information from the target text data and obtain the target text features.

[0082] Step 606: Input the target text feature into the personalized diffusion decoder to obtain a first target Mel spectrum.

[0083] After obtaining the target text features, the target text features may be input into a personalized diffusion decoder to convert the text into a Mel spectrum with a target timbre added thereto, so as to obtain a first target Mel spectrum.

[0084] Step 607: Input the first target Mel-spectrogram into the trained vocoder to obtain personalized synthesized speech, where the timbre of the personalized synthesized speech matches the timbre of the target personalized speech.

[0085] After generating the first target Mel spectrum, the first target Mel spectrum is further input into the trained vocoder, which converts it into a personalized synthesized speech output, thereby converting the text data into a personalized synthesized speech with a timbre matching the target personalized speech.

[0086] Figure 7 This is a schematic diagram of the application process of the speech synthesis model provided in the embodiment of the present application. Figure 7 As shown, in the application process of the speech synthesis model, the target text data 701 is input into the text encoder 702 to obtain the target text feature 703, and then the target text feature 703 is input into the personalized diffusion decoder 704 to obtain the first target Mel spectrum 705, and finally the first target Mel spectrum 705 is input into the vocoder 706 to obtain the personalized synthesized speech 707.

[0087] Step 608: Determine the trained speech content recognition network, personalized diffusion decoder, and trained vocoder as a timbre conversion model.

[0088] The timbre conversion model is used to extract content features from the target speech data, combine these features with the timbre of the target personalized speech, and output the personalized converted speech. The corresponding timbre conversion model may include a speech content recognition network, a personalized decoder, and a vocoder. The trained speech content recognition network, personalized diffusion decoder, and trained vocoder can be identified as the timbre conversion model.

[0089] Step 609: Input the target speech data into the trained speech content recognition network to obtain target content features.

[0090] The target content features include the content information contained in the target speech data, as well as the phoneme features, semantic features, word and sentence division features corresponding to the content information, and the duration features of the sample speech data corresponding to different words.

[0091] Step 610: Input the target content feature into a personalized diffusion decoder to obtain a second target Mel spectrum.

[0092] Step 611: Input the second target Mel-spectrogram into the trained vocoder to obtain a personalized converted speech, where the timbre of the personalized converted speech matches the timbre of the target personalized speech.

[0093] In the application process of the timbre conversion model, the target speech data is input into the trained speech content recognition network to obtain the target content features corresponding to the target speech data. The target content features are then passed through a personalized diffusion decoder to obtain a second target Mel spectrum. Finally, the second target Mel spectrum is input into the trained vocoder to obtain personalized converted speech with the target personalized speech timbre.

[0094] Figure 8 Schematic diagram of the application process of the timbre conversion model provided in the embodiment of the present application. Figure 8 As shown, during the application of the timbre conversion model, the target speech data 801 is input into the speech content recognition network 802, which outputs the target content features 803. The target content features 803 are then input into the personalized diffusion decoder 804 to obtain the second target Mel spectrum 805. Finally, the second target Mel spectrum 805 is input into the vocoder 806 to obtain the personalized converted speech 807.

[0095] In this embodiment, since the speech feature encoding network is only used to fine-tune the diffusion decoder during training, it is not required during model application. Based on the functions of the trained modules and the purposes of speech synthesis and timbre conversion, during model application, the trained speech content recognition network, vocoder, and personalized diffusion decoder can be combined to generate a timbre conversion model for personalized timbre conversion; and the trained text encoder, personalized diffusion decoder, and vocoder can be combined to generate a speech synthesis model for personalized speech synthesis. This achieves the goal of simultaneously training two functional models and improves model training efficiency.

[0096] The speech content recognition network may include a speech feature extraction network and a content feature encoding network. The speech feature extraction network is used to extract content features and voiceprint features from sample speech data; the content feature encoding network is used to separate voiceprint features from the features extracted by the speech feature extraction network to obtain content features corresponding to the sample speech data. In this embodiment, the speech feature extraction network and the content feature encoding network can be trained separately to improve the performance of the speech feature extraction network.

[0097] Figure 9 This is a flowchart of another joint training method for a speech synthesis model and a timbre conversion model provided in an embodiment of the present application. This method is described using a computer device as an example. The method includes:

[0098] Step 901: Train a speech feature extraction network based on sample speech data and sample text data to obtain a trained speech feature extraction network. The speech feature extraction network is used to extract content features and voiceprint features from the sample speech data.

[0099] Speech data contains text features, voiceprint features, and some noise. During speech processing, it is necessary to extract useful information from the speech data. In one possible implementation, a speech feature extraction network can be trained based on sample speech data and sample text data to obtain a trained speech feature extraction network.

[0100] Figure 10 This is a schematic diagram of the process of speech feature extraction network training provided by the embodiment of the present application. Figure 10 As shown, the specific training process of the speech feature extraction network is: input the sample speech data 1001 into the speech feature extraction network 1002, output the sample content feature 1003 and the sample voiceprint feature 1004 corresponding to the sample speech data 1001, determine the loss function according to the sample content feature 1003 and the standard content feature 1006 of the sample text data 1005 corresponding to the sample speech data 1001, and based on the loss function, update the parameters of the speech feature extraction network 1002 to obtain the trained speech feature extraction network 1002.

[0101] Optionally, the speech feature extraction network can adopt an end-to-end speech recognition framework. In order to make the speech feature extraction network better applicable to subsequent timbre conversion, the convolutional network is not used for downsampling operation in the encoding part of the speech feature extraction network.

[0102] Step 902: training the text encoder, diffusion decoder, and vocoder based on the sample speech data and the sample text data to obtain trained text encoder, diffusion decoder, and vocoder.

[0103] The implementation of step 902 can refer to the above embodiment, and will not be described in detail in this embodiment.

[0104] It should be noted that step 901 and step 902 may be performed simultaneously, or step 901 may be performed first and then step 902; or step 902 may be performed first and then step 901. This embodiment does not limit the order of performing step 901 and step 902.

[0105] Step 903 : Based on the trained diffusion decoder, the trained speech feature extraction network and the sample speech data, the content feature coding network in the speech content recognition network is trained to obtain a trained content feature coding network.

[0106] Among them, the content feature encoding network is used to further extract the content features and voiceprint features output by the speech feature extraction network, strip the voiceprint features, and extract higher-dimensional context information to obtain the content features of the sample speech data.

[0107] The speech feature extraction network removes some noise from the sample speech data and obtains sample text features and sample voiceprint features corresponding to the sample speech data. However, during the timbre conversion process, what needs to be obtained are the text features in the speech data. A corresponding content feature encoding network is also provided to further remove voiceprint information and extract higher-dimensional contextual information. In one possible implementation, after the diffusion decoder and speech feature extraction network are trained, the parameters of the diffusion decoder and speech feature extraction network are fixed, the sample speech data is input into the trained speech feature extraction network, the sample text features and sample voiceprint features in the sample speech data are output, the content features and voiceprint features are input into the content feature encoding network to obtain content features, and the content features are then input into the trained diffusion decoder to obtain the corresponding mel spectrum. The mel spectrum loss is determined based on the mel spectrum and the mel spectrum corresponding to the sample speech features. The content encoding network is trained based on the mel spectrum loss to obtain a trained content feature encoding network.

[0108] In an exemplary example, step 903 may include steps 903A to 903E.

[0109] 903A, input the sample speech data into the trained speech feature extraction network to obtain a first sample speech feature, which includes a sample text feature and a sample voiceprint feature.

[0110] 903B, input the sample text feature into the content feature encoding network to obtain the second sample speech feature.

[0111] The second sample speech feature is the sample text feature after stripping the sample voiceprint feature, and the second sample speech feature has higher-dimensional context information than the first sample speech feature.

[0112] 903C: Input the second sample speech feature into the trained diffusion decoder to obtain the first sample Mel spectrum.

[0113] 903D, determine a first model loss based on the first sample Mel spectrum and the standard Mel spectrum corresponding to the sample speech data.

[0114] 903E: Based on the first model loss, update the model parameters of the content feature coding network to obtain a trained content feature coding network.

[0115] During the training process of the content feature coding network, the model parameters of the speech feature extraction network and the diffusion decoder are fixed, and the sample speech data is input into the trained speech feature extraction network to output a first sample speech feature. The first sample speech feature includes a sample text feature and a sample voiceprint feature. The sample text feature is then input into the content feature coding network, which further removes the voiceprint information from the sample text feature and extracts higher-dimensional context information to obtain a second sample speech feature. Furthermore, the second sample speech feature is input into the trained diffusion decoder to obtain a first sample Mel spectrum. During the loss determination process, a first model loss can be determined based on the first sample Mel spectrum and the standard Mel spectrum corresponding to the sample speech data, and the model parameters of the content feature coding network are updated to obtain a trained content feature coding network. Optionally, the first model loss can also be determined based on the Mel spectrum loss, the duration loss, and the diffusion loss to train the content feature coding network.

[0116] Figure 11 This is a schematic diagram of the process of content feature coding network training provided by the embodiment of this application. Figure 11 As shown, during the training of the content feature coding network, the sample speech data 1101 is input into the speech feature extraction network 1102 to obtain a first sample speech feature 1103, wherein the first sample speech feature 1103 includes a sample text feature 1104 and a sample voiceprint feature 1105, and then the sample text feature 1104 is input into the content feature coding network 1106 to obtain a second sample speech feature 1107, and the second sample speech feature 1107 is input into the diffusion decoder 1108 to obtain a first sample Mel spectrum 1109, and based on the first sample Mel spectrum 1109 and the standard Mel spectrum 1110 corresponding to the sample speech data 1101, the first model loss is determined for training the content feature coding network 1106.

[0117] Step 904: Based on the trained diffusion decoder and the sample speech data, the speech feature coding network is trained to obtain a trained speech feature coding network.

[0118] Similar to the training process for the content feature coding network, the training of the speech feature coding network also requires fixing the model parameters of the trained diffusion decoder to train the speech feature coding network. Sample speech data is fed into the speech feature coding network to extract speech features. Once the speech features are obtained, the trained diffusion decoder is fed to generate the mel-spectrogram corresponding to the sample speech data. The model parameters of the speech feature coding network are then updated based on the mel-spectrogram loss.

[0119] In an exemplary example, step 904 may further include steps 904A to 904D.

[0120] Step 904A: Input the sample speech data into the speech feature coding network to obtain the third sample speech feature.

[0121] The third sample voice feature at least includes a voiceprint feature corresponding to the sample voice data, and may also include a content feature corresponding to the sample voice data.

[0122] Step 904B: Input the third sample speech feature into the trained diffusion decoder to obtain a second sample Mel spectrum.

[0123] Step 904C: Determine a second model loss based on the second sample Mel spectrum and the standard Mel spectrum corresponding to the sample speech data.

[0124] Step 904D: Based on the second model loss, update the model parameters of the speech feature coding network to obtain a trained speech feature coding network.

[0125] During the training process of the speech feature coding network, the model parameters of the diffusion decoder are fixed, the sample speech data is input into the speech feature coding network, and a third sample speech feature corresponding to the sample speech data is obtained. The third sample speech feature is then input into the trained diffusion decoder to obtain a second sample Mel spectrum output by the diffusion decoder. Based on the second sample Mel spectrum and the standard Mel spectrum corresponding to the sample speech data, a second model loss is determined, and the model parameters of the speech feature coding network are updated based on the second model loss to obtain a trained speech feature coding network. Optionally, the duration loss and diffusion loss can also be calculated, and the sum of the duration loss, diffusion loss, and Mel spectrum loss can be determined as the second model loss to update the model parameters of the speech feature coding network.

[0126] Figure 12 This is a schematic diagram of the process of speech feature coding network training provided by the embodiment of the present application. Figure 12 As shown, during the training of the speech feature coding network, the sample speech data 1201 is input into the speech feature coding network 1202 to obtain a third sample speech feature 1203, the third sample speech feature 1203 is input into the diffusion decoder 1204, and a second sample Mel spectrum 1205 is output. The second model loss is determined based on the second sample Mel spectrum 1205 and the standard Mel spectrum 1206 corresponding to the sample speech data 1201 to train the speech feature coding network 1202.

[0127] Step 905 : Based on the trained speech feature coding network and the target personalized speech, the trained diffusion decoder is trained to obtain a personalized diffusion decoder.

[0128] In order to enable the diffusion decoder to process the input content and output speech with personalized timbre, in one possible implementation, based on the trained diffusion decoder, the diffusion decoder is adjusted in combination with the target personalized speech to obtain a personalized diffusion decoder.

[0129] In an exemplary example, step 905 may further include steps 905A to 905D.

[0130] Step 905A: Input the target personalized speech into the trained speech feature coding network to obtain a fourth sample speech feature.

[0131] The fourth sample voice feature at least includes a voiceprint feature corresponding to the target personalized voice.

[0132] Step 905B: Input the fourth sample speech feature into the trained diffusion decoder to obtain the third sample Mel spectrum.

[0133] Step 905C: Determine a third model loss based on the third sample Mel spectrum and the standard Mel spectrum corresponding to the target personalized speech.

[0134] Step 905D: Based on the third model loss, update the model parameters of the trained diffusion decoder to obtain a personalized diffusion decoder.

[0135] To enable the diffusion decoder to output personalized speech, during the training process of the personalized diffusion decoder, the model parameters of the speech feature encoding network are fixed, the target personalized speech is input into the trained speech feature encoding network, and the voiceprint features of the target personalized speech are obtained. The voiceprint features are then input into the trained diffusion decoder to obtain a third sample Mel spectrum. Based on the third sample Mel spectrum and the standard Mel spectrum corresponding to the target personalized speech, a third model loss is determined. The third model loss is then used to update the model parameters of the trained diffusion decoder to obtain the personalized diffusion decoder. Optionally, a duration loss and diffusion loss can also be calculated, and the sum of the duration loss, diffusion loss, and Mel spectrum loss is determined as the third model loss to update the model parameters of the diffusion decoder.

[0136] Figure 13 This is a schematic diagram of the process of training a personalized diffusion decoder provided by an embodiment of the present application. Figure 13As shown, during the training of the personalized diffusion decoder, the target personalized speech 1301 is input into the speech feature encoding network 1302 to obtain a fourth sample speech feature 1303, the fourth sample speech feature 1303 is input into the diffusion decoder 1304, and a third sample Mel spectrum 1305 is output. The third model loss is determined based on the third sample Mel spectrum 1305 and the standard Mel spectrum 1306 corresponding to the sample speech data 1301 to train the diffusion decoder 1304 and obtain a personalized diffusion decoder.

[0137] Step 906 : Generate a speech synthesis model and a timbre conversion model based on the trained text encoder, the personalized diffusion decoder, the trained vocoder, and the trained speech content recognition network.

[0138] Among them, the speech synthesis model is used to process the input text into personalized speech with a target timbre, and the timbre conversion model is used to process the input speech into personalized speech with a target timbre, where the target timbre is the timbre of the target personalized speech.

[0139] The implementation of step 906 can refer to the above embodiment, and will not be described in detail in this embodiment.

[0140] In this embodiment, fine-tuning the diffusion decoder through the speech feature coding network allows inputting only the target personalized speech without the need for text data. This reduces the time required for manual annotation of text data and further improves model training efficiency. Furthermore, during model training, corresponding model losses are provided and used to update model parameters, so that the model approaches the standard prediction results, thereby improving model training accuracy.

[0141] Figure 14 This is a structural block diagram of a joint training device for a speech synthesis model and a timbre conversion model provided in an embodiment of the present application, the device comprising:

[0142] A first training module 1401 is configured to train a text encoder, a diffusion decoder, and a vocoder based on sample speech data and sample text data to obtain trained text encoders, diffusion decoders, and vocoders;

[0143] A second training module 1402 is configured to train a speech content recognition network and a speech feature coding network based on the trained diffusion decoder and the sample speech data, respectively, to obtain a trained speech content recognition network and a trained speech feature coding network. The speech content recognition network is configured to extract content features from the sample speech data, and the speech feature coding network is configured to extract content features and voiceprint features from the sample speech data.

[0144] The third training module 1403 is configured to train the trained diffusion decoder based on the trained speech feature coding network and the target personalized speech to obtain a personalized diffusion decoder;

[0145] The model generation module 1404 is used to generate a speech synthesis model and a timbre conversion model based on the trained text encoder, personalized diffusion decoder, trained vocoder, and trained speech content recognition network. The speech synthesis model is used to process the input text into personalized speech with a target timbre, and the timbre conversion model is used to process the input speech into personalized speech with a target timbre, where the target timbre is the timbre of the target personalized speech.

[0146] In an optional embodiment, the model generation module 1404 is further configured to:

[0147] The trained text encoder, personalized diffusion decoder, and trained vocoder are determined as a speech synthesis model;

[0148] The trained speech content recognition network, personalized diffusion decoder, and trained vocoder are determined as a timbre conversion model.

[0149] In an optional embodiment, the device further comprises:

[0150] The first acquisition module is used to input the target text data into the trained text encoder to obtain the target text features;

[0151] The second acquisition module is used to input the target text feature into the personalized diffusion decoder to obtain a first target Mel spectrum;

[0152] The third acquisition module is used to input the first target Mel spectrum into the trained vocoder to obtain personalized synthesized speech, and the timbre of the personalized synthesized speech matches the timbre of the target personalized speech.

[0153] In an optional embodiment, the device further comprises:

[0154] A fourth acquisition module is used to input the target speech data into the trained speech content recognition network to obtain target content features;

[0155] A fifth acquisition module is configured to input the target content feature into a personalized diffusion decoder to obtain a second target Mel spectrum;

[0156] The sixth acquisition module is configured to input the second target Mel spectrum into the trained vocoder to obtain a personalized converted speech, wherein the timbre of the personalized converted speech matches the timbre of the target personalized speech.

[0157] In an optional embodiment, the speech content recognition network includes a speech feature extraction network and a content feature encoding network;

[0158] The device also includes:

[0159] A fourth training module is used to train the speech feature extraction network based on the sample speech data and the sample text data to obtain a trained speech feature extraction network, which is used to extract content features and voiceprint features from the sample speech data;

[0160] The second training module 1402 is further configured to:

[0161] Based on the trained diffusion decoder, the trained speech feature extraction network, and the sample speech data, the content feature encoding network in the speech content recognition network is trained to obtain a trained content feature encoding network;

[0162] Based on the trained diffusion decoder and sample speech data, the speech feature coding network is trained to obtain a trained speech feature coding network.

[0163] In an optional embodiment, the second training module 1402 is further configured to:

[0164] Inputting the sample speech data into the trained speech feature extraction network to obtain a first sample speech feature, where the first sample speech feature includes a sample text feature and a sample voiceprint feature;

[0165] Inputting the sample text features into the content feature encoding network to obtain the second sample speech features;

[0166] Input the second sample speech feature into the trained diffusion decoder to obtain the first sample Mel spectrum;

[0167] Determining a first model loss based on the first sample Mel spectrum and a standard Mel spectrum corresponding to the sample speech data;

[0168] Based on the first model loss, the model parameters of the content feature encoding network are updated to obtain a trained content feature encoding network.

[0169] In an optional embodiment, the second training module 1402 is further configured to:

[0170] Inputting the sample speech data into a speech feature coding network to obtain a third sample speech feature;

[0171] Input the third sample speech feature into the trained diffusion decoder to obtain the second sample Mel spectrum;

[0172] Determining a second model loss based on the second sample Mel spectrum and a standard Mel spectrum corresponding to the sample speech data;

[0173] Based on the second model loss, the model parameters of the speech feature coding network are updated to obtain a trained speech feature coding network.

[0174] In an optional embodiment, the third training module 1403 is further configured to:

[0175] Inputting the target personalized speech into the trained speech feature encoding network to obtain a fourth sample speech feature;

[0176] Inputting the fourth sample speech feature into the trained diffusion decoder to obtain the third sample Mel spectrum;

[0177] determining a third model loss based on the third sample Mel spectrum and a standard Mel spectrum corresponding to the target personalized voice;

[0178] Based on the third model loss, the model parameters of the trained diffusion decoder are updated to obtain a personalized diffusion decoder.

[0179] Figure 15 It is a structural diagram of a computer device provided in an embodiment of the present application.

[0180] For example, Figure 15 As shown, the computer device 1500 includes: a memory 1501 and a processor 1502, wherein the memory 1501 stores an executable program code 1503, and the processor 1502 is used to call and execute the executable program code 1503 to perform a joint training method for a speech synthesis model and a timbre conversion model.

[0181] In addition, an embodiment of the present application also protects a device, which may include a memory and a processor, wherein the memory stores executable program code, and the processor is used to call and execute the executable program code to perform a joint training method for a speech synthesis model and a timbre conversion model provided in an embodiment of the present application.

[0182] In this embodiment, the device can be divided into functional modules based on the above-described method examples. For example, each functional module can be mapped to a specific functional module, or two or more functions can be integrated into a single processing module. The integrated module can be implemented in hardware. It should be noted that the module division in this embodiment is illustrative and represents only a logical functional division. In actual implementation, other division methods may be used.

[0183] In the case of dividing the functional modules into corresponding modules, the device may further include a first acquisition module, a first control module, a second control module, etc. It should be noted that all relevant contents of the steps involved in the above method embodiment can be referred to the functional description of the corresponding functional modules and will not be repeated here.

[0184] It should be understood that the device provided in this embodiment is used to execute the above-mentioned joint training method of a speech synthesis model and a timbre conversion model, and thus can achieve the same effect as the above-mentioned implementation method.

[0185] In the case of an integrated unit, the device may include a processing module and a storage module. When the device is applied to a computer device, the processing module may be used to control and manage the operation of the computer device. The storage module may be used to support the computer device in executing relevant program codes, etc.

[0186] The processing module may be a processor or controller that implements or executes the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a digital signal processing (DSP) and a microprocessor, and the storage module may be a memory.

[0187] In addition, the device provided in the embodiments of the present application can specifically be a chip, component or module, and the chip may include a connected processor and memory; wherein the memory is used to store instructions, and when the processor calls and executes the instructions, the chip can execute a joint training method for a speech synthesis model and a timbre conversion model provided in the above embodiment.

[0188] This embodiment also provides a computer-readable storage medium, which stores computer program code. When the computer program code runs on a computer, the computer executes the above-mentioned related method steps to implement a joint training method for a speech synthesis model and a timbre conversion model provided in the above embodiment.

[0189] This embodiment also provides a computer program product. When the computer program product runs on a computer, it enables the computer to execute the above-mentioned related steps to implement a joint training method for a speech synthesis model and a timbre conversion model provided in the above embodiment.

[0190] Among them, the device, computer-readable storage medium, computer program product or chip provided in this embodiment are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods provided above, and will not be repeated here.

[0191] Through the description of the above implementation methods, technical personnel in the relevant field can understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0192] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0193] The above content is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A joint training method for a speech synthesis model and a timbre conversion model, characterized in that: The method comprises: Training a text encoder, a diffusion decoder, and a vocoder based on the sample speech data and the sample text data to obtain trained text encoders, diffusion decoders, and vocoders; Based on the trained diffusion decoder and the sample speech data, respectively training a speech content recognition network and a speech feature coding network to obtain trained speech content recognition network and speech feature coding network, wherein the speech content recognition network is used to extract content features from the sample speech data, and the speech feature coding network is used to extract content features and voiceprint features from the sample speech data; Based on the trained speech feature coding network and the target personalized speech, the trained diffusion decoder is trained to obtain a personalized diffusion decoder; Based on the trained text encoder, the personalized diffusion decoder, the trained vocoder, and the trained speech content recognition network, a speech synthesis model and a timbre conversion model are generated. The speech synthesis model is used to process the input text into personalized speech with a target timbre, and the timbre conversion model is used to process the input speech into personalized speech with the target timbre, where the target timbre is the timbre of the target personalized speech.

2. The method according to claim 1, characterized in that The generating of a speech synthesis model and a timbre conversion model based on the trained text encoder, the personalized diffusion decoder, the trained vocoder, and the trained speech content recognition network includes: Determining the trained text encoder, the personalized diffusion decoder, and the trained vocoder as the speech synthesis model; The trained speech content recognition network, the personalized diffusion decoder, and the trained vocoder are determined as the timbre conversion model.

3. The method according to claim 2, characterized in that The method further comprises: Inputting target text data into the trained text encoder to obtain target text features; Inputting the target text feature into the personalized diffusion decoder to obtain a first target Mel spectrum; The first target Mel-spectrogram is input into the trained vocoder to obtain a personalized synthesized speech, wherein the timbre of the personalized synthesized speech matches the timbre of the target personalized speech.

4. The method according to claim 2, characterized in that The method further comprises: Inputting the target speech data into the trained speech content recognition network to obtain target content features; Inputting the target content feature into the personalized diffusion decoder to obtain a second target Mel spectrum; The second target Mel-spectrogram is input into the trained vocoder to obtain a personalized converted speech, wherein the timbre of the personalized converted speech matches the timbre of the target personalized speech.

5. The method according to any one of claims 1 to 4, characterized in that: The speech content recognition network includes a speech feature extraction network and a content feature encoding network; The method further comprises: Training the speech feature extraction network based on the sample speech data and the sample text data to obtain a trained speech feature extraction network, wherein the speech feature extraction network is used to extract content features and voiceprint features from the sample speech data; The method of training a speech content recognition network and a speech feature coding network based on the trained diffusion decoder and the sample speech data to obtain the trained speech content recognition network and speech feature coding network comprises: Based on the trained diffusion decoder, the trained speech feature extraction network, and the sample speech data, the content feature coding network in the speech content recognition network is trained to obtain a trained content feature coding network; The speech feature coding network is trained based on the trained diffusion decoder and the sample speech data to obtain the trained speech feature coding network.

6. The method according to claim 5, characterized in that The step of training the content feature coding network in the speech content recognition network based on the trained diffusion decoder and the sample speech data to obtain a trained content feature coding network includes: Inputting the sample speech data into the trained speech feature extraction network to obtain a first sample speech feature, wherein the first sample speech feature includes a sample text feature and a sample voiceprint feature; Inputting the sample text feature into the content feature encoding network to obtain a second sample speech feature; Inputting the second sample speech feature into the trained diffusion decoder to obtain a first sample Mel spectrum; Determining a first model loss based on the first sample Mel spectrum and a standard Mel spectrum corresponding to the sample speech data; Based on the first model loss, the model parameters of the content feature encoding network are updated to obtain the trained content feature encoding network.

7. The method according to claim 5, characterized in that The step of training the speech feature coding network based on the trained diffusion decoder and the sample speech data to obtain the trained speech feature coding network comprises: Inputting the sample speech data into the speech feature coding network to obtain a third sample speech feature; Inputting the third sample speech feature into the trained diffusion decoder to obtain a second sample Mel spectrum; Determining a second model loss based on the second sample Mel spectrum and a standard Mel spectrum corresponding to the sample speech data; Based on the second model loss, the model parameters of the speech feature coding network are updated to obtain the trained speech feature coding network.

8. The method according to any one of claims 1 to 4, characterized in that: The step of training the trained diffusion decoder based on the trained speech feature coding network and the target personalized speech to obtain a personalized diffusion decoder includes: Inputting the target personalized speech into the trained speech feature coding network to obtain a fourth sample speech feature; Inputting the fourth sample speech feature into the trained diffusion decoder to obtain a third sample Mel spectrum; Determining a third model loss based on the third sample Mel spectrum and a standard Mel spectrum corresponding to the target personalized voice; Based on the third model loss, the model parameters of the trained diffusion decoder are updated to obtain the personalized diffusion decoder.

9. A joint training device for a speech synthesis model and a timbre conversion model, characterized in that: The device comprises: A first training module is used to train the text encoder, diffusion decoder and vocoder based on sample speech data and sample text data to obtain trained text encoder, diffusion decoder and vocoder; a second training module, configured to train a speech content recognition network and a speech feature coding network based on the trained diffusion decoder and the sample speech data, respectively, to obtain a trained speech content recognition network and a trained speech feature coding network, wherein the speech content recognition network is configured to extract content features from the sample speech data, and the speech feature coding network is configured to extract content features and voiceprint features from the sample speech data; A third training module is configured to train the trained diffusion decoder based on the trained speech feature coding network and the target personalized speech to obtain a personalized diffusion decoder; A model generation module is used to generate a speech synthesis model and a timbre conversion model based on the trained text encoder, the personalized diffusion decoder, the trained vocoder, and the trained speech content recognition network. The speech synthesis model is used to process input text into personalized speech with a target timbre, and the timbre conversion model is used to process input speech into personalized speech with the target timbre, where the target timbre is the timbre of the target personalized speech.

10. A computer device, characterized in that: The computer device includes: a processor and a memory, the memory stores a computer program, and the computer program is loaded and executed by the processor to implement the joint training method of the speech synthesis model and the timbre conversion model as described in any one of claims 1 to 8.