A speech synthesis model training, speech synthesis method and related device

By building a speech synthesis model that includes voiceprint network and tone support network, the problem of poor tone fit in the existing technology is solved, high-quality tone fit and low-threshold personalized TTS system are realized, and training costs and labeling costs are reduced.

CN114187891BActive Publication Date: 2025-09-02BIGO TECH PTE LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210040692.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-14
Publication Date
2025-09-02
Estimated Expiration
2042-01-14

AI Technical Summary

Technical Problem

In the prior art, the timbre similarity between the spectrum signal fitted by voiceprints and the target spectrum signal is low, resulting in poor tone fitting effect of the personalized TTS system.

Method used

By constructing a speech synthesis model, including voiceprint network, tone support network and acoustic network, we obtain the original spectrum signal and encode the voiceprint features in the voiceprint network, encode the tone supplementary features in the tone support network, and fuse it into the tone total feature, and use the tone embedding feature to correct the total feature to train the acoustic network to achieve high-quality fitting of the tone.

Benefits of technology

It improves the tone similarity between the spectrum signal and the target spectrum signal, reduces training costs and thresholds, facilitates user operations, reduces manual labeling costs and time costs, and improves the tone fitting quality of the personalized TTS system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114187891B_ABST
    Figure CN114187891B_ABST
Patent Text Reader

Abstract

The present invention provides a training method for a speech synthesis model, a speech synthesis method, and related devices. The method comprises: obtaining an original spectral signal and a speaker's timbre embedded feature, wherein the original spectral signal is converted from an original speech signal recorded when the speaker speaks according to text information; encoding the original spectral signal into a voiceprint feature in a voiceprint network, wherein the voiceprint feature is used to verify the speaker's identity; encoding the original spectral signal into a timbre supplementary feature in a timbre support network, wherein the timbre supplementary feature is a feature missing from the voiceprint feature in timbre; fusing the voiceprint feature and the timbre supplementary feature into a total timbre feature; and training an acoustic network and a timbre support network based on the total timbre feature and the original spectral signal, under the condition that the timbre embedded feature corrects the total timbre feature. This ensures the comprehensiveness of the feature in timbre, thereby fitting a high-quality spectral signal and improving the timbre similarity between the fitted spectral signal and the target spectral signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech processing, and in particular to a speech synthesis model training, a speech synthesis method and related devices. Background Art

[0002] TTS (Text To Speech) aims to convert text into speech. It is part of human-computer dialogue and enables machines to speak. Personalized TTS is that the user records one or several voice clips, and the machine extracts the timbre from the voice and imitates it. That is, any text can be input and the machine will output the corresponding content and voice with similar timbre. It is widely used in scenarios such as audiobooks (simulating the voice of parents telling stories to children), navigation broadcasts (using one's own voice to broadcast navigation instructions), and personalized idols.

[0003] Currently, one way to implement personalized TTS requires users to record a voice, learn the voiceprint of the voice, and use the voiceprint and arbitrary text to generate a spectral signal with the corresponding timbre. However, the timbre similarity between the spectral signal fitted with the voiceprint and the target spectral signal is low. Summary of the Invention

[0004] The present invention proposes a speech synthesis model training, speech synthesis method and related devices to solve the problem of low timbre similarity between a spectral signal fitted by voiceprint and a target spectral signal.

[0005] In a first aspect, an embodiment of the present invention provides a method for training a speech synthesis model, wherein the speech synthesis model includes a voiceprint network, a timbre support network, and an acoustic network, and the method includes:

[0006] Obtaining an original spectrum signal and a speaker's timbre embedding feature, wherein the original spectrum signal is converted from an original speech signal recorded when the speaker speaks according to the text information;

[0007] In the voiceprint network, the original spectrum signal is encoded into a voiceprint feature, and the voiceprint feature is used to verify the identity of the speaker;

[0008] In the timbre support network, the original spectrum signal is encoded into a timbre supplementary feature, where the timbre supplementary feature is a feature missing from the voiceprint feature in timbre;

[0009] Merging the voiceprint feature and the timbre supplementary feature into a timbre total feature;

[0010] Under the condition that the timbre embedding feature modifies the timbre total amount feature, the acoustic network and the timbre support network are trained according to the timbre total amount feature and the original spectrum signal.

[0011] In a second aspect, an embodiment of the present invention further provides a speech synthesis method, comprising:

[0012] Loading a speech synthesis model, the speech synthesis model including a voiceprint network, a timbre support network, and an acoustic network;

[0013] Determine the text information and the speaker's original speech signal;

[0014] Converting the original speech signal into an original spectrum signal;

[0015] In the voiceprint network, the original spectrum signal is encoded into a voiceprint feature, and the voiceprint feature is used to verify the identity of the speaker;

[0016] In the timbre support network, the original spectrum signal is encoded into a timbre supplementary feature, where the timbre supplementary feature is a feature missing from the voiceprint feature in timbre;

[0017] Merging the voiceprint feature and the timbre supplementary feature into a timbre total feature;

[0018] In the acoustic network, a target spectrum signal having the text information and the total amount of timbre characteristics is fitted.

[0019] In a third aspect, an embodiment of the present invention further provides a training device for a speech synthesis model, wherein the speech synthesis model includes a voiceprint network, a timbre support network, and an acoustic network, and the device includes:

[0020] A data set acquisition module is used to acquire an original spectrum signal and a speaker's timbre embedding feature, wherein the original spectrum signal is converted from an original speech signal recorded when the speaker speaks according to the text information;

[0021] a voiceprint feature encoding module, configured to encode the original spectrum signal into a voiceprint feature in the voiceprint network, wherein the voiceprint feature is used to verify the identity of the speaker;

[0022] a timbre supplementary feature encoding module, configured to encode the original spectrum signal into a timbre supplementary feature in the timbre support network, wherein the timbre supplementary feature is a feature missing from the voiceprint feature in timbre;

[0023] A timbre total feature fusion module, configured to fuse the voiceprint feature and the timbre supplementary feature into a timbre total feature;

[0024] The network correction training module is used to train the acoustic network and the timbre support network according to the timbre total amount feature and the original spectrum signal under the condition that the timbre embedding feature corrects the timbre total amount feature.

[0025] In a fourth aspect, an embodiment of the present invention further provides a speech synthesis device, comprising:

[0026] A speech synthesis model loading module is used to load a speech synthesis model, wherein the speech synthesis model includes a voiceprint network, a timbre support network, and an acoustic network;

[0027] A learning object determination module is used to determine text information and the speaker's original speech signal;

[0028] An original spectrum signal conversion module, used to convert the original speech signal into an original spectrum signal;

[0029] A voiceprint feature encoding module, configured to encode the original spectrum signal into a voiceprint feature in the voiceprint network, wherein the voiceprint feature is used to verify the identity of the speaker;

[0030] a timbre supplementary feature encoding module, configured to encode the original spectrum signal into a timbre supplementary feature in the timbre support network, wherein the timbre supplementary feature is a feature missing from the voiceprint feature in timbre;

[0031] A timbre total feature fusion module, configured to fuse the voiceprint feature and the timbre supplementary feature into a timbre total feature;

[0032] The target spectrum signal fitting module is used to fit the target spectrum signal whose content is the text information and has the timbre total amount characteristics in the acoustic network.

[0033] In a fifth aspect, an embodiment of the present invention further provides a computer device, comprising:

[0034] one or more processors;

[0035] a memory for storing one or more programs,

[0036] When the one or more programs are executed by the one or more processors, the one or more processors implement the training method of the speech synthesis model as described in the first aspect or the speech synthesis method as shown in the second aspect.

[0037] In the sixth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the training method of the speech synthesis model as described in the first aspect or the speech synthesis method as shown in the second aspect.

[0038] In this embodiment, the original spectrum signal and the speaker's timbre embedding feature are obtained. The original spectrum signal is converted from the original speech signal recorded when the speaker speaks according to the text information. In the voiceprint network, the original spectrum signal is encoded into a voiceprint feature for verifying the speaker's identity. In the timbre support network, the original spectrum signal is encoded into a feature that the voiceprint feature lacks in timbre, which serves as a timbre supplementary feature. The voiceprint feature and the timbre supplementary feature are fused into a timbre total feature. Under the condition that the timbre total feature is corrected by the timbre embedding feature, the acoustic network and the timbre support network are trained based on the timbre total feature and the original spectrum signal. This embodiment uses voiceprint-based timbre learning to avoid tedious operations such as fine-tuning some or all parameters according to personalized timbre after speech synthesis model training is completed and saving the fine-tuned parameters. It also allows for recording of a short voice signal with no content restrictions, thereby maintaining a low threshold for personalized TTS and facilitating user operation. The voiceprint network can reuse pre-trained models, reducing the amount of computation required during training and the labor and time costs associated with labeling. Furthermore, considering that the primary function of the voiceprint network is to identify the speaker's identity rather than learning the speaker's timbre, the voiceprint features it generates, while possessing some timbre characteristics, also lack others. By supplementing the missing timbre features of the voiceprint features through the timbre support network, the comprehensiveness of the features in timbre can be ensured, thereby fitting a high-quality spectral signal and improving the timbre similarity between the fitted spectral signal and the target spectral signal. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 A flowchart of a method for training a speech synthesis model provided in Example 1 of the present invention;

[0040] Figure 2 A schematic diagram of a process for training a timbre support network and an acoustic network provided in the first embodiment of the present invention;

[0041] Figure 3 A schematic diagram of the structure of a timbre support network provided in the first embodiment of the present invention;

[0042] Figure 4 This is a flowchart of a method for training a speech synthesis model provided in the second embodiment of the present invention;

[0043] Figure 5 This is a flow chart of a speech synthesis method provided by Embodiment 3 of the present invention;

[0044] Figure 6 A flowchart of speech synthesis provided in the third embodiment of the present invention;

[0045] Figure 7 This is a flow chart of a speech synthesis method provided by the fourth embodiment of the present invention;

[0046] Figure 8 A schematic diagram of the structure of a speech synthesis model training device provided in a fifth embodiment of the present invention;

[0047] Figure 9 A schematic diagram of the structure of a speech synthesis device provided in Example 6 of the present invention;

[0048] Figure 10 A structural diagram of a computer device provided in Example 7 of the present invention. DETAILED DESCRIPTION

[0049] The present invention will be further described in detail below with reference to the accompanying drawings and examples. It will be understood that the specific embodiments described herein are intended only to illustrate the present invention and are not intended to limit the present invention. It should also be noted that, for ease of description, the accompanying drawings only illustrate portions relevant to the present invention, not all structures.

[0050] Example 1

[0051] Figure 1 This is a flowchart of a method for training a speech synthesis model provided in Example 1 of the present invention. This embodiment is applicable to the case of supplementing the missing timbre of a voiceprint on the basis of a voiceprint to thereby train a personalized speech synthesis model. The method can be performed by a training device for a speech synthesis model. The training device for a speech synthesis model can be implemented by software and / or hardware and can be configured in a computer device, such as a server, workstation, personal computer, etc. In this embodiment, the speech synthesis model includes a voiceprint network (Speaker Verification Encoder), a timbre support network (Tone Support Encoder), and an acoustic network. The method specifically includes the following steps:

[0052] Step 101: Obtain the original spectrum signal and the speaker's timbre embedding features.

[0053] like Figure 2 As shown, in order to facilitate the collection of a sufficient number of training sets, this embodiment can collect the voice signal recorded when the speaker speaks according to the established text information from common channels such as open source data sets and / or open source projects, and record it as the original voice signal, that is, the content of the original voice signal is the text information, and the text information can be obtained by performing voice recognition on the original voice signal.

[0054] The original speech signal can be converted into a spectrum signal, such as a Mel Spectrogram, by a method such as a fast Fourier transform (FFT), and recorded as an original spectrum signal. That is, the original spectrum signal is converted from the original speech signal recorded when the speaker speaks according to the text information.

[0055] In addition, the timbre features (vectors) of each speaker are learned in advance in an embedding table. The timbre features of the speaker can be found by looking up the table (Lookup Table), which is recorded as the timbre embedding feature (SpeakerInner Embedding).

[0056] Step 102: In the voiceprint network, encode the original spectrum signal into voiceprint features.

[0057] The voiceprint network can be used for SVF (Spearkr Verification) tasks, that is, the voiceprint network can learn the voiceprint features of the speaker, and the voiceprint features can be used to verify the identity of the speaker.

[0058] The voiceprint network is an encoder that can reuse pre-trained voiceprint networks from third parties or other projects, such as the GE2E network, to reduce the amount of computation required during training and the labor and time costs associated with labeling.

[0059] Of course, in addition to reusing the trained voiceprint network, you can also use the pre-trained voiceprint network for fine tuning (Fine Tuning), or you can independently train a new voiceprint network and then connect it to the speech synthesis model in this embodiment after training. This embodiment does not limit this.

[0060] In this embodiment, if Figure 2 As shown, the original spectrum signal can be input into the voiceprint network, and the voiceprint network performs encoding operations on the original spectrum signal and outputs voiceprint features for verifying the identity of the speaker.

[0061] Step 103: In the timbre support network, the original spectrum signal is encoded into timbre supplementary features.

[0062] There are two main ways to achieve personalized TTS. One of them is to fine-tune the TTS network based on the timbre (FineTuning), thereby realizing a personalized TTS network.

[0063] Specifically, a relatively large TTS network is trained in advance based on a relatively large data set, and then some or all parameters in the TTS network are fine-tuned using a small amount of speech signals of the timbre to be learned, and the fine-tuned parameters for the timbre to be learned are stored, which results in a relatively high learning cost.

[0064] In addition, this method has high requirements for the voice data used to learn the timbre. It generally requires users to record 20 voice signals in a quiet and noise-free environment. Moreover, the recorded voice signals are for specific text information and the pronunciation is error-free. The learning threshold is relatively high and the user usage remains unchanged.

[0065] Another method is to train the TTS network based on the personalized voiceprint-learnable timbre. This method generally requires recording a voice signal of about 5 seconds instead of 20 voice signals. In addition, there is no restriction on the content of the recorded voice signal (i.e. text information), which lowers the learning threshold and facilitates user use.

[0066] Since the main function of the voiceprint network is to identify the identity of the speaker, not to learn the speaker's timbre, the voiceprint features it generates have some timbre characteristics, but also lack some timbre characteristics. As a result, the timbre of the speech signal fitted by the TTS network is somewhat different from the timbre of the original speech signal, and the similarity is low.

[0067] In this embodiment, if Figure 2 As shown, an independent timbre support network can be additionally configured for the voiceprint network. The original spectrum signal is input into the timbre support network. The voiceprint network performs encoding operations on the original spectrum signal and outputs the features that are missing from the voiceprint feature in timbre, which are recorded as timbre supplementary features. That is, the timbre supplementary features are the features that are missing from the voiceprint feature in timbre.

[0068] The timbre support network belongs to the encoder. Due to its role as a timbre supplement, it is generally a lightweight network. Its structure is not limited to artificially designed neural networks. It can also be a neural network optimized by a model quantization method, a neural network searched for speech characteristics by a NAS (Neural Architecture Search) method, and so on. This embodiment does not impose any restrictions on this.

[0069] In one embodiment of the present invention, Figure 3 As shown, the timbre support network includes a first convolutional layer, multiple convolutional blocks, and a long short-term memory network (LSTM). In this embodiment, step 103 may include the following steps:

[0070] Step 1031: Input the original spectrum signal into the first convolutional layer to perform a convolution operation to obtain a first spectrum feature.

[0071] In this embodiment, if Figure 3 As shown, the original spectrum signal is input into the first convolution layer. The first convolution layer is composed of several convolution units, which can perform a one-dimensional (1D) convolution operation on the original spectrum signal and output a low-dimensional first spectrum feature.

[0072] Step 1032: Perform a first-level normalization operation on the first spectrum feature to obtain a second spectrum feature.

[0073] like Figure 3 As shown, in the timbre support network, different time steps share parameters, so that sequences of different lengths can be processed. For each time step in the first spectral feature, the first layer normalization operation (LayerNormalization, LN) can be performed to horizontally normalize different time steps, that is, each time step has its own distribution, and the second spectral feature is output, so that a single sample and variable-length sequence can be processed.

[0074] Step 1033: Input the second spectrum feature into multiple convolution blocks in sequence to perform convolution operations to obtain a third spectrum feature.

[0075] In this embodiment, each convolution block is an encapsulation and abstraction of some structures containing convolution, so as to facilitate the reuse of the structure and reduce the cost of research and development for technical personnel.

[0076] Furthermore, the number of convolution blocks can be set according to business requirements, and the structures of the convolution blocks can be the same or different, which is not limited in this embodiment.

[0077] Multiple convolution blocks are arranged in sequence, the output of the previous convolution block is the input of the next convolution block, the second spectrum features are input into the multiple convolution blocks in sequence to perform convolution operations, extract high-dimensional features, and output the third spectrum features.

[0078] In one example, if Figure 3 As shown, it is assumed that the structure of each convolution block is the same, and each convolution block includes a second convolutional layer.

[0079] Then, in this example, the first candidate feature input to each convolution block Block is determined. If the convolution block Block is ranked first, the first candidate feature is the second spectral feature. If the convolution block Block is not ranked first, the first candidate feature is the third candidate feature output by the previous convolution block Block.

[0080] In each convolution block, the first candidate feature is input into the second convolution block to perform a one-dimensional (1D) convolution operation to obtain the second candidate feature, and the second layer normalization operation LN under the self-attention mechanism is performed on the second candidate feature to obtain the third candidate feature.

[0081] Among them, the output sequence length of the self-attention mechanism (i.e., the second candidate feature) is the same as the length of the input sequence, and the corresponding output vector takes into account the information of the entire input sequence.

[0082] If the convolution block is not ranked last, the third candidate feature is output to the next convolution block.

[0083] If the convolution block is ranked last, the third candidate feature is output as the third spectral feature to the long short-term memory network.

[0084] To better understand this embodiment, the following describes the processing method of multiple convolution blocks in this embodiment through specific examples.

[0085] In this example, the timbre support network is configured with four convolution blocks, which are denoted as the first convolution block Block_1, the second convolution block Block_2, the third convolution block Block_3, and the fourth convolution block Block_4.

[0086] 1. In the first convolutional block Block_1:

[0087] Input the second spectrum feature Vectors_11.

[0088] The second spectrum feature Vectors_11 is input into the second convolution block to perform a convolution operation to obtain the second candidate feature Vectors_12.

[0089] Perform the second-layer normalization operation under the self-attention mechanism on the second candidate feature Vectors_12 to obtain the third candidate feature Vectors_13.

[0090] Output the third candidate feature Vectors_13 to the second convolution block Block_2.

[0091] 2. In the second convolution block Block_2:

[0092] Input the first candidate feature Vectors_21, which is the third candidate feature Vectors_13 output by the first convolution block Block_1.

[0093] The first candidate feature Vectors_21 is input into the second convolution block to perform a convolution operation to obtain the second candidate feature Vectors_22.

[0094] The second-layer normalization operation under the self-attention mechanism is performed on the second candidate feature Vectors_22 to obtain the third candidate feature Vectors_23.

[0095] Output the third candidate feature Vectors_23 to the third convolution block Block_3.

[0096] 3. In the third convolution block Block_3:

[0097] Input the first candidate feature Vectors_31, which is the third candidate feature Vectors_23 output by the second convolution block Block_2.

[0098] The first candidate feature Vectors_31 is input into the second convolution block to perform a convolution operation to obtain the second candidate feature Vectors_32.

[0099] The second-layer normalization operation under the self-attention mechanism is performed on the second candidate feature Vectors_32 to obtain the third candidate feature Vectors_33.

[0100] Output the third candidate feature Vectors_33 to the fourth convolution block Block_4.

[0101] 4. In the fourth convolution block Block_4:

[0102] Input the first candidate feature Vectors_41, which is the third candidate feature Vectors_33 output by the third convolution block Block_3.

[0103] The first candidate feature Vectors_41 is input into the second convolution block to perform a convolution operation to obtain the second candidate feature Vectors_42.

[0104] The second-layer normalization operation under the self-attention mechanism is performed on the second candidate feature Vectors_42 to obtain the third spectral feature Vectors_43.

[0105] Output the third spectral feature Vectors_43 to the long short-term memory network.

[0106] Step 1034: Input the third spectrum feature into the long short-term memory network for processing to obtain the features that are missing from the voiceprint feature in terms of timbre as supplementary features of the timbre.

[0107] In general, for long short-term memory networks, the following key variables can be divided:

[0108] Input: h t-1 (hidden layer at time t-1) and x t (eigenvector at time t)

[0109] Output: h t (Adding activation functions such as softmax can be used as the real output, otherwise it is used as a hidden layer)

[0110] Main line / memory: c t-1 and c t

[0111] There are three main stages within the LSTM network:

[0112] 1. Forgetting stage

[0113] The forgetting stage is mainly to selectively forget the input from the previous node, specifically by calculating the z f (f means forget) is used as a forget gate to control the c of the previous state t-1 What to keep and what to forget.

[0114] 2. Select the memory stage

[0115] The selective memory stage selectively "memorizes" the input, mainly the input x t Select memory, record the important ones and remember less important ones. The current input content is represented by the z calculated above. The selected gate signal is represented by z i (i stands for information) to control.

[0116] Adding the results of the above two stages, we can get c transmitted to the next state t .

[0117] 3. Output stage

[0118] The output stage will determine what will be considered as the output of the current state. Mainly through z o To control. And also to the c obtained in the previous stage o The image is scaled (transformed via a tanh activation function).

[0119] In this embodiment, if Figure 3 As shown, the third spectral feature is input into the long short-term memory network. The long short-term memory network processes the third spectral feature, controls the transmission state through the gate state, extracts more important information in time sequence, and forgets unimportant information, thereby obtaining the features that are missing in the timbre of the voiceprint feature, which are recorded as timbre supplementary features.

[0120] Of course, the above-described timbre support network is merely an example. When implementing the embodiments of the present invention, other timbre support networks may be configured based on actual circumstances. For example, a new convolutional layer may be added after the first convolutional layer, or the first convolutional layer may be replaced with an RNN (Recurrent Neural Network) layer, and so on. The embodiments of the present invention are not limited in this regard. Furthermore, in addition to the above-described timbre support network, those skilled in the art may also adopt other timbre support networks based on actual needs, and the embodiments of the present invention are not limited in this regard.

[0121] Step 104: Merge the voiceprint feature and the timbre supplementary feature into a timbre total feature.

[0122] The timbre supplementary feature is a supplement to the timbre dimension in the voiceprint feature. Therefore, Figure 2 As shown in FIG, by fusing the voiceprint feature with the timbre supplementary feature, a more comprehensive timbre feature can be obtained, which is recorded as the total timbre feature.

[0123] In a specific implementation, the addition operation Add can be performed on the voiceprint feature and the timbre supplementary feature to obtain the timbre total amount feature. Since the addition operation Add maintains the dimensions (channels) of the voiceprint feature and the timbre supplementary feature, and the amount of information (channels) in each dimension is increased, the new feature (timbre total amount feature) obtained by the addition operation Add can reflect some characteristics of the original feature (voiceprint feature and timbre supplementary feature). In this process, some information of the original feature may be lost.

[0124] Step 105: Under the condition that the timbre embedding feature modifies the timbre total amount feature, the acoustic network and the timbre support network are trained according to the timbre total amount feature and the original spectrum signal.

[0125] Experiments have shown that the multi-timbre TTS network trained with timbre embedding features can generate speech with a timbre very close to the original timbre. Figure 2 As shown in the figure, the timbre embedding features can be used as target features (Target Vectors) for learning timbre, and the total timbre features can be modified. Based on this, the total timbre features and the original spectrum signals are used to train the acoustic network and the timbre support network.

[0126] In one embodiment of the present invention, step 105 may include the following steps:

[0127] Step 1051: fuse part of the timbre embedding features and part of the timbre total amount features into a timbre correction feature.

[0128] The timbre embedding feature is used as the target feature to modify the total timbre feature, such as Figure 2 As shown, on the one hand, it can be reflected in the training of acoustic networks. Then, part of the timbre embedding features can be fused with part of the timbre total features to obtain new features, which are recorded as timbre correction features.

[0129] If the parameters in the timbre complement network are designed correctly, then after the partial timbre embedding features are fused with the partial timbre total features, the timbre correction features can be used to fit the spectrum signal with the correct timbre.

[0130] In the specific implementation, the batch size can be determined. The batch size is a hyperparameter that defines the number of samples (original audio signals, original spectrum signals, and text information) processed in a batch during the training of the acoustic network and the timbre support network.

[0131] In order to ensure the normal training of the acoustic network and the timbre support network, the timbre embedding features and the timbre total amount features meet the batch number. Therefore, the timbre embedding features can be divided into the first sub-embedding features and the second sub-embedding features according to the batch number, that is, the sum of the number of the first sub-embedding features and the second sub-embedding features is the batch number.

[0132] In general, the timbre embedding feature can be divided equally, that is, the ratio between the first sub-embedding feature and the second sub-embedding feature is 1:1, and both belong to half of the timbre embedding feature.

[0133] Of course, in addition to the equal division, the timbre embedding features may also be divided in other proportions, for example, 4:6, 3:7, 6:4, 7:3, etc., which is not limited in this embodiment.

[0134] In addition, the timbre total amount feature can be divided into a first sub-total amount feature and a second sub-total amount feature according to the batch quantity, that is, the sum of the quantities of the first sub-total amount feature and the second sub-total amount feature is the batch quantity.

[0135] In general, the timbre total amount feature can be divided evenly, that is, the ratio between the first sub-total amount feature and the second sub-total amount feature is 1:1, and both belong to half of the timbre total amount feature.

[0136] Of course, in addition to the equal division, other proportions may be used to divide the total timbre characteristics, for example, 4:6, 3:7, 6:4, 7:3, etc. This embodiment does not impose any limitation on this.

[0137] When fusing the partial timbre embedding feature with the partial timbre total amount feature, if the first sub-embedded feature and the second sub-timbre feature meet the batch quantity, that is, the first sub-embedded feature and the second sub-timbre feature are complementary, then a concatenation operation can be performed on the first sub-embedded feature and the second sub-timbre feature to obtain a timbre correction feature. Since the concatenation operation increases the dimension (channel) of the partial timbre embedding feature and the partial timbre total amount feature, the information under the dimension (channel) of the partial timbre embedding feature and the partial timbre total amount feature is not increased. Therefore, the new feature (timbre correction feature) obtained by the concatenation operation can reflect the characteristics of the original features (partial timbre embedding feature and partial timbre total amount feature). In this process, the information of the original features is not lost.

[0138] Step 1052: In the acoustic network, fit a target spectrum signal whose content is text information and has a timbre correction feature.

[0139] The acoustic model is used to convert text information into spectral signals that match a specified timbre. It can reuse pre-trained acoustic networks from third parties or other projects for fine-tuning, such as the FastSpeech network, Tacotron network, DeepVoice network, DurIAN network, and Transformer network, to reduce the amount of computation required during training and lower the labor and time costs associated with labeling.

[0140] Of course, in addition to reusing the pre-trained voiceprint network, a new acoustic network can also be built for training, which is not limited in this embodiment.

[0141] In this embodiment, if Figure 2 As shown, text information (represented in the form of phonemes, vectors, etc.) and timbre correction features are input into the acoustic network, and the acoustic network fits the target spectrum information. The content of the target spectrum information is the text information and has a timbre represented by the timbre correction features, that is, the content of the spoken words is the text information and the timbre features are the timbre correction features.

[0142] Taking the FastSpeech network as an example, to speed up inference, the FastSpeech network adopts a non-auto-regressive seq-to-seq (encoding-decoding) model. It does not rely on the input of the previous time step, allowing the entire FastSpeech network to be parallelized.

[0143] The text (phoneme) passes through the encoder to produce the encoder output. Since the length of the phoneme is often shorter than the mel-spectrogram, the FastSpeech network has a length regulator to adjust the decoder input length. This padded encoder output is padded to the same length as the mel-spectrogram, and then directly used as the decoder input. It is worth noting that the FastSpeech network uses 1D Conv layers instead of the fully connected network in the Transformer network.

[0144] The Duration Predictor is responsible for predicting the number of times each vector needs to be waited based on the Encoder Output. Here, in order to change the implicit Alignment of the traditional seq-to-seq model, the FastSpeech network adds an explicit Alignment label, which is provided by the Attention mechanism of Transformer-TTS (Transformer-based speech synthesis system). Transformer is a Multi-Head Attention mechanism, and one Head can be selected as the Alignment. The Duration Extractor extracts the target from the Attention Matrix of this Head. The prediction result of the Duration Predictor is compared with the target to perform MSE Loss (mean square error loss), and backpropagation is performed throughout the training process.

[0145] In addition, the speed of the synthesized speech can be controlled by human intervention in the prediction results of Duration Predictor.

[0146] Step 1053: Calculate the difference between the original spectrum signal and the target spectrum signal to obtain a first loss value.

[0147] In this embodiment, if Figure 2 As shown, the original spectrum signal and the target spectrum signal are substituted into the preset loss function to calculate the difference between the original spectrum signal and the target spectrum signal, and obtain the loss value LOSS, which is recorded as the first loss value. Then, the first loss value represents the loss of the speech synthesis model on the spectrum signal.

[0148] For example, the first loss value LOSS1 is calculated as follows:

[0149]

[0150] Where T is the number of frames, i is a positive integer, i∈T, the number of frames of the original spectrum signal is equal to the number of frames of the target spectrum signal, y i is the original spectrum signal, y' i is the target spectrum signal.

[0151] Step 1054: Calculate the difference between the partial timbre embedding feature and the partial timbre total feature as the second loss value.

[0152] In this embodiment, the partial timbre embedding feature is aligned with the partial timbre total amount feature so that the dimension of the partial timbre embedding feature is the same as the dimension of the partial timbre total amount feature. Figure 2 As shown, the partial timbre embedding features and the partial timbre total amount features are substituted into a preset loss function, such as the L1 norm, etc., so as to calculate the difference between the partial timbre embedding features and the partial timbre total amount features, and obtain the loss value LOSS, which is recorded as the second loss value. Then, the second loss value represents the loss of the speech synthesis model in the timbre features.

[0153] In a specific implementation, if the timbre total amount feature is divided into a first sub-total amount feature and a second sub-total amount feature according to the batch number, and the timbre total amount feature is divided into a first sub-total amount feature and a second sub-total amount feature according to the batch number, and a concatenation operation Concate is performed on the first sub-embedded feature and the second sub-timbre feature to obtain the timbre correction feature, then the difference between the second sub-embedded feature and the first sub-total amount feature can be calculated as the second loss value.

[0154] Taking the L1 norm as an example, the second loss value LOSS2 is calculated as follows:

[0155]

[0156] Among them, m is the number of dimensions, j is a positive integer, j∈m, the number of dimensions of the second sub-embedding feature is equal to the number of dimensions of the first sub-total feature, y (j) is the second sub-embedding feature, y' (j) It is the first sub-total characteristic.

[0157] The L1 norm, also known as the minimum absolute deviation (LAD) or minimum absolute error (LAE), minimizes the sum of the absolute differences between the target value (the second sub-embedded feature) and the estimated value (the first sub-total feature).

[0158] Step 1055: Merge the first loss value and the second loss value into a third loss value.

[0159] In this embodiment, if Figure 2 As shown, when updating the acoustic network and the timbre support network, the first loss value LOSS1 and the second loss value LOSS2 are comprehensively considered, and the first loss value LOSS1 and the second loss value LOSS2 are merged into the third loss value LOSS3, that is, the third loss value LOSS3 represents the comprehensive loss of the speech synthesis model in the spectral signal and timbre characteristics.

[0160] In a specific implementation, the fusion method can be linear fusion or nonlinear fusion, which is not limited in this embodiment.

[0161] For linear fusion, appropriate weights can be configured for the first loss value LOSS1 and the second loss value LOSS2 respectively according to business needs, so as to calculate the sum of the first loss value LOSS1 and the second loss value LOSS2 after the configured weights (i.e., weighted) as the third loss value LOSS3, i.e., LOSS3 = α*LOSS1+β*LOSS2, where α and β are weights.

[0162] In some services, if the spectrum signal is more important than the timbre feature, the weight of the first loss value LOSS1 is greater than the weight of the second loss value LOSS2.

[0163] In some businesses, if the importance of the spectral signal is similar to that of the timbre feature, then the weight of the first loss value LOSS1 is equal to the weight of the second loss value LOSS2. At this time, the first loss value LOSS1 can be added to the second loss value LOSS2 to obtain the third loss value LOSS3, that is, LOSS3 = LOSS1 + LOSS2.

[0164] In some services, if the timbre feature is more important than the spectrum signal, then the weight of the first loss value LOSS1 is greater than the weight of the second loss value LOSS2.

[0165] Step 1056: Update the acoustic network and the timbre support network according to the third loss value.

[0166] After completing forward propagation in the speech synthesis model, backpropagation can be performed on the acoustic network and the timbre support network. The third loss value can be substituted into optimization algorithms such as SGD (stochastic gradient descent) and Adam (Adaptive momentum) to calculate the first gradient of the parameters in the acoustic network and the second gradient of the parameters in the timbre support network, and then the parameters in the acoustic network and the timbre support network are updated according to the first gradient and the second gradient, respectively.

[0167] Step 1057 , determine whether the preset first training condition is met; if so, execute step 1058 ; if not, return to execute step 1012 .

[0168] Step 1058: Determine whether the timbre support network has completed training.

[0169] In this embodiment, a first training condition can be set in advance as a condition for stopping training the acoustic network and the timbre support network. For example, the number of iterations reaches a threshold, the amplitude of change of the third loss value for multiple consecutive times is less than a certain threshold, and so on. In each round of iterative training, it is determined whether the first training condition is met.

[0170] If the first training condition is met, the training of the acoustic network and the timbre support network can be considered complete. At this time, the parameters in the acoustic network and the timbre support network are output respectively and persisted in the database.

[0171] If the first training condition is not met, the next round of iterative training can be entered, and steps 102 to 105 can be executed again, and the iterative training can be repeated until the training of the acoustic network and the timbre support network is completed.

[0172] In this embodiment, the original spectrum signal and the speaker's timbre embedding feature are obtained. The original spectrum signal is converted from the original speech signal recorded when the speaker speaks according to the text information. In the voiceprint network, the original spectrum signal is encoded into a voiceprint feature for verifying the speaker's identity. In the timbre support network, the original spectrum signal is encoded into a feature that the voiceprint feature lacks in timbre, which serves as a timbre supplementary feature. The voiceprint feature and the timbre supplementary feature are fused into a timbre total feature. Under the condition that the timbre total feature is corrected by the timbre embedding feature, the acoustic network and the timbre support network are trained based on the timbre total feature and the original spectrum signal. This embodiment uses voiceprint-based timbre learning to avoid tedious operations such as fine-tuning some or all parameters according to personalized timbre after speech synthesis model training is completed and saving the fine-tuned parameters. It also allows for recording of a short voice signal with no content restrictions, thereby maintaining a low threshold for personalized TTS and facilitating user operation. The voiceprint network can reuse pre-trained models, reducing the amount of computation required during training and the labor and time costs associated with labeling. Furthermore, considering that the primary function of the voiceprint network is to identify the speaker's identity rather than learning the speaker's timbre, the voiceprint features it generates, while possessing some timbre characteristics, also lack others. By supplementing the missing timbre features of the voiceprint features through the timbre support network, the comprehensiveness of the features in timbre can be ensured, thereby fitting a high-quality spectral signal and improving the timbre similarity between the fitted spectral signal and the target spectral signal.

[0173] Example 2

[0174] Figure 4 This is a flowchart of a method for training a speech synthesis model provided in Example 2 of the present invention. This embodiment, based on the previous embodiment, further adds an operation for training a vocoder. The speech synthesis model includes a voiceprint network, a timbre support network, an acoustic network, and a vocoder. The method specifically includes the following steps:

[0175] Step 401: Obtain the original spectrum signal and the speaker's timbre embedding features.

[0176] The original spectrum signal is converted from the original speech signal recorded when the speaker speaks according to the text information.

[0177] Step 402: In the voiceprint network, encode the original spectrum signal into voiceprint features.

[0178] Among them, the voiceprint feature is used to verify the identity of the speaker.

[0179] Step 403: In the timbre support network, the original spectrum signal is encoded into timbre supplementary features.

[0180] Among them, the timbre supplementary feature is the feature that the voiceprint feature lacks in timbre.

[0181] Step 404: Merge the voiceprint feature and the timbre supplementary feature into a timbre total feature.

[0182] Step 405: Under the condition that the timbre embedding feature modifies the timbre total feature, the acoustic network and the timbre support network are trained according to the timbre total feature and the original spectrum signal.

[0183] Step 406: Acquire candidate speech signals and convert the candidate speech signals into candidate spectrum signals.

[0184] Step 407: Train the vocoder using the candidate spectrum signal as a sample and the candidate speech signal as a label.

[0185] The vocoder is used to fit spectral signals into speech signals. It can reuse pre-trained vocoders from third parties or other projects, such as the HiFiGAN network (Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis), MelGan network, PWGan network, VocGan network, etc., to reduce the amount of computation during training and the labor and time costs associated with labeling.

[0186] Of course, in addition to reusing the trained vocoder, you can also use a pre-trained vocoder for fine-tuning, or you can independently train a new voiceprint network and then connect it to the speech synthesis model in this embodiment after training. This embodiment does not limit this.

[0187] To improve the adaptability of the vocoder to the service, the pre-trained vocoder can be fine-tuned. When fine-tuning the vocoder, in order to facilitate the collection of a sufficient number of training sets, voice signals recorded when the speaker speaks according to the established text information can be collected from common channels such as open source data sets and / or open source projects, and recorded as candidate voice signals. That is, the content of the candidate voice signal is the text information, and voice recognition of the candidate voice signal can obtain the text information.

[0188] The candidate speech signal can be converted into a spectrum signal, such as MelSpectrogram, by means of fast Fourier transform or the like, and recorded as a candidate spectrum signal.

[0189] Furthermore, the original speech signal and the candidate speech signal may be the same or different, and the original spectrum signal and the candidate spectrum signal may be the same or different, which is not limited in this embodiment.

[0190] For a certain speaker, the candidate spectrum signal is used as the training sample and the candidate speech signal is used as the label Tag. Under the supervision of the candidate speech signal, the vocoder is trained in a supervised manner.

[0191] In one embodiment of the present invention, step 407 may include the following steps:

[0192] Step 4071: In the vocoder, convert the candidate spectrum signal into a reference speech signal.

[0193] In this embodiment, the candidate spectrum signal is input into the vocoder, and the vocoder predicts the speech signal according to the spectrum signal, which is recorded as the reference speech signal.

[0194] There are two main challenges in training the vocoder in the speech synthesis model: noisy speech signals and a limited number of samples in the dataset. In one example, to address these two challenges, HiFiGAN, which has high synthesis quality and fast speed, can be selected as the vocoder in the speech synthesis model.

[0195] HiFiGAN consists of a generator and two discriminators. Each discriminator has a sub-discriminator to generate an audio signal with a fixed period. The discriminators are a scale detector and a multi-period detector.

[0196] The generator is a convolutional neural network whose input is a candidate spectral signal (such as a Mel-spectrogram) and which increases the sampling until the number of output frames is the same as the specified duration.

[0197] Speech signals are composed of many sinusoidal signals of different periods. HiFiGAN can improve audio quality by modeling audio periodic patterns. In addition, HiFiGAN generates speech signals quickly.

[0198] Step 4072: Calculate the difference between the candidate speech signal and the reference speech signal to obtain a fourth loss value.

[0199] In this embodiment, the predicted speech signal (i.e., the reference speech signal) and the true speech signal (i.e., the candidate speech signal) are substituted into a preset loss function, such as the L1 norm, the L2 norm, etc., so as to calculate the difference between the predicted speech signal (i.e., the reference speech signal) and the true speech signal (i.e., the candidate speech signal), and obtain a loss value LOSS, which is recorded as the fourth loss value. Then, the fourth loss value represents the loss of the speech synthesis model on the speech signal.

[0200] Step 4073: Update the vocoder according to the fourth loss value.

[0201] After completing forward propagation in the vocoder, the vocoder can be back-propagated, and the fourth loss value can be substituted into the optimization algorithm such as SGD and Adam to calculate and update the third gradient of the parameters in the vocoder, and the parameters in the vocoder are updated according to the third gradient.

[0202] Step 4074: Determine whether the preset second training condition is met; if so, execute step 4075; if not, return to execute step 4071.

[0203] Step 4075: Determine whether the vocoder has completed training.

[0204] In this embodiment, a second training condition can be pre-set as a condition for stopping training the vocoder, for example, the number of iterations reaches a threshold, the amplitude of change of the fourth loss value is less than a certain threshold for multiple consecutive times, etc. In each round of iterative training, it is determined whether the second training condition is met.

[0205] If the second training condition is met, the vocoder training is considered complete. At this point, the parameters in the vocoder are output and persisted in the database.

[0206] If the second training condition is not met, the next round of iterative training can be entered, and steps 4077 to 4073 can be executed again, and the iterative training is repeated in this way until the vocoder training is completed.

[0207] Example 3

[0208] Figure 5This is a flowchart of a speech synthesis method provided in Example 1 of the present invention. This embodiment is applicable to the case where a missing timbre is supplemented based on a voiceprint to perform speech synthesis. The method can be performed by a speech synthesis device. The speech synthesis device can be implemented by software and / or hardware and can be configured in a computer device, such as a server, workstation, personal computer, mobile terminal (such as a mobile phone, tablet computer), wearable device (such as a watch, glasses, etc.), smart toy (such as a storytelling machine), vehicle-mounted terminal, etc. The method specifically includes the following steps:

[0209] Step 501: Load the speech synthesis model.

[0210] In this embodiment, the operating systems in the computer device include Windows, Android, iOS, etc., and these operating systems can support running speech synthesis applications, such as navigation applications, storytelling applications, novel applications, news applications, live broadcast applications, short video quotations, instant messaging tools, conference applications, etc.

[0211] While the application is running, the speech synthesis model and its parameters can be loaded into the memory and run, waiting for speech synthesis.

[0212] Among them, Figure 4 As shown in Figure 1, the speech synthesis model includes a voiceprint network, a timbre support network, and an acoustic network.

[0213] In one embodiment of the present invention, the training method of the timbre support network and the acoustic network is as follows:

[0214] Obtaining an original spectrum signal and a speaker's timbre embedding feature, where the original spectrum signal is converted from an original speech signal recorded when the speaker speaks according to the text information;

[0215] In the voiceprint network, the original spectrum signal is encoded into voiceprint features, which are used to verify the identity of the speaker;

[0216] In the timbre support network, the original spectrum signal is encoded into timbre supplementary features, which are the features that the voiceprint features lack in timbre;

[0217] The voiceprint feature and the timbre supplementary feature are combined into the total timbre feature;

[0218] Under the condition that the timbre embedding feature modifies the timbre total amount feature, the acoustic network and the timbre support network are trained according to the timbre total amount feature and the original spectrum signal.

[0219] In this embodiment, since the training methods of the timbre support network and the acoustic network are basically similar to those used in the first embodiment, the description is relatively simple. For relevant details, please refer to the partial description of the first embodiment, and this embodiment will not be described in detail here.

[0220] Step 502: Determine text information and the speaker's original voice signal.

[0221] like Figure 4 As shown, on the one hand, the user acts as a speaker and provides a speech signal of the timbre to be learned in the application by recording, uploading a file, etc. during the first speech synthesis, which is recorded as the original speech signal. After the characteristics of the speaker's original speech signal (total timbre characteristics) are learned, the original speech signal is not necessarily uploaded during non-first speech synthesis.

[0222] On the other hand, the user selects text information to be synthesized into speech in the application, such as the content of a story, navigation, novel, news, web page, etc.

[0223] Furthermore, for the speech synthesis model based on the voiceprint network, the user only needs to input an original speech signal of about 5 seconds. Moreover, the text content is arbitrary, and the original speech signal and the text content are not necessarily related.

[0224] Step 503: Convert the original speech signal into an original spectrum signal.

[0225] In this embodiment, if Figure 4 As shown in FIG, the original speech signal can be converted into a spectrum signal, such as Mel Spectrogram, by means of fast Fourier transform and the like, and is recorded as the original spectrum signal.

[0226] Step 504: In the voiceprint network, encode the original spectrum signal into voiceprint features.

[0227] In this embodiment, if Figure 4 As shown, the original spectrum signal can be input into the voiceprint network, which performs encoding operations on the original spectrum signal and outputs voiceprint features, which are used to verify the identity of the speaker.

[0228] Step 505: In the timbre support network, encode the original spectrum signal into timbre supplementary features.

[0229] In this embodiment, if Figure 4 As shown, the original spectrum signal is input into the timbre support network, and the voiceprint network performs encoding operations on the original spectrum signal and outputs timbre supplementary features. The timbre supplementary features are the features that are missing from the voiceprint features in timbre.

[0230] In one embodiment of the present invention, the timbre support network includes a first convolutional layer, multiple convolutional blocks, and a long short-term memory network; in this embodiment, step 505 may include the following steps:

[0231] Step 5051: Input the original spectrum signal into the first convolution layer to perform a convolution operation to obtain a first spectrum feature;

[0232] Step 5052: Perform a first-level normalization operation on the first spectrum feature to obtain a second spectrum feature;

[0233] Step 5053: input the second spectrum feature into multiple convolution blocks in sequence to perform convolution operations to obtain a third spectrum feature;

[0234] Step 5054: Input the third spectrum feature into the long short-term memory network for processing to obtain the features that are missing from the voiceprint feature in the timbre as the timbre supplementary features.

[0235] Furthermore, each convolution block includes a second convolution layer; step 5053 may include the following steps:

[0236] Determine the first candidate feature input to each convolution block. If the convolution block is ranked first, the first candidate feature is the second spectrum feature. If the convolution block is not ranked first, the first candidate feature is the third candidate feature output by the previous convolution block.

[0237] In each convolution block, the first candidate feature is input into the second convolution block to perform a convolution operation to obtain the second candidate feature;

[0238] Perform the second-layer normalization operation under the self-attention mechanism on the second candidate feature to obtain the third candidate feature;

[0239] If the convolution block is not ranked last, the third candidate feature is output to the next convolution block;

[0240] If the convolution block is ranked last, the third candidate feature is output as the third spectral feature to the long short-term memory network.

[0241] In this embodiment, since the application of step 505 is basically similar to that of step 103, the description is relatively simple. For relevant details, please refer to the partial description of step 103, and this embodiment will not be described in detail here.

[0242] Step 506: Merge the voiceprint feature and the timbre supplementary feature into a timbre total feature.

[0243] The timbre supplement feature is a supplement to the timbre dimension in the voiceprint feature. Therefore, by fusing the voiceprint feature with the timbre supplement feature, a more comprehensive timbre feature can be obtained, which is recorded as the total timbre feature.

[0244] In a specific implementation, an addition operation Add may be performed on the voiceprint feature and the timbre supplementary feature to obtain the timbre total amount feature.

[0245] Step 507: In the acoustic network, fit a target spectrum signal having text information and total timbre characteristics.

[0246] In this embodiment, if Figure 4 As shown, text information (represented in the form of phonemes, vectors, etc.) and timbre total amount features are input into the acoustic network, and the acoustic network fits the target spectrum information. The content of the target spectrum information is the text information and has a timbre represented by the timbre total amount features, that is, the content of the spoken words is the text information and the timbre features are the timbre total amount features.

[0247] In this embodiment, a speech synthesis model is loaded. The speech synthesis model includes a voiceprint network, a timbre support network, and an acoustic network. The text information and the speaker's original speech signal are determined, and the original speech signal is converted into an original spectrum signal. In the voiceprint network, the original spectrum signal is encoded as a voiceprint feature for verifying the speaker's identity; in the timbre support network, the original spectrum signal is encoded as a feature that is missing from the voiceprint feature in timbre, as a timbre supplementary feature; the voiceprint feature and the timbre supplementary feature are fused into a timbre total feature; in the acoustic network, a target spectrum signal whose content is text information and has a timbre total feature is fitted. This embodiment uses voiceprint-based timbre learning to avoid tedious operations such as fine-tuning some or all parameters according to personalized timbre after speech synthesis model training is completed and saving the fine-tuned parameters. It also allows for recording of a short voice signal with no content restrictions, thereby maintaining a low threshold for personalized TTS and facilitating user operation. The voiceprint network can reuse pre-trained models, reducing the amount of computation required during training and the labor and time costs associated with labeling. Furthermore, considering that the primary function of the voiceprint network is to identify the speaker's identity rather than learning the speaker's timbre, the voiceprint features it generates, while possessing some timbre characteristics, also lack others. By supplementing the missing timbre features of the voiceprint features through the timbre support network, the comprehensiveness of the features in timbre can be ensured, thereby fitting a high-quality spectral signal and improving the timbre similarity between the fitted spectral signal and the target spectral signal.

[0248] Example 4

[0249] Figure 7 This is a flowchart of a speech synthesis method provided in Embodiment 4 of the present invention. This embodiment is based on the previous embodiment and further adds an operation of converting a speech signal by a vocoder. The method specifically includes the following steps:

[0250] Step 701: Load the speech synthesis model.

[0251] Among them, the speech synthesis model includes voiceprint network, timbre support network, acoustic network, and vocoder.

[0252] Step 702: Determine text information and the speaker's original voice signal.

[0253] Step 703: Convert the original speech signal into an original spectrum signal.

[0254] Step 704: In the voiceprint network, encode the original spectrum signal into voiceprint features.

[0255] Among them, the voiceprint feature is used to verify the identity of the speaker.

[0256] Step 705: In the timbre support network, encode the original spectrum signal into timbre supplementary features.

[0257] Among them, the timbre supplementary feature is the feature that the voiceprint feature lacks in timbre.

[0258] Step 706: Merge the voiceprint feature and the timbre supplementary feature into a timbre total feature.

[0259] Step 707: In the acoustic network, fit a target spectrum signal having text information and total timbre characteristics.

[0260] Step 708: In the vocoder, convert the target spectrum signal into a target speech signal.

[0261] In this embodiment, the target spectrum signal is input into the vocoder, and the vocoder predicts the speech signal according to the spectrum signal, which is recorded as the target speech signal.

[0262] The vocoder may be a pre-trained network or may be fine-tuned based on a pre-trained network.

[0263] In one embodiment of the present invention, for fine-tuning, the vocoder training method is as follows:

[0264] Acquire a candidate speech signal and a candidate spectrum signal converted from the candidate speech signal;

[0265] The vocoder is trained using the candidate spectrum signals as samples and the candidate speech signals as labels.

[0266] In this embodiment, since the training method of the vocoder is basically similar to the application of the second embodiment, the description is relatively simple. For relevant details, please refer to the partial description of the second embodiment, and this embodiment will not be described in detail here.

[0267] In this embodiment, since the timbre similarity between the fitted spectrum signal and the target spectrum signal is improved, the timbre similarity between the speech signal generated based on the fitted spectrum signal and the target speech signal can be further improved.

[0268] It should be noted that for the sake of simplicity, the method embodiments are described as a series of actions. However, those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.

[0269] Example 5

[0270] Figure 8 This is a block diagram of a training device for a speech synthesis model provided in Example 5 of the present invention. The speech synthesis model includes a voiceprint network, a timbre support network, and an acoustic network. The device may specifically include the following modules:

[0271] The data set acquisition module 801 is used to acquire the original spectrum signal and the speaker's timbre embedding features, wherein the original spectrum signal is converted from the original speech signal recorded when the speaker speaks according to the text information;

[0272] A voiceprint feature encoding module 802 is configured to encode the original spectrum signal into a voiceprint feature in the voiceprint network, wherein the voiceprint feature is used to verify the identity of the speaker;

[0273] A timbre supplementary feature encoding module 803 is configured to encode the original spectrum signal into a timbre supplementary feature in the timbre support network, wherein the timbre supplementary feature is a feature that is missing from the voiceprint feature in timbre;

[0274] The timbre total feature fusion module 804 is used to fuse the voiceprint feature and the timbre supplementary feature into a timbre total feature;

[0275] The network correction training module 805 is used to train the acoustic network and the timbre support network according to the timbre total amount feature and the original spectrum signal under the condition that the timbre embedding feature corrects the timbre total amount feature.

[0276] In one embodiment of the present invention, the timbre support network includes a first convolutional layer, a plurality of convolutional blocks, and a long short-term memory network;

[0277] The timbre supplementary feature encoding module 803 includes:

[0278] a first convolution operation module, configured to input the original spectrum signal into the first convolution layer to perform a convolution operation to obtain a first spectrum feature;

[0279] A first-layer normalization operation module, configured to perform a first-layer normalization operation on the first spectrum feature to obtain a second spectrum feature;

[0280] a block convolution module, configured to sequentially input the second spectrum feature into a plurality of the convolution blocks to perform a convolution operation to obtain a third spectrum feature;

[0281] The time series processing module is used to input the third spectrum feature into the long short-term memory network for processing, and obtain the feature missing from the voiceprint feature in timbre as a timbre supplementary feature.

[0282] In one embodiment of the present invention, each of the convolutional blocks includes a second convolutional layer;

[0283] The block convolution module includes:

[0284] a candidate feature determination module, configured to determine a first candidate feature input to each of the convolution blocks, where if the convolution block is ranked first, the first candidate feature is the second spectrum feature; and if the convolution block is not ranked first, the first candidate feature is the third candidate feature output by the previous convolution block;

[0285] a second convolution operation module, configured to, in each of the convolution blocks, input the first candidate feature into the second convolution block to perform a convolution operation to obtain a second candidate feature;

[0286] A second-layer normalization operation module is used to perform a second-layer normalization operation under the self-attention mechanism on the second candidate feature to obtain a third candidate feature;

[0287] A candidate feature output module, configured to output the third candidate feature to the next convolution block if the convolution block is not ranked last;

[0288] A spectrum feature output module is configured to output the third candidate feature as a third spectrum feature to the long short-term memory network if the convolution block is ranked last.

[0289] In one embodiment of the present invention, the timbre total feature fusion module 804 includes:

[0290] The addition operation execution module is used to perform an addition operation Add on the voiceprint feature and the timbre supplementary feature to obtain a timbre total amount feature.

[0291] In one embodiment of the present invention, the network correction training module 805 includes:

[0292] A timbre correction feature fusion module, configured to fuse part of the timbre embedding feature and part of the timbre total feature into a timbre correction feature;

[0293] a target spectrum signal fitting module, configured to fit a target spectrum signal having the text information and the timbre correction feature in the acoustic network;

[0294] a first loss value calculation module, configured to calculate a difference between the original spectrum signal and the target spectrum signal to obtain a first loss value;

[0295] a second loss value calculation module, configured to calculate a difference between a portion of the timbre embedding feature and a portion of the timbre total feature as a second loss value;

[0296] A third loss value fusion module, configured to fuse the first loss value and the second loss value into a third loss value;

[0297] a network updating module, configured to update the acoustic network and the timbre support network according to the third loss value;

[0298] A first training condition determination module is configured to determine whether a preset first training condition is met; if so, the training completion determination module is called; if not, the voiceprint feature encoding module 802 is called back;

[0299] The training completion determination module is used to determine whether the timbre support network has completed training.

[0300] In one embodiment of the present invention, the network correction training module 805 further includes:

[0301] Batch quantity determination module, used to determine the batch quantity;

[0302] a timbre embedding feature division module, configured to divide the timbre embedding feature into a first sub-embedding feature and a second sub-embedding feature according to the batch number;

[0303] a timbre total amount feature division module, configured to divide the timbre total amount feature into a first sub-total amount feature and a second sub-total amount feature according to the batch number;

[0304] The timbre correction feature fusion module includes:

[0305] a concatenation operation execution module, configured to perform a concatenation operation on the first sub-embedding feature and the second sub-timbre feature if the first sub-embedding feature and the second sub-timbre feature meet the batch quantity, so as to obtain a timbre correction feature;

[0306] The second loss value calculation module includes:

[0307] A feature difference calculation module is used to calculate the difference between the second sub-embedding feature and the first sub-total feature as a second loss value.

[0308] In one embodiment of the present invention, the third loss value fusion module includes:

[0309] A loss adding module is used to add the first loss value and the second loss value to obtain a third loss value.

[0310] In one embodiment of the present invention, the speech synthesis model further includes a vocoder, and the apparatus further includes:

[0311] A training set acquisition module, configured to acquire candidate speech signals and convert the candidate speech signals into candidate spectrum signals;

[0312] The vocoder training module is used to train the vocoder using the candidate spectrum signal as a sample and the candidate speech signal as a label.

[0313] In one embodiment of the present invention, the vocoder training module includes:

[0314] a reference speech signal conversion module, configured to convert the candidate spectrum signal into a reference speech signal in the vocoder;

[0315] a fourth loss value calculation module, configured to calculate a difference between the candidate speech signal and the reference speech signal to obtain a fourth loss value;

[0316] a vocoder updating module, configured to update the vocoder according to the fourth loss value;

[0317] A second training condition judgment module is used to judge whether a preset second training condition is met; if so, calling the vocoder completion determination module; if not, returning to calling the reference speech signal conversion module;

[0318] The vocoder completion determination module is used to determine whether the vocoder has completed training.

[0319] The training device for the speech synthesis model provided in the embodiment of the present invention can execute the training method for the speech synthesis model provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0320] Example 6

[0321] Figure 9 This is a structural block diagram of a speech synthesis device provided in Example 6 of the present invention, which may specifically include the following modules:

[0322] The speech synthesis model loading module 901 is used to load the speech synthesis model, wherein the speech synthesis model includes a voiceprint network, a timbre support network, and an acoustic network;

[0323] A learning object determination module 902 is used to determine text information and the speaker's original speech signal;

[0324] The original spectrum signal conversion module 903 is used to convert the original speech signal into an original spectrum signal;

[0325] A voiceprint feature encoding module 904 is configured to encode the original spectrum signal into a voiceprint feature in the voiceprint network, wherein the voiceprint feature is used to verify the identity of the speaker;

[0326] A timbre supplementary feature encoding module 905 is configured to encode the original spectrum signal into a timbre supplementary feature in the timbre support network, wherein the timbre supplementary feature is a feature missing from the voiceprint feature in timbre;

[0327] The timbre total feature fusion module 906 is used to fuse the voiceprint feature and the timbre supplementary feature into a timbre total feature;

[0328] The target spectrum signal fitting module 907 is configured to fit a target spectrum signal containing the text information and having the timbre total amount characteristics in the acoustic network.

[0329] In one embodiment of the present invention, the training methods of the acoustic network and the timbre support network are as follows:

[0330] Obtaining an original spectrum signal and a speaker's timbre embedding feature, wherein the original spectrum signal is converted from an original speech signal recorded when the speaker speaks according to the text information;

[0331] In the voiceprint network, encoding the original spectrum signal into a voiceprint feature for verifying the speaker's identity;

[0332] In the timbre support network, the original spectrum signal is encoded into the features missing from the voiceprint feature in timbre as timbre supplementary features;

[0333] Merging the voiceprint feature and the timbre supplementary feature into a timbre total feature;

[0334] Under the condition that the timbre embedding feature modifies the timbre total amount feature, the acoustic network and the timbre support network are trained according to the timbre total amount feature and the original spectrum signal.

[0335] In one embodiment of the present invention, the speech synthesis model further includes a vocoder, and the training method of the vocoder is as follows:

[0336] Acquire a candidate speech signal and a candidate spectrum signal converted from the candidate speech signal;

[0337] The vocoder is trained using the candidate spectrum signal as a sample and the candidate speech signal as a label.

[0338] In one embodiment of the present invention, the speech synthesis model further includes a vocoder, and the apparatus further includes:

[0339] The target speech signal conversion module is used to convert the target spectrum signal into a target speech signal in the vocoder.

[0340] In one embodiment of the present invention, the timbre supplementary feature encoding module 905 includes:

[0341] a first convolution operation module, configured to input the original spectrum signal into the first convolution layer to perform a convolution operation to obtain a first spectrum feature;

[0342] A first-layer normalization operation module, configured to perform a first-layer normalization operation on the first spectrum feature to obtain a second spectrum feature;

[0343] a block convolution module, configured to sequentially input the second spectrum feature into a plurality of the convolution blocks to perform a convolution operation to obtain a third spectrum feature;

[0344] The time series processing module is used to input the third spectrum feature into the long short-term memory network for processing, and obtain the feature missing from the voiceprint feature in timbre as a timbre supplementary feature.

[0345] In one embodiment of the present invention, each of the convolutional blocks includes a second convolutional layer;

[0346] The block convolution module includes:

[0347] a candidate feature determination module, configured to determine a first candidate feature input to each of the convolution blocks, where if the convolution block is ranked first, the first candidate feature is the second spectrum feature; and if the convolution block is not ranked first, the first candidate feature is the third candidate feature output by the previous convolution block;

[0348] a second convolution operation module, configured to, in each of the convolution blocks, input the first candidate feature into the second convolution block to perform a convolution operation to obtain a second candidate feature;

[0349] A second-layer normalization operation module is used to perform a second-layer normalization operation under the self-attention mechanism on the second candidate feature to obtain a third candidate feature;

[0350] A candidate feature output module, configured to output the third candidate feature to the next convolution block if the convolution block is not ranked last;

[0351] A spectrum feature output module is configured to output the third candidate feature as a third spectrum feature to the long short-term memory network if the convolution block is ranked last.

[0352] In one embodiment of the present invention, the timbre total feature fusion module 906 includes:

[0353] The addition operation execution module is used to perform an addition operation Add on the voiceprint feature and the timbre supplementary feature to obtain a timbre total amount feature.

[0354] The speech synthesis device provided in the embodiment of the present invention can execute the speech synthesis method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0355] Example 7

[0356] Figure 10 A structural diagram of a computer device provided in Example 7 of the present invention. Figure 10 A block diagram of an exemplary computer device 12 suitable for use in implementing embodiments of the present invention is shown. Figure 10 The computer device 12 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.

[0357] like Figure 10 As shown, computer device 12 is implemented as a general-purpose computing device. Components of computer device 12 may include, but are not limited to, one or more processors or processing units 16, system memory 28, and a bus 18 that connects various system components (including system memory 28 and processing unit 16).

[0358] Bus 18 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor, or a local bus using any of a variety of bus architectures. Examples of these architectures include, but are not limited to, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MAC) bus, an Enhanced ISA bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus.

[0359] The computer device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the computer device 12, including volatile and non-volatile media, removable and non-removable media.

[0360] System memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. Computer device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 may be configured to read and write non-removable, non-volatile magnetic media ( Figure 10Not shown, often called a "hard drive"). Although Figure 10 Not shown, a magnetic disk drive for reading and writing to a removable non-volatile magnetic disk (e.g., a "floppy disk"), and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to bus 18 via one or more data media interfaces. Memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of various embodiments of the present invention.

[0361] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data, each of which, or some combination thereof, may include an implementation of a network environment. Program modules 42 generally implement the functions and / or methods of the embodiments described herein.

[0362] The computer device 12 can also communicate with one or more external devices 14 (e.g., a keyboard, pointing device, display 24, etc.), one or more devices that enable a user to interact with the computer device 12, and / or any device that enables the computer device 12 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). Such communication can occur via an input / output (I / O) interface 22. Furthermore, the computer device 12 can communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network such as the Internet) via a network adapter 20. As shown, the network adapter 20 communicates with the other modules of the computer device 12 via a bus 18. It should be understood that, although not shown, other hardware and / or software modules can be used in conjunction with the computer device 12, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0363] The processing unit 16 executes various functional applications and data processing by running programs stored in the system memory 28, such as implementing the training method of the speech synthesis model or the speech synthesis method provided in the embodiment of the present invention.

[0364] Example 8

[0365] Embodiment 8 of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the training method of the above-mentioned speech synthesis model or each process of the speech synthesis method, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0366] Among them, computer-readable storage media can include, for example, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or components, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or device.

[0367] Note that the above are only preferred embodiments of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and that various obvious changes, readjustments, and substitutions can be made by those skilled in the art without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments and may include many other equivalent embodiments without departing from the concept of the present invention. The scope of the present invention is determined by the scope of the appended claims.

Claims

1. A method for training a speech synthesis model, characterized in that: The speech synthesis model includes a voiceprint network, a timbre support network, and an acoustic network, and the method includes: Obtaining an original spectrum signal and a speaker's timbre embedding feature, wherein the original spectrum signal is converted from an original speech signal recorded when the speaker speaks according to the text information; In the voiceprint network, the original spectrum signal is encoded into a voiceprint feature, and the voiceprint feature is used to verify the identity of the speaker; In the timbre support network, the original spectrum signal is encoded into a timbre supplementary feature, where the timbre supplementary feature is a feature missing from the voiceprint feature in timbre; Merging the voiceprint feature and the timbre supplementary feature into a timbre total feature; Under the condition that the timbre embedding feature modifies the timbre total amount feature, the acoustic network and the timbre support network are trained according to the timbre total amount feature and the original spectrum signal; Wherein, under the condition that the timbre embedding feature corrects the timbre total amount feature, training the acoustic network and the timbre support network according to the timbre total amount feature and the original spectrum signal includes: fusing part of the timbre embedding feature and part of the timbre total amount feature into a timbre correction feature; In the acoustic network, fitting a target spectrum signal having the text information and the timbre correction feature; Calculating a difference between the original spectrum signal and the target spectrum signal to obtain a first loss value; Calculating a difference between part of the timbre embedding feature and part of the timbre total amount feature as a second loss value; Merging the first loss value and the second loss value into a third loss value; updating the acoustic network and the timbre support network according to the third loss value; Determine whether a preset first training condition is met; if so, determine that the timbre support network has completed training; if not, return to executing the encoding of the original spectrum signal into a voiceprint feature in the voiceprint network.

2. The method according to claim 1, characterized in that The timbre support network includes a first convolutional layer, a plurality of convolutional blocks, and a long short-term memory network; The step of encoding the original spectrum signal into a timbre supplementary feature in the timbre support network includes: Inputting the original spectrum signal into the first convolution layer to perform a convolution operation to obtain a first spectrum feature; performing a first-layer normalization operation on the first spectrum feature to obtain a second spectrum feature; Inputting the second spectrum feature into the plurality of convolution blocks in sequence to perform convolution operations to obtain a third spectrum feature; The third spectrum feature is input into the long short-term memory network for processing to obtain the feature missing from the voiceprint feature in terms of timbre as a timbre supplementary feature.

3. The method according to claim 2, characterized in that Each of the convolutional blocks includes a second convolutional layer; The step of sequentially inputting the second spectrum feature into a plurality of convolution blocks to perform a convolution operation to obtain a third spectrum feature includes: Determine a first candidate feature input to each of the convolution blocks, if the convolution block is ranked first, then the first candidate feature is the second spectrum feature; if the convolution block is not ranked first, then the first candidate feature is the third candidate feature output by the previous convolution block; In each of the convolution blocks, the first candidate feature is input into the second convolution layer to perform a convolution operation to obtain a second candidate feature; Performing a second-layer normalization operation under the self-attention mechanism on the second candidate feature to obtain a third candidate feature; If the convolution block is not ranked last, outputting the third candidate feature to the next convolution block; If the convolution block is ranked last, the third candidate feature is output as the third spectral feature to the long short-term memory network.

4. The method according to claim 1, wherein The method further comprises: training the acoustic network and the timbre support network according to the timbre total amount feature and the original spectrum signal under the condition that the timbre embedding feature modifies the timbre total amount feature; Determine the batch quantity; Dividing the timbre embedding feature into a first sub-embedding feature and a second sub-embedding feature according to the batch number; dividing the timbre total amount feature into a first sub-total amount feature and a second sub-total amount feature according to the batch number; The step of fusing part of the timbre embedding feature and part of the timbre total amount feature into a timbre correction feature includes: If the first sub-embedded feature and the second sub-total feature meet the batch quantity, performing a splicing operation on the first sub-embedded feature and the second sub-total feature to obtain a timbre correction feature; The calculating, as a second loss value, a difference between part of the timbre embedding feature and part of the timbre total feature includes: The difference between the second sub-embedded feature and the first sub-total feature is calculated as a second loss value.

5. The method according to any one of claims 1 to 4, characterized in that The speech synthesis model further includes a vocoder, and the method further includes: Acquire a candidate speech signal and convert the candidate speech signal into a candidate spectrum signal; The vocoder is trained using the candidate spectrum signal as a sample and the candidate speech signal as a label.

6. The method according to claim 5, characterized in that The step of training the vocoder using the candidate spectrum signal as a sample and the candidate speech signal as a label includes: In the vocoder, converting the candidate spectrum signal into a reference speech signal; Calculating a difference between the candidate speech signal and the reference speech signal to obtain a fourth loss value; updating the vocoder according to the fourth loss value; Determine whether a preset second training condition is met; if so, determine that the vocoder has completed training; if not, return to executing the conversion of the candidate spectrum signal into a reference speech signal in the vocoder.

7. A speech synthesis method, characterized in that: include: Loading a speech synthesis model, the speech synthesis model including a voiceprint network, a timbre support network, and an acoustic network; Determine the text information and the speaker's original speech signal; Converting the original speech signal into an original spectrum signal; In the voiceprint network, the original spectrum signal is encoded into a voiceprint feature, and the voiceprint feature is used to verify the identity of the speaker; In the timbre support network, the original spectrum signal is encoded into a timbre supplementary feature, where the timbre supplementary feature is a feature missing from the voiceprint feature in timbre; Merging the voiceprint feature and the timbre supplementary feature into a timbre total feature; In the acoustic network, a target spectrum signal having the text information and the total amount of timbre characteristics is fitted.

8. The method according to claim 7, characterized in that The speech synthesis model further includes a vocoder, and the method further includes: In the vocoder, the target spectrum signal is converted into a target speech signal.

9. A training device for a speech synthesis model, characterized in that: The speech synthesis model includes a voiceprint network, a timbre support network, and an acoustic network, and the device includes: A data set acquisition module is used to acquire an original spectrum signal and a speaker's timbre embedding feature, wherein the original spectrum signal is converted from an original speech signal recorded when the speaker speaks according to the text information; a voiceprint feature encoding module, configured to encode the original spectrum signal into a voiceprint feature in the voiceprint network, wherein the voiceprint feature is used to verify the identity of the speaker; a timbre supplementary feature encoding module, configured to encode the original spectrum signal into a timbre supplementary feature in the timbre support network, wherein the timbre supplementary feature is a feature missing from the voiceprint feature in timbre; A timbre total feature fusion module, configured to fuse the voiceprint feature and the timbre supplementary feature into a timbre total feature; A network correction training module is used to train the acoustic network and the timbre support network according to the timbre total amount feature and the original spectrum signal under the condition that the timbre embedding feature corrects the timbre total amount feature; Wherein, the network correction training module includes: A timbre correction feature fusion module, configured to fuse part of the timbre embedding feature and part of the timbre total feature into a timbre correction feature; a target spectrum signal fitting module, configured to fit a target spectrum signal having the text information and the timbre correction feature in the acoustic network; a first loss value calculation module, configured to calculate a difference between the original spectrum signal and the target spectrum signal to obtain a first loss value; a second loss value calculation module, configured to calculate a difference between a portion of the timbre embedding feature and a portion of the timbre total feature as a second loss value; A third loss value fusion module, configured to fuse the first loss value and the second loss value into a third loss value; a network updating module, configured to update the acoustic network and the timbre support network according to the third loss value; A first training condition judgment module is used to judge whether a preset first training condition is met; if so, calling the training completion determination module; if not, returning to calling the voiceprint feature encoding module; The training completion determination module is used to determine whether the timbre support network has completed training.

10. A speech synthesis device, characterized in that: include: A speech synthesis model loading module is used to load a speech synthesis model, wherein the speech synthesis model includes a voiceprint network, a timbre support network, and an acoustic network; A learning object determination module is used to determine text information and the speaker's original speech signal; An original spectrum signal conversion module, used to convert the original speech signal into an original spectrum signal; a voiceprint feature encoding module, configured to encode the original spectrum signal into a voiceprint feature in the voiceprint network, wherein the voiceprint feature is used to verify the identity of the speaker; a timbre supplementary feature encoding module, configured to encode the original spectrum signal into a timbre supplementary feature in the timbre support network, wherein the timbre supplementary feature is a feature missing from the voiceprint feature in timbre; A timbre total feature fusion module, configured to fuse the voiceprint feature and the timbre supplementary feature into a timbre total feature; The target spectrum signal fitting module is used to fit the target spectrum signal whose content is the text information and has the timbre total amount characteristics in the acoustic network.

11. A computer device, characterized in that: The computer device comprises: one or more processors; a memory for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the training method of the speech synthesis model as described in any one of claims 1-6 or the speech synthesis method as described in any one of claims 7-8.

12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the training method for a speech synthesis model as described in any one of claims 1 to 6 or the speech synthesis method as described in any one of claims 7 to 8.

Citation Information

Patent Citations

  • Voice generation method, and device, equipment and computer readable medium

    CN111785247A

  • Voice conversion method and device, corresponding model training method and device, equipment and storage medium

    CN112466275A