Speech synthesis model training method, speech synthesis method, device, and product

By encoding training spectrogram information and timbre features in the speech synthesis model and adjusting the model parameters, the problem of inaccurate timbre was solved, and higher timbre similarity and speech synthesis accuracy were achieved.

CN114566140BActive Publication Date: 2025-12-09TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210157576.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-21
Publication Date
2025-12-09
Estimated Expiration
2042-02-21

AI Technical Summary

Technical Problem

The synthesized timbre produced by existing technology still differs from the speaker's timbre, resulting in inaccurate speech timbre.

Method used

By acquiring the training spectrogram information of the training speech object, inputting it into the encoding module of the speech synthesis model, encoding the obtained encoding vector and determining the vector distribution parameters, combining the timbre features of the training object and text information to generate synthesized speech, and adjusting the model parameters according to the differences to obtain the pre-trained speech synthesis model.

Benefits of technology

It improves the similarity between the synthesized speech timbre and the speaker's timbre, enhancing the accuracy and natural flow of speech synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114566140B_ABST
    Figure CN114566140B_ABST
Patent Text Reader

Abstract

The application relates to a speech synthesis model training method, a speech synthesis method, equipment and products. The speech synthesis model training method comprises the following steps: obtaining training speech spectrum information corresponding to a training speech sample of a training speech object; inputting the training speech spectrum information into a first coding module in a speech synthesis model to be trained, obtaining a first coding vector corresponding to the training speech spectrum information through the first coding module, and determining training vector distribution parameters corresponding to the first coding vector; obtaining a training object timbre feature corresponding to the training speech object based on the training vector distribution parameters, obtaining first synthesized speech according to the training object timbre feature and training text information corresponding to the training speech sample; obtaining a first model loss value according to the difference between the first synthesized speech and the training speech sample; and adjusting model parameters according to the first model loss value to obtain a pre-trained speech synthesis model, which can effectively improve the similarity between the timbre of the synthesized speech and the timbre of the speaker.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of speech synthesis, and in particular to a speech synthesis model training method, a speech synthesis method, a computer device, and a computer program product. BACKGROUND

[0002] With the development of artificial intelligence technology, speech synthesis is increasingly widely used.

[0003] In speech synthesis, speech corpus data corresponding to a speaker is obtained, and the speech corpus data is further analyzed and processed to obtain a corresponding speech synthesis model. Through the model, a sound with a target timbre of the speaker can be synthesized to achieve timbre transfer. However, there is still a certain difference between the synthesized timbre and the speaker's timbre in related technologies, and there is a problem of inaccurate timbre of the synthesized speech. SUMMARY

[0004] Therefore, it is necessary to provide a speech synthesis model training method, a speech synthesis method, a computer device, and a computer program product to solve the above technical problems.

[0005] In a first aspect, the present application provides a speech synthesis model training method. The method comprises:

[0006] obtaining training spectrogram information corresponding to a training speech sample of a training speech object;

[0007] inputting the training spectrogram information into a first encoding module in a to-be-trained speech synthesis model, encoding the training spectrogram information through the first encoding module to obtain a first encoding vector corresponding to the training spectrogram information, and determining a training vector distribution parameter corresponding to the first encoding vector;

[0008] obtaining a training object timbre feature corresponding to the training speech object based on the training vector distribution parameter, and obtaining a first synthesized speech based on the training object timbre feature and training text information corresponding to the training speech sample;

[0009] obtaining a first model loss value based on a difference between the first synthesized speech and the training speech sample;

[0010] adjusting a model parameter of the to-be-trained speech synthesis model according to the first model loss value to obtain a pre-trained speech synthesis model.

[0011] In one embodiment, the first model loss value is obtained based on a difference between the first synthesized speech and the training speech sample, comprising:

[0012] obtaining a speech loss value based on a difference between the first synthesized speech and the training speech sample;

[0013] inputting the training target voice timbre feature into a target classification network to obtain a target classification result;

[0014] obtaining a target classification loss value based on the target classification result, the target classification loss value being used to indicate a difference between the target classification result and a standard classification result;

[0015] obtaining the first model loss value based on the speech loss value and the target classification loss value.

[0016] In one embodiment, the first synthesized speech is obtained according to the training target voice timbre feature and training text information corresponding to the training speech sample, and the method comprises:

[0017] based on the training speech sample, obtaining training acoustic attribute information corresponding to training text information of the training speech sample, and determining a second encoding vector corresponding to the training acoustic attribute information;

[0018] inputting the training speech spectrum information, the training target voice timbre feature and the second encoding vector into a decoding network in the to-be-trained speech synthesis model to obtain a first synthesized speech obtained by decoding.

[0019] In one embodiment, the training acoustic attribute information comprises speech pause results and emotional feature information, and the obtaining of the training acoustic attribute information corresponding to the training text information of the training speech sample comprises:

[0020] obtaining training text information corresponding to the training speech sample and speech pause results corresponding to the training text information;

[0021] based on the training speech sample, obtaining emotional feature information corresponding to the training text information.

[0022] In a second aspect, the present application further provides a speech synthesis method. The method comprises:

[0023] obtaining target speech spectrum information corresponding to a target speech sample of a target speech object;

[0024] training a pre-trained speech synthesis model based on the target speech spectrum information to obtain a target speech synthesis model of the target speech object; the pre-trained speech synthesis model is obtained according to the training method of the speech synthesis model according to any one of claims 1-4, and the target speech synthesis model is used to synthesize speech with the voice timbre feature of the target speech object;

[0025] obtaining target text of to-be-synthesized speech;

[0026] input the target text into the target speech synthesis model to obtain target synthesized speech corresponding to the target text output by the target speech synthesis model.

[0027] In one of the embodiments, the pre-trained speech synthesis model is trained based on a plurality of training spectrogram information, and the training spectrogram information is associated with an object identifier of a corresponding training speech object; the pre-trained speech synthesis model is trained based on the target spectrogram information to obtain a target speech synthesis model of the target speech object, including:

[0028] obtaining a training sample set corresponding to the pre-trained speech synthesis model; the training sample set includes training spectrogram information corresponding to each of the plurality of training speech objects;

[0029] determining a target training speech object from the plurality of training speech objects; the training object timbre feature of the target training speech object matches the sample timbre feature of the target speech sample;

[0030] associating a target object identifier corresponding to the target training speech object with the target spectrogram information;

[0031] training the pre-trained speech synthesis model based on the target spectrogram information associated with the target object identifier to obtain a target speech synthesis model of the target speech object.

[0032] In one of the embodiments, the pre-trained speech synthesis model is trained based on the target spectrogram information associated with the target object identifier to obtain the target speech synthesis model of the target speech object, including:

[0033] inputting the target spectrogram information associated with the target object identifier into a first encoding module in the pre-trained speech synthesis model;

[0034] encoding the target spectrogram information by the first encoding module to obtain a third encoding vector corresponding to the target spectrogram information, and determining a training vector distribution parameter of the target training speech object based on the target object identifier;

[0035] adjusting the training vector distribution parameter of the target training speech object based on the third encoding vector to obtain a target vector distribution parameter;

[0036] obtaining a target object timbre feature corresponding to the target speech object based on the target vector distribution parameter, and obtaining a second synthesized speech based on the target object timbre feature and target text information corresponding to the target speech sample;

[0037] obtaining a second model loss value based on the difference between the second synthesized speech and the target speech sample;

[0038] adjust a model parameter of the pre-trained speech synthesis model according to the second model loss value, to obtain a target speech synthesis model of the target speech object.

[0039] In one of the embodiments, the determining the target training speech object from the plurality of training speech objects comprises:

[0040] obtaining a first target voiceprint feature corresponding to the target speech sample and a first training voiceprint feature corresponding to the training speech sample, and determining a first timbre similarity between the target speech object and the training speech object according to the first target voiceprint feature and the first training voiceprint feature; the first target voiceprint feature and the first training voiceprint feature are extracted by a first voiceprint recognition model;

[0041] obtaining a second target voiceprint feature corresponding to the target speech sample and a second training voiceprint feature corresponding to the training speech sample, and determining a second timbre similarity between the target speech object and the training speech object according to the second target voiceprint feature and the second training voiceprint feature; the second target voiceprint feature and the second training voiceprint feature are extracted by a second voiceprint recognition model;

[0042] determining a timbre similarity between each training speech object and the target speech object based on the first timbre similarity and the second timbre similarity;

[0043] determining the target training speech object based on the timbre similarity.

[0044] In a third aspect, the present application further provides a computer device. The computer device comprises a memory and a processor, the memory stores a computer program, and the processor implements the steps of the method according to any one of the above aspects when executing the computer program.

[0045] In a fourth aspect, the present application further provides a computer program product. The computer program product comprises a computer program, and the computer program implements the steps of the method according to any one of the above aspects when executed by a processor.

[0046] The training method of the speech synthesis model, the speech synthesis method, the computer device and the computer program product can obtain training spectrogram information corresponding to a training speech sample of a training speech object, input the training spectrogram information into a first encoding module in a speech synthesis model to be trained, obtain a first encoding vector corresponding to the training spectrogram information through the first encoding module, determine a training vector distribution parameter corresponding to the first encoding vector, obtain a training object vocal characteristics corresponding to the training speech object based on the training vector distribution parameter, obtain a first synthesized speech based on the training object vocal characteristics and training text information corresponding to the training speech sample, obtain a first model loss value based on a difference between the first synthesized speech and the training speech sample, adjust a model parameter of the speech synthesis model to be trained based on the first model loss value, and obtain a pre-trained speech synthesis model. In the present application, the training object vocal characteristics with the speaker vocal characteristics can be determined based on the training vector distribution parameter, and the training object vocal characteristics can be used to assist speech synthesis based on the spectrogram information and the text information, thereby effectively improving the similarity between the synthesized speech and the speaker vocal characteristics. BRIEF DESCRIPTION OF DRAWINGS

[0047] Figure 1 An application environment diagram of the training method of the speech synthesis model and the speech synthesis method in one embodiment;

[0048] Figure 2 A flowchart of the training method of the speech synthesis model in one embodiment;

[0049] Figure 3 A flowchart of the step of determining the first model loss value in one embodiment;

[0050] Figure 4 A flowchart of the speech synthesis method in one embodiment;

[0051] Figure 5 A flowchart of the speech synthesis method in another embodiment;

[0052] Figure 6 A structure diagram of the speech synthesis model in one embodiment;

[0053] Figure 7 A structure block diagram of the training device of the speech synthesis model in one embodiment;

[0054] Figure 8 A structure block diagram of the speech synthesis device in one embodiment;

[0055] Figure 9 An internal structure diagram of the computer device in one embodiment;

[0056] Figure 10 An internal structure diagram of the computer device in another embodiment. DETAILED DESCRIPTION

[0057] In order to make the purposes, technical solutions and advantages of the present application clearer, further detailed description will be made to the present application in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and do not limit the present application.

[0058] Figure 1 The application environment of the training method of the speech synthesis model and the speech synthesis method provided by an embodiment of the present application is shown in FIG. 1, which can include a terminal 110 and a server 120, and the speech synthesis model to be trained is trained by the server 120 to obtain a pre-trained speech synthesis model. After obtaining the pre-trained speech synthesis model, the server 120 can store it to perform speech synthesis by using the pre-trained speech synthesis model, or can send the pre-trained speech synthesis model to a corresponding device after receiving a model loading request sent by the terminal 110 or other servers. For example, the server 120 can deploy the pre-trained speech synthesis model in a speech synthesis application, and the terminal 110 can install the speech synthesis application to further train the pre-trained speech synthesis model to obtain a target speech synthesis model, and synthesize a speech with a specified tone by using the target speech synthesis model. Figure 1

[0059] The terminal 110 can be, but is not limited to, various personal computers, notebook computers, smart phones, tablet computers, Internet of Things devices and portable wearable devices, and the portable wearable device can be a smart watch, a smart bracelet, a head-mounted device, etc.

[0060] It can be understood that the above application scenario is only an example, and cannot constitute a limitation on the speech synthesis method provided by the embodiments of the present application. The method can also be applied to a server, and can also be applied to a system including a terminal and a server, and can be realized through the interaction of the terminal and the server. The server 120 can be realized by an independent server or a server cluster composed of multiple servers, which can be a physical server or a cloud server providing basic cloud computing services such as cloud server, cloud database, cloud storage and CDN.

[0061] In an embodiment, as shown in FIG. 2, a training method of a speech synthesis model is provided, which is applied to the server 120 in FIG. 1 for example. Of course, the method can also be applied to the terminal 110 to train the model by the terminal 110. Specifically, the method can include the following steps: Figure 2 Figure 1

[0062] ​​​In step S210, training spectrogram information corresponding to the training speech sample of the training speech object is obtained.

[0063] As an example, the training speech object can be a speaker who provides training corpus in the speech synthesis model training process, and the speech uttered by the training speech object can be used as the training speech sample. The training speech sample can be a speech sample with semantics, which can express specific information.

[0064] The training spectrogram information can be the spectrogram information corresponding to the training speech sample. The training spectrogram information can be information representing the original signal characteristics of the training speech sample. For example, the frequency spectrum information of the time-domain sound signal of the training speech sample can be extracted as the spectrogram information. When the frequency spectrum is a mel frequency spectrum, the spectrogram information can also include mel frequency cepstral coefficients.

[0065] In actual applications, the training speech sample uttered by the training speech object can be obtained, and the training spectrogram information corresponding to the training speech sample can be extracted.

[0066] In step S220, the training spectrogram information is input into a first encoding module in the speech synthesis model to be trained, and a first encoding vector corresponding to the training spectrogram information is obtained by encoding the training spectrogram information through the first encoding module. The training vector distribution parameter corresponding to the first encoding vector is determined.

[0067] As an example, the first encoding module can be a module for obtaining a distribution parameter. The vector distribution parameter can also be referred to as a normal distribution parameter, which includes a mean and a variance. The training vector distribution parameter is the vector distribution parameter corresponding to the first encoding vector.

[0068] Specifically, the first encoding module can be a variational auto-encoder (VAE). The variational auto-encoder is a neural network module that obtains the distribution parameter of the real sample and makes the result generated based on the distribution parameter consistent with the real sample through training model parameters.

[0069] In actual applications, the speech synthesis model to be trained can be obtained in advance. The speech synthesis model can include the first encoding module. After obtaining the training spectrogram information, the training spectrogram information can be input into the first encoding module of the speech synthesis model to be trained. The input training spectrogram information can be encoded by the first encoding module to obtain the first encoding vector corresponding to the training spectrogram information and determine the training vector distribution parameter corresponding to the first encoding vector. The training vector distribution parameter corresponding to each training object can be obtained by the first encoding module. The training vector distribution parameter obtained based on the training spectrogram information can have a one-to-one correspondence with the training object.

[0070] In an embodiment, after receiving the training spectrogram information, the first encoding module can encode the input training spectrogram information through a two-dimensional convolution and a recurrent neural network in the first encoding module to obtain a first encoding vector, and input the first encoding vector into a fully connected network inside the first encoding module to trigger the fully connected network to process the input first encoding vector, thereby obtaining a training vector distribution parameter covering the pronunciation characteristics of the training subject.

[0071] Step S230, obtaining a training subject timbre feature corresponding to the training speech object based on the training vector distribution parameter, and obtaining a first synthesized speech based on the training subject timbre feature and training text information corresponding to the training speech sample.

[0072] As an example, the training subject timbre feature can be information representing the timbre characteristics of the training subject, and the training subject timbre feature can be one-to-one corresponding to the training subject.

[0073] The training text information can be the text expressed by the speech content in the training speech sample, i.e., the speaking content in the training speech sample. The text can be a Chinese text or an English text, or a text corresponding to other types of natural languages.

[0074] After obtaining the training vector distribution parameter, the training subject timbre feature corresponding to the training speech object can be obtained based on the training vector distribution parameter, and the training text information corresponding to the training speech sample can be obtained, and the training subject timbre feature and the training text information can be used for speech synthesis to obtain the first synthesized speech.

[0075] Step S240, obtaining a first model loss value according to the difference between the first synthesized speech and the training speech sample.

[0076] After obtaining the first synthesized speech, the difference between the first synthesized speech and the training speech sample can be obtained, and the first model loss value of the speech synthesis model to be trained can be determined according to the difference.

[0077] Step S250, adjusting the model parameters of the speech synthesis model to be trained according to the first model loss value to obtain a pre-trained speech synthesis model.

[0078] After determining the first model loss value, the model parameters of the speech synthesis model to be trained can be adjusted based on this value. When adjusting the model parameters based on the first model loss value, gradient descent can be used to adjust the model parameters in the direction that decreases the corresponding loss value. Specifically, since the first synthesized speech is the speech output by the speech synthesis model to be trained, simulating the timbre of the training object's voice, the smaller the difference between the first synthesized speech and the training speech samples, the better, so that the speech synthesis model to be trained can more accurately synthesize speech with the timbre of the training object.

[0079] Therefore, after determining the first model loss value based on the difference between the first synthesized speech and the training speech samples, the model parameters of the speech synthesis model can be adjusted in the direction of reducing the first model loss value. During the training process of the speech synthesis model to be trained, the model parameters are gradually adjusted until the training termination condition is met, resulting in a pre-trained speech synthesis model that can perform speech synthesis to produce speech with the timbre of the training target.

[0080] In this embodiment, training spectrogram information corresponding to training speech samples of the training speech object can be obtained. This training spectrogram information is input into the first encoding module of the speech synthesis model to be trained. The first encoding module encodes the training spectrogram information to obtain a first encoding vector. Training vector distribution parameters corresponding to the first encoding vector are determined. Based on the training vector distribution parameters, the timbre features of the training speech object are obtained. A first synthesized speech is obtained based on the timbre features of the training object and the training text information corresponding to the training speech samples. A first model loss value is obtained based on the difference between the first synthesized speech and the training speech samples. The model parameters of the speech synthesis model to be trained are adjusted based on the first model loss value to obtain a pre-trained speech synthesis model. In this embodiment, the timbre features of the training object with speaker timbre characteristics can be determined based on the training vector distribution parameters. By combining the timbre features of the training object with spectrogram information and text information to assist speech synthesis, the similarity between the timbre of the synthesized speech and the speaker's timbre is effectively improved.

[0081] In one embodiment, such as Figure 3 As shown, in step S240, the first model loss value is obtained based on the difference between the first synthesized speech and the training speech samples, which may include:

[0082] Step S241: Obtain the speech loss value based on the difference between the first synthesized speech and the training speech sample.

[0083] In actual application, after the first synthesized speech is obtained, the first synthesized speech can be compared with the training speech sample to determine the difference between the first synthesized speech and the training speech sample, and according to the difference, the speech loss value is determined.

[0084] In step S242, the training target voice characteristic is input into the target classification network to obtain a target classification result.

[0085] In addition, the training target voice characteristic can also be input into the target classification network, and according to the output of the target classification network, the target classification result is obtained.

[0086] Specifically, the target classification network can be part of the speech synthesis model to be trained. After the training vector distribution parameter is obtained, the first encoding module can input the training vector distribution parameter into the full connection network in the speech synthesis model to be trained, and obtain the training target voice characteristic through the full connection network. The training target voice characteristic can be input into the target classification network by the full connection network. The target classification network can be trained together with each network in the speech synthesis model to be trained, or it can be a pre-trained target classification network.

[0087] In step S243, a target classification loss value is obtained based on the target classification result, and the target classification loss value is used to indicate the difference between the target classification result and a standard classification result.

[0088] After the target classification result is obtained, the target classification loss value can be determined according to the target classification result.

[0089] Specifically, the target classification result can be the association probability of the training target voice characteristic and each training voice target, that is, the probability of the target classification network predicting that the training target voice characteristic is the voice characteristic of the corresponding training voice target. The standard classification result can be the label corresponding to the training target voice characteristic, which can be determined according to the object number corresponding to the training voice target. In specific implementation, after the association probability is obtained, the association probability and the corresponding label can be compared, for example, the association probability and the probability corresponding to the label are compared, and the corresponding target classification loss value is determined according to the comparison result.

[0090] In step S244, a first model loss value is obtained based on the speech loss value and the target classification loss value.

[0091] After the speech loss value and the target classification loss value are determined, the first model loss value can be determined according to the speech loss value and the target classification loss value.

[0092] In actual application, the speech loss value and the object classification loss value can be positively correlated with the first model loss value, that is, the smaller the speech loss value and the object classification loss value are, the smaller the first model loss value is, and the closer the first synthesized speech obtained based on the current speech synthesis model is to the speech emitted by the corresponding training speech object.

[0093] In the embodiment, the speech loss value can be obtained according to the difference between the first synthesized speech and the training speech sample, the object classification result can be obtained by inputting the training object timbre feature into the object classification network, the object classification loss value can be obtained based on the object classification result, and the first model loss value can be obtained based on the speech loss value and the object classification loss value. In the scheme of the embodiment, by determining the first model loss value based on the speech loss value and the object classification loss value, the synthesized speech output by the speech synthesis model can be more and more similar to the speech emitted by the training speech object, and can be distinguished from the timbre of other training speech objects, thereby effectively improving the accuracy of the synthesized speech generated by the speech synthesis model.

[0094] In one embodiment, in step S230, obtaining the first synthesized speech according to the training object timbre feature and the training text information corresponding to the training speech sample can include:

[0095] Based on the training speech sample, training acoustic attribute information corresponding to the training text information of the training speech sample is obtained, and a second encoding vector corresponding to the training acoustic attribute information is determined; the training spectrogram information, the object timbre feature and the second encoding vector are input into a decoding network in the speech synthesis model to be trained to obtain the first synthesized speech.

[0096] As an example, the acoustic attribute information can represent the processing manner of the pronunciation corresponding to the characters in the text information, and / or the processing manner of the pronunciation between the plurality of characters in the text information.

[0097] By way of example, the acoustic attribute information can include prosodic features, which are a kind of phonological structure of language, and are related to other linguistic structures such as syntax and discourse structure, information structure, etc.; the prosodic features can include three elements: intonation, time domain distribution and stress, which are realized by suprasegmental features. Suprasegmental features include pitch, intensity and temporal characteristics, which are loaded by phonemes or phoneme groups. Prosody is a typical feature of human natural language, and has many common characteristics across languages, such as: pitch declination, stress, pause, etc. are universally present in different languages, and prosodic features are one of the important forms of language and emotional expression. The training acoustic attribute information is the acoustic attribute information corresponding to the training text information.

[0098] In actual application, after obtaining the training speech sample, the training text information corresponding to the training speech sample can be obtained, and by analyzing the voice data corresponding to the training speech sample, the reading manner of the training speech object when reading the training text information can be determined, the acoustic attribute information corresponding to the training text information can be obtained, and then the training acoustic attribute information can be determined.

[0099] After obtaining the training acoustic attribute information, the server 120 can determine the second encoding vector corresponding to the training acoustic attribute information, and input the training spectrogram information, the object timbre feature, and the second encoding vector into the decoding network in the to-be-trained speech synthesis model, and obtain the first synthesized speech based on the output result corresponding to the decoding network.

[0100] In one embodiment, after obtaining the training acoustic attribute information, the server can input the training acoustic attribute information into the second encoding module in the to-be-trained speech synthesis model, and the second encoding module can be composed of a decoding network and an attention network; after being processed by the encoding network and the attention network, the training acoustic attribute information can output the corresponding second encoding vector by the attention network.

[0101] In this embodiment, the server 120 obtains the training acoustic attribute information corresponding to the training text information of the training speech sample, determines the second encoding vector corresponding to the training acoustic attribute information, inputs the training spectrogram information, the training object timbre feature, and the second encoding vector into the decoding network in the to-be-trained speech synthesis model, and obtains the first synthesized speech by decoding, which can generate the first synthesized speech by imitating the processing manner of the training object to the reading of the text information, so that the first synthesized speech is more natural and fluent, conforms to human language habits, and improves the accuracy of speech synthesis.

[0102] In one embodiment, the training acoustic attribute information includes a voice pause result and emotional feature information, and the obtaining of the training acoustic attribute information corresponding to the training text information of the training speech sample based on the training speech sample can include:

[0103] Obtaining the training text information corresponding to the training speech sample and the voice pause result corresponding to the training text information; and obtaining the emotional feature information corresponding to the training text information based on the training speech sample.

[0104] As an example, the voice pause result can represent the pause positions of each character in the training text information, for example, the voice pause result can be the segmentation result obtained after the segmentation processing of the training text information. Through the voice pause result, the synthesized speech can reflect the effect of ups and downs in different texts such as words or phrases.

[0105] The emotional feature information can be information used to control the emotion embodied in the synthesized speech. Specifically, the same text information can embody different emotions when read in different ways. In an example, the emotional feature information can include at least one of tone information or prosody information.

[0106] The tone information is used to indicate the pitch of the speech corresponding to the character, and the prosody information represents the reading rhythm of the character in the training text information. By obtaining the prosody information, the speed of the reading of the training speech object in reading the training text information can be determined, so that the characters in the text can have corresponding pauses when the speech synthesis is performed.

[0107] In this embodiment, the corresponding training text information and the speech pause result corresponding to the training text information can be obtained after obtaining the training speech sample. In actual application, after obtaining the training text information, the training text information can be converted into corresponding phonemes, such as converting Chinese characters into pinyin or converting English words into international phonetic alphabet.

[0108] In addition, the emotional feature information corresponding to the training text information can be obtained based on the training speech sample. Further, the training acoustic attribute information corresponding to the training text information can be obtained based on the speech pause result and the emotional feature information.

[0109] In this embodiment, the speech pause result and the emotional feature information corresponding to the training text information can be used as the corresponding training acoustic attribute information. Compared with the way of speech synthesis by phonemes in the related art, the scheme of this embodiment can make the speech synthesis model to be trained obtain various acoustic attributes involved in the text reading process, effectively enhance the speech synthesis effect of the model, and obtain a synthesized speech more similar to the training speech object.

[0110] In the related art, when synthesizing a sound with a target timbre, the speaker often needs to provide sufficient corpus data. If the corpus data provided by the speaker is insufficient, it is difficult to accurately synthesize a sound with a target timbre, and there is a problem of needing to rely on a large amount of corpus data. To at least solve the above problems, as shown in Figure 4 In one embodiment, the present application provides a speech synthesis method, which can be applied to the terminal 110 or the server 120 in Figure 1 Specifically, the method can include the following steps:

[0111] Step S410, obtaining target speech spectrum information corresponding to a target speech sample of a target speech object.

[0112] As an example, the target speech object can be a new speech object different from the training speech object, i.e., a new speaker, i.e., a speaker who does not appear in the training of the speech synthesis model to be trained, to obtain the pre-trained speech synthesis model. The speech uttered by the target speech object can be a target speech sample. The target speech sample can be a speech sample with semantics, capable of expressing specific information.

[0113] The target spectrogram information can be spectrogram information corresponding to the target speech sample, wherein the target spectrogram information can be information representing the original signal characteristics of the target speech sample, such as a mel spectrum.

[0114] In actual applications, the target speech sample provided by the target speech object can be obtained. For example, a user can record a speech sample on the terminal 110, and send the speech sample to the server 120 through the terminal 110. The server 120 can receive the speech sample as a target speech sample, and obtain the target spectrogram information corresponding to the target speech sample. In an example, the sample amount of the target speech sample can be less than a preset sample amount threshold, which is the sample amount required from training the initial model to obtaining a speech synthesis model capable of synthesizing the specified object.

[0115] In step S420, the pre-trained speech synthesis model is trained based on the target spectrogram information to obtain a target speech synthesis model of the target speech object.

[0116] The pre-trained speech synthesis model can be obtained according to the training method of the speech synthesis model described above, and the target speech synthesis model obtained by training can be used to synthesize speech with the timbre characteristics of the target speech object.

[0117] In a specific implementation, since the pre-trained speech synthesis model pre-trained based on the training spectrogram information can be obtained, after obtaining the target spectrogram information corresponding to the target speech sample of the target speech object, the pre-trained speech synthesis model can be further trained using the target spectrogram information to obtain the corresponding target speech synthesis model.

[0118] It can be understood that, on the one hand, in the process of obtaining the pre-trained speech synthesis model, the training object voice characteristics with the speaker voice characteristics are determined based on the training vector distribution parameters, so as to help improve the similarity between the synthesized voice and the speaker voice, and increase the reliability of the finally obtained pre-trained speech synthesis model. On the other hand, the pre-trained speech synthesis model can be based on a large amount of training spectrogram information to adjust the model parameters of the speech synthesis model to be trained, and in this process, the pre-trained speech synthesis model can learn the mapping relationship between the spectrogram information and the corresponding speech synthesis model which has commonality and universality. When obtaining new target spectrogram information different from the training spectrogram information, by training the pre-trained speech synthesis model using the target spectrogram information, the corresponding target speech synthesis model can be quickly obtained.

[0119] In step S430, the target text of the speech to be synthesized is obtained.

[0120] As an example, the target text can include text expressing natural language, which can be Chinese, English, etc., and can include at least one of words, phrases, and sentences. Of course, in another example, the target text can also include individual characters without semantics, such as numbers or punctuation marks, etc.

[0121] In actual application, in response to the detected input operation, the target text of the speech to be synthesized can be obtained.

[0122] When inputting the target text, one or more ways can be used for input, for example, in response to detecting a text input operation, the terminal can obtain the target text input directly by the user. Alternatively, the target text can be input by voice, for example, after detecting the currently input voice, the text corresponding to the currently input voice can be obtained by automatic speech recognition (ASR) as the target text, wherein the speaker corresponding to the currently input voice and the target speech object can be the same object or different objects.

[0123] In step S440, the target text is input into the target speech synthesis model to obtain the target synthesized speech corresponding to the target text output by the target speech synthesis model.

[0124] After obtaining the target text, the target text can be input into the target speech synthesis model, and the target text can be converted into the target synthesized speech with the target speech object voice characteristics by the target speech synthesis model, and the target synthesized speech finally output by the target speech synthesis model can be obtained.

[0125] In the voice synthesis method, target voice spectrum information corresponding to the target voice sample of the target voice object can be obtained, the pre-trained voice synthesis model is trained based on the target voice spectrum information, and a target voice synthesis model of the target voice object is obtained. The target voice synthesis model can be used to synthesize voice with a timbre feature of the target voice object. Then, the target text to be synthesized can be obtained, and the target text is input into the target voice synthesis model to obtain target synthesized voice corresponding to the target text output by the target voice synthesis model. In this embodiment, the target voice object timbre feature representing the timbre feature of the target voice object can be quickly obtained by obtaining the target voice spectrum information of the target voice object and training the pre-trained voice synthesis model based on the target voice spectrum information. The voice with the timbre of the target voice object can be synthesized without relying on a large amount of training corpus data. The timbre migration of a large number of users can be quickly realized while ensuring the synthesis accuracy.

[0126] In this embodiment, the pre-trained voice synthesis model is repeatedly trained on the initial voice synthesis model by using the training voice samples provided by the multiple training reference objects. Then, the pre-trained voice synthesis model is trained again by using the target voice sample corresponding to the target frequency spectrum information and the target text information provided by the target training object. In this way, the target voice synthesis model associated with the target voice object can be obtained while the model converges quickly. Therefore, the target voice synthesis model corresponding to the target voice object can be quickly obtained in the case of a small amount of corpus.

[0127] In one embodiment, the pre-trained voice synthesis model is trained based on multiple training voice spectrum information. Each training voice spectrum information is associated with an object identifier of a corresponding training voice object. For example, N (N≥2) training voice objects can be used to train the training voice spectrum information corresponding to each training voice object. The object identifiers of the training voice objects include ID1, ID2, ID3,..., IDN. The training voice spectrum information of the training voice object with the object identifier ID1 is associated with the object identifier ID1.

[0128] In step S420, the pre-trained voice synthesis model is trained based on the target voice spectrum information to obtain a target voice synthesis model of the target voice object. The method can include the following steps:

[0129] In step S421, a training sample set corresponding to the pre-trained voice synthesis model is obtained. The training sample set includes multiple training voice spectrum information corresponding to multiple training voice objects.

[0130] In step S422, a target training voice object is determined from the multiple training voice objects. The training object timbre feature of the target training voice object matches the sample timbre feature of the target voice sample.

[0131] In a specific implementation, the speech synthesis model to be trained can be trained by using a training sample set, the training sample set including training spectrogram information corresponding to each of a plurality of training speech objects. By using the training spectrogram information of each of the plurality of training speech objects to train the model, the pre-trained speech synthesis model obtained finally can synthesize speech with the timbre characteristics of any training speech object.

[0132] If the target speech sample corresponding to the target speech object is received, the plurality of training speech objects corresponding to the set can be further obtained, and based on the timbre characteristics corresponding to each of the plurality of training speech objects, a training speech object whose corresponding timbre characteristics match the sample timbre of the target speech sample can be determined from the plurality of training speech objects as the target training speech object.

[0133] In step S423, the target object identifier corresponding to the target training speech object is associated with the target spectrogram information.

[0134] In step S424, the pre-trained speech synthesis model is trained based on the target spectrogram information associated with the target object identifier, to obtain a target speech synthesis model of the target speech object.

[0135] After the target training speech object is determined, the object identifier corresponding to the target training speech object can be used as the target object identifier, and the target object identifier is associated with the target spectrogram information. Then, the pre-trained speech synthesis model can be trained based on the target spectrogram information associated with the target object identifier, to obtain a target speech synthesis model of the target speech object.

[0136] Specifically, since the pre-trained speech synthesis model is trained based on the training spectrogram information of the plurality of training speech objects, for each training speech object, the pre-trained speech synthesis model constructs a mapping relationship between the input text and the synthesized speech (i.e., the synthesized speech containing the semantic of the text and having the timbre characteristics of the training speech object). Then, after the target spectrogram information of the target speech object is obtained, by associating the target spectrogram information of the target speech object with the object identifier corresponding to the training speech object whose timbre characteristics match, and inputting to the pre-trained speech synthesis model for training, the pre-trained speech synthesis model can use the current input target spectrogram information as the training material of the target training speech object, and adjust the relevant parameters in the pre-trained speech synthesis model to make the timbre characteristics of the output speech more similar to the timbre characteristics of the target speech object.

[0137] And, by selecting a training speech object with a matching timbre feature from the multiple training speech objects as the target training speech object, the model parameters associated with the training speech objects with a more similar timbre can be used as the adjustment basis, and the target speech spectrum information is not associated with the identification of other training speech objects with a large timbral gap, which can effectively shorten the training time, reduce the required training materials, and quickly obtain a target speech synthesis model that can synthesize a target speech object with a timbral feature.

[0138] In the embodiment, by obtaining a training sample set corresponding to the pre-trained speech synthesis model, determining a target training speech object from the multiple training speech objects corresponding to the training sample set, associating a target object identifier corresponding to the target training speech object with the target speech spectrum information, and training the pre-trained speech synthesis model based on the target speech spectrum information associated with the target object identifier, a target speech synthesis model of the target speech object is obtained, which can effectively shorten the acquisition time of the target speech synthesis model, and an accurate target speech synthesis model can be quickly obtained with less corpus.

[0139] In one embodiment, in step S424, training the pre-trained speech synthesis model based on the target speech spectrum information associated with the target object identifier to obtain a target speech synthesis model of the target speech object can include the following steps:

[0140] Step S4241, inputting the target speech spectrum information associated with the target object identifier to a first encoding module in the pre-trained speech synthesis model.

[0141] The pre-trained speech synthesis model is trained based on the training spectrum information of the multiple training speech objects, in other words, the pre-trained speech synthesis model can include model parameters corresponding to each training speech object, that is, the pre-trained speech synthesis model can include model parameters with common characteristics and model parameters with characteristics of each training speech object, and then when receiving text, the model parameters of the specified training speech object can be used to synthesize a synthesized speech with the timbral characteristics of the specified training speech object.

[0142] When the pre-trained model is further trained based on the target spectrogram information, since the training speech object whose timbre matches the timbre of the target speech object can be determined, the model parameters corresponding to the training speech object are adjusted based on the target spectrogram information, which effectively accelerates the training speed of the speech synthesis model of the target speech object. When the pre-trained speech synthesis model is trained using the target spectrogram information associated with the target object identifier, the process of associating the target object identifier with the target spectrogram information can be understood as treating (or disguising) the target spectrogram information as the spectrogram information of the target training speech object. In this way, the pre-trained speech synthesis model can adjust the model parameters corresponding to the target training speech object in the direction of matching the timbre of the target speech object.

[0143] Specifically, the target spectrogram information associated with the target object identifier can be input into the first encoding module in the pre-trained speech synthesis model, and the target spectrogram information is encoded by the first encoding module.

[0144] In step S4242, a third encoding vector corresponding to the target spectrogram information is obtained by encoding the target spectrogram information by the first encoding module, and the training vector distribution parameter of the target training speech object is determined based on the target object identifier.

[0145] After the target spectrogram information is encoded by the first encoding module, the third encoding vector corresponding to the target spectrogram information can be obtained, and the training vector distribution parameter corresponding to the target training speech object trained by the pre-trained model during the pre-training process can be determined based on the target object identifier associated with the target spectrogram information.

[0146] In step S4243, the training vector distribution parameter of the target training speech object is adjusted based on the third encoding vector to obtain a target vector distribution parameter.

[0147] After the training vector distribution parameter corresponding to the target training speech object is determined, the training vector distribution parameter can be adjusted based on the current third encoding vector to obtain a target vector distribution parameter.

[0148] In step S4244, the target object timbre feature corresponding to the target speech object is obtained based on the target vector distribution parameter, and the second synthesized speech is obtained according to the target text information corresponding to the target speech sample and the target object timbre feature.

[0149] After obtaining the target vector distribution parameter, a target object timbre feature corresponding to the target speech object can be obtained based on the target vector distribution parameter, for example, the target vector distribution parameter is taken as the target object timbre feature, or the target object timbre feature can also be obtained based on the target vector distribution parameter in combination with other information. Further, speech synthesis can be performed according to the target object timbre feature and the target speech sample, for example, the target object timbre feature is input into a decoding network, and the decoding network decodes based on the input target object timbre feature and the encoding information corresponding to the target text information to obtain a second synthesized speech.

[0150] In step S4245, a second model loss value is obtained according to the difference between the second synthesized speech and the target speech sample.

[0151] In step S4246, the model parameters of the pre-trained speech synthesis model are adjusted according to the second model loss value to obtain a target speech synthesis model of the target speech object.

[0152] After obtaining the second synthesized speech, the difference between the second synthesized speech and the target speech sample can be obtained, and the second model loss value of the pre-trained speech synthesis model can be determined according to the difference. Further, the model parameters of the pre-trained speech synthesis model can be adjusted according to the second model loss value, for example, the model parameters associated with the target training speech object in the pre-trained speech synthesis model can be adjusted according to the second model loss value. In the adjustment, the gradient descent method can be used to adjust the model parameters in the direction of reducing the second model loss value.

[0153] In the embodiment, the target speech spectrum information associated with the target object identifier can be input into the first encoding module of the pre-trained speech synthesis model, the third encoding vector corresponding to the target speech spectrum information can be obtained by encoding through the first encoding module, the training vector distribution parameter of the target training speech object can be determined based on the target object identifier, the training vector distribution parameter of the target training speech object can be adjusted based on the third encoding vector to obtain the target vector distribution parameter, the target object timbre feature corresponding to the target speech object can be obtained based on the target vector distribution parameter, and the second synthesized speech can be obtained according to the target object timbre feature and the target text information corresponding to the target speech sample. Adjusting the model parameters of the pre-trained speech synthesis model can make the model converge quickly while obtaining the target speech synthesis model associated with the target speech object, so that the target speech synthesis model corresponding to the target speech object can be quickly obtained in the case of less corpus.

[0154] In one embodiment, in step S422, the target training speech object is determined from the plurality of training speech objects, which can include:

[0155] The first target voiceprint feature corresponding to the target voice sample and the first training voiceprint feature corresponding to the training voice sample are acquired, and the first timbre similarity corresponding to the target voice object and the training voice object is determined according to the first target voiceprint feature and the first training voiceprint feature; the second target voiceprint feature corresponding to the target voice sample and the second training voiceprint feature corresponding to the training voice sample are acquired, and the second timbre similarity corresponding to the target voice object and the training voice object is determined according to the second target voiceprint feature and the second training voiceprint feature; the timbre similarity between each training voice object and the target voice object is determined based on the first timbre similarity and the second timbre similarity; and the target training voice object is determined based on the timbre similarity.

[0156] The first target voiceprint feature and the first training voiceprint feature are extracted by a first voiceprint recognition model; and the second target voiceprint feature and the second training voiceprint feature are extracted by a second voiceprint recognition model. The first voiceprint recognition model and the second voiceprint recognition model are different voiceprint recognition models.

[0157] In a specific implementation, the first voiceprint recognition model can be used to extract voiceprint features of the target voice sample and the training voice sample respectively to obtain the first target voiceprint feature and the first training voiceprint feature, and the first timbre similarity corresponding to the target voice object and the training voice object is determined according to the first target voiceprint feature and the first training voiceprint feature.

[0158] Meanwhile, the second voiceprint recognition model can be used to extract voiceprint features of the target voice sample and the training voice sample respectively to obtain the second target voiceprint feature and the second training voiceprint feature, and the second timbre similarity corresponding to the target voice object and the training voice object is determined according to the second target voiceprint feature and the second training voiceprint feature.

[0159] For each training voice object, after the corresponding first timbre similarity and the second timbre similarity are determined, the timbre similarity between the training voice object and the target voice object can be determined based on the first timbre similarity and the second timbre similarity, and then the target training voice object can be determined from the plurality of training voice objects according to the timbre similarity.

[0160] In one embodiment, the first voiceprint recognition model can be a voiceprint recognition model based on a Gaussian mean hyper vector, such as a Gaussian mixture model (GMM) or a GMM-UBM (Universal Background Model) model. The second voiceprint recognition model can be a neural network model, which can be composed of at least one type of network, such as a convolutional neural network, a recurrent neural network, and a deep neural network.

[0161] In the timbre matching, the timbre similarity can be determined based on the following formula:

[0162]

[0163] where ID is a training object number corresponding to the training voice object; a is a weight parameter corresponding to the first voiceprint recognition model, and β is a weight parameter corresponding to the second voiceprint recognition model, used to represent the influence degree of the two voiceprint recognition models on the timbre similarity result, and exemplarily, a and β can be configured as 0.25 and 0.5 respectively.

[0164] Vec n is the first training voiceprint feature corresponding to the nth training voice object in the N training voice objects, and Vec target is the first target voiceprint feature of the target voice object, and cosine is the cosine distance between two vectors. The result extracted by the first voiceprint recognition model can also be referred to as an i-vector (Identity-vector).

[0165] Emb n is the second training voiceprint feature corresponding to the nth training voice object in the N training voice objects, and Emb target is the second target voiceprint feature of the target voice object, and softmax is the normalized exponential of two vectors.

[0166] In the embodiment, the first timbre similarity is obtained by the first voiceprint recognition model, and the second timbre similarity is obtained by the second voiceprint recognition model, and the timbre similarity between the training voice object and the target voice object is determined based on the first timbre similarity and the second timbre similarity, which can avoid the recognition error caused by a single voiceprint recognition model, effectively improve the accuracy and reliability of the target training voice object, and provide a basis for subsequent rapid acquisition of a target voice recognition model associated with the target voice object.

[0167] The voice synthesis method provided by the embodiment of the present application will be described below with a specific example. Figure 5

[0168] 1. Train the voice synthesis model to be trained.

[0169] In a specific implementation, the voice synthesis model to be trained can also be referred to as a basic codec network, which is specifically a neural network structure. Exemplarily, the model structure of the voice synthesis model to be trained is as shown in Figure 6 may include a variational autoencoder network, a fully connected network, a classification network, an encoding network, an attention network, and a decoding network.

[0170] ​After obtaining the to-be-trained speech synthesis model, a plurality of training speech samples corresponding to a plurality of training speech objects can be obtained, and training spectrogram information and training text information corresponding to the training speech samples can be obtained.

[0171] Specifically, after obtaining the training text information, phonemes, word segmentation results, tones, prosody, and training object numbers corresponding to the training text information can be obtained, and the phonemes, word segmentation results, tones, prosody, and training object numbers are input into the encoding network. For the training object numbers, the training object numbers can be processed before being input into the encoding network to obtain corresponding training object vectors, which are then input into the decoding network to represent the training speech object corresponding to the plurality of information input at present and the training object vector.

[0172] Meanwhile, after obtaining the training spectrogram information, the server 120 can input the training spectrogram information into the variational autoencoder network. After determining the encoding vector corresponding to the training spectrogram information, the variational autoencoder network can obtain the mean and variance corresponding to the encoding vector as the normal distribution parameters (i.e., training vector distribution parameters), and input the normal distribution parameters into the fully connected network to generate the corresponding speaker vector (i.e., training object voice characteristics) based on the input normal distribution parameters. The variational autoencoder can model different training speech objects based on the input training spectrogram information to obtain the normal distribution parameters corresponding to each training speech object.

[0173] After determining the speaker vector, the fully connected network can input the speaker vector into the decoding network. After inputting the phonemes, word segmentation results, tones, prosody, and training object numbers into the encoding network, the output of the encoding network can be input into the decoding network after being processed by the attention network. The decoding network can perform decoding processing based on this, in combination with the speaker vector and the training spectrogram information, output the corresponding spectrogram, and perform speech synthesis on the training text information using the spectrogram to obtain a first synthesized speech. After obtaining the first synthesized speech, the speech loss value can be determined according to the difference between the first synthesized speech and the training sample speech.

[0174] In addition, the speaker vector output by the fully connected network can be input into the classification network. The classification network can determine the prediction probability of each of the plurality of training speech objects corresponding to the speaker vector based on the input speaker vector, i.e., predict which training speech object the speaker vector corresponds to. Then, the difference between the probability and the true probability can be determined, the object classification loss value corresponding to the difference can be determined based on the difference, and the object loss value can be fed back to the basic encoding and decoding network. Then, the model parameters can be adjusted according to the speech loss value and the object classification loss value.

[0175] The pre-trained speech synthesis model is obtained by training the basic coding and decoding network based on the training spectrogram information and the training text information of each training object. The pre-trained model can be used to synthesize speech of any training object. For example, if N training objects are used to train the speech synthesis model, the pre-trained model obtained is an acoustic model corresponding to the N training objects, which can be used to synthesize speech of any of the N training objects.

[0176] 2. Determine a target training object from the plurality of training objects, which has a similar voice tone to the target voice object.

[0177] In step 1, the pre-trained speech synthesis model corresponding to the plurality of training objects is obtained. When the target voice sample provided by the new speaker (i.e., the target voice object) is obtained, the target voice object is matched with the plurality of training objects in terms of voice tone, and a target training object having a similar voice tone to the target voice object is determined from the plurality of training objects. For example, a user records a voice sample on a terminal according to a provided text.

[0178] 3. Load the pre-trained speech synthesis model, and train the pre-trained speech synthesis model based on the target spectrogram information and the target text information of the target voice sample of the target training object.

[0179] In this step, after the target training object is determined, the pre-trained speech synthesis model can be trained based on the target spectrogram information and the text information of the target training object. Specifically, the training sample set can include training spectrogram information and training text information corresponding to the plurality of training objects, and the server 120 can replace the training spectrogram information and the training text information corresponding to the target training object with the target spectrogram information and the text information. This process can be understood as disguising the target spectrogram information as the spectrogram information corresponding to the target training object, so that the target spectrogram information can be associated with the target training object. Furthermore, when the pre-trained model is trained, the model parameters of the target training object in the pre-trained speech synthesis model can be adjusted based on the target spectrogram information, and a speech synthesis model of the target voice object is obtained. This speech synthesis model can synthesize sounds with the voice tone of the target voice object, as well as sounds with the voice tones of the other N-1 training objects.

[0180] 4. Synthesize speech based on the obtained target speech synthesis model.

[0181] After the user synthesizes the target speech synthesis model corresponding to his own voice through the server 120, the user can input the text he wants to hear and send it to the server through the terminal 110. The server 120 can process the text according to the target object voice characteristics corresponding to the target speech synthesis model to obtain the corresponding target synthesized speech, and return the target synthesized speech to the terminal 110 for the user to listen to.

[0182] Of course, in another example, the target speech synthesis model can also be loaded into the terminal 110. After the terminal obtains the text input by the user, it can use local computing resources (such as mobile phone cpu, gpu) to obtain the synthesized speech according to the target object voice characteristics corresponding to the target speech synthesis model, and play it to the user.

[0183] In the case of no network, the terminal device is used, but this scheme requires less computation and fewer model parameters, and the general effect is worse than that of the server processing result. It is recommended to use the server synthesis scheme in the case of smooth network.

[0184] It should be understood that although each step in the flowchart involved in each of the above embodiments is displayed in sequence according to the arrow, these steps are not necessarily executed in sequence according to the arrow. Unless otherwise stated herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other orders. Moreover, as described above, at least part of the steps in the flowchart involved in each of the above embodiments can include multiple steps or stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least part of other steps or steps or stages in other steps.

[0185] Based on the same inventive concept, the embodiments of the present application also provide a speech synthesis device for implementing the above-mentioned speech synthesis method. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme described in the above method, so the specific limitations in one or more speech synthesis device embodiments provided below can refer to the limitations of the speech synthesis method in the above text, which will not be repeated here.

[0186] In one embodiment, as shown in Figure 7 A speech synthesis model training device 700 is provided, comprising:

[0187] A training spectrogram information acquisition module 701 is configured to acquire training spectrogram information corresponding to a training speech sample of a training speech object;

[0188] The distribution parameter obtaining module 702 is configured to input the training spectrogram information into a first encoding module in the to-be-trained speech synthesis model, encode the training spectrogram information by using the first encoding module to obtain a first encoding vector corresponding to the training spectrogram information, and determine a training vector distribution parameter corresponding to the first encoding vector.

[0189] The synthesis module 703 is configured to obtain a training object timbre feature corresponding to the training speech object based on the training vector distribution parameter, and obtain first synthesized speech based on the training object timbre feature and training text information corresponding to the training speech sample.

[0190] The loss value obtaining module 704 is configured to obtain a first model loss value based on a difference between the first synthesized speech and the training speech sample.

[0191] The parameter adjusting module 705 is configured to adjust a model parameter of the to-be-trained speech synthesis model based on the first model loss value to obtain a pre-trained speech synthesis model.

[0192] In one of the embodiments, the loss value obtaining module 704 is specifically configured to:

[0193] obtain a speech loss value based on a difference between the first synthesized speech and the training speech sample;

[0194] input the training object timbre feature into an object classification network to obtain an object classification result;

[0195] obtain an object classification loss value based on the object classification result, the object classification loss value being used to indicate a difference between the object classification result and a standard classification result;

[0196] obtain the first model loss value based on the speech loss value and the object classification loss value.

[0197] In one of the embodiments, the synthesis module 703 includes:

[0198] A second encoding vector obtaining sub-module is configured to obtain training acoustic attribute information corresponding to training text information of the training speech sample based on the training speech sample, and determine a second encoding vector corresponding to the training acoustic attribute information.

[0199] A decoding sub-module is configured to input the training spectrogram information, the training object timbre feature, and the second encoding vector into a decoding network in the to-be-trained speech synthesis model to obtain the first synthesized speech.

[0200] In one of the embodiments, the training acoustic attribute information includes speech pause results and emotion feature information, and the second encoding vector obtaining submodule is specifically configured to:

[0201] obtain training text information corresponding to the training speech sample and speech pause results corresponding to the training text information;

[0202] obtain emotion feature information corresponding to the training text information based on the training speech sample.

[0203] In one of the embodiments, as shown in Figure 8 a speech synthesis device 800 is provided, and the device includes:

[0204] a target speech spectrum information obtaining module 801 configured to obtain target speech spectrum information corresponding to a target speech sample of a target speech object;

[0205] a target speech synthesis model training module 802 configured to train a pre-trained speech synthesis model based on the target speech spectrum information to obtain a target speech synthesis model of the target speech object; the pre-trained speech synthesis model is obtained according to the training method of any one of the speech synthesis models, and the target speech synthesis model is used to synthesize speech with timbre characteristics of the target speech object;

[0206] a target text obtaining module 803 configured to obtain target text to be synthesized;

[0207] a target synthesized speech obtaining module 804 configured to input the target text into the target speech synthesis model to obtain target synthesized speech corresponding to the target text output by the target speech synthesis model.

[0208] In one of the embodiments, the pre-trained speech synthesis model is trained based on a plurality of training speech spectrum information, and the training speech spectrum information is associated with an object identifier of a corresponding training speech object; the target speech synthesis model training module 802 includes:

[0209] a training sample set obtaining submodule configured to obtain a training sample set corresponding to the pre-trained speech synthesis model; the training sample set includes training speech spectrum information corresponding to each of the plurality of training speech objects;

[0210] a target training speech object determining submodule configured to determine a target training speech object from the plurality of training speech objects; the training object timbre characteristics of the target training speech object match the sample timbre characteristics of the target speech sample;

[0211] The object identification association submodule is configured to associate a target object identifier corresponding to the target training voice object with the target spectrogram information.

[0212] The target voice synthesis model acquisition submodule is configured to train the pre-trained voice synthesis model based on the target spectrogram information associated with the target object identifier, to obtain a target voice synthesis model of the target voice object.

[0213] In one of the embodiments, the target voice synthesis model acquisition submodule is specifically configured to:

[0214] input the target spectrogram information associated with the target object identifier into a first encoding module in the pre-trained voice synthesis model;

[0215] obtain a third encoding vector corresponding to the target spectrogram information through the first encoding module, and determine a training vector distribution parameter of the target training voice object based on the target object identifier;

[0216] adjust the training vector distribution parameter of the target training voice object based on the third encoding vector, to obtain a target vector distribution parameter;

[0217] obtain a target object timbre feature corresponding to the target voice object based on the target vector distribution parameter, and obtain a second synthesized voice based on the target object timbre feature and target text information corresponding to the target voice sample;

[0218] obtain a second model loss value based on a difference between the second synthesized voice and the target voice sample;

[0219] adjust a model parameter of the pre-trained voice synthesis model based on the second model loss value, to obtain a target voice synthesis model of the target voice object.

[0220] In one of the embodiments, the target training voice object determination submodule is specifically configured to:

[0221] obtain a first target voiceprint feature corresponding to the target voice sample and a first training voiceprint feature corresponding to the training voice sample, and determine a first timbre similarity corresponding to the target voice object and the training voice object based on the first target voiceprint feature and the first training voiceprint feature; the first target voiceprint feature and the first training voiceprint feature are extracted through a first voiceprint recognition model;

[0222] obtaining second target voiceprint features corresponding to the target speech sample and second training voiceprint features corresponding to the training speech sample, and determining second timbre similarities between the target speech object and the training speech objects according to the second target voiceprint features and the second training voiceprint features; the second target voiceprint features and the second training voiceprint features are extracted by a second voiceprint recognition model;

[0223] determining timbre similarities between each training speech object and the target speech object based on the first timbre similarity and the second timbre similarity;

[0224] determining a target training speech object based on the timbre similarities.

[0225] Each of the above apparatuses can be implemented by software, hardware, or a combination thereof. Each of the above apparatuses can be embedded in or independent of a processor in a computer device in a hardware form, or stored in a memory in a computer device in a software form, so as to be called and executed by a processor to perform operations corresponding to each of the above apparatuses.

[0226] In one embodiment, a computer device is provided, which can be a server, and an internal structure diagram of the computer device can be as shown in Figure 9 The computer device includes a processor, a memory, and a network interface connected through a system bus. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The database of the computer device can be used to store spectrogram information (training spectrogram information and / or target spectrogram information) and a speech synthesis model (a speech synthesis model to be trained and / or a pre-trained speech synthesis model). The network interface of the computer device is configured to communicate with an external terminal through a network connection. The computer program is executed by the processor to implement the method in the above embodiments.

[0227] In one embodiment, a computer device is provided, which can be a terminal, and an internal structure diagram of the computer device can be as shown in Figure 10As shown. The computer device includes a processor, memory, communication interface, display screen, and input device connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements the methods described in the above embodiments. The display screen can be a liquid crystal display (LCD) or an e-ink display. The input device can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the computer device casing, or an external keyboard, touchpad, or mouse.

[0228] Those skilled in the art will understand that Figure 9 or Figure 10 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0229] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:

[0230] Obtain the training spectrogram information corresponding to the training speech samples of the training speech object;

[0231] The training spectrogram information is input into the first encoding module of the speech synthesis model to be trained, and the first encoding module encodes the first encoding vector corresponding to the training spectrogram information to determine the training vector distribution parameters corresponding to the first encoding vector.

[0232] Based on the training vector distribution parameters, the training object timbre features corresponding to the training speech object are obtained, and the first synthesized speech is obtained according to the training object timbre features and the training text information corresponding to the training speech sample.

[0233] The first model loss value is obtained based on the difference between the first synthesized speech and the training speech sample;

[0234] The model parameters of the speech synthesis model to be trained are adjusted based on the first model loss value to obtain a pre-trained speech synthesis model.

[0235] In one embodiment, the processor, when executing the computer program, also implements the steps in the above method embodiments.

[0236] In one embodiment, a computer device is provided, comprising a memory and a processor, the memory storing a computer program, and the processor, when executing the computer program, implements the following steps:

[0237] obtaining target spectrogram information corresponding to a target speech sample of a target speech object;

[0238] training a pre-trained speech synthesis model based on the target spectrogram information to obtain a target speech synthesis model of the target speech object; the pre-trained speech synthesis model is obtained according to the training method of the speech synthesis model in any one of the above, and the target speech synthesis model is used to synthesize speech with vocal characteristics of the target speech object;

[0239] obtaining target text to be synthesized;

[0240] inputting the target text into the target speech synthesis model to obtain target synthesized speech corresponding to the target text output by the target speech synthesis model.

[0241] In one embodiment, a computer readable storage medium is provided, which stores a computer program, and the computer program, when executed by a processor, implements the following steps:

[0242] obtaining training spectrogram information corresponding to a training speech sample of a training speech object;

[0243] inputting the training spectrogram information into a first encoding module in a speech synthesis model to be trained, obtaining a first encoding vector corresponding to the training spectrogram information through encoding of the first encoding module, and determining training vector distribution parameters corresponding to the first encoding vector;

[0244] obtaining training object vocal characteristics corresponding to the training speech object based on the training vector distribution parameters, and obtaining first synthesized speech according to the training object vocal characteristics and training text information corresponding to the training speech sample;

[0245] obtaining a first model loss value according to the difference between the first synthesized speech and the training speech sample;

[0246] adjusting model parameters of the speech synthesis model to be trained according to the first model loss value to obtain a pre-trained speech synthesis model.

[0247] In one embodiment, the computer program, when executed by the processor, also implements the steps in the above method embodiments.

[0248] In one embodiment, a computer readable storage medium is provided, and the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the following steps:

[0249] obtaining target spectrogram information corresponding to a target speech sample of a target speech object;

[0250] training a pre-trained speech synthesis model based on the target spectrogram information to obtain a target speech synthesis model of the target speech object; the pre-trained speech synthesis model is obtained according to the training method of the speech synthesis model in any one of the above, and the target speech synthesis model is used to synthesize speech with timbre characteristics of the target speech object;

[0251] obtaining target text to be synthesized;

[0252] inputting the target text into the target speech synthesis model to obtain target synthesized speech corresponding to the target text output by the target speech synthesis model.

[0253] In one embodiment, a computer program product is provided, and the computer program product includes a computer program, and the computer program is executed by a processor to implement the following steps:

[0254] obtaining training spectrogram information corresponding to a training speech sample of a training speech object;

[0255] inputting the training spectrogram information into a first encoding module in a speech synthesis model to be trained, encoding the training spectrogram information through the first encoding module to obtain a first encoding vector corresponding to the training spectrogram information, and determining a training vector distribution parameter corresponding to the first encoding vector;

[0256] obtaining a training object timbre characteristic corresponding to the training speech object based on the training vector distribution parameter, and obtaining a first synthesized speech according to the training object timbre characteristic and training text information corresponding to the training speech sample;

[0257] obtaining a first model loss value according to a difference between the first synthesized speech and the training speech sample;

[0258] adjusting a model parameter of the speech synthesis model to be trained according to the first model loss value to obtain a pre-trained speech synthesis model.

[0259] In one embodiment, the computer program is executed by the processor to further implement the steps in the above method embodiments.

[0260] In one embodiment, a computer program product is provided, and the computer program product includes a computer program, and the computer program is executed by a processor to implement the following steps:

[0261] obtaining target spectrogram information corresponding to a target speech sample of a target speech object;

[0262] training a pre-trained speech synthesis model based on the target spectrogram information to obtain a target speech synthesis model of the target speech object; the pre-trained speech synthesis model is obtained according to the training method of the speech synthesis model in any one of the above, and the target speech synthesis model is used to synthesize speech with the timbre characteristics of the target speech object;

[0263] obtaining target text to be synthesized;

[0264] inputting the target text into the target speech synthesis model to obtain target synthesized speech corresponding to the target text output by the target speech synthesis model.

[0265] It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties.

[0266] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (Read-Only Memory, ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive memory (ReRAM), magnetoresistive random access memory (Magnetoresistive Random Access Memory, MRAM), ferroelectric memory (Ferroelectric Random Access Memory, FRAM), phase change memory (Phase Change Memory, PCM), graphene memory, etc. Volatile memory can include random access memory (Random Access Memory, RAM) or external cache memory, etc. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (Static Random Access Memory, SRAM) or dynamic random access memory (Dynamic Random Access Memory, DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without being limited thereto.

[0267] Any combination of the technical features of the above embodiments can be made. In order to make the description simple, all possible combinations of the technical features in the above embodiments are not described, however, as long as the combination of the technical features does not exist contradictory, it should be considered as the scope of the present application.

[0268] The above embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of the patent of the present application. It should be pointed out that for ordinary skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are within the scope of protection of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A method of training a speech synthesis model, the method comprising: The method comprises: obtaining training speech samples of a plurality of training speech objects respectively corresponding to training spectrogram information; inputting a plurality of the training spectrogram information into a first encoding module in a to-be-trained speech synthesis model, obtaining a first encoding vector corresponding to the training spectrogram information through the first encoding module, and determining training vector distribution parameters corresponding to the first encoding vector; generating training object timbre features corresponding to the training speech objects based on the training vector distribution parameters, and obtaining first synthesized speech according to the training object timbre features and training text information corresponding to the training speech samples; obtaining a first model loss value according to the difference between the first synthesized speech and the training speech samples; adjusting model parameters of the to-be-trained speech synthesis model according to the first model loss value to obtain a pre-trained speech synthesis model; the pre-trained speech synthesis model comprises a plurality of training vector distribution parameters respectively corresponding to the training spectrogram information; when target spectrogram information of a target speech object different from each of the training speech objects is obtained, a target speech synthesis model corresponding to the target speech object is obtained by adjusting the pre-trained speech synthesis model according to a second model loss value, the second model loss value is determined according to the difference between second synthesized speech and target speech samples provided by the target speech object, the second synthesized speech is generated according to target object timbre features and target text information of the target speech samples, the target object timbre features are determined based on the adjustment result of the training vector distribution parameters of a target training speech object in the pre-trained speech synthesis model; the timbre of the target training speech object matches the timbre of the target speech object.

2. The method of claim 1, wherein, The first model loss value is obtained according to the difference between the first synthesized speech and the training speech samples, comprising: obtaining a speech loss value according to the difference between the first synthesized speech and the training speech samples; inputting the training object timbre features into an object classification network to obtain an object classification result; obtaining an object classification loss value based on the object classification result, the object classification loss value being used to indicate the difference between the object classification result and a standard classification result; obtaining the first model loss value based on the speech loss value and the object classification loss value.

3. The method of claim 1, wherein, The first synthesized speech is obtained according to the training object timbre features and the training text information corresponding to the training speech samples, comprising: based on the training speech samples, obtaining training acoustic attribute information corresponding to the training text information of the training speech samples, and determining a second encoding vector corresponding to the training acoustic attribute information; inputting the training spectrogram information, the training object timbre features and the second encoding vector into a decoding network in the to-be-trained speech synthesis model to obtain the first synthesized speech.

4. The method of claim 3, wherein, The training acoustic attribute information comprises speech pause results and emotional feature information, and the training acoustic attribute information corresponding to the training text information of the training speech samples is obtained based on the training speech samples, comprising: obtain training text information corresponding to the training speech sample and a speech pause result corresponding to the training text information; obtain emotional feature information corresponding to the training text information based on the training speech sample.

5. A speech synthesis method characterized by, The method comprises: obtain target spectrogram information corresponding to a target speech sample of a target speech object; train a pre-trained speech synthesis model based on the target spectrogram information to obtain a target speech synthesis model of the target speech object; the pre-trained speech synthesis model is obtained according to the training method of the speech synthesis model in any one of claims 1-4, and the target speech synthesis model is used to synthesize speech with a timbre feature of the target speech object; obtain target text of speech to be synthesized; input the target text into the target speech synthesis model to obtain target synthesized speech corresponding to the target text output by the target speech synthesis model.

6. The method of claim 5, wherein, The pre-trained speech synthesis model is trained based on a plurality of training spectrogram information, and the training spectrogram information is associated with an object identifier of a corresponding training speech object; The training of the pre-trained speech synthesis model based on the target spectrogram information to obtain the target speech synthesis model of the target speech object comprises: obtain a training sample set corresponding to the pre-trained speech synthesis model; the training sample set comprises training spectrogram information corresponding to each of the plurality of training speech objects; determine a target training speech object from the plurality of training speech objects; the training object timbre feature of the target training speech object matches a sample timbre feature of the target speech sample; associate a target object identifier corresponding to the target training speech object with the target spectrogram information; train the pre-trained speech synthesis model based on the target spectrogram information associated with the target object identifier to obtain the target speech synthesis model of the target speech object.

7. The method of claim 6, wherein, The training of the pre-trained speech synthesis model based on the target spectrogram information associated with the target object identifier to obtain the target speech synthesis model of the target speech object comprises: input the target spectrogram information associated with the target object identifier into a first encoding module in the pre-trained speech synthesis model; encode the target spectrogram information through the first encoding module to obtain a third encoding vector corresponding to the target spectrogram information, and determine a training vector distribution parameter of the target training speech object based on the target object identifier; adjust the training vector distribution parameter of the target training speech object based on the third encoding vector to obtain a target vector distribution parameter; obtain a target object timbre feature corresponding to the target speech object based on the target vector distribution parameter, and obtain a second synthesized speech based on the target object timbre feature and target text information corresponding to the target speech sample; obtain a second model loss value according to the difference between the second synthesized speech and the target speech sample; adjust a model parameter of the pre-trained speech synthesis model according to the second model loss value to obtain the target speech synthesis model of the target speech object.

8. The method of claim 6, wherein, The determining the target training voice object from the plurality of training voice objects comprises: obtaining a first target voiceprint feature corresponding to the target voice sample and a first training voiceprint feature corresponding to the training voice sample, and determining a first timbre similarity corresponding to the target voice object and the training voice object according to the first target voiceprint feature and the first training voiceprint feature; the first target voiceprint feature and the first training voiceprint feature are extracted by a first voiceprint recognition model; obtaining a second target voiceprint feature corresponding to the target voice sample and a second training voiceprint feature corresponding to the training voice sample, and determining a second timbre similarity corresponding to the target voice object and the training voice object according to the second target voiceprint feature and the second training voiceprint feature; the second target voiceprint feature and the second training voiceprint feature are extracted by a second voiceprint recognition model; determining a timbre similarity between each training voice object and the target voice object based on the first timbre similarity and the second timbre similarity; determining a target training voice object based on the timbre similarity. 9.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-8 when the computer program is executed by the processor. The processor executes the computer program to realize the steps of the method in any one of claims 1 to 8.

10. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to realize the steps of the method in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Voice generation method, and device, equipment and computer readable medium

    CN111785247A

  • Audio generation method and device, storage medium and electronic equipment

    CN113205793A

  • Song synthesis method and device, equipment, medium and product

    CN113808555A