Training Method, Device, Equipment and Storage Medium of Speech Synthesis Model

Through the end-to-end training of speech synthesis model, combined with text encoder, duration predictor and decoder, the cumbersome training process of speech synthesis model is solved, achieving a more efficient and natural speech synthesis effect.

CN115294962BActive Publication Date: 2025-07-04TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210919946.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-01
Publication Date
2025-07-04
Estimated Expiration
2042-08-01

AI Technical Summary

Technical Problem

In the prior art, the training process of speech synthesis models is cumbersome and fusion training cannot be achieved, resulting in the synthetic audio being stiff and not very natural.

Method used

The end-to-end training method is adopted, through the combination of text encoder, duration predictor and decoder, and the first pronunciation time and acoustic features are used as supervision to train the speech synthesis model to solve the problem of mismatch between the sample text length and the acoustic feature length.

Benefits of technology

The training process of speech synthesis model is simplified, the efficiency and quality of speech synthesis are improved, and more natural audio is generated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115294962B_ABST
    Figure CN115294962B_ABST
Patent Text Reader

Abstract

The present application discloses a training method, device, equipment and storage medium for a speech synthesis model, relating to the field of artificial intelligence. The method includes: obtaining a sample hidden text representation through a text encoder; obtaining a first pronunciation duration and a first predicted acoustic feature through a first decoder based on the sample hidden text representation and the sample acoustic feature; obtaining a second pronunciation duration through a duration predictor based on the sample hidden text feature; obtaining a second predicted acoustic feature through a second decoder based on a sample hidden text extended representation obtained by performing upsampling processing on the first pronunciation duration; training the text encoder, the duration predictor, the first decoder and the second decoder based on the first pronunciation duration, the second pronunciation duration, the sample acoustic feature, the first predicted acoustic feature and the second predicted acoustic feature; and constructing a speech synthesis model based on the trained text encoder, duration predictor and second decoder. This solution helps to improve the training effect of the speech synthesis model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of artificial intelligence, and particularly to a method, device, equipment and storage medium for training a speech synthesis model. Background Art

[0002] Speech synthesis refers to the process of converting text into audio. In this process, a speech synthesis model is usually used for speech synthesis.

[0003] In the related art, when training a speech synthesis model, multiple independent trainings are required, which makes the training process of the speech synthesis model fragmented, the training process cumbersome and unable to realize the advantages of fusion training, and the synthesized audio is relatively rigid and has low naturalness. Summary of the Invention

[0004] The embodiments of the present application provide a method, device, equipment and storage medium for training a speech synthesis model. The technical solutions are as follows:

[0005] On the one hand, the embodiments of the present application provide a method for training a speech synthesis model, the method comprising:

[0006] Encoding the sample text through a text encoder to obtain a sample hidden text representation;

[0007] Based on the sample hidden text representation and the sample acoustic features corresponding to the sample text, performing duration prediction and decoding through a first decoder to obtain a first pronunciation duration and a first predicted acoustic feature corresponding to the sample hidden text representation;

[0008] Performing duration prediction through a duration predictor based on the sample hidden text features to obtain a second pronunciation duration;

[0009] Performing upsampling processing on the sample hidden text representation based on the first pronunciation duration to obtain a sample hidden text extended representation;

[0010] Decoding the sample hidden text extended representation through a second decoder to obtain a second predicted acoustic feature;

[0011] Using the first pronunciation duration as the supervision of the second pronunciation duration, and using the sample acoustic features as the supervision of the first predicted acoustic feature and the second predicted acoustic feature, and training the text encoder, the duration predictor, the first decoder and the second decoder in an end-to-end manner;

[0012] Constructing a speech synthesis model based on the trained text encoder, duration predictor and second decoder.

[0013] On the other hand, an embodiment of the present application provides a training device for a speech synthesis model, the device comprising:

[0014] A text encoding module, configured to encode sample text through a text encoder to obtain a sample hidden text representation;

[0015] A first decoding module, configured to perform duration prediction and decoding through a first decoder based on the sample hidden text representation and the sample acoustic features corresponding to the sample text, to obtain a first pronunciation duration and a first predicted acoustic feature corresponding to the sample hidden text representation;

[0016] A duration prediction module, configured to perform duration prediction through a duration predictor based on the sample hidden text features to obtain a second pronunciation duration;

[0017] An upsampling module, configured to perform upsampling processing on the sample hidden text representation based on the first pronunciation duration to obtain a sample hidden text extended representation;

[0018] A second decoding module, configured to decode the sample hidden text extended representation through a second decoder to obtain a second predicted acoustic feature;

[0019] A training module, configured to use the first pronunciation duration as the supervision for the second pronunciation duration, and use the sample acoustic features as the supervision for the first predicted acoustic feature and the second predicted acoustic feature, and train the text encoder, the duration predictor, the first decoder, and the second decoder in an end-to-end manner;

[0020] A model construction module, configured to construct a speech synthesis model based on the trained text encoder, duration predictor, and second decoder.

[0021] On the other hand, an embodiment of the present application provides a computer device, the computer device comprising a processor and a memory, and at least one program is stored in the memory, and the at least one program is loaded and executed by the processor to implement the speech synthesis model training method as described in the above aspect.

[0022] On the other hand, an embodiment of the present application provides a computer-readable storage medium, and at least one program is stored in the computer-readable storage medium, and the at least one program is loaded and executed by a processor to implement the speech synthesis model training method as described in the above aspect.

[0023] On the other hand, an embodiment of the present application provides a computer program product, which includes computer instructions stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to cause the earphone to execute the training method of the voice synthesis model provided in the above aspect.

[0024] The beneficial effects brought by the technical solution provided by the embodiment of the present application at least include:

[0025] In the embodiment of the present application, the computer device first encodes the sample text through a text encoder to obtain a sample hidden text representation, and then based on the sample hidden text representation and the sample acoustic features corresponding to the sample text, performs duration prediction and decoding through a first decoder to obtain a first pronunciation duration and a first predicted acoustic feature corresponding to the sample hidden text representation, and based on the sample hidden text features, performs duration prediction through a duration predictor to obtain a second pronunciation duration. The sample hidden text representation is upsampled based on the first pronunciation duration to obtain an extended sample hidden text representation. Further, the extended sample hidden text representation is decoded through a second decoder to obtain a second predicted acoustic feature. Finally, the first pronunciation duration is used as the supervision for the second pronunciation duration, and the sample acoustic features are used as the supervision for the first predicted acoustic feature and the second predicted acoustic feature. The text encoder, duration predictor, first decoder, and second decoder are trained in an end-to-end manner, and a voice synthesis model is constructed based on the trained text encoder, duration predictor, and second decoder; by using the solution provided by the embodiment of the present application, the sample hidden text representation can be upsampled based on the first pronunciation duration and then input into the second decoder, which solves the problem of the length mismatch between the sample text length and the corresponding sample acoustic features, and further simplifies the structure of the acoustic model and improves the efficiency and conversion quality of the acoustic model in converting the text to be synthesized into acoustic features by adopting an end-to-end training method. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0027] Figure 1 Shows a schematic diagram of the implementation environment of the voice synthesis model provided by an exemplary embodiment of the present application;

[0028] Figure 2 Shows a flowchart of the training method of the voice synthesis model provided by an exemplary embodiment of the present application;

[0029] Figure 3 shows an implementation schematic diagram of the training process of a speech synthesis model provided by an exemplary embodiment of the present application;

[0030] Figure 4 shows a flowchart of the duration prediction and decoding process of the first decoder provided by an exemplary embodiment of the present application;

[0031] Figure 5 shows an implementation schematic diagram of the duration prediction and decoding process of the first decoder provided by an exemplary embodiment of the present application;

[0032] Figure 6 shows a flowchart of the training method of a speech synthesis model provided by another exemplary embodiment of the present application;

[0033] Figure 7 shows an implementation schematic diagram of the training process of a speech synthesis model provided by another exemplary embodiment of the present application;

[0034] Figure 8 shows a flowchart of the speech synthesis process provided by an exemplary embodiment of the present application;

[0035] Figure 9 shows an implementation schematic diagram of the speech synthesis process provided by an exemplary embodiment of the present application;

[0036] Figure 10 is a structural block diagram of a training device of a speech synthesis model provided by an exemplary embodiment of the present application;

[0037] Figure 11 shows a structural schematic diagram of a computer device provided by an exemplary embodiment of the present application. Detailed implementation manners

[0038] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0039] As used herein, "a plurality of" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0040] For ease of understanding, the following explains the nouns involved in the embodiments of the present application.

[0041] Speech synthesis: Also known as Text to Speech (TTS), its function is to convert the text information generated by the computer itself or input externally into understandable and fluent speech and read it out.

[0042] Vocoder: Derived from the abbreviation of Voice Encoder, also known as a speech signal analysis and synthesis system, its function is to convert acoustic features into sound.

[0043] Convolutional Neural Network (CNN): A type of feedforward neural network whose neurons can respond to units within the receptive field. CNNs usually consist of multiple convolutional layers and a fully connected layer at the top. By sharing parameters, they reduce the number of model parameters and are widely used in image and speech recognition.

[0044] Recurrent Neural Network (RNN): A class of recursive neural networks that take sequence data as input, perform recursion in the evolution direction of the sequence, and all nodes (recurrent units) are connected in a chain. Its structure consists of an input layer, a hidden layer, and an output layer. The current output of the RNN is related to the previous output. Specifically, the RNN remembers the previous information and applies it to the calculation of the current output, that is, the nodes between the hidden layers are no longer unconnected but connected, and the input to the hidden layer includes not only the output of the input layer but also the output of the previous hidden layer.

[0045] Loss function: Also known as the cost function, it is a function used to evaluate the degree of difference between the predicted value and the true value of a neural network model. The smaller the function value of the loss function, the better the performance of the neural network model. The training process of the model is to adjust the model parameters to minimize the value of the loss function. For different neural network models, different loss functions are used. Common loss functions include the 0-1 loss function, absolute value loss function, logarithmic loss function, exponential loss function, perceptron loss function, cross-entropy loss function, Kullback-Leibler divergence loss function, Triplet Loss function, and so on.

[0046] Artificial Intelligence (AI) uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence. It is a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.

[0047] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0048] The key technologies of speech technology include Automatic Speech Recognition (ASR), Text To Speech (TTS), speech enhancement technology, and voiceprint recognition technology, etc. Enabling the computer to listen, see, speak, and feel is the future development direction of human-computer interaction, and speech has become one of the most promising human-computer interaction methods in the future.

[0049] The training method of the speech synthesis model involved in the embodiments of this application, that is, the application of speech synthesis technology in the field of Natural Language Processing (NLP), can reduce the training process while improving the training effect of the speech synthesis model through end-to-end fusion training, thereby enhancing the accuracy of the synthesis result of the trained speech synthesis model.

[0050] Please refer to Figure 1 , which shows a schematic diagram of the implementation environment of the speech synthesis model provided by the exemplary embodiments of this application. The implementation environment may include: a terminal 110 and a server 120.

[0051] The terminal 110 is an electronic device with speech synthesis function.

[0052] Among them, the terminal 110 includes but is not limited to smart phones, tablet computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, laptop computers, desktop computers, and so on.

[0053] A client that provides a text-to-speech function can run on the terminal 110. The client can be an instant messaging application, a music playback application, a reading application, etc. The embodiments of the present application do not limit the specific type of the client.

[0054] The server 120 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. In the embodiments of the present application, the server is the background server of the client that provides the text-to-speech function in the terminal 110 and can convert text into speech.

[0055] Among them, data communication is carried out between the terminal 110 and the server 120 through a communication network. Optionally, the communication network can be a wired network or a wireless network, and the communication network can be at least one of a local area network, a metropolitan area network, and a wide area network.

[0056] In a possible implementation manner, during the training stage of the text-to-speech model, when the terminal 110 receives a training instruction for the text-to-speech model, it uploads the sample text and the sample acoustic features corresponding to the sample text to the server 120. The server 120 trains the text-to-speech model based on the sample text and the sample acoustic features corresponding to the sample text, so as to obtain a trained text-to-speech model. Among them, since the sample text and the sample acoustic features corresponding to the sample text can be the private data of the user, the user can train a text-to-speech model with the desired timbre to achieve personalized training.

[0057] Further, during the usage stage of the text-to-speech model, when the terminal 110 receives a text-to-speech operation, it uploads the target text 121 to the server 120. The server 120 predicts the acoustic features of the target text 121 through the trained text-to-speech model 122 to obtain the target acoustic features 123 corresponding to the target text.

[0058] Further, the server 120 inputs the target features into the vocoder 124. The vocoder 124 performs text-to-speech based on the target acoustic features 123 and transmits the synthesized target audio 125 back to the terminal 110 for the terminal 110 to play, realizing the text-to-speech function.

[0059] In other possible implementation manners, the text-to-speech model 122 and the vocoder 124 can also be deployed in the terminal 110, so that the terminal 110 realizes the text-to-speech function locally, reducing the processing pressure on the server 120. This embodiment does not limit this.

[0060] In addition, the above vocoder can be trained by the server 120 or deployed on the server 120 side after being trained by other devices. For the convenience of description, in the following embodiments, it is assumed that the speech synthesis method is applied to a computer device (which can be the Figure 1 server or terminal in), and the training of the speech synthesis model is performed by the computer device as an example for illustration.

[0061] Please refer to Figure 2 , which shows a flowchart of a method for training a speech synthesis model provided by an exemplary embodiment of the present application. This embodiment is described by taking this method as an example for a computer device. The method includes the following steps.

[0062] Step 201: Encode the sample text through a text encoder to obtain a sample hidden text representation.

[0063] In a possible implementation manner, after the computer device obtains the sample text, it first performs a text preprocessing operation on the sample text, converts the sample text into phonemes, and obtains a sample phoneme sequence corresponding to the sample text.

[0064] Optionally, the text preprocessing operation may include text regularization processing, prosody analysis processing, polyphone analysis processing, etc. The embodiments of the present application do not limit the specific text preprocessing operation.

[0065] Optionally, the sample phoneme sequence contains multiple phonemes. A phoneme is the smallest speech unit divided according to the natural attributes of speech. Taking Mandarin Chinese as an example, phonemes may include initials, finals, tones, etc. The phonemes corresponding to different languages may be different. For example, the phonemes of Mandarin Chinese corresponding to the text are different from those of dialects, or the phonemes of Chinese corresponding to the text are different from the corresponding English phonemes.

[0066] Taking Mandarin Chinese as an example, phonemes may include initials, finals, tones, etc. For example, when the sample text is "The weather is really nice today", the corresponding sample phoneme sequence may be "jin1 tian1 de1 tian1 qi4 zhen1 hao3". The sample phoneme sequence may be the phoneme sequence of the language to be synthesized for the sample text.

[0067] Further, the computer device encodes the sample phoneme sequence through a text encoder to obtain a sample hidden text representation.

[0068] Optionally, the sample hidden text representation output by the text encoder can be expressed in the form of a vector or a matrix. Since the sample hidden text representation of the sample text output by the encoder can be considered as the output in the intermediate processing process of the model, it may not be interpretable.

[0069] In the embodiments of the present application, the sample hidden text representation output by the text encoder is expressed in matrix form. Each row / column of the matrix corresponds to a phoneme in the sample phoneme sequence. It can also be understood that the sample hidden text representations corresponding to the respective phonemes in the sample phoneme sequence exist in the form of a vector in the sample hidden text representation matrix corresponding to the overall sample phoneme sequence. For ease of distinction, the sample hidden text representation matrix corresponding to the sample phoneme sequence can be referred to as the sample hidden text representation, and the sample hidden text representation vectors corresponding to the respective phonemes in the sample phoneme sequence are referred to as sample hidden text sub-representations.

[0070] Step 202: Based on the sample hidden text representation and the sample acoustic features corresponding to the sample text, perform duration prediction and decoding through the first decoder to obtain the first pronunciation duration and the first predicted acoustic features corresponding to the sample hidden text representation.

[0071] In a possible implementation manner, after obtaining the sample hidden text representation output by the text encoder, the first decoder performs duration prediction and decoding based on the sample hidden text representation and the sample acoustic features corresponding to the sample text to obtain the first pronunciation duration and the first predicted acoustic features corresponding to the sample hidden text representation.

[0072] Optionally, the acoustic features are used to represent the spectral features of speech. The sample acoustic features are the spectral features corresponding to the sample speech corresponding to the sample text, which can be mel-spectrogram, Mel-scale Frequency Cepstral Coefficients (MFCC), Linear Prediction Cepstral Coefficients (LPCC), Perceptual Linear Predictive (PLP), etc.

[0073] Optionally, the first decoder is a two-layer RNN structure. The first layer of RNN obtains the first pronunciation duration corresponding to the sample hidden text representation based on the sample acoustic features and the sample hidden text representation. Then, the second layer of RNN obtains the predicted acoustic features corresponding to the current moment based on the output of the first layer of RNN at the current moment and the sample hidden text representation. Finally, the first decoder obtains the first predicted acoustic features corresponding to the sample hidden text representation based on the predicted acoustic features output by the second layer of RNN at each moment.

[0074] Optionally, the first decoder is only used during the training process to obtain the first pronunciation duration and the first predicted acoustic features of the sample hidden text representation, and uses the obtained first pronunciation duration as a label to train the duration predictor.

[0075] Step 203: Based on the sample hidden text features, perform duration prediction through a duration predictor to obtain a second pronunciation duration.

[0076] In a possible implementation, the computer device can use a duration predictor to predict the pronunciation duration corresponding to the sample hidden text features, so as to improve the accuracy of the character pronunciation duration in the synthesized speech and improve the naturalness of the voice.

[0077] Optionally, after the computer device obtains the sample hidden text representation output by the text encoder, it inputs the sample hidden text representation into the duration predictor. The duration predictor can predict the duration of the acoustic features corresponding to each phoneme in the sample text based on the sample hidden text representation, that is, the pronunciation duration of each phoneme.

[0078] Step 204: Perform upsampling processing on the sample hidden text representation based on the first pronunciation duration to obtain a sample hidden text extended representation.

[0079] Since the lengths of the sample hidden text representation and the sample acoustic features corresponding to the sample text are different, in order to facilitate the subsequent second decoder to decode the sample hidden text representation to obtain the second predicted acoustic features, in a possible implementation, after the computer device obtains the first pronunciation duration output by the first decoder, it performs upsampling processing on the sample hidden text representation based on the number of pronunciation frames of the acoustic features corresponding to different sample hidden text sub-representations in the sample hidden text representation indicated by the first pronunciation duration, and extends each phoneme-corresponding sample hidden text sub-representation in the sample hidden text representation to the length corresponding to its pronunciation duration, so as to obtain a sample hidden text extended representation.

[0080] Optionally, the length of the sample hidden text sub-representation is positively correlated with its corresponding pronunciation duration, that is, the longer the pronunciation duration, the longer the length of the corresponding sample hidden text sub-representation.

[0081] Optionally, the computer device performs representation replication on the sample hidden text sub-representation in the sample hidden text representation based on the number of pronunciation frames indicated by the first pronunciation duration to obtain a sample hidden text extended representation.

[0082] For example, if the sample hidden text representation is "[abc]", and the first pronunciation duration shows that the number of pronunciation frames corresponding to phoneme a is 3, the number of pronunciation frames corresponding to phoneme b is 4, and the number of pronunciation frames corresponding to phoneme c is 2, then based on the pronunciation duration of the target phoneme to be synthesized, the sample hidden text extended representation "[aaabbbbcc]" after upsampling the target hidden text representation can be obtained.

[0083] Step 205: Decode the sample hidden text extended representation through a second decoder to obtain a second predicted acoustic feature.

[0084] In a possible implementation, the computer device inputs the sample hidden text extended representation into a second decoder. Based on the characteristics such as the part-of-speech of characters, the context association of characters, the character duration, and the tones, intonations, stresses, rhythms, etc. implied by the prosody features in the sample hidden text extended representation, the second decoder generates a second predicted acoustic feature corresponding to the sample text.

[0085] Optionally, the second decoder is one of a transformer structure or a CNN structure, and can finally obtain a second predicted acoustic feature corresponding to the sample text through multiple non-linear transformations.

[0086] In the embodiment of the present application, the second decoder is a parallel decoder. After the second decoder obtains the upsampled sample hidden text extended representation, since the length of the sample hidden text extended representation is the same as the length of the sample acoustic feature, the second decoder can directly perform parallel decoding on the sample hidden text sub-representations corresponding to different acoustic features in the input sample hidden text extended representation.

[0087] Optionally, the computer device performs representation division on the sample hidden text extended representation based on the number of pronunciation frames indicated by the first pronunciation duration to obtain a representation division result, and then performs parallel decoding on the representation division result through the second decoder to obtain predicted acoustic sub-features corresponding to different representation division results, and finally generates a second predicted acoustic feature based on the predicted acoustic sub-features.

[0088] Optionally, when the second decoder obtains the sample hidden text extended representation, it can obtain the first pronunciation duration and perform representation division on the sample hidden text extended representation based on the number of pronunciation frames indicated by the first pronunciation duration to obtain a representation division result.

[0089] For example, if the sample hidden text extended representation is "[aaabbbbcc]", and the first pronunciation duration shows that the number of pronunciation frames corresponding to the phoneme a is 3, the number of pronunciation frames corresponding to the phoneme b is 4, and the number of pronunciation frames corresponding to the phoneme c is 2, then the second decoder can perform representation division on the sample hidden text extended representation based on the number of pronunciation frames of each phoneme indicated by the first pronunciation duration to obtain a representation division result of "aaa", "bbbb", "cc", and then perform parallel decoding on them through the second decoder, that is, parallel decode "aaa", "bbbb", "cc", and then obtain a second predicted acoustic feature corresponding to the input sample text.

[0090] Step 206: Use the first pronunciation duration as the supervision for the second pronunciation duration, and use the sample acoustic feature as the supervision for the first predicted acoustic feature and the second predicted acoustic feature, and train the text encoder, the duration predictor, the first decoder, and the second decoder in an end-to-end manner.

[0091] In a possible implementation, the computer device trains a text encoder, a duration predictor, a first decoder, and a second decoder in an end-to-end manner based on a first pronunciation duration, a second pronunciation duration, sample acoustic features, a first predicted acoustic feature, and a second predicted acoustic feature.

[0092] Optionally, the end-to-end training method can solve the problem in the related art that each part of the task is independent and cannot be jointly optimized, making the training of the speech synthesis model simpler, and as the number of model layers increases and the training data becomes larger, the accuracy is higher.

[0093] Step 207: Construct a speech synthesis model based on the trained text encoder, duration predictor, and second decoder.

[0094] In a possible implementation, after the computer device trains the text encoder, duration predictor, first decoder, and second decoder in an end-to-end manner, it constructs a speech synthesis model based on the text encoder, duration predictor, and second decoder.

[0095] Optionally, when using the speech synthesis model to perform speech synthesis on the target text, first obtain the target hidden text representation through the text encoder, and based on the target hidden text feature, perform duration prediction through the duration predictor to obtain the target pronunciation duration, and then perform upsampling processing on the target hidden text representation based on the target pronunciation duration to obtain the target hidden text extended representation, and finally decode the target hidden text extended representation through the second decoder to obtain the target acoustic feature.

[0096] Schematically, as Figure 3 shown, after the computer device obtains the sample text 301, it outputs the sample hidden text representation 303 corresponding to the sample text 301 through the text encoder 310. The sample hidden text representation 303 will be sent to the first decoder 320, the duration predictor 330, and the upsampling processing 340 respectively, so that the first decoder 320 generates the first pronunciation duration 304 and the first predicted acoustic feature 305 based on the sample hidden text representation 303 and the sample acoustic feature 302 corresponding to the sample text 301.

[0097] The first pronunciation duration 304 will be sent to the duration predictor 330 and the upsampling processing 340 respectively. The duration predictor 330 generates the second pronunciation duration 306 based on the sample hidden text representation 303. The upsampling processing 340 performs upsampling processing on the sample hidden text representation 303 based on the first pronunciation duration 304 to obtain the sample hidden text extended representation 307, and then the second decoder 350 decodes the sample hidden text extended representation 307 to obtain the second predicted acoustic feature 308.

[0098] Finally, based on the first pronunciation duration 304, the second pronunciation duration 306, the sample acoustic features 302, the first predicted acoustic features 305, and the second predicted acoustic features 308, the computer device trains the text encoder 310, the duration predictor 330, the first decoder 320, and the second decoder 350 in an end-to-end manner, and constructs a speech synthesis model based on the trained text encoder 310, duration predictor 330, and second decoder 350.

[0099] In summary, in the embodiment of the present application, the computer device first encodes the sample text through the text encoder to obtain a sample hidden text representation. Then, based on the sample hidden text representation and the sample acoustic features corresponding to the sample text, the first decoder performs duration prediction and decoding to obtain the first pronunciation duration and the first predicted acoustic features corresponding to the sample hidden text representation. Based on the sample hidden text features, the duration predictor performs duration prediction to obtain the second pronunciation duration. The sample hidden text representation is upsampled based on the first pronunciation duration to obtain an extended sample hidden text representation. Further, the second decoder decodes the extended sample hidden text representation to obtain the second predicted acoustic features. Finally, the first pronunciation duration is used as the supervision for the second pronunciation duration, and the sample acoustic features are used as the supervision for the first predicted acoustic features and the second predicted acoustic features. The text encoder, duration predictor, first decoder, and second decoder are trained in an end-to-end manner, and a speech synthesis model is constructed based on the trained text encoder, duration predictor, and second decoder. By using the solution provided in the embodiment of the present application, the sample hidden text representation can be upsampled based on the first pronunciation duration and then input into the second decoder, which solves the problem of the mismatch in length between the sample text length and the corresponding sample acoustic features. By adopting an end-to-end training method, the structure of the acoustic model is further simplified, and the efficiency and conversion quality of converting the text to be synthesized into acoustic features by the acoustic model are improved.

[0100] In a possible implementation manner, when training the text encoder, duration predictor, first decoder, and second decoder, after determining the similarity between the first pronunciation duration and the second pronunciation duration, the feature similarity between the sample acoustic features and the first predicted acoustic features, and the feature similarity between the sample acoustic features and the second predicted acoustic features, these similarities are used as the loss function to adjust the parameters in the text encoder, duration predictor, first decoder, and second decoder, so that the acoustic features output by the speech synthesis model can continuously approach the sample acoustic features.

[0101] Optionally, the computer device determines a duration prediction loss based on the first pronunciation duration and the second pronunciation duration, determines a first acoustic feature prediction loss based on the sample acoustic features and the first predicted acoustic features, determines a second acoustic feature prediction loss based on the sample acoustic features and the second predicted acoustic features, and then trains the text encoder, the duration predictor, the first decoder, and the second decoder in an end-to-end manner based on the duration prediction loss, the first acoustic feature prediction loss, and the second acoustic feature prediction loss.

[0102] Optionally, the computer device sends the first pronunciation duration output by the first decoder to the duration predictor, and then uses the first pronunciation duration as a label to obtain a duration prediction loss based on the first pronunciation duration and the second pronunciation duration, and then performs supervised training on the duration predictor based on the duration prediction loss.

[0103] Optionally, the computer device calculates the distance between the first predicted acoustic features and the sample acoustic features and uses it as the first acoustic feature prediction loss, and then performs first acoustic feature prediction loss training based on the first acoustic feature prediction loss. Among them, during the first acoustic feature prediction loss training, the computer device only trains the text encoder and the first decoder.

[0104] Optionally, the computer device calculates the distance between the second predicted acoustic features and the sample acoustic features and uses it as the second acoustic feature prediction loss, and then performs second acoustic feature prediction loss training based on the second acoustic feature prediction loss. Among them, during the second acoustic feature prediction loss training, the computer device trains the text encoder, the first decoder, and the second decoder.

[0105] Optionally, the computer device can use the gradient descent or backpropagation (BP) algorithm to perform loss training and adjust the model parameters of the speech synthesis model based on the duration prediction loss, the first acoustic feature prediction loss, and the second acoustic feature prediction loss obtained each time.

[0106] Thereafter, if the similarity between the first pronunciation duration and the second pronunciation duration, the feature similarity between the first predicted acoustic features and the sample acoustic features, and the feature similarity between the second predicted acoustic features and the sample acoustic features meet the preset conditions, it can be considered that the acoustic model training is completed.

[0107] Optionally, the preset condition may be that the similarity between the first pronunciation duration and the second pronunciation duration, the feature similarity between the first predicted acoustic feature and the sample acoustic feature, and the feature similarity between the second predicted acoustic feature and the sample acoustic feature are all higher than a preset threshold, the similarity between the first pronunciation duration and the second pronunciation duration, the feature similarity between the first predicted acoustic feature and the sample acoustic feature, and the feature similarity between the second predicted acoustic feature and the sample acoustic feature are all basically unchanged, etc. The embodiments of the present application do not limit this.

[0108] In the embodiments of the present application, the computer device may train the text encoder, the duration predictor, the first decoder, and the second decoder in the same round of training process based on the duration prediction loss, the first acoustic feature prediction loss, and the second acoustic feature prediction loss, which simplifies the training process and reduces the time required to train the speech synthesis model.

[0109] In a possible implementation manner, the first decoder includes an attention mechanism and a sub-decoder. The computer device obtains the first pronunciation duration through the attention mechanism in the first decoder and obtains the first predicted acoustic feature through the sub-decoder in the first decoder. The following is a schematic description of the specific methods for duration prediction and decoding of the first decoder.

[0110] Please refer to Figure 4 , which shows a flowchart of the duration prediction and decoding process of the first decoder provided by an exemplary embodiment of the present application. This embodiment is described by taking this method as an example for a computer device. The method includes the following steps.

[0111] Step 401, based on the sample hidden text representation and the sample acoustic feature, determine the alignment matrix and the attention weight between the sample hidden text representation and the sample acoustic feature through the attention mechanism.

[0112] In a possible implementation manner, an RNN is included in the attention mechanism. The RNN can obtain the hidden state output by the hidden layer at the current moment based on the hidden state output by the hidden layer at the previous moment and the sample acoustic feature at the previous moment.

[0113] The attention mechanism then determines the alignment matrix and the attention weight between the sample hidden text representation and the sample acoustic feature based on the hidden state output by the hidden layer at the current moment and the sample hidden text representation. Among them, the attention mechanism can determine which sample hidden text representations will be used in each decoding step in the first decoder.

[0114] Optionally, after the RNN obtains the sample acoustic features and the hidden state of the previous moment, it obtains the hidden state of the hidden layer output at the current moment. The attention mechanism then calculates the distance value between the hidden state of the hidden layer output at the current moment and each sample hidden text sub-representation in the hidden text representation, and compares each distance value to determine that the phoneme corresponding to the sample hidden text sub-representation with the smallest distance value is the phoneme corresponding to the current moment.

[0115] For example, the sample text is "我", and the sample phoneme sequence obtained after preprocessing is "wo3". Since the sample phoneme sequence contains 3 phonemes, the hidden text representation of the sample phoneme sequence after being encoded by the text encoder is a three-row matrix [X1, X2, X3] T , X1, X2, X3 are the three sample hidden text sub-representations corresponding to the three phonemes. Furthermore, the attention mechanism combines the hidden state of the current hidden layer output with the input sample hidden text representation [X1, X2, X3] T The distances of different sample hidden text sub-representations X1, X2, and X3 are calculated to obtain three values ​​d1, d2, and d3, and then the sizes of these three values ​​are compared. If d1 is the smallest at the current moment, it can be determined that the sample hidden text sub-representation corresponding to the current moment is the sample hidden text sub-representation of the phoneme "w", that is, the phoneme corresponding to the current moment is "w".

[0116] Optionally, the attention mechanism calculates the distance value between the hidden state of the hidden layer output at the current moment and each sample hidden text sub-representation in the hidden text representation through the softmax function to obtain the probability distribution of the acoustic feature at the current moment being matched to each phoneme. This probability distribution is the attention weight between the sample hidden text representation and the sample acoustic feature.

[0117] Optionally, since the number of frames of the sample acoustic feature is much larger than the number of frames of the corresponding sample hidden text representation, after the first decoder obtains the phonemes corresponding to each moment, it determines the alignment matrix between the sample hidden text representation and the sample acoustic feature through the attention mechanism. The alignment matrix can reflect which frame sample acoustic features each phoneme corresponds to. It should be noted that each row / column in the alignment matrix determined by the attention mechanism corresponds to a phoneme, and each column / row corresponds to the acoustic feature of the frame sample at the current moment. Then each value in the alignment matrix represents the probability that the phoneme corresponding to the row / column corresponds to the acoustic feature of the frame sample corresponding to the column / row at the current moment. Obviously, there will be an almost diagonal value in the alignment matrix that is very large, and the rest of the values ​​are very small.

[0118] For example, the number of frames of the sample acoustic features corresponding to the sample text "I" is 200 frames. Among these 200 frames of acoustic features, the phoneme "w" corresponds to 160 frames of sample acoustic features, and both the phoneme "o" and the phoneme "3" correspond to 20 frames of sample acoustic features. If the first row of the alignment matrix determined by the attention mechanism corresponds to the phoneme "w", the second row corresponds to the phoneme "o", and the third row corresponds to the phoneme "3", and each column corresponds to this frame of sample acoustic features at the current moment. For the sake of convenient expression, taking 1 to indicate that this phoneme is likely to correspond to this frame of sample acoustic features and 0 to indicate that this phoneme is unlikely to correspond to this frame of sample acoustic features as an example, a 3×200 alignment matrix can be obtained. In this alignment matrix, only the values corresponding to the first column to the 160th column in the first row are 1, and the values corresponding to the remaining columns are 0. In the second row, only the values corresponding to the 160th column to the 180th column are 1, and the values corresponding to the remaining columns are 0. In the third row, only the values corresponding to the 180th column to the 200th column are 1, and the values corresponding to the remaining columns are 0.

[0119] Step 402: Based on the alignment matrix, determine the first pronunciation duration corresponding to the sample hidden text representation.

[0120] In a possible implementation manner, after the computer device obtains the alignment matrix output by the attention mechanism, it obtains each probability value in the alignment matrix that represents the probability that the phoneme corresponding to this row / column corresponds to this frame of sample acoustic features at the current moment corresponding to this column / row, and counts the number of large probability values corresponding to each frame of sample acoustic features at the current moment among the phonemes corresponding to each row / column. This number is used as the pronunciation duration of this phoneme, and then the first pronunciation duration corresponding to the sample phoneme sequence is obtained.

[0121] Optionally, the computer device can preset a probability threshold. A probability value greater than this probability threshold is a large probability value, and a probability value less than this probability threshold is a small probability value. The embodiments of the present application do not limit the specific method for determining whether a probability value is a large probability or a small probability.

[0122] For example, if the first row in the alignment matrix determined by the attention mechanism corresponds to the phoneme "w", the second row corresponds to the phoneme "o", the third row corresponds to the phoneme "3", and in the alignment matrix, only the values corresponding to the 1st to 160th columns in the first row are high-probability values, and the values corresponding to the remaining columns are low-probability values; only the values corresponding to the 160th to 180th columns in the second row are high-probability values, and the values corresponding to the remaining columns are low-probability values; only the values corresponding to the 180th to 200th columns in the third row are high-probability values, and the values corresponding to the remaining columns are low-probability values, then the computer device can obtain that the pronunciation duration of the phoneme "w" corresponding to the first row is 160 frames, the pronunciation duration of the phoneme "o" corresponding to the second row is 20 frames, the pronunciation duration of the phoneme "3" corresponding to the third row is 20 frames, and further obtain the first pronunciation duration corresponding to the sample phoneme sequence "wo3".

[0123] In another possible implementation manner, optionally, the first decoder can also directly determine the sample hidden text sub-representation corresponding to the current moment based on the distance values between the hidden state output by the hidden layer at the current moment obtained by the attention mechanism and different sample hidden text sub-representations in the input sample hidden text representation, that is, the phoneme corresponding to the current moment can be determined.

[0124] Since the number of frames of the sample acoustic features is much larger than the number of frames of the corresponding sample hidden text representation, the attention mechanism can directly determine that there are multiple consecutive frames of acoustic features corresponding to the same sample hidden text sub-representation. Whenever the attention mechanism detects that the sample hidden text sub-representation corresponding to the current moment is the same as the sample hidden text sub-representation corresponding to the previous moment, the duration information is accumulated starting from 1, and finally the obtained duration information is used as the number of frames of the acoustic features corresponding to the sample hidden text sub-representation, that is, the pronunciation duration of the phoneme corresponding to the sample hidden text sub-representation, and further the first pronunciation duration corresponding to the sample phoneme sequence is obtained.

[0125] For example, if the attention mechanism determines that the sample hidden text sub-representation corresponding to the 1st frame is the sample hidden text sub-representation of the phoneme "w", and until the sample hidden text sub-representation corresponding to the 160th frame is the sample hidden text sub-representation of the phoneme "w", then the duration information is accumulated starting from 1, and the accumulated result is that there are 160 frames of acoustic features corresponding to the sample hidden text sub-representation of the phoneme "w", that is, the duration information corresponding to the phoneme "w" is 160 frames. Similarly, the attention mechanism determines that the duration information corresponding to the phonemes "o" and "3" is both 20 frames, and further the first pronunciation duration corresponding to the sample phoneme sequence "wo3" is obtained.

[0126] Step 403: Based on the attention weights, the sample hidden text representation, and the sample acoustic features, perform decoding through a sub-decoder to obtain the first predicted acoustic features.

[0127] In a possible implementation, the sub-decoder also includes an RNN, which can obtain the first predicted acoustic feature corresponding to the sample hidden text representation based on the hidden state output by the RNN hidden layer in the attention mechanism at the current moment (obtained based on the hidden state output by the hidden layer at the previous moment and the sample acoustic feature at the previous moment), the attention weight, and the sample hidden text representation.

[0128] Optionally, the sub-decoder performs attention calculation on the sample hidden text representation based on the attention weight at the t-th moment to obtain the context feature at the t-th moment.

[0129] In the embodiment of the present application, the sub-decoder further obtains the context feature context at the t-th moment through weighted summation based on the attention weight at the t-th moment. For example, the hidden text representation corresponding to the sample phoneme sequence "wo3" is a three-row matrix [X1, X2, X3] T , where X1, X2, and X3 are three sample hidden text sub-representations corresponding to three phonemes. The sub-decoder obtains the attention weight corresponding to the phoneme "w" at the current moment as 0.8, the attention weight corresponding to the phoneme "o" as 0.1, and the attention weight corresponding to the phoneme "3" as 0.1. Then, the context at the current moment can be obtained through weighted summation: context = 0.8X1 + 0.1X2 + 0.1X3.

[0130] Furthermore, the computer device decodes through the sub-decoder based on the context feature at the t-th moment and the sample acoustic feature at the (t - 1)-th moment to obtain the predicted acoustic sub-feature at the t-th moment, and generates the first predicted acoustic feature based on the predicted acoustic sub-features at each moment.

[0131] In the embodiment of the present application, the sub-decoder combines the output state output by the RNN output layer in the attention mechanism at the current moment (obtaining the hidden state output by the hidden layer at the current moment based on the hidden state output by the hidden layer at the previous moment and the sample acoustic feature at the previous moment, and then obtaining the output state output by the output layer at the current moment based on the hidden state output by the hidden layer at the current moment) with the context feature context at the current moment through the concat function to obtain a vector, and then uses this vector as the input of the RNN input layer in the sub-decoder. Decoding is performed through this RNN to output the predicted acoustic sub-feature at the current moment, and then the first sample predicted acoustic feature is obtained.

[0132] Schematically, as Figure 5 shown, the attention mechanism 510 and the sub-decoder 520 in the first decoder are both RNN structures. The RNN in the attention mechanism 510 obtains the sample acoustic feature Y at the previous moment t-1 and the hidden state H at the previous moment t-1After that, the hidden state H of the hidden layer at the current moment is obtained t and the output state O of the output layer t , and then after determining the alignment matrix and attention weights between the sample hidden text representation and the sample acoustic features through the attention mechanism 510, the sub-decoder 520 performs attention calculation on the sample hidden text representation based on the attention weights at the current moment to obtain the context feature context at the current moment. The sub-decoder 520 combines the output state O t of the RNN output layer in the attention mechanism 510 at the current moment with the context feature context at the current moment through the concat function to obtain a vector vector, so that the sub-decoder 520 uses this vector vector as the input of the RNN input layer in the sub-decoder 520. The sub-decoder 520 obtains the vector vector at the current moment and the hidden state h t-1 at the previous moment, and then obtains the acoustic feature y t at the current moment, and further obtains the first sample predicted acoustic feature corresponding to the sample text.

[0133] In the embodiments of the present application, the first decoder determines the first pronunciation duration corresponding to the sample hidden text representation through the attention mechanism based on the sample hidden text representation and the sample acoustic features, and obtains the first predicted acoustic feature through the sub-decoder, so that the computer device trains the duration predictor based on the first pronunciation duration, and there is no need to separately train a duration extractor, reducing the training steps, and enabling the speech synthesis model to complete the training more simply and efficiently while ensuring the training quality.

[0134] Combining the above various embodiments, Figure 6 a flow of the training method of the speech synthesis model is provided:

[0135] Step 601, encoding the sample text through a text encoder to obtain a sample hidden text representation.

[0136] Step 602, determining the alignment matrix and attention weights between the sample hidden text representation and the sample acoustic features through the attention mechanism based on the sample hidden text representation and the sample acoustic features.

[0137] Step 603, determining the first pronunciation duration corresponding to the sample hidden text representation based on the alignment matrix.

[0138] Step 604, performing attention calculation on the sample hidden text representation based on the attention weights at the t-th moment to obtain the context feature at the t-th moment.

[0139] Step 605, decoding through the sub-decoder based on the context feature at the t-th moment and the sample acoustic feature at the (t - 1)-th moment to obtain the predicted acoustic sub-feature at the t-th moment.

[0140] Step 606: Generate a first predicted acoustic feature based on the predicted acoustic sub-features at each moment.

[0141] Step 607: Based on the sample hidden text feature, perform duration prediction through a duration predictor to obtain a second pronunciation duration.

[0142] Step 608: Upsample the sample hidden text representation based on the first pronunciation duration to obtain an extended sample hidden text representation.

[0143] Step 609: Based on the number of pronunciation frames indicated by the first pronunciation duration, perform representation partitioning on the extended sample hidden text representation to obtain a representation partitioning result.

[0144] Step 610: Parallelly decode the representation partitioning result through a second decoder to obtain the predicted acoustic sub-features corresponding to different representation partitioning results.

[0145] Step 611: Generate a second predicted acoustic feature based on the predicted acoustic sub-features.

[0146] Step 612: Determine a duration prediction loss based on the first pronunciation duration and the second pronunciation duration.

[0147] Step 613: Determine a first acoustic feature prediction loss based on the sample acoustic feature and the first predicted acoustic feature.

[0148] Step 614: Determine a second acoustic feature prediction loss based on the sample acoustic feature and the second predicted acoustic feature.

[0149] Step 615: Based on the duration prediction loss, the first acoustic feature prediction loss, and the second acoustic feature prediction loss, train a text encoder, a duration predictor, a first decoder, and a second decoder in an end-to-end manner.

[0150] Step 616: Construct a speech synthesis model based on the trained text encoder, duration predictor, and second decoder.

[0151] Schematically, as Figure 7 shown, after the computer device obtains the sample text 7001, it outputs the sample hidden text representation 7003 corresponding to the sample text 7001 through the text encoder 710. This sample hidden text representation 7003 will be sent to the first decoder 720, the duration predictor 730, and the upsampling process 740 respectively, so that the attention mechanism 721 in the first decoder 720 determines the alignment matrix 7004 and the attention weights 7005 between the sample hidden text representation 7003 and the sample acoustic feature 7002 corresponding to the sample text 7001.

[0152] The first decoder 720 further determines a first pronunciation duration 7006 corresponding to the sample hidden text representation 7003 based on the alignment matrix 7004. Then, the sub-decoder 722 obtains a first predicted acoustic feature 7007 based on the attention weight 7005, the sample hidden text representation 7003, and the sample acoustic feature 7002.

[0153] The first pronunciation duration 7006 determined by the attention mechanism 721 is respectively sent to the duration predictor 730 and the upsampling process 740. The duration predictor 730 generates a second pronunciation duration 7008 based on the sample hidden text representation 7003. The upsampling process 740 performs upsampling on the sample hidden text representation 7003 based on the first pronunciation duration 7006 to obtain a sample hidden text extended representation 7009. Then, the second decoder 750 decodes the sample hidden text extended representation 7009 to obtain a second predicted acoustic feature 7010.

[0154] Finally, the computer device determines a duration prediction loss based on the first pronunciation duration 7006 and the second pronunciation duration 7008, determines a first acoustic feature prediction loss based on the sample acoustic feature 7002 and the first predicted acoustic feature 7007, determines a second acoustic feature prediction loss based on the sample acoustic feature 7002 and the second predicted acoustic feature 7010. Then, based on these three losses, the text encoder 710, the duration predictor 730, the first decoder 720, and the second decoder 750 are trained in an end-to-end manner, and a speech synthesis model is constructed based on the trained text encoder 710, duration predictor 730, and second decoder 750.

[0155] The training method for the speech synthesis model provided by the embodiments of this application is widely applicable and can be applied to cloud services, enabling users to perform privatized deployment services on a private cloud. In a possible implementation, the computer device, in response to a model training request from a target object, obtains a sample text and a sample audio, where the sample audio belongs to the target object, generates a sample acoustic feature corresponding to the sample text based on the sample audio, and then constructs a speech synthesis model corresponding to the target object based on the trained text encoder, duration predictor, and second decoder.

[0156] For example, the computer device, in response to a model training request from a user, obtains the user's private data (private sample audio and corresponding private sample text), then generates a private sample acoustic feature corresponding to the sample text based on the user's private sample audio, and then constructs a private speech synthesis model with the user's desired timbre based on the trained text encoder, duration predictor, and second decoder.

[0157] In the embodiments of the present application, a user can train a voice synthesis model based on private data, and then construct a personalized voice synthesis model with a specific timbre to implement a privatized deployment service.

[0158] In a possible implementation manner, after a computer device trains a text encoder, a duration predictor, a first decoder, and a second decoder, it constructs a voice synthesis model based on the trained text encoder, duration predictor, and second decoder, enabling the user to perform voice synthesis through this voice synthesis model.

[0159] Please refer to Figure 8 , which shows a flowchart of a voice synthesis process provided by an exemplary embodiment of the present application. This embodiment is described by taking the method being used in a computer device as an example. The method includes the following steps.

[0160] Step 801: Encode the target text through a text encoder to obtain a target hidden text representation.

[0161] In a possible implementation manner, after the computer device obtains the target text, it first performs a text preprocessing operation on the target text, converts the target text into phonemes to obtain a target phoneme sequence corresponding to the target text, and then encodes the target phoneme sequence through the text encoder to obtain a target hidden text representation.

[0162] Step 802: Perform duration prediction through a duration predictor based on the target hidden text feature to obtain a target pronunciation duration.

[0163] In a possible implementation manner, after the computer device obtains the target hidden text representation output by the text encoder, it inputs the target hidden text representation into the duration predictor. The duration predictor can predict the duration of the acoustic features corresponding to each phoneme in the target text based on the target hidden text representation, and then obtain the target pronunciation duration corresponding to the target hidden text representation.

[0164] Step 803: Perform upsampling processing on the target hidden text representation based on the target pronunciation duration to obtain a target hidden text extended representation.

[0165] In a possible implementation manner, after the computer device obtains the target pronunciation duration output by the duration predictor, it performs upsampling processing on the target hidden text representation based on the number of pronunciation frames of the acoustic features corresponding to different target hidden text sub-representations in the target hidden text representation indicated by the target pronunciation duration, and extends each target hidden text sub-representation corresponding to each phoneme in the target hidden text representation to the length corresponding to its pronunciation duration, thereby obtaining a sample target hidden text extended representation.

[0166] Step 804: Decode the target hidden text extended representation through a second decoder to obtain the target acoustic features.

[0167] In a possible implementation, the computer device inputs the target hidden text extended representation into the second decoder, and the second decoder directly performs parallel decoding on the target hidden text sub-representations corresponding to different acoustic features in the input target hidden text extended representation to obtain the target acoustic features.

[0168] Step 805: Perform voice conversion on the target acoustic features through a vocoder to obtain the target audio corresponding to the target text.

[0169] In a possible implementation, after obtaining the target acoustic features corresponding to the target text, the computer device can input the target acoustic features corresponding to the target text into a vocoder, and the vocoder generates text speech based on the acoustic features to complete the speech synthesis of the text to be synthesized.

[0170] Optionally, the vocoder is used to convert the target acoustic features into a playable voice waveform, that is, to restore the target acoustic features to the target audio.

[0171] Optionally, the vocoder can be a neural network-based vocoder, such as WaveNet, HIFIGAN, or MelGAN, etc. The specific structure of the vocoder is not limited in this embodiment.

[0172] Schematically, as Figure 9 shown, after the computer device obtains the target text 901, it outputs the target hidden text representation 902 corresponding to the target text 901 through the text encoder 910. The target hidden text representation 902 will be sent to the duration predictor 920 and the upsampling process 930 respectively, so that the duration predictor 920 generates the target pronunciation duration 903 based on the target hidden text representation 902, and the upsampling process 930 performs upsampling processing on the target hidden text representation 902 based on the target pronunciation duration 903 to obtain the target hidden text extended representation 904. Then, the second decoder 940 decodes the target hidden text extended representation 904 to obtain the target acoustic features 905. Finally, the vocoder 950 decodes the target acoustic features 905 to obtain the target audio 906 corresponding to the target text 901.

[0173] In the embodiments of the present application, the target acoustic features corresponding to the input target text can be obtained based on a trained speech synthesis model, and then the target audio corresponding to the target text can be obtained based on the vocoder, which can improve the efficiency of speech synthesis while ensuring the quality of speech synthesis.

[0174] Please refer to Figure 10, which shows a structural block diagram of a training device for a speech synthesis model provided by an exemplary embodiment of the present application. The device includes:

[0175] A first text encoding module 1001, configured to encode a sample text through a text encoder to obtain a sample hidden text representation;

[0176] A first decoding module 1002, configured to perform duration prediction and decoding through a first decoder based on the sample hidden text representation and the sample acoustic features corresponding to the sample text, to obtain a first pronunciation duration and a first predicted acoustic feature corresponding to the sample hidden text representation;

[0177] A first duration prediction module 1003, configured to perform duration prediction through a duration predictor based on the sample hidden text features to obtain a second pronunciation duration;

[0178] A first upsampling module 1004, configured to perform upsampling processing on the sample hidden text representation based on the first pronunciation duration to obtain a sample hidden text extended representation;

[0179] A second decoding module 1005, configured to decode the sample hidden text extended representation through a second decoder to obtain a second predicted acoustic feature;

[0180] A training module 1006, configured to use the first pronunciation duration as the supervision of the second pronunciation duration, and use the sample acoustic features as the supervision of the first predicted acoustic feature and the second predicted acoustic feature, and train the text encoder, the duration predictor, the first decoder, and the second decoder in an end-to-end manner;

[0181] A model construction module 1007, configured to construct a speech synthesis model based on the trained text encoder, duration predictor, and second decoder.

[0182] Optionally, the first decoder includes an attention mechanism and a sub-decoder;

[0183] The first decoding module 1002 is configured to:

[0184] Based on the sample hidden text representation and the sample acoustic features, determine an alignment matrix and attention weights between the sample hidden text representation and the sample acoustic features through the attention mechanism;

[0185] Based on the alignment matrix, determine the first pronunciation duration corresponding to the sample hidden text representation;

[0186] Based on the attention weights, the sample hidden text representation, and the sample acoustic features, decoding is performed through the sub-decoder to obtain the first predicted acoustic features.

[0187] Optionally, the first decoding module 1002 is specifically configured to:

[0188] Perform attention calculation on the sample hidden text representation based on the attention weights at the t-th moment to obtain the context features at the t-th moment;

[0189] Based on the context features at the t-th moment and the sample acoustic features at the (t - 1)-th moment, decoding is performed through the sub-decoder to obtain the predicted acoustic sub-features at the t-th moment;

[0190] Generate the first predicted acoustic features based on the predicted acoustic sub-features at each moment.

[0191] Optionally, the training module 1006 is used to:

[0192] Determine the duration prediction loss based on the first pronunciation duration and the second pronunciation duration;

[0193] Determine the first acoustic feature prediction loss based on the sample acoustic features and the first predicted acoustic features;

[0194] Determine the second acoustic feature prediction loss based on the sample acoustic features and the second predicted acoustic features;

[0195] Train the text encoder, the duration predictor, the first decoder, and the second decoder in an end-to-end manner based on the duration prediction loss, the first acoustic feature prediction loss, and the second acoustic feature prediction loss.

[0196] Optionally, the first pronunciation duration is used to indicate the number of pronunciation frames of the acoustic features corresponding to different sample hidden text sub-representations in the sample hidden text representation;

[0197] The first upsampling module 1004 is used to:

[0198] Based on the number of pronunciation frames indicated by the first pronunciation duration, perform representation replication on the sample hidden text sub-representations in the sample hidden text representation to obtain the sample hidden text extended representation.

[0199] Optionally, the second decoding module 1005 is used to:

[0200] Based on the number of pronunciation frames indicated by the first pronunciation duration, perform representation partitioning on the sample hidden text extended representation to obtain a representation partitioning result;

[0201] The second decoder performs parallel decoding on the representation division result to obtain predicted acoustic sub - features corresponding to different representation division results;

[0202] Generate the second predicted acoustic feature based on the predicted acoustic sub - features.

[0203] Optionally, the device further includes:

[0204] An acquisition module, configured to acquire the sample text and the sample audio in response to a model training request of a target object, where the sample audio belongs to the target object;

[0205] An acoustic feature generation module, configured to generate the sample acoustic feature corresponding to the sample text based on the sample audio;

[0206] The model construction module 1007 is configured to:

[0207] Construct the speech synthesis model corresponding to the target object based on the trained text encoder, the duration predictor, and the second decoder.

[0208] Optionally, the device further includes:

[0209] A second text encoding module, configured to encode the target text through the text encoder to obtain a target hidden text representation;

[0210] A second duration prediction module, configured to perform duration prediction through the duration predictor based on the target hidden text feature to obtain a target pronunciation duration;

[0211] A second up - sampling module, configured to perform up - sampling processing on the target hidden text representation based on the target pronunciation duration to obtain a target hidden text extended representation;

[0212] A third decoding module, configured to decode the target hidden text extended representation through the second decoder to obtain a target acoustic feature;

[0213] A voice conversion module, configured to perform voice conversion on the target acoustic feature through a vocoder to obtain a target audio corresponding to the target text.

[0214] In summary, in the embodiment of the present application, the computer device first encodes the sample text through a text encoder to obtain a sample hidden text representation. Then, based on the sample hidden text representation and the sample acoustic features corresponding to the sample text, the first decoder is used for duration prediction and decoding to obtain the first pronunciation duration corresponding to the sample hidden text representation and the first predicted acoustic features. Based on the sample hidden text features, the duration predictor is used for duration prediction to obtain the second pronunciation duration. The sample hidden text representation is upsampled based on the first pronunciation duration to obtain an extended sample hidden text representation. Further, the second decoder decodes the extended sample hidden text representation to obtain the second predicted acoustic features. Finally, the first pronunciation duration is used as the supervision for the second pronunciation duration, and the sample acoustic features are used as the supervision for the first predicted acoustic features and the second predicted acoustic features. The text encoder, duration predictor, first decoder, and second decoder are trained in an end-to-end manner, and a speech synthesis model is constructed based on the trained text encoder, duration predictor, and second decoder. By using the solution provided in the embodiment of the present application, the sample hidden text representation can be upsampled based on the first pronunciation duration and then input into the second decoder, which solves the problem of the length mismatch between the sample text length and the corresponding sample acoustic features. By adopting an end-to-end training method, the structure of the acoustic model is further simplified, and the efficiency and quality of converting the text to be synthesized into acoustic features by the acoustic model are improved.

[0215] Please refer to Figure 11 , which shows a schematic structural diagram of a computer device provided by an exemplary embodiment of the present application. Specifically: The computer device 1100 includes a central processing unit (CPU) 1101, a system memory 1104 including a random access memory 1102 and a read-only memory 1103, and a system bus 1105 connecting the system memory 1104 and the central processing unit 1101. The computer device 1100 may further include a basic input / output system (Input / Output, I / O system) 1106 for facilitating the transfer of information between various components within the computer, and a mass storage device 1107 for storing an operating system 1113, application programs 1114, and other program modules 1115.

[0216] In some embodiments, the basic input / output system 1106 may include a display 1108 for displaying information and input devices 1109 such as a mouse, keyboard, etc. for user input of information. Both the display 1108 and the input devices 1109 are connected to the central processing unit 1101 through an input / output controller 1110 connected to the system bus 1105. The basic input / output system 1106 may also include an input / output controller 1110 for receiving and processing inputs from a plurality of other devices such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 1110 also provides outputs to a display screen, printer, or other types of output devices.

[0217] The mass storage device 1107 is connected to the central processing unit 1101 through a mass storage controller (not shown) connected to the system bus 1105. The mass storage device 1107 and its associated computer-readable medium provide non-volatile storage for the computer device 1100. That is, the mass storage device 1107 may include a computer-readable medium (not shown) such as a hard disk or a drive.

[0218] Without loss of generality, the computer-readable medium may include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes random access memory (RAM), read-only memory (ROM), flash memory or other solid-state storage technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic tape cartridges, tapes, magnetic disk storage or other magnetic storage devices. Of course, those skilled in the art will know that the computer storage media is not limited to the above several. The above system memory 1104 and mass storage device 1107 may be collectively referred to as memory.

[0219] The memory stores one or more programs, the one or more programs are configured to be executed by one or more central processing units 1101, the one or more programs contain instructions for implementing the above method, and the central processing unit 1101 executes the one or more programs to implement the methods provided by the above various method embodiments.

[0220] According to various embodiments of the present application, the computer device 1100 can also run on a remote computer on the network through a network such as the Internet. That is, the computer device 1100 can be connected to the network 1112 through the network interface unit 1111 connected to the system bus 1105. Or rather, the network interface unit 1111 can also be used to connect to other types of networks or remote computer systems (not shown).

[0221] The memory further includes one or more programs, and the one or more programs are stored in the memory. The one or more programs include steps for performing the method provided by the embodiments of the present application that are executed by the computer device.

[0222] Embodiments of the present application further provide a computer-readable storage medium. The computer-readable storage medium stores at least one segment of program, and the at least one segment of program is loaded and executed by the processor to implement the training method of the speech synthesis model as described in the above various embodiments.

[0223] Embodiments of the present application further provide a computer program product. The computer program product includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the earphone executes the training method of the speech synthesis model provided in various alternative implementations of the above aspects.

[0224] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium. The computer-readable storage medium can be the computer-readable storage medium included in the memory in the above embodiments; it can also exist separately and be a computer-readable storage medium not assembled into the terminal. The computer-readable storage medium stores at least one instruction, at least one segment of program, a code set or an instruction set, and the at least one instruction, the at least one segment of program, the code set or the instruction set are loaded and executed by the processor to implement the speech synthesis method described in any of the above method embodiments.

[0225] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disc, etc.

[0226] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included within the protection scope of the present application.

Claims

1. A training method for a speech synthesis model, characterized in that, The method includes: Encoding the sample text through a text encoder to obtain a sample hidden text representation; Based on the sample hidden text representation and the sample acoustic features corresponding to the sample text, determining an alignment matrix and attention weights between the sample hidden text representation and the sample acoustic features through the attention mechanism of a first decoder; Determining a first pronunciation duration corresponding to the sample hidden text representation based on the alignment matrix; Decoding through a sub-decoder of the first decoder based on the attention weights, the sample hidden text representation, and the sample acoustic features to obtain a first predicted acoustic feature; Performing duration prediction through a duration predictor based on the sample hidden text representation to obtain a second pronunciation duration; Performing upsampling processing on the sample hidden text representation based on the first pronunciation duration to obtain an extended sample hidden text representation; Decoding the extended sample hidden text representation through a second decoder to obtain a second predicted acoustic feature; Using the first pronunciation duration as the supervision for the second pronunciation duration, and using the sample acoustic features as the supervision for the first predicted acoustic feature and the second predicted acoustic feature, and training the text encoder, the duration predictor, the first decoder, and the second decoder in an end-to-end manner; Constructing a speech synthesis model based on the trained text encoder, duration predictor, and second decoder.

2. The method according to claim 1, wherein The decoding through the sub-decoder based on the attention weights and the sample hidden text representation to obtain a first predicted acoustic feature includes: Performing attention calculation on the sample hidden text representation based on the attention weights at the t-th moment to obtain a context feature at the t-th moment; Decoding through the sub-decoder based on the context feature at the t-th moment and the sample acoustic features at the (t - 1)-th moment to obtain a predicted acoustic sub-feature at the t-th moment; Generating the first predicted acoustic feature based on the predicted acoustic sub-features at each moment.

3. The method according to claim 1, characterized in that, The using the first pronunciation duration as the supervision for the second pronunciation duration, and using the sample acoustic features as the supervision for the first predicted acoustic feature and the second predicted acoustic feature, and training the text encoder, the duration predictor, the first decoder, and the second decoder in an end-to-end manner includes: Determining a duration prediction loss based on the first pronunciation duration and the second pronunciation duration; Determining a first acoustic feature prediction loss based on the sample acoustic features and the first predicted acoustic feature; Determining a second acoustic feature prediction loss based on the sample acoustic features and the second predicted acoustic feature; Training the text encoder, the duration predictor, the first decoder, and the second decoder in an end-to-end manner based on the duration prediction loss, the first acoustic feature prediction loss, and the second acoustic feature prediction loss.

4. The method according to claim 1, wherein The first pronunciation duration is used to indicate the number of pronunciation frames of the acoustic features corresponding to different sample hidden text sub-representations in the sample hidden text representation; Performing upsampling processing on the sample hidden text representation based on the first pronunciation duration to obtain a sample hidden text extended representation, including: Performing representation replication on the sample hidden text sub-representation in the sample hidden text representation based on the number of pronunciation frames indicated by the first pronunciation duration to obtain the sample hidden text extended representation.

5. The method according to claim 4, wherein Decoding the sample hidden text extended representation through a second decoder to obtain a second predicted acoustic feature, including: Performing representation partitioning on the sample hidden text extended representation based on the number of pronunciation frames indicated by the first pronunciation duration to obtain a representation partitioning result; Performing parallel decoding on the representation partitioning result through the second decoder to obtain predicted acoustic sub-features corresponding to different representation partitioning results; Generating the second predicted acoustic feature based on the predicted acoustic sub-features.

6. The method according to any one of claims 1 to 5, characterized in that, The method further includes: In response to a model training request of a target object, obtaining the sample text and the sample audio, where the sample audio belongs to the target object; Generating the sample acoustic feature corresponding to the sample text based on the sample audio; Constructing a speech synthesis model based on the trained text encoder, the duration predictor, and the second decoder, including: Constructing the speech synthesis model corresponding to the target object based on the trained text encoder, the duration predictor, and the second decoder.

7. According to the method described in any one of claims 1 to 5, characterized in that, The method further includes: Encoding a target text through the text encoder to obtain a target hidden text representation; Performing duration prediction through the duration predictor based on the target hidden text feature to obtain a target pronunciation duration; Performing upsampling processing on the target hidden text representation based on the target pronunciation duration to obtain a target hidden text extended representation; Decoding the target hidden text extended representation through the second decoder to obtain a target acoustic feature; Performing voice conversion on the target acoustic feature through a vocoder to obtain a target audio corresponding to the target text.

8. A training device for a speech synthesis model, characterized in that, The device includes: A first text encoding module, configured to encode a sample text through a text encoder to obtain a sample hidden text representation; A first decoding module, configured to determine an alignment matrix and attention weights between the sample hidden text representation and the sample acoustic feature through an attention mechanism of a first decoder based on the sample hidden text representation and the sample acoustic feature corresponding to the sample text; determining a first pronunciation duration corresponding to the sample hidden text representation based on the alignment matrix; decoding through a sub-decoder of the first decoder based on the attention weights, the sample hidden text representation, and the sample acoustic feature to obtain a first predicted acoustic feature; A first duration prediction module, configured to perform duration prediction through a duration predictor based on the sample hidden text representation to obtain a second pronunciation duration; A first upsampling module, configured to perform upsampling processing on the sample hidden text representation based on the first pronunciation duration to obtain a sample hidden text extended representation; A second decoding module, configured to decode the sample hidden text expansion representation through a second decoder to obtain a second predicted acoustic feature; A training module, configured to use the first pronunciation duration as supervision for the second pronunciation duration, and use the sample acoustic feature as supervision for the first predicted acoustic feature and the second predicted acoustic feature, and train the text encoder, the duration predictor, the first decoder, and the second decoder in an end-to-end manner; A model construction module, configured to construct a speech synthesis model based on the trained text encoder, the duration predictor, and the second decoder.

9. A computer device, characterized in that, The computer device includes a processor and a memory, and at least one program is stored in the memory, and the at least one program is loaded and executed by the processor to implement the training method of the speech synthesis model according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, At least one program is stored in the computer-readable storage medium, and the at least one program is loaded and executed by a processor to implement the training method of the speech synthesis model according to any one of claims 1 to 7.

11. A computer program product, characterized in that, The computer program product includes computer instructions, the computer instructions are stored in a computer-readable storage medium, a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to implement the training method of the speech synthesis model according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Speech synthesis method and system for new tone generation

    CN112802448A

  • Speech synthesis model training method, speech synthesis method and device

    CN113823260A