A training method for a voice synthesis model, a voice synthesis method and apparatus

By combining speech recognition error and spectrum error to evaluate the speech synthesis model, the problem of insufficient synthesized speech accuracy in the prior art is solved, and higher speech synthesis accuracy is achieved.

CN113393828BActive Publication Date: 2025-07-11TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011336173.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-24
Publication Date
2025-07-11
Estimated Expiration
2040-11-24

AI Technical Summary

Technical Problem

During the training process, the use of mean square error loss value of existing end-to-end speech synthesis models may cause blurred pronunciation or background noise, making it difficult to ensure the accuracy of synthesized speech.

Method used

Combining speech recognition error and spectrum error, the speech synthesis model is comprehensively evaluated, by obtaining the text and audio to be trained in the pair of samples to be trained, the speech synthesis model is used to obtain the Mel spectrum and the speech recognition model to obtain the phoneme sequence, and the model parameters are updated to improve the accuracy of synthesized speech.

Benefits of technology

By comprehensively evaluating speech recognition error and spectrum error, a speech synthesis model with better prediction results was trained, which improved the accuracy of synthesized speech.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113393828B_ABST
    Figure CN113393828B_ABST
Patent Text Reader

Abstract

The present application discloses a method for training a speech synthesis model implemented based on artificial intelligence technology, specifically relating to the field of speech processing technology. The present application includes: obtaining a training sample pair to be trained; obtaining a first Mel spectrogram through a speech synthesis model based on the text to be trained; obtaining a first phoneme sequence through a speech recognition model based on the first Mel spectrogram; and updating the model parameters of the speech synthesis model according to the loss value between the first Mel spectrogram and the true Mel spectrogram, and the loss value between the first phoneme sequence and the labeled phoneme sequence. The embodiments of the present application also provide a method and device for speech synthesis, which can comprehensively evaluate the speech synthesis model by combining speech recognition error and spectral error, thereby facilitating the training of a speech synthesis model with better prediction effect and improving the accuracy of the synthesized speech.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of speech processing, and in particular, to a method for training a speech synthesis model, a method and apparatus for speech synthesis. Background Art

[0002] Speech is a very common way of daily communication among people. With the development of Artificial Intelligence (AI) technology, Text To Speech (TTS) technology has attracted more and more attention. Using TTS technology, any text information can be converted into corresponding speech, making the synthesized speech intelligible, clear, natural and expressive.

[0003] Currently, the mainstream end-to-end speech synthesis model is usually adopted to implement speech synthesis. First, the text to be synthesized is converted into a phoneme sequence, and then the phoneme sequence is input into the speech synthesis model, and the synthesized speech is output through the speech synthesis model.

[0004] In the process of training the speech synthesis model, the mean-square error (MSE) loss value is used to determine whether the training is completed. However, since extremely small errors in the spectrum may lead to unclear pronunciation or background noise, and the MSE loss value only considers the spectrum error, the trained speech synthesis model is difficult to ensure the accuracy of the synthesized speech. Summary of the Invention

[0005] Embodiments of the present application provide a method for training a speech synthesis model, a method and apparatus for speech synthesis, which can comprehensively evaluate the speech synthesis model by combining speech recognition error and spectrum error, thereby facilitating the training of a speech synthesis model with better prediction effect and improving the accuracy of the synthesized speech.

[0006] In view of this, one aspect of the present application provides a method for training a speech synthesis model, including:

[0007] Obtaining a pair of samples to be trained, where the pair of samples to be trained includes a text to be trained and an audio to be trained with a corresponding relationship, the text to be trained corresponds to a labeled phoneme sequence, and the audio to be trained corresponds to a true Mel spectrum;

[0008] Based on the text to be trained, obtaining a first Mel spectrum through a speech synthesis model;

[0009] Based on the first Mel spectrum, obtaining a first phoneme sequence through a speech recognition model;

[0010] Updating the model parameters of the speech synthesis model according to the loss value between the first Mel spectrum and the true Mel spectrum, and the loss value between the first phoneme sequence and the labeled phoneme sequence.

[0011] On the other hand, this application provides a method for speech synthesis, including:

[0012] Obtain the text to be synthesized;

[0013] Based on the text to be synthesized, obtain the target Mel spectrogram through a speech synthesis model, where the speech synthesis model is trained according to the training methods of the above aspects;

[0014] Generate the target synthesized speech according to the target Mel spectrogram.

[0015] On the other hand, this application provides a speech synthesis model training device, including:

[0016] An obtaining module, configured to obtain a pair of samples to be trained, where the pair of samples to be trained includes a text to be trained and an audio to be trained with a corresponding relationship, the text to be trained corresponds to an annotated phoneme sequence, and the audio to be trained corresponds to a true Mel spectrogram;

[0017] The obtaining module is further configured to obtain a first Mel spectrogram through a speech synthesis model based on the text to be trained;

[0018] The obtaining module is further configured to obtain a first phoneme sequence through a speech recognition model based on the first Mel spectrogram;

[0019] A training module, configured to update the model parameters of the speech synthesis model according to the loss value between the first Mel spectrogram and the true Mel spectrogram, and the loss value between the first phoneme sequence and the annotated phoneme sequence.

[0020] In a possible design, in another implementation manner of the other aspect of the embodiments of this application, the audio to be trained is from a first object, and the first object corresponds to a first identity identifier;

[0021] The obtaining module is specifically configured to obtain a first Mel spectrogram through a speech synthesis model based on the text to be trained and the first identity identifier.

[0022] In a possible design, in another implementation manner of the other aspect of the embodiments of this application,

[0023] The training module is specifically configured to determine a mean square error loss value according to the first Mel spectrogram and the true Mel spectrogram;

[0024] Determine a first cross-entropy loss value according to the first phoneme sequence and the annotated phoneme sequence;

[0025] Determine a first target loss value according to the mean square error loss value and the first cross-entropy loss value;

[0026] Update the model parameters of the speech synthesis model according to the first target loss value.

[0027] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,

[0028] The training module is specifically configured to obtain an M-frame predicted frequency amplitude vector corresponding to the first Mel spectrogram, where each frame of the predicted frequency amplitude vector in the M-frame predicted frequency amplitude vector corresponds to a frame of audio signal in the audio to be trained, and M is an integer greater than or equal to 1;

[0029] Obtain an M-frame labeled frequency amplitude vector corresponding to the true Mel spectrogram, where each frame of the labeled frequency amplitude vector in the M-frame labeled frequency amplitude vector corresponds to a frame of audio signal in the audio to be trained;

[0030] Determine the average value of the predicted frequency amplitude according to the M-frame predicted frequency amplitude vector;

[0031] Determine the average value of the labeled frequency amplitude according to the M-frame labeled frequency amplitude vector;

[0032] Determine the M-frame frequency amplitude difference according to the average value of the predicted frequency amplitude and the average value of the labeled frequency amplitude;

[0033] Perform an averaging process on the M-frame frequency amplitude difference to obtain the mean squared error loss value.

[0034] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,

[0035] The training module is specifically configured to obtain an M-frame predicted phoneme vector corresponding to the first phoneme sequence, where each frame of the predicted phoneme vector in the M-frame predicted phoneme vector corresponds to a frame of audio signal in the audio to be trained, and M is an integer greater than or equal to 1;

[0036] Obtain an M-frame labeled phoneme vector corresponding to the labeled phoneme sequence, where each frame of the labeled phoneme vector in the M-frame labeled phoneme vector corresponds to a frame of audio signal in the audio to be trained;

[0037] Determine the cross-entropy loss value of the M-frame phonemes according to the M-frame predicted phoneme vector and the M-frame labeled phoneme vector;

[0038] Perform an averaging process on the cross-entropy loss value of the M-frame phonemes to obtain the first cross-entropy loss value.

[0039] In a possible design, in another implementation manner of another aspect of the embodiments of the present application, the voice synthesis model training device further includes a determination module;

[0040] The obtaining module is further configured to obtain a text to be tested and a second identity identifier corresponding to the text to be tested after the training module updates the model parameters of the speech synthesis model according to the loss value between the first mel-spectrogram and the real mel-spectrogram, and the loss value between the first phoneme sequence and the labeled phoneme sequence, where the second identity identifier corresponds to a second object;

[0041] The obtaining module is further configured to obtain a second mel-spectrogram through the speech synthesis model based on the text to be tested;

[0042] The obtaining module is further configured to obtain a predicted identity identifier through the object recognition model based on the second mel-spectrogram;

[0043] The obtaining module is further configured to obtain a second phoneme sequence through the speech recognition model based on the second mel-spectrogram;

[0044] The obtaining module is further configured to obtain a weight matrix through the speech synthesis model based on the text to be tested;

[0045] The determining module is configured to determine a target phoneme sequence according to the weight matrix;

[0046] The training module is further configured to update the model parameters of the speech synthesis model according to the loss value between the second identity identifier and the predicted identity identifier, and the loss value between the second phoneme sequence and the target phoneme sequence.

[0047] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,

[0048] The training module is specifically configured to determine a second cross-entropy loss value according to the second identity identifier and the predicted identity identifier;

[0049] Determine a third cross-entropy loss value according to the second phoneme sequence and the target phoneme sequence;

[0050] Determine a second target loss value according to the second cross-entropy loss value and the third cross-entropy loss value;

[0051] Update the model parameters of the speech synthesis model according to the second target loss value.

[0052] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,

[0053] The training module is specifically configured to obtain a labeled identity vector corresponding to the second identity identifier;

[0054] Obtain a predicted identity vector corresponding to the predicted identity identifier;

[0055] Determine a second cross-entropy loss value according to the labeled identity vector and the predicted identity vector.

[0056] In a possible design, in another implementation of another aspect of the embodiments of the present application,

[0057] The training module is specifically configured to obtain N-frame predicted phoneme vectors corresponding to the second phoneme sequence, where each frame of the N-frame predicted phoneme vectors corresponds to a frame of audio signal, and N is an integer greater than or equal to 1;

[0058] Obtain N-frame phoneme vectors corresponding to the target phoneme sequence, where each frame of the N-frame phoneme vectors corresponds to a frame of audio signal;

[0059] Determine the cross-entropy loss value of the N-frame phonemes according to the N-frame predicted phoneme vectors and the N-frame phoneme vectors;

[0060] Perform an averaging process on the cross-entropy loss value of the N-frame phonemes to obtain a third cross-entropy loss value.

[0061] In a possible design, in another implementation of another aspect of the embodiments of the present application,

[0062] The training module is further configured to update the model parameters of the speech recognition model according to the loss value between the first Mel spectrogram and the true Mel spectrogram, and the loss value between the first phoneme sequence and the labeled phoneme sequence.

[0063] Another aspect of the present application provides a speech synthesis device, including:

[0064] An acquisition module, configured to acquire the text to be synthesized;

[0065] The acquisition module is further configured to obtain a target Mel spectrogram through a speech synthesis model based on the text to be synthesized, where the speech synthesis model is trained according to the training methods of the above aspects;

[0066] A generation module, configured to generate a target synthesized speech according to the target Mel spectrogram.

[0067] In a possible design, in another implementation of another aspect of the embodiments of the present application,

[0068] The acquisition module is further configured to acquire a target identity identifier;

[0069] The acquisition module is specifically configured to obtain a target Mel spectrogram through a speech synthesis model based on the text to be synthesized and the target identity identifier.

[0070] Another aspect of the present application provides a computer device, including: a memory, a processor, and a bus system;

[0071] Wherein, the memory is used to store programs;

[0072] The processor is used to execute programs in the memory, and the processor is used to execute the methods of the above aspects according to the instructions in the program code;

[0073] The bus system is used to connect the memory and the processor to enable communication between the memory and the processor.

[0074] Another aspect of the present application provides a computer-readable storage medium, in which instructions are stored. When it runs on a computer, it causes the computer to execute the methods of the above aspects.

[0075] Another aspect of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to execute the methods provided in the above aspects.

[0076] From the above technical solutions, it can be seen that the embodiments of the present application have the following advantages:

[0077] In the embodiments of the present application, a method for training a speech synthesis model is provided. First, a to-be-trained sample pair is obtained, then based on the to-be-trained text, a first Mel spectrogram is obtained through the speech synthesis model, then based on the first Mel spectrogram, a first phoneme sequence is obtained through the speech recognition model, and finally, according to the loss value between the first Mel spectrogram and the real Mel spectrogram, and the loss value between the first phoneme sequence and the labeled phoneme sequence, the model parameters of the speech synthesis model are updated. Through the above method, the pre-trained speech recognition model is introduced into the model training framework, which can recognize the Mel spectrogram output by the to-be-trained speech synthesis model, determine the speech recognition error according to the recognized phoneme sequence and the labeled phoneme sequence, and determine the spectral error according to the predicted Mel spectrogram and the real Mel spectrogram. Combining the speech recognition error and the spectral error to comprehensively evaluate the speech synthesis model is beneficial to training a speech synthesis model with better prediction effect and improving the accuracy of the synthesized speech. Description of the Drawings

[0078] Figure 1 It is a schematic diagram of an application scenario of the speech synthesis method in the embodiments of the present application;

[0079] Figure 2 It is a schematic diagram of an architecture of the speech synthesis system in the embodiments of the present application;

[0080] Figure 3 It is a schematic diagram of an embodiment of the method for training a speech synthesis model in the embodiments of the present application;

[0081] Figure 4A schematic diagram of a framework for training a speech synthesis model based on supervised learning in an embodiment of the present application;

[0082] Figure 5 Another schematic diagram of a framework for training a speech synthesis model based on supervised learning in an embodiment of the present application;

[0083] Figure 6 A schematic diagram of a structure of a speech synthesis model in an embodiment of the present application;

[0084] Figure 7 Another schematic diagram of a structure of a speech synthesis model in an embodiment of the present application;

[0085] Figure 8 A schematic diagram of a framework for training a speech synthesis model based on self-supervised learning in an embodiment of the present application;

[0086] Figure 9 A schematic diagram of an embodiment of a speech synthesis method in an embodiment of the present application;

[0087] Figure 10 A schematic diagram of a speech synthesis interface in an embodiment of the present application;

[0088] Figure 11 Another schematic diagram of a speech synthesis interface in an embodiment of the present application;

[0089] Figure 12 A schematic diagram of an embodiment of a speech synthesis model training device in an embodiment of the present application;

[0090] Figure 13 A schematic diagram of an embodiment of a speech synthesis device in an embodiment of the present application;

[0091] Figure 14 A schematic diagram of a structure of a server in an embodiment of the present application;

[0092] Figure 15 A schematic diagram of a structure of a terminal device in an embodiment of the present application. Detailed implementation manners

[0093] Embodiments of the present application provide a method for training a speech synthesis model, a method and device for speech synthesis, which can comprehensively evaluate a speech synthesis model by combining speech recognition error and spectrum error, thereby facilitating training a speech synthesis model with better prediction effect and improving the accuracy of synthesized speech.

[0094] In the description and claims of this application and the above-mentioned drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "corresponding to" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0095] Text To Speech (TTS) technology and speech recognition technology are two key technologies necessary for realizing human-machine voice communication and establishing a spoken language system with the ability to listen and speak. Enabling computer devices to have the ability to speak similar to humans is an important competitive market in the information industry today. TTS technology, also known as text-to-speech conversion technology, can convert any text information into standard and fluent speech for reading aloud, which is equivalent to installing an artificial mouth on the machine. It involves multiple disciplinary technologies such as acoustics, linguistics, digital signal processing, and computer science. It is a cutting-edge technology in the field of Chinese information processing. The main problem it solves is how to convert text information into audible sound information, that is, to make the machine speak like a human.

[0096] TTS technology is used in various services. For example, automatic responses in call centers, voice announcements in public transportation, car navigation, electronic dictionaries, smart phones, smart speakers, voice assistants, entertainment robots, TV programs, community broadcasts, e-book reading, etc. In addition, TTS technology can also be used for people with speech impairments or reading disabilities. For example, people who have difficulty speaking due to illness can use synthetic speech to replace their own voices.

[0097] TTS technology belongs to the speech technology of Artificial Intelligence (AI) technology. Among them, AI is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results of theory, method, technology, and application system. In other words, AI is a comprehensive technology of computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. AI is also to study the design principles and implementation methods of various intelligent machines to enable the machine to have the functions of perception, reasoning, and decision-making.

[0098] AI technology is a comprehensive discipline that involves a wide range of fields, including both hardware-level and software-level technologies. AI basic technologies generally include technologies such as sensors, dedicated AI chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. AI software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0099] The key technologies of speech technology are ASR technology, TTS technology, and voiceprint recognition technology. Enabling computers to listen, see, speak, and feel is the future development direction of human-computer interaction, and among them, speech has become one of the most promising human-computer interaction methods in the future.

[0100] Taking the intelligent customer service scenario as an example, more and more companies are trying to gradually replace the positions of human customer service by connecting intelligent customer service robots, and the application of customer service robots has become more and more extensive. The intelligent customer service system is an industry-oriented application developed on the basis of large-scale knowledge processing, applicable to technologies such as large-scale knowledge processing, natural language understanding, knowledge management, automatic question answering systems, and reasoning. The intelligent customer service not only provides enterprises with fine-grained knowledge management technology, but also establishes a fast and effective technical means based on natural language for communication between enterprises and a large number of users. At the same time, it can also provide statistical analysis information required for refined management of enterprises. For ease of understanding, please refer to Figure 1 , Figure 1 This is a schematic diagram of an application scenario of the speech synthesis method in the embodiments of this application. As shown in the figure, exemplarily, when the user enters the application interface, taking the interface of a certain shopping application as an example, the user can enter "I am 180 cm tall. What size of clothes do I fit?" on the interface, and call the function of the Natural Language Processing (NLP) Software Development Kit (SDK) to detect the question input by the user, so as to determine the user's needs. Then, in combination with the knowledge database, the text to be synthesized is determined. For example, "You fit large-sized clothes". Then, the function of the TTSSDK is called to convert the text to be synthesized into the target synthesized voice, and the target synthesized voice is fed back by the customer service robot.

[0101] Exemplarily, when the user enters the application interface, taking the interface of a certain weather application as an example, the user can input a piece of speech through the microphone of the terminal device. For example, "What's the weather like today?", and call the function of the Automatic Speech Recognition (ASR) SDK to detect the speech spoken by the user, so as to determine the user's needs, and then combine the knowledge database to determine the text to be synthesized. For example, "Sunny turning to cloudy". Then call the function of the TTS SDK to convert the text to be synthesized into the target synthesized speech, and the customer service robot feeds back the target synthesized speech.

[0102] In order to synthesize more accurate and clearer speech in the above scenario, the present application proposes a training method for a speech synthesis model and a method for speech synthesis, both of which can be applied to Figure 2 the speech synthesis system shown. Please refer to Figure 2 , Figure 2 which is a schematic architecture diagram of the speech synthesis system in an embodiment of the present application. As shown in the figure, the speech synthesis system may include a server and a terminal device, and the client is deployed on the terminal device. The server involved in the present application may be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal device may be a smart phone, a tablet computer, a laptop computer, a handheld computer, a personal computer, a smart TV, a smart watch, etc., but is not limited thereto. The terminal device and the server may be directly or indirectly connected through wired or wireless communication methods, and the present application does not make any restrictions here. The number of servers and terminal devices is also not restricted.

[0103] The model training can be divided into two stages. The first stage is the supervised learning stage, and the second stage is the self-supervised learning stage. In the supervised learning stage, the server can obtain a large number of pairs of samples to be trained. Each pair of samples to be trained includes a text to be trained and an audio to be trained. Among them, it is necessary to pre-annotate the text to be trained to obtain the annotated phoneme sequence. At the same time, the true Mel spectrogram can be obtained according to the audio to be trained. During the model training process, first input each text to be trained into the speech synthesis model to be trained, and thus predict the first Mel spectrogram. Then input the predicted first Mel spectrogram into the speech recognition model, and thus predict the first phoneme sequence. Based on this, the server updates the model parameters of the speech synthesis model according to the loss value between the first Mel spectrogram and the true Mel spectrogram, and the loss value between the first phoneme sequence and the annotated phoneme sequence.

[0104] After the supervised learning phase ends, the performance of the speech synthesis model can be further improved in the self-supervised learning phase. In the self-supervised learning phase, the server can obtain a large number of texts to be tested and the identity identifiers labeled for each text to be tested. However, there is no audio corresponding to the text to be tested at this time. Therefore, it is necessary to simulate the text to be tested through the speech synthesis model to obtain the corresponding target phoneme sequence. During the model training process, each text to be tested is input into the speech synthesis model to be optimized, and thus the second Mel spectrogram is predicted. Then, the predicted second Mel spectrogram is input into the speech recognition model, and thus the second phoneme sequence is predicted. Based on this, the server tunes the model parameters of the speech synthesis model according to the loss value between the labeled identity identifier and the predicted identity identifier, and the loss value between the second phoneme sequence and the target phoneme sequence.

[0105] The model training process involves Machine Learning (ML). Among them, ML is an interdisciplinary subject in multiple fields, involving multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. ML is the core of AI and the fundamental way to make computers intelligent, and its applications cover all fields of AI. ML and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, and inductive learning.

[0106] After completing the training of the speech synthesis model, the server can store the speech synthesis model locally or send it to the terminal device. Exemplarily, if the speech synthesis model is stored on the server side, the terminal device needs to send the text to be synthesized to the server. The server inputs the text to be synthesized into the speech synthesis model, and the corresponding target Mel spectrogram is output through the speech synthesis model. The server converts the target Mel spectrogram into a speech signal, that is, generates the target synthesized speech. Finally, the server sends the target synthesized speech to the terminal device, and the terminal device plays the target synthesized speech.

[0107] Exemplarily, if the speech synthesis model is stored on the terminal device side, after the terminal device obtains the text to be synthesized, it can directly call the local speech synthesis model to output the corresponding target Mel spectrogram. Then, the terminal device converts the target Mel spectrogram into a speech signal, that is, generates the target synthesized speech. Finally, the terminal device plays the target synthesized speech.

[0108] This application involves relevant professional terms. For the convenience of understanding, these professional terms will be introduced separately below.

[0109] 1. Mel spectrogram: Mel spectrum is a spectrum obtained by performing Fourier transform on the acoustic signal and then transforming it through Mel scale. The spectrogram is often a large image. In order to obtain the sound features of appropriate size, the spectrogram can be transformed into a Mel spectrum after passing through the Mel scale filter bank.

[0110] 2. GT mel: Ground Truth (GT) mel spectrum, that is, the real mel spectrum involved in this application.

[0111] 3. GTA training: Ground Truth Autoregressive (GTA) training, that is, using the real Mel spectrum as the input of the decoder, performing autoregression and predicting the new spectrum.

[0112] 4. Free Running: Free training, that is, no real Mel spectrum is provided, only text (test text) is provided, so that the speech synthesis model can autoregressively predict the new spectrum (the second Mel spectrum).

[0113] 5. Linguistic Features, including but not limited to Chinese phonemes, English phonemes, Chinese vowel tones, word boundaries, phrase boundaries, sentence boundaries, etc. The texts to be synthesized, the texts to be trained, and the texts to be tested in this application all belong to linguistic features.

[0114] 6. Phone: It is the smallest unit of speech divided according to the natural properties of speech. It is analyzed based on the pronunciation action in the syllable, and one action constitutes a phoneme. Phonemes are divided into two categories: vowels and consonants. For example, the Chinese syllable "啊(ā)" has only one phoneme, "爱(ài)" has two phonemes, and "代(dài)" has three phonemes.

[0115] 7. Speaker Identification: It is to determine whether the audio belongs to a certain speaker based on a segment of Mel spectrum.

[0116] 8. Content Encoder: Maps the original phoneme sequence to a distributed vector representation that contains contextual information.

[0117] 9. Autoregressive Decoder: Each step of predicting the Mel spectrum depends on the Mel spectrum predicted in the previous step.

[0118] 10. Attention Mechanism: Used to provide the decoder with the contextual information required for each decoding step.

[0119] 11. Spectrogram Postnet: That is, it post-processes the mel spectrogram predicted by the autoregressive decoder to make the mel spectrogram smoother and of better quality.

[0120] 12. Speaker Identity: Usually, a set of vectors is used to represent the unique identifier of a speaker, and this identifier is the identity identifier involved in this application.

[0121] 13. Self-supervised learning: Without given input and output data pairs, the model can obtain reasonable labels based on arbitrary input and perform self-correction to improve the model performance.

[0122] 14. Mean Squared Error (MSE): Also known as the mean squared difference error function. Here, it refers to using the squared difference between the mel spectrogram predicted by the model and the true mel spectrogram as the objective for model optimization. The smaller the squared difference, the more accurate the mel spectrogram predicted by the model.

[0123] 15. Cross Entropy (CE): It measures the gap between the predicted distribution and the true distribution. In this application, it includes the error between the phoneme distribution predicted by the speech recognition model and the true phoneme distribution, and the error between the identity vector predicted by the object recognition model and the true identity vector.

[0124] 16. Loss function: In machine learning, it refers to the objective that the model training needs to minimize.

[0125] Combined with the above introduction, the training method of the speech synthesis model in this application will be introduced below. Please refer to Figure 3 One embodiment of the training method of the speech synthesis model in the embodiments of this application includes:

[0126] 101. Obtain the training sample pairs to be trained, where the training sample pairs to be trained include the text to be trained and the audio to be trained with a corresponding relationship. The text to be trained corresponds to the labeled phoneme sequence, and the audio to be trained corresponds to the true mel spectrogram.

[0127] In this embodiment, the speech synthesis model training model obtains the training sample pairs to be trained. In actual training, usually a large number of training sample pairs need to be obtained. For the convenience of description, one of the training sample pairs will be taken as an example for introduction below.

[0128] Specifically, a pair of samples to be trained includes two parts, namely, the text to be trained and the audio to be trained. The text to be trained is represented by linguistic features. Taking the original text "speech synthesis" as an example, the corresponding text to be trained is represented as "v3 in1 h e2 ch eng2", where "v" represents the phoneme of the word "language", and "3" represents that the tone of the word "language" is the third tone. "in" represents the phoneme of the word "sound", and "1" represents that the tone of the word "sound" is the first tone. "h" and "e" are both phonemes of the word "combination", and the first "2" represents that the tone of the word "combination" is the second tone. "ch" and "eng" are both phonemes of the word "combination", and the second "2" represents that the tone of the word "combination" is the second tone.

[0129] The audio to be trained refers to an audio of the original text. For example, subject A reads the four words "speech synthesis" and records it to obtain an audio to be trained (i.e., a speech signal). Since the high-frequency signal in the audio to be trained is weak, it is necessary to increase the high-frequency signal through pre-emphasis to balance the high-frequency and low-frequency signals. This can avoid the numerical operation problem in Fourier transform and also improve the signal-to-noise ratio (SNR). After pre-emphasis filtering of the training audio, it is also necessary to perform a sliding window Fourier transform on the signal in the time domain. Before the Fourier transform, in order to prevent energy leakage, a window function (e.g., Hanning window function) can be used. After short-time Fourier transform (STFT) processing, the linear spectrum of the training audio can be obtained. The linear spectrum dimension is usually high, for example, n_fft = 1024, hop size - 240, where n_fft = 1024 means that the input is sampled using a window of size 1024, and hop size = 240 means that 240 sampling points are staggered between two adjacent windows. Based on this, take the entire spectrum and divide it into equally spaced frequencies of n_mels = 80, where the equally spaced frequencies here refer to the distance heard by the human ear. Finally, when generating the true Mel spectrum, for each window, the amplitude of the signal in its component corresponds to the frequency in the mel ratio.

[0130] The real Mel spectrum obtains several sampling points according to the frame division. Assuming that each frame is 5 milliseconds (ms), if the audio to be trained is 1.56 seconds (i.e. 1560ms), the audio to be trained is divided into 312 frames. Based on this, it is also necessary to annotate each frame of audio with phonemes, thereby obtaining a real annotated phoneme sequence. The annotation method can be machine annotation or manual annotation, which is not limited here.

[0131] It should be noted that the speech synthesis model training model is deployed on a computer device, which can be a server or a terminal device. In this application, the speech synthesis model training model is deployed on a server as an example, but this should not be construed as a limitation of this application.

[0132] 102. Based on the text to be trained, obtain the first Mel spectrogram through the speech synthesis model;

[0133] In this embodiment, the speech synthesis model training model inputs the text to be trained into the speech synthesis model to be trained, and the speech synthesis model outputs the first Mel spectrogram. Here, the first Mel spectrogram is the predicted result.

[0134] 103. Based on the first Mel spectrogram, obtain the first phoneme sequence through the speech recognition model;

[0135] In this embodiment, the speech synthesis model training model inputs the predicted first Mel spectrogram into the pre-trained speech recognition model, and the speech recognition model predicts the first phoneme sequence. It should be noted that the first phoneme sequence has a corresponding relationship with the labeled phoneme sequence. For example, if the labeled phoneme sequence is a 312-frame labeled phoneme sequence, then the first phoneme sequence is a 312-frame predicted phoneme sequence. That is to say, each frame in the audio to be trained corresponds to a labeled phoneme and a predicted phoneme.

[0136] 104. Update the model parameters of the speech synthesis model according to the loss value between the first Mel spectrogram and the true Mel spectrogram, and the loss value between the first phoneme sequence and the labeled phoneme sequence.

[0137] In this embodiment, after the speech synthesis model training model obtains the first Mel spectrogram and the first phoneme sequence, it can calculate the loss value between the first Mel spectrogram and the true Mel spectrogram. For example, this loss value is L1. It can also calculate the loss value between the first phoneme sequence and the labeled phoneme sequence. For example, this loss value is L2. Based on this, the following method can be used to calculate the comprehensive loss value:

[0138] L = a * L1 + b * L2;

[0139] Where L represents the comprehensive loss value, a represents a weight value, b represents another weight value, L1 represents the loss value between the first Mel spectrogram and the true Mel spectrogram, and L2 represents the loss value between the first phoneme sequence and the labeled phoneme sequence. Finally, taking the minimization of the comprehensive loss value as the training objective, the model parameters of the speech synthesis model are optimized through the stochastic gradient descent (SGD) algorithm.

[0140] For ease of understanding, please refer to Figure 4 ,Figure 4 FIG. Figure 4 is a schematic framework diagram for training a speech synthesis model based on supervised learning in an embodiment of the present application. As shown in the figure, specifically, first, based on the text to be trained and the audio to be trained, a true Mel spectrum and an annotated phoneme sequence are obtained. Then, the text to be trained is input into the speech synthesis model, and a first Mel spectrum is output through the speech synthesis model. Based on the first Mel spectrum and the true Mel spectrum, a loss value is calculated. Next, the first Mel spectrum is input into the speech recognition model, and a first phoneme sequence is output through the speech recognition model. Based on the first phoneme sequence and the annotated phoneme sequence, another loss value is calculated. Finally, the two loss values are combined to update the model parameters of the speech synthesis model. After multiple iterations, a speech synthesis model with better performance can be trained.

[0141] In an embodiment of the present application, a method for training a speech synthesis model is provided. First, a pair of samples to be trained is obtained. Then, based on the text to be trained, a first Mel spectrum is obtained through the speech synthesis model. Next, based on the first Mel spectrum, a first phoneme sequence is obtained through the speech recognition model. Finally, according to the loss value between the first Mel spectrum and the true Mel spectrum, and the loss value between the first phoneme sequence and the annotated phoneme sequence, the model parameters of the speech synthesis model are updated. By the above method, the pre-trained speech recognition model is introduced into the model training framework, which can recognize the Mel spectrum output by the speech synthesis model to be trained, determine the speech recognition error according to the recognized phoneme sequence and the annotated phoneme sequence, and determine the spectral error according to the predicted Mel spectrum and the true Mel spectrum. Combining the speech recognition error and the spectral error to comprehensively evaluate the speech synthesis model is beneficial to training a speech synthesis model with better prediction effect and improving the accuracy of the synthesized speech.

[0142] Optionally, based on the above Figure 3 In another optional embodiment of the speech synthesis model training method provided in the embodiment of the present application corresponding to the above embodiment, the audio to be trained is from a first object, and the first object corresponds to a first identity identifier;

[0143] Based on the text to be trained, obtaining the first Mel spectrum through the speech synthesis model specifically includes the following steps:

[0144] Based on the text to be trained and the first identity identifier, the first Mel spectrum is obtained through the speech synthesis model.

[0145] In this embodiment, a method of introducing the speaker identity for training in the GTA training stage is introduced. In order to make the predicted synthesized speech closer to the real speech of a certain speaker, during the model training process, the first identity identifier can also be added, where the first identity identifier is the identifier of the first object, and the first object represents the speaker corresponding to the audio to be trained.

[0146] Specifically, during the training process, a large number of training sample pairs are often used. Each training sample pair includes a text to be trained and an audio to be trained. Some of the audios to be trained may come from the same object, and some may come from different objects. Please refer to Table 1, which is a schematic illustration of the relationship between the audios to be trained and the identity identifiers.

[0147] Table 1

[0148] Audio to be trained Object Identity identifier Identity vector 1st audio to be trained Tom 001 (1,0,0,0) 2nd audio to be trained Tom 001 (1,0,0,0) 3rd audio to be trained Tom 001 (1,0,0,0) 4th audio to be trained Tom 001 (1,0,0,0) 5th audio to be trained Tom 001 (1,0,0,0) 6th audio to be trained Jack 002 (0,1,0,0) 7th audio to be trained Jack 002 (0,1,0,0) 8th audio to be trained Jack 002 (0,1,0,0) 9th audio to be trained Anna 003 (0,0,1,0) 10th audio to be trained Anna 003 (0,0,1,0) 11th audio to be trained Betty 004 (0,0,0,1) 12th audio to be trained Betty 004 (0,0,0,1)

[0149] As can be seen from Table 1, assuming there are a total of 12 training sample pairs, that is, including 12 audios to be trained, these audios to be trained come from 4 speakers, namely "Tom", "Jack", "Anna" and "Betty". Each object has an identity identifier, and the identity identifiers of different objects are inconsistent. Taking 4 objects as an example, the identity vectors respectively include 4 elements, and the identity identifiers are encoded in the form of one-hot vectors, that is, the position corresponding to each element indicates an object. For example, the first element being "1" indicates the object is "Tom", the second element being "1" indicates the object is "Jack", and so on, which will not be elaborated here.

[0150] Based on this, assuming the audio to be trained in this application is the 2nd audio to be trained, then the first object is "Tom" and the first identity identifier is "001".

[0151] Combined with the above introduction, for the sake of easy understanding, please refer to Figure 5 , Figure 5 which is another schematic diagram of the framework for training a speech synthesis model based on supervised learning in the embodiments of this application. As shown in the figure, first, based on the text to be trained and the audio to be trained, the true Mel spectrogram and the labeled phoneme sequence are obtained, and at the same time, the first identity identifier corresponding to the audio to be trained is obtained. Then, the text to be trained and the identity vector corresponding to the first identity identifier are jointly input into the speech synthesis model, and the first Mel spectrogram is output through the speech synthesis model. Based on the first Mel spectrogram and the true Mel spectrogram, a loss value is calculated. Then, the first Mel spectrogram is input into the speech recognition model, and the first phoneme sequence is output through the speech recognition model. Based on the first phoneme sequence and the labeled phoneme sequence, another loss value is calculated. Finally, the two loss values are combined to update the model parameters of the speech synthesis model. After multiple iterations, a speech synthesis model with better performance can be trained.

[0152] Secondly, in the embodiments of the present application, a method of introducing speaker identity for training in the GTA training stage is provided. Through the above method, it is possible to more specifically train the speech belonging to a certain speaker, so that the finally synthesized speech is closer to the real speech of a certain speaker, thereby improving the model performance and enhancing the effect of speech personalization.

[0153] Optionally, on the basis of the above Figure 3 corresponding embodiment, in another optional embodiment of the speech synthesis model training method provided by the embodiments of the present application, according to the loss value between the first Mel spectrogram and the real Mel spectrogram, and the loss value between the first phoneme sequence and the labeled phoneme sequence, the model parameters of the speech synthesis model are updated, which specifically includes the following steps:

[0154] Determine the mean squared error loss value according to the first Mel spectrogram and the real Mel spectrogram;

[0155] Determine the first cross-entropy loss value according to the first phoneme sequence and the labeled phoneme sequence;

[0156] Determine the first target loss value according to the mean squared error loss value and the first cross-entropy loss value;

[0157] Update the model parameters of the speech synthesis model according to the first target loss value.

[0158] In this embodiment, a method of jointly training the speech synthesis model using the cross-entropy loss value and the MSE loss value in the GTA training stage is introduced. Here, two loss values are used. One is the MSE calculated for the first Mel spectrogram and the real Mel spectrogram, and the other is the CE loss value calculated for the first phoneme sequence and the labeled phoneme sequence. Based on this, the first target loss value can be calculated in the following way:

[0159] The first target loss value = w1 * LMSE + w2 * CE1;

[0160] wherein, w1 represents the first weight value, w2 represents the second weight value, LMSE represents the MSE loss value, and CE1 represents the first CE loss value. Finally, taking the minimization of the first target loss value as the training objective, the model parameters of the speech synthesis model are optimized through the SGD algorithm. If the first target loss value reaches convergence, or the number of training iterations reaches the iteration number threshold, it is determined that the model training condition has been satisfied, and the speech synthesis model can be output. It should be noted that the speech synthesis model can also adopt different types of network structures, for example, tacotron, tacotron 2, clarinet, and Deepvoice, etc. For the sake of understanding, the network structures of two speech synthesis models will be introduced separately below.

[0161] Exemplarily, please refer to Figure 6 , Figure 6 which is a schematic structural diagram of a speech synthesis model in an embodiment of the present application. As shown in the figure, the speech synthesis model includes four modules, namely, a Content Encoder, an Attention Mechanism, an Autoregressive Decoder, and a Spectrogram Postnet. Among them, the Content Encoder can convert the input text to be trained into context-related hidden features. The Content Encoder is usually composed of models with forward and backward relevance (for example, convolutional filter banks, highway networks, and bidirectional gated recurrent units). The features output by the Content Encoder have the ability to model the context.

[0162] The Attention Mechanism can combine the current state of the decoder to generate corresponding content context information for the decoder to better predict the next frame of spectrum. Speech synthesis is a task of establishing a monotonic mapping from a text sequence to a spectrum sequence. Therefore, when generating each frame of mel-spectrum, only a small part of phoneme content needs to be focused on, and this part of phoneme content is generated through the Attention Mechanism. The Attention Mechanism adopted in the present application can be location-sensitive attention, that is, the weight vector of the previous step is included in the calculation range of the current step context vector.

[0163] The Autoregressive Decoder generates the current frame of spectrum based on the content information generated by the current Attention Mechanism and the spectrum predicted in the previous frame. Since it needs to rely on the output of the previous frame, it is called an Autoregressive Decoder. Also because of the autoregressive nature, in an actual production environment, if the sequence is long, pronunciation errors may occur due to error accumulation.

[0164] The Spectrogram Postnet can smooth the spectrum predicted by the decoder to obtain a higher-quality spectrum, that is, output the first mel-spectrum. It can be seen that the previously trained speech recognition model is connected to the Spectrogram Postnet to classify each frame of phonemes and calculate the cross-entropy between the class distribution predicted by the speech recognition network and the label distribution corresponding to the true phonemes. At this stage, the model parameters of the speech synthesis network will be jointly updated by the mel-spectrum reconstruction error and the phoneme classification CE.

[0165] Exemplarily, please refer to Figure 7 , Figure 7Another structural schematic diagram of the speech synthesis model in the embodiments of the present application is shown in the figure. After the text to be trained is input into the speech synthesis model, the duration can be predicted first, that is, the pronunciation time of these phonemes when speaking needs to be considered. Since the duration of phonemes is determined based on the context, the pronunciation duration of each phoneme can be predicted by understanding it. Next, the fundamental frequency needs to be predicted, that is, in order to make the pronunciation as close to the human voice as possible, the pitch and intonation of each phoneme also need to be predicted. Since the same sound has completely different meanings when pronounced with different pitches and stresses. Predicting the frequency of each phoneme helps to pronounce each phoneme well because the frequency will tell the system what pitch and tone each phoneme should be pronounced. In addition, some phonemes are not fully voiced, which means that the vocal cords do not need to vibrate every time these sounds are pronounced. Finally, the text to be trained, the duration, and the frequency are combined to output the audio, and then the audio is converted into a Mel spectrogram, that is, the first Mel spectrogram is obtained.

[0166] Secondly, in the embodiments of the present application, a method for jointly training a speech synthesis model using the cross-entropy loss value and the MSE loss value during the GTA training phase is provided. Through the above method, only judging whether a model reaches the optimal from the MSE loss value is not sufficient to guarantee the pronunciation accuracy of the model. Therefore, the cross-entropy loss value between phoneme sequences can also be combined, so as to be able to reflect the pronunciation accuracy of the model, thereby improving the accuracy of the synthesized speech.

[0167] Optionally, based on the above Figure 3 In another optional embodiment of the speech synthesis model training method provided in the embodiments of the present application on the basis of the corresponding embodiment, the mean square error loss value is determined according to the first Mel spectrogram and the true Mel spectrogram, which specifically includes the following steps:

[0168] Obtain the M-frame predicted frequency amplitude vectors corresponding to the first Mel spectrogram, where each frame of the predicted frequency amplitude vectors in the M-frame predicted frequency amplitude vectors corresponds to a frame of audio signal in the audio to be trained, and M is an integer greater than or equal to 1;

[0169] Obtain the M-frame labeled frequency amplitude vectors corresponding to the true Mel spectrogram, where each frame of the labeled frequency amplitude vectors in the M-frame labeled frequency amplitude vectors corresponds to a frame of audio signal in the audio to be trained;

[0170] Determine the average predicted frequency amplitude according to the M-frame predicted frequency amplitude vectors;

[0171] Determine the average labeled frequency amplitude according to the M-frame labeled frequency amplitude vectors;

[0172] Determine the M-frame frequency amplitude difference according to the average predicted frequency amplitude and the average labeled frequency amplitude;

[0173] Average the difference in frequency and amplitude of the M frames to obtain the mean squared error loss value.

[0174] In this embodiment, a method for determining the MSE loss value in the GTA training stage is introduced. In the GTA stage, the teacher-student framework is mainly used for training. Based on the reconstruction error between the first mel-spectrogram and the real mel-spectrogram, the entire model is trained to obtain a relatively stable effect. Taking Figure 6 the shown speech synthesis model as an example, if the attention mechanism is replaced with an explicit duration model (e.g., Long Short-Term Memory (LSTM)), the alignment stability can be further improved. In addition, the mel-spectrogram reconstruction loss can adopt Dynamic Time Warping (DTW), which can further improve the quality of the predicted mel-spectrogram.

[0175] Specifically, assume that the mel-spectrogram includes M frames of audio signals. For example, M in [M, D] represents the number of frames of the audio signal, and D represents the mel-level frequency components. The specific values represent the amplitude. Based on this, the MSE loss value is calculated in the following way:

[0176]

[0177] where MSE represents the MSE loss value, M represents the number of frames of the audio signal, m represents the m-th frame of the audio signal, and y m represents the predicted frequency amplitude vector of the m-th frame of the audio signal, and this predicted frequency amplitude vector can include D values (D can be 80), represents the labeled frequency amplitude vector of the m-th frame of the audio signal. D represents the number of values included in each frequency amplitude vector. Therefore, y m / D represents the average predicted frequency amplitude, represents the average labeled frequency amplitude.

[0178] Based on this, represents the difference in frequency and amplitude of the M frames. After averaging the difference in frequency and amplitude of the M frames, the MSE loss value is obtained.

[0179] Again, in the embodiment of the present application, a method for determining the MSE loss value in the GTA training stage is provided. Through the above method, the predicted first mel-spectrogram and the labeled real mel-spectrogram can be effectively utilized to calculate the MSE loss value between the two. The MSE loss value can measure the average difference between the two mel-spectrograms, so as to minimize the difference between the mel-spectrograms as much as possible during the training process.

[0180] Optionally, in the above Figure 3Based on the corresponding embodiments, in another alternative embodiment of the speech synthesis model training method provided by the embodiments of the present application, to determine the first cross-entropy loss value according to the first phoneme sequence and the labeled phoneme sequence, the following steps are specifically included:

[0181] Obtain M-frame predicted phoneme vectors corresponding to the first phoneme sequence, where each frame of the predicted phoneme vectors in the M-frame predicted phoneme vectors corresponds to a frame of audio signal in the audio to be trained, and M is an integer greater than or equal to 1;

[0182] Obtain M-frame labeled phoneme vectors corresponding to the labeled phoneme sequence, where each frame of the labeled phoneme vectors in the M-frame labeled phoneme vectors corresponds to a frame of audio signal in the audio to be trained;

[0183] Determine the cross-entropy loss value of the M-frame phonemes according to the M-frame predicted phoneme vectors and the M-frame labeled phoneme vectors;

[0184] Perform an averaging process on the cross-entropy loss value of the M-frame phonemes to obtain the first cross-entropy loss value.

[0185] In this embodiment, a method for determining the first CE loss value in the GTA training stage is introduced. Since the audio to be trained contains the phonemes represented by each frame, the true labeled phoneme sequence is extracted from the audio to be trained, and then combined with the probability distribution corresponding to the first phoneme sequence predicted by the speech recognition network to calculate the CE.

[0186] Specifically, assume that the mel spectrogram includes M frames of audio signals, each frame of audio signal corresponds to a phoneme vector (i.e., a probability distribution vector), and each frame of audio signal corresponds to a phoneme. Taking a total of 50 phonemes as an example, a phoneme vector is represented as a 50-dimensional vector. Based on this, the following method is used to calculate the CE loss value of the M-frame phonemes:

[0187]

[0188] where CE1 represents the CE loss value of the M-frame phonemes, M represents the number of frames of the audio signal, m represents the m-th frame of the audio signal, represents the labeled phoneme vector of the m-th frame of the audio signal in the labeled phoneme sequence, and p m represents the predicted phoneme vector of the m-th frame of the audio signal in the first phoneme sequence.

[0189] Finally, average the CE loss value of the M-frame phonemes, that is, divide by M, thereby obtaining the first CE loss value.

[0190] Again, in an embodiment of the present application, a method for determining a first CE loss value in the GTA training phase is provided. Through the above method, the predicted first phoneme sequence and the labeled phoneme sequence can be effectively utilized to calculate the CE loss between the two. The CE loss can be taken as a frame, and the classification difference between the corresponding phonemes of each frame can be predicted, so that the difference between the phonemes can be minimized as much as possible during the training process.

[0191] Optionally, in the above Figure 3 On the basis of the corresponding embodiment, in another optional embodiment of the speech synthesis model training method provided in the embodiment of the present application, after the model parameters of the speech synthesis model are updated according to the loss value between the first mel spectrum and the true mel spectrum, and the loss value between the first phoneme sequence and the labeled phoneme sequence, the following steps are also included:

[0192] Acquire the text to be tested and a second identity identifier corresponding to the text to be tested, wherein the second identity identifier corresponds to the second object;

[0193] Based on the text to be tested, a second mel spectrum is obtained through a speech synthesis model;

[0194] Based on the second mel spectrum, the predicted identity is obtained through the object recognition model;

[0195] Based on the second mel-spectrogram, a second phoneme sequence is obtained through a speech recognition model;

[0196] Based on the text to be tested, a weight matrix is ​​obtained through a speech synthesis model;

[0197] Determine a target phoneme sequence according to a weight matrix;

[0198] Model parameters of the speech synthesis model are updated according to the loss value between the second identity identifier and the predicted identity identifier, and the loss value between the second phoneme sequence and the target phoneme sequence.

[0199] In this embodiment, a method for training a speech synthesis model in the free running training stage is introduced. In the free running stage, since there is no real mel spectrum, the real phoneme sequence is unknown, and the phoneme probability distribution corresponding to each frame of mel spectrum can be approximated based on the weight matrix in the attention mechanism, that is, the target phoneme sequence is obtained. With the real label for calculating the CE loss value, combined with the second phoneme sequence predicted by the speech recognition model, the cross entropy between the two phoneme sequences can be calculated.

[0200] After the GTA training stage is completed, a large number of texts and obscure and unsmooth sentences that are not within the training set can be collected and input into the speech synthesis model to obtain the predicted Mel spectrogram. Since there is no true Mel spectrogram, the reconstruction error between the predicted Mel spectrogram and the true Mel spectrogram cannot be calculated. However, the alignment relationship between the spectrogram and phonemes can be extracted from the weight matrix in the attention mechanism. Based on this alignment relationship, when the Mel spectrogram is input into the speech recognition model, the cross-entropy between phoneme distributions can be calculated and passed to the speech synthesis model to further improve the pronunciation stability.

[0201] For ease of introduction, please refer to Figure 8 , Figure 8 which is a schematic framework diagram for training a speech synthesis model based on self-supervised learning in an embodiment of the present application. As shown in the figure, specifically, first, the text to be tested and the first identity identifier are obtained, and then the text to be tested is input into the speech synthesis model. The weight matrix is output through the attention network in the speech synthesis model, and the target phoneme sequence is obtained after aligning the weight matrix with the frames. The second Mel spectrogram is output through the speech synthesis model, and then the second Mel spectrogram is input into the speech recognition model, and the second phoneme sequence is output through the speech recognition model. Based on the target phoneme sequence and the second phoneme sequence, a loss value is calculated. The second Mel spectrogram is input into the object recognition model (for example, Speaker Identification) to obtain the predicted identity identifier. Based on the predicted identity identifier and the first identity identifier, another loss value is calculated. Finally, the two loss values are combined to update the model parameters of the speech synthesis model. After multiple iterations, a speech synthesis model with better performance can be trained.

[0202] Secondly, in the embodiment of the present application, a method for training a speech synthesis model in the free running training stage is provided. Through the above method, speech recognition technology and speaker recognition technology are applied to the model training task based on the attention mechanism. Through staged training, the speech synthesis model can maintain more accurate pronunciation ability and higher similarity on less corpus or single-language corpus. The advantages of self-supervised learning are fully utilized, and the dependence of adaptive speech synthesis technology on data diversity is significantly reduced, so that the model maintains stronger robustness. In addition, combining the ASR error can effectively improve the problem that the evaluation cost of existing models is too high. Since the effects of existing models can only be auditioned by human ears and the number of artificial test sentences is limited, a comprehensive understanding of the model effects cannot be obtained, while the present application can effectively solve this problem.

[0203] Optionally, in the above Figure 3Based on the corresponding embodiments, in another optional embodiment of the speech synthesis model training method provided by the embodiments of the present application, according to the loss value between the second identity identifier and the predicted identity identifier, and the loss value between the second phoneme sequence and the target phoneme sequence, the model parameters of the speech synthesis model are updated, which specifically includes the following steps:

[0204] Determine the second cross-entropy loss value according to the second identity identifier and the predicted identity identifier;

[0205] Determine the third cross-entropy loss value according to the second phoneme sequence and the target phoneme sequence;

[0206] Determine the second target loss value according to the second cross-entropy loss value and the third cross-entropy loss value;

[0207] Update the model parameters of the speech synthesis model according to the second target loss value.

[0208] In this embodiment, a method of jointly training a speech synthesis model using two cross-entropy loss values in the free running training stage is introduced. Two loss values are used here. One is the third CE loss value calculated for the second phoneme sequence and the target phoneme sequence, and the other is the CE loss value calculated for the first phoneme sequence and the labeled phoneme sequence. Based on this, the second target loss value can be calculated in the following way:

[0209] Second target loss value = w3 * CE2 + w4 * CE3;

[0210] Among them, w3 represents the third weight value, w4 represents the fourth weight value, CE2 represents the second CE loss value, and CE3 represents the third CE loss value. Finally, taking the minimization of the second target loss value as the training objective, the model parameters of the speech synthesis model are optimized by the SGD algorithm. If the second target loss value reaches convergence, or the number of training iterations reaches the iteration number threshold, it is determined that the model training condition has been met, and the speech synthesis model can be output.

[0211] Again, in the embodiments of the present application, a method of jointly training a speech synthesis model using two cross-entropy loss values in the free running training stage is provided. Through the above method, combined with the self-supervised learning of any text, the model has more texts in different fields and different difficulties in the training stage, thus reducing the requirements for the quantity and content of the recorded corpus. At the same time, the present application incorporates the phoneme accuracy of accurately reading each frame into the CE loss function, which can significantly reduce the probability of errors of the existing speech synthesis system on unknown texts.

[0212] Optionally, in the above Figure 3Based on the corresponding embodiments, in another alternative embodiment of the speech synthesis model training method provided by the embodiments of the present application, the second cross-entropy loss value is determined according to the second identity identifier and the predicted identity identifier, which specifically includes the following steps:

[0213] Obtain the labeled identity vector corresponding to the second identity identifier;

[0214] Obtain the predicted identity vector corresponding to the predicted identity identifier;

[0215] Determine the second cross-entropy loss value according to the labeled identity vector and the predicted identity vector.

[0216] In this embodiment, a method for determining the second CE loss value in the free running training stage is introduced. In the free running stage, in order to prevent the voice actor's timbre from deviating from the original timbre due to unstable model updates, an object recognition model is added. The mel spectrogram is used as the input of the object recognition model to obtain the error of the timbre distribution, and this error is passed to the speech synthesis model to constrain the parameters of the model, so as to ensure a high similarity between the audio synthesized by the model and the original voice actor. During the model training process, a second identity identifier can also be added, where the second identity identifier is the identifier of the second object, and the second object represents the speaker corresponding to a certain audio to be trained in the GAT training stage.

[0217] Specifically, each frame of identity identifier corresponds to an identity vector (i.e., a probability distribution vector). Taking a total of 500 objects as an example, an identity vector is represented as a 500-dimensional vector. Based on this, the second CE loss value can be calculated in the following way:

[0218]

[0219] Among them, CE2 represents the second CE loss value, K represents the total dimension of the identity vector, k represents the k-th feature element in the identity vector, represents the k-th feature element in the labeled identity vector, p k represents the k-th feature element in the predicted identity vector.

[0220] Furthermore, in the embodiments of the present application, a method for determining the second CE loss value in the free running training stage is provided. By the above method, the speaker verification technology is integrated into the speech synthesis model, which can effectively prevent the deviation of the speaker's timbre due to parameter update, and further improve the effect and stability of speech synthesis. In the free running stage, the network is trained only with text without audio, eliminating the dependence on recorded audio. And in the free running stage, a large number of rare text corpora can be used to enhance the effect of the speech synthesis model.

[0221] Optionally, based on the above Figure 3 corresponding embodiment, in another optional embodiment of the speech synthesis model training method provided by the embodiments of the present application, a third cross-entropy loss value is determined according to the second phoneme sequence and the target phoneme sequence, which specifically includes the following steps:

[0222] Obtain the N-frame predicted phoneme vectors corresponding to the second phoneme sequence, where each frame of the predicted phoneme vectors in the N-frame predicted phoneme vectors corresponds to a frame of audio signal, and N is an integer greater than or equal to 1;

[0223] Obtain the N-frame phoneme vectors corresponding to the target phoneme sequence, where each frame of the phoneme vectors in the N-frame phoneme vectors corresponds to a frame of audio signal;

[0224] Determine the cross-entropy loss value of the N-frame phonemes according to the N-frame predicted phoneme vectors and the N-frame phoneme vectors;

[0225] Perform an averaging process on the cross-entropy loss value of the N-frame phonemes to obtain the third cross-entropy loss value.

[0226] In this embodiment, a method for determining the third CE loss value in the free running training stage is introduced. Since the test text contains the phonemes represented by each frame, therefore, based on the target phoneme sequence estimated from the test text, and then combined with the probability distribution corresponding to the second phoneme sequence predicted by the speech recognition network to calculate the CE.

[0227] Specifically, assume that the Mel spectrum includes N frames of audio signals, and each frame of audio signal corresponds to a phoneme vector (i.e., a probability distribution vector). Taking 50 phonemes as an example, a phoneme vector is represented as a 50-dimensional vector. Based on this, the following method is used to calculate the CE loss value of the N-frame phonemes:

[0228]

[0229] where CE3 represents the CE loss value of the N-frame phonemes, N represents the number of frames of the audio signal, n represents the nth frame of the audio signal, The phoneme vector of the n-th frame audio signal in the target phoneme sequence, p m represents the predicted phoneme vector of the n-th frame audio signal in the second phoneme sequence.

[0230] Finally, the CE loss values of N frames of phonemes are averaged, that is, divided by N, to obtain the third CE loss value.

[0231] Furthermore, in the embodiments of the present application, a method for determining the third CE loss value in the free running training stage is provided. Through the above method, the phonemes represented by each frame are included in the text to be tested. Therefore, a true target phoneme sequence can be obtained based on the text to be tested, and then the CE can be calculated by combining with the probability distribution corresponding to the second phoneme sequence predicted by the speech recognition network.

[0232] Optionally, on the basis of the above Figure 3 corresponding embodiment, in another optional embodiment of the speech synthesis model training method provided by the embodiments of the present application, the following steps may further be included:

[0233] Update the model parameters of the speech recognition model according to the loss value between the first Mel spectrogram and the true Mel spectrogram, and the loss value between the first phoneme sequence and the labeled phoneme sequence.

[0234] In this embodiment, a method for optimizing the speech recognition model in the GTA training stage is introduced. Here, two loss values are used. One is the MSE calculated for the first Mel spectrogram and the true Mel spectrogram, and the other is the CE loss value calculated for the first phoneme sequence and the labeled phoneme sequence. Based on this, the first target loss value is obtained. Finally, taking minimizing the first target loss value as the training objective, the model parameters of the speech recognition model are optimized by the SGD algorithm.

[0235] In the free running stage, the model parameters of the speech recognition model can also be optimized. Here, two loss values are also used. One is the CE loss calculated for the second identity identifier and the predicted identity identifier, and the other is the CE loss value calculated for the second phoneme sequence and the target phoneme sequence. Based on this, the second target loss value is obtained. Finally, taking minimizing the second target loss value as the training objective, the model parameters of the speech recognition model are optimized by the SGD algorithm.

[0236] It should be noted that the speech recognition model involved in this application can specifically be an ASR model. The ASR model can adopt the structure of a hybrid model, such as a Gaussian Mixture Model (GMM) and a Hidden Markov Model (HMM), a Deep Neural Network (DNN) and an HMM, an LSTM and an HMM, a Convolutional Neural Networks (CNN) and an HMM, a Recurrent Neural Network (RNN) and an HMM. The ASR model can also adopt a single model, such as an LSTM, a DNN, a CNN, an HMM, and an RNN, etc., which is not limited here.

[0237] Secondly, in the embodiments of this application, a method for optimizing the speech recognition model during the GTA training stage is provided. Through the above method, during the process of supervised learning, not only can the speech synthesis model be trained, but also the trained speech recognition model can be optimized. In this way, the speech recognition model can output a more accurate phoneme sequence, thereby further improving the performance of the model.

[0238] Combined with the above introduction, after the speech synthesis model is trained, it is possible to augment data using the speech synthesis model, which is highly generalizable. This application can be applied to products with speech synthesis capabilities, including but not limited to intelligent speakers, screen speakers, smart watches, smartphones, smart homes, and smart cars and other intelligent devices. It can also be applied to intelligent robots, AI customer service, and TTS cloud services, etc. Their usage scenarios can all strengthen the pronunciation stability and reduce the dependence on training data through the self-supervised learning algorithm proposed in this application. Based on this, the speech synthesis method in this application will be introduced below. Please refer to Figure 9 , an embodiment of the speech synthesis method in the embodiments of this application includes:

[0239] 201. Obtain the text to be synthesized;

[0240] In this embodiment, the speech synthesis device obtains the text to be synthesized, and the text to be synthesized is represented as linguistic features. Taking the original text as "speech synthesis" as an example, its corresponding text to be synthesized is represented as "v3 in1 h e2 ch eng2".

[0241] It should be noted that the speech synthesis device is deployed on a computer device. The computer device can be a terminal device or a server. This application takes the speech synthesis device being deployed on a terminal device as an example for illustration, but this should not be construed as a limitation to this application.

[0242] 202. Obtain a target Mel spectrogram through a speech synthesis model based on the text to be synthesized, where the speech synthesis model is trained according to the training method described in the above embodiments;

[0243] In this embodiment, the speech synthesis device will call the trained speech synthesis model to process the text to be synthesized and obtain the target Mel spectrogram.

[0244] 203. Generate a target synthesized speech according to the target Mel spectrogram.

[0245] In this embodiment, the speech synthesis device can inverse-transform the target Mel spectrogram into a time-domain waveform sample. Specifically, the WaveNet model can be used to transform the target Mel spectrogram into a time-domain waveform sample, and the target synthesized speech can be obtained according to the time-domain waveform sample. The Mel spectrogram is related to the STFT spectrogram. Relative to language and sound features, the Mel spectrogram is a relatively simple and low-level representation, and the WaveNet model can directly generate audio through this representation. It should be noted that other methods can also be used to convert the target Mel spectrogram into the target synthesized speech. This is only an example and should not be construed as a limitation of the present application.

[0246] Specifically, taking a speech synthesis product as an example for introduction, please refer to Figure 10 , Figure 10 which is a schematic diagram of the speech synthesis interface in the embodiment of the present application. As shown in the figure, the user can directly input the original text in the speech synthesis interface. For example, "speech synthesis", and then the original text input by the user can be seen in the text preview box. Or, the user can click the "Upload" button to select a piece of original text and upload it directly. Based on the original text input by the user or the uploaded original text, the corresponding text to be synthesized can be automatically generated. When the user clicks the "Synthesize" button, the text to be synthesized can be uploaded to the server, and the server will call the speech synthesis model to process the text to be synthesized to obtain the target Mel spectrogram. Or, when the user clicks the "Synthesize" button, the terminal device will call the local speech synthesis model to process the text to be synthesized to obtain the target Mel spectrogram. Finally, the target synthesized speech is generated according to the target Mel spectrogram. When the user clicks the "Preview" button, the target synthesized speech can be played through the terminal device.

[0247] In an embodiment of the present application, a speech synthesis method is provided. First, the text to be synthesized is obtained. Then, based on the text to be synthesized, a target Mel spectrogram is obtained through a speech synthesis model. Finally, a target synthesized speech is generated according to the target Mel spectrogram. By the above method, a pre-trained speech recognition model is introduced into the model training framework, which can recognize the Mel spectrogram output by the speech synthesis model to be trained, determine the speech recognition error according to the recognized phoneme sequence and the labeled phoneme sequence, and determine the spectral error according to the predicted Mel spectrogram and the true Mel spectrogram. Combining the speech recognition error and the spectral error to comprehensively evaluate the speech synthesis model is beneficial to training a speech synthesis model with better prediction effect, thereby improving the accuracy of the synthesized speech.

[0248] Optionally, based on the above Figure 9 corresponding embodiment, in another optional embodiment of the speech synthesis method provided by the embodiment of the present application, the following steps may further be included:

[0249] Obtain a target identity identifier;

[0250] Based on the text to be synthesized, obtaining the target Mel spectrogram through the speech synthesis model specifically includes the following steps:

[0251] Based on the text to be synthesized and the target identity identifier, obtain the target Mel spectrogram through the speech synthesis model.

[0252] In this embodiment, a method for synthesizing the speech of a certain object is introduced. Since in the process of model training, in order to make the predicted synthesized speech closer to the real speech of a certain speaker, an identity identifier can also be added as an input to the model. Based on this, in the process of model prediction, the identifier of the object to be simulated can also be added, that is, the target identity identifier is input. The speech synthesis model outputs the target Mel spectrogram according to the target identity identifier and the text to be synthesized, and finally converts the target Mel spectrogram into the target synthesized speech.

[0253] Specifically, taking a speech synthesis product as an example for introduction, please refer to Figure 11 , Figure 11Another schematic diagram of the speech synthesis interface in the embodiment of the present application is shown as follows. As shown in the figure, the user can directly input the original text in the speech synthesis interface. For example, "speech synthesis". Then, the original text input by the user can be seen in the text preview box. Or, the user can click the "Upload" button to select a piece of original text for direct upload. In addition, the user can also select the object to be synthesized on the speech synthesis interface. For example, the user can select to synthesize the voice of a cross-talk actor, that is, trigger the selection instruction for the cross-talk actor. The target identity identifier carried in the selection instruction is the identity identifier of the cross-talk actor, for example, 006. Based on the original text input by the user or the uploaded original text, the corresponding text to be synthesized can be automatically generated. When the user clicks the "Synthesize" button, the text to be synthesized and the target identity identifier selected by the user can be uploaded to the server, and the server calls the speech synthesis model to process the text to be synthesized and the target identity identifier to obtain the target Mel spectrogram. Or, when the user clicks the "Synthesize" button, the terminal device calls the local speech synthesis model to process the text to be synthesized and the target identity identifier selected by the user to obtain the target Mel spectrogram. Finally, the target synthesized speech is generated according to the target Mel spectrogram. When the user clicks the "Preview" button, the target synthesized speech can be played through the terminal device.

[0254] Since the self-supervised algorithm proposed in the present application can be used to enhance the speech synthesis effect, on the one hand, it can improve the effect of the speech synthesis model, and on the other hand, it can reduce the data acquisition cost. Based on these two advantages, it can be used for the customization of star voices. Since stars usually have a tight schedule and less clean corpus can be obtained. At the same time, it can also be used for the customization of the voices of teachers in online education. Since there are a large number of online education teachers and the answering work is very cumbersome, the present application can realize the customization of the voices of teachers only with a small amount of recorded audio of teachers, reduce the burden on teachers, and make the answering voice more anthropomorphic.

[0255] Secondly, in the embodiment of the present application, a method for synthesizing the voice of a certain object is provided. Through the above method, the target identity identifier can also be added. The target identity identifier is the identity identifier of the target object. Therefore, the synthesized target synthesized speech is more in line with the voice characteristics of the target object, thereby improving the speech synthesis effect.

[0256] The speech synthesis model training device in the present application will be described in detail below. Please refer to Figure 12 , Figure 12 An embodiment schematic diagram of the speech synthesis model training device in the embodiment of the present application is shown. The speech synthesis model training device 30 includes:

[0257] An acquisition module 301 is configured to acquire pairs of samples to be trained, where a pair of samples to be trained includes a text to be trained and an audio to be trained with a corresponding relationship. The text to be trained corresponds to an annotated phoneme sequence, and the audio to be trained corresponds to a true Mel spectrogram.

[0258] The acquisition module 301 is further configured to obtain a first Mel spectrogram based on the text to be trained through a speech synthesis model.

[0259] The acquisition module 301 is further configured to obtain a first phoneme sequence based on the first Mel spectrogram through a speech recognition model.

[0260] A training module 302 is configured to update the model parameters of the speech synthesis model according to the loss value between the first Mel spectrogram and the true Mel spectrogram, and the loss value between the first phoneme sequence and the annotated phoneme sequence.

[0261] In an embodiment of the present application, a speech synthesis model training device is provided. By using the above device and introducing a pre-trained speech recognition model into the model training framework, it is possible to recognize the Mel spectrogram output by the speech synthesis model to be trained, determine the speech recognition error according to the recognized phoneme sequence and the annotated phoneme sequence, and determine the spectral error according to the predicted Mel spectrogram and the true Mel spectrogram. By combining the speech recognition error and the spectral error, the speech synthesis model is comprehensively evaluated, which is beneficial to training a speech synthesis model with better prediction effect and improving the accuracy of the synthesized speech.

[0262] Optionally, on the basis of the corresponding embodiment above Figure 12 In another embodiment of the speech synthesis model training device 30 provided in the embodiment of the present application, the audio to be trained is from a first object, and the first object corresponds to a first identity identifier.

[0263] The acquisition module 301 is specifically configured to obtain a first Mel spectrogram based on the text to be trained and the first identity identifier through a speech synthesis model.

[0264] In an embodiment of the present application, a speech synthesis model training device is provided. By using the above device, it is possible to more specifically train the speech belonging to a certain speaker, so that the finally synthesized speech is closer to the true speech of a certain speaker, thereby improving the model performance and enhancing the effect of speech personalization.

[0265] Optionally, on the basis of the corresponding embodiment above Figure 12 In another embodiment of the speech synthesis model training device 30 provided in the embodiment of the present application

[0266] The training module 302 is specifically configured to determine a mean square error loss value according to the first Mel spectrogram and the true Mel spectrogram.

[0267] Determine a first cross-entropy loss value according to the first phoneme sequence and the labeled phoneme sequence;

[0268] Determine a first target loss value according to the mean squared error loss value and the first cross-entropy loss value;

[0269] Update the model parameters of the speech synthesis model according to the first target loss value.

[0270] In an embodiment of the present application, a speech synthesis model training device is provided. By using the above device, only judging whether a model is optimal from the MSE loss value is not sufficient to guarantee the pronunciation accuracy of the model. Therefore, the cross-entropy loss value between phoneme sequences can also be combined, so as to be able to reflect the pronunciation accuracy of the model, thereby improving the accuracy of the synthesized speech.

[0271] Optionally, on the basis of the corresponding embodiment above, Figure 12 In another embodiment of the speech synthesis model training device 30 provided in the embodiment of the present application,

[0272] The training module is specifically configured to obtain an M-frame predicted frequency amplitude vector corresponding to the first Mel spectrogram, where each frame of the predicted frequency amplitude vector in the M-frame predicted frequency amplitude vector corresponds to a frame of audio signal in the audio to be trained, and M is an integer greater than or equal to 1;

[0273] Obtain an M-frame labeled frequency amplitude vector corresponding to the true Mel spectrogram, where each frame of the labeled frequency amplitude vector in the M-frame labeled frequency amplitude vector corresponds to a frame of audio signal in the audio to be trained;

[0274] Determine the average predicted frequency amplitude according to the M-frame predicted frequency amplitude vector;

[0275] Determine the average labeled frequency amplitude according to the M-frame labeled frequency amplitude vector;

[0276] Determine the M-frame frequency amplitude difference according to the average predicted frequency amplitude and the average labeled frequency amplitude;

[0277] Perform an averaging process on the M-frame frequency amplitude difference to obtain the mean squared error loss value.

[0278] In an embodiment of the present application, a speech synthesis model training device is provided. By using the above device,

[0279] In an embodiment of the present application, a speech synthesis model training device is provided. By using the above device, the predicted first Mel spectrogram and the labeled true Mel spectrogram can be effectively utilized to calculate the MSE loss value between the two. The MSE loss value can measure the average difference between the two Mel spectrograms, so as to minimize the difference between the Mel spectrograms as much as possible during the training process.

[0280] Optionally, based on the above Figure 12 In another embodiment of the voice synthesis model training device 30 provided by the embodiments of the present application, on the basis of the corresponding embodiment,

[0281] The training module 302 is specifically configured to obtain M-frame predicted phoneme vectors corresponding to the first phoneme sequence, where each frame of the predicted phoneme vectors in the M-frame predicted phoneme vectors corresponds to a frame of audio signal in the audio to be trained, and M is an integer greater than or equal to 1;

[0282] Obtain M-frame labeled phoneme vectors corresponding to the labeled phoneme sequence, where each frame of the labeled phoneme vectors in the M-frame labeled phoneme vectors corresponds to a frame of audio signal in the audio to be trained;

[0283] Determine the cross-entropy loss value of the M-frame phonemes according to the M-frame predicted phoneme vectors and the M-frame labeled phoneme vectors;

[0284] Perform an averaging process on the cross-entropy loss value of the M-frame phonemes to obtain the first cross-entropy loss value.

[0285] In the embodiments of the present application, a voice synthesis model training device is provided. Using the above device,

[0286] In the embodiments of the present application, a voice synthesis model training device is provided. Using the above device, the effectively utilized predicted first phoneme sequence and labeled phoneme sequence can be used to calculate the CE loss between the two. The CE loss can predict the classification difference between the corresponding phonemes of each frame in units of frames, so that the difference between phonemes can be minimized as much as possible during the training process.

[0287] Optionally, based on the above Figure 12 In another embodiment of the voice synthesis model training device 30 provided by the embodiments of the present application, on the basis of the corresponding embodiment, the voice synthesis model training device 30 further includes a determination module 303;

[0288] The acquisition module 301 is further configured to obtain the text to be tested and the second identity identifier corresponding to the text to be tested after the training module updates the model parameters of the voice synthesis model according to the loss value between the first mel spectrogram and the true mel spectrogram, and the loss value between the first phoneme sequence and the labeled phoneme sequence, where the second identity identifier corresponds to the second object;

[0289] The acquisition module 301 is further configured to obtain a second mel spectrogram through the voice synthesis model based on the text to be tested;

[0290] The acquisition module 301 is further configured to obtain a predicted identity identifier through the object recognition model based on the second mel spectrogram;

[0291] The obtaining module 301 is further configured to obtain a second phoneme sequence based on the second Mel spectrogram through a speech recognition model;

[0292] The obtaining module 301 is further configured to obtain a weight matrix based on the text to be tested through a speech synthesis model;

[0293] The determining module 303 is configured to determine a target phoneme sequence according to the weight matrix;

[0294] The training module 302 is further configured to update the model parameters of the speech synthesis model according to the loss value between the second identity identifier and the predicted identity identifier, and the loss value between the second phoneme sequence and the target phoneme sequence.

[0295] In the embodiments of the present application, a speech synthesis model training device is provided. By using the above device, speech recognition technology and speaker recognition technology are applied to the model training task based on the attention mechanism. Through staged training, the speech synthesis model can maintain more accurate pronunciation ability and higher similarity on less corpus or single-language corpus. The advantages of self-supervised learning are fully utilized, and the dependence of adaptive speech synthesis technology on data diversity is significantly reduced, thereby making the model more robust. In addition, combining the ASR error can effectively improve the problem that the evaluation cost of existing models is too high. Since the effects of existing models can only be auditioned by human ears and the number of artificial test sentences is limited, it is impossible to have a more comprehensive understanding of the model effects, while the present application can effectively solve this problem.

[0296] Optionally, based on the corresponding embodiments above, in another embodiment of the speech synthesis model training device 30 provided by the embodiments of the present application, Figure 12 The training module 302 is specifically configured to determine a second cross-entropy loss value according to the second identity identifier and the predicted identity identifier;

[0297] Determine a third cross-entropy loss value according to the second phoneme sequence and the target phoneme sequence;

[0298] Determine a second target loss value according to the second cross-entropy loss value and the third cross-entropy loss value;

[0299] Update the model parameters of the speech synthesis model according to the second target loss value.

[0300] Update the model parameters of the speech synthesis model according to the second target loss value.

[0301] In the embodiment of the present application, a speech synthesis model training device is provided. By using the above device and combining it with self-supervised learning of any text, the model has more texts of different fields and different difficulties during the training phase, which reduces the requirements for the quantity and content of the recording corpus. At the same time, the present application incorporates the accuracy of the phonemes of each frame into the CE loss function, which can significantly reduce the probability of errors in the existing speech synthesis system on unknown texts.

[0302] Optionally, in the above Figure 12 On the basis of the corresponding embodiment, in another embodiment of the speech synthesis model training device 30 provided in the embodiment of the present application,

[0303] The training module 302 is specifically used to obtain a labeled identity vector corresponding to the second identity identifier;

[0304] Obtaining a predicted identity vector corresponding to the predicted identity identifier;

[0305] A second cross entropy loss value is determined based on the labeled identity vector and the predicted identity vector.

[0306] In the embodiment of the present application, a speech synthesis model training device is provided. By using the above device, speaker verification technology is integrated into the speech synthesis model, which can effectively prevent the deviation of the speaker's timbre due to parameter update, and further improve the effect and stability of speech synthesis. In the free running stage, only text is used to train the network without audio, eliminating the dependence on recorded audio, and in the free running stage, a large amount of uncommon text corpus can be used to enhance the effect of the speech synthesis model.

[0307] Optionally, in the above Figure 12 On the basis of the corresponding embodiment, in another embodiment of the speech synthesis model training device 30 provided in the embodiment of the present application,

[0308] The training module 302 is specifically configured to obtain N frames of predicted phoneme vectors corresponding to the second phoneme sequence, wherein each frame of the predicted phoneme vectors in the N frames corresponds to a frame of audio signal, and N is an integer greater than or equal to 1;

[0309] Obtaining N frames of phoneme vectors corresponding to the target phoneme sequence, wherein each frame of the N frames of phoneme vectors corresponds to a frame of audio signal;

[0310] Determine the cross entropy loss value of the N-frame phonemes according to the N-frame predicted phoneme vectors and the N-frame phoneme vectors;

[0311] The cross entropy loss values ​​of the N frames of phonemes are averaged to obtain a third cross entropy loss value.

[0312] In the embodiments of the present application, a voice synthesis model training device is provided. By using the above device, the test text contains the phonemes represented by each frame. Therefore, a true target phoneme sequence can be obtained based on the test text, and then the CE can be calculated by combining with the probability distribution corresponding to the second phoneme sequence predicted in the speech recognition network.

[0313] Optionally, based on the corresponding embodiments above, in another embodiment of the voice synthesis model training device 30 provided in the embodiments of the present application, Figure 12 the training module 302 is further configured to update the model parameters of the speech recognition model according to the loss value between the first Mel spectrum and the true Mel spectrum, and the loss value between the first phoneme sequence and the labeled phoneme sequence.

[0314]

[0315] In the embodiments of the present application, a voice synthesis model training device is provided. By using the above device, in the process of supervised learning, not only can the voice synthesis model be trained, but also the trained speech recognition model can be optimized. In this way, the speech recognition model can output a more accurate phoneme sequence, thereby further improving the performance of the model.

[0316] Figure 13 The following will describe the voice synthesis device in the present application in detail. Please refer to Figure 13 , which is a schematic diagram of an embodiment of the voice synthesis device in the embodiments of the present application. The voice synthesis device 40 includes:

[0317] An acquisition module 401, configured to acquire the text to be synthesized;

[0318] The acquisition module 401 is further configured to obtain a target Mel spectrum through the voice synthesis model based on the text to be synthesized, where the voice synthesis model is trained according to the training method provided in the above embodiments;

[0319] A generation module 402, configured to generate a target synthesized voice according to the target Mel spectrum.

[0320] In the embodiments of the present application, a voice synthesis device is provided. By using the above device, introducing a pre-trained speech recognition model into the model training framework can identify the Mel spectrum output by the voice synthesis model to be trained, determine the speech recognition error according to the recognized phoneme sequence and the labeled phoneme sequence, and determine the spectral error according to the predicted Mel spectrum and the true Mel spectrum. Combining the speech recognition error and the spectral error to comprehensively evaluate the voice synthesis model is beneficial to training a voice synthesis model with better prediction effect, thereby improving the accuracy of the synthesized voice.

[0321] Figure 13 Optionally, based on the above Figure 13Based on the corresponding embodiments, in another embodiment of the speech synthesis device 40 provided in the embodiments of the present application,

[0322] The acquisition module 401 is further configured to acquire a target identity identifier;

[0323] Specifically, the acquisition module 401 is configured to obtain a target Mel spectrogram through a speech synthesis model based on the text to be synthesized and the target identity identifier.

[0324] In the embodiments of the present application, a speech synthesis device is provided. By using the above device, a target identity identifier can also be added. The target identity identifier is the identity identifier of the target object. Therefore, the synthesized target speech is more in line with the speech characteristics of the target object, thereby improving the speech synthesis effect.

[0325] The speech synthesis model training device and the speech synthesis device provided in the present application can be deployed on a server. Please refer to Figure 14 , Figure 14 FIG. is a schematic structural diagram of a server provided in an embodiment of the present application. The server 500 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 522 (for example, one or more processors) and a memory 532, and one or more storage media 530 (for example, one or more mass storage devices) for storing application programs 542 or data 544. Among them, the memory 532 and the storage media 530 may be transient storage or persistent storage. The program stored in the storage media 530 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 522 may be configured to communicate with the storage media 530 and execute a series of instruction operations in the storage media 530 on the server 500.

[0326] The server 500 may further include one or more power supplies 526, one or more wired or wireless network interfaces 550, one or more input / output interfaces 558, and / or one or more operating systems 541, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM and so on.

[0327] In the embodiments of the present application, the CPU 522 included in the server further has the following functions:

[0328] Obtain pairs of samples to be trained, where each pair of samples to be trained includes a text to be trained and an audio to be trained with a corresponding relationship, the text to be trained corresponds to an annotated phoneme sequence, and the audio to be trained corresponds to a true Mel spectrogram;

[0329] Based on the text to be trained, obtain a first Mel spectrogram through a speech synthesis model;

[0330] Based on the first Mel spectrogram, obtain a first phoneme sequence through a speech recognition model;

[0331] Update the model parameters of the speech synthesis model according to the loss value between the first Mel spectrogram and the true Mel spectrogram, and the loss value between the first phoneme sequence and the annotated phoneme sequence.

[0332] In the embodiments of the present application, the CPU 522 included in the server further has the following functions:

[0333] Obtain the text to be synthesized;

[0334] Based on the text to be synthesized, obtain a target Mel spectrogram through a speech synthesis model;

[0335] Generate a target synthesized speech according to the target Mel spectrogram.

[0336] The steps performed by the server in the above embodiments can be based on the Figure 14 shown server structure.

[0337] The speech synthesis model training device and speech synthesis device provided by the present application can be deployed on a server. As Figure 15 shown, for ease of illustration, only the parts related to the embodiments of the present application are shown. For specific technical details not disclosed, please refer to the method part of the embodiments of the present application. In the embodiments of the present application, a smart phone is used as an example of a terminal device for illustration:

[0338] Figure 15 Shown is a block diagram of a part of the structure of a smart phone related to the terminal device provided by the embodiments of the present application. Referring to Figure 15 , a smart phone includes: a radio frequency (RF) circuit 610, a memory 620, an input unit 630, a display unit 640, a sensor 650, an audio circuit 660, a wireless fidelity (WiFi) module 670, a processor 680, and a power supply 690, etc. Those skilled in the art can understand that Figure 15 the smart phone structure shown in does not constitute a limitation on the smart phone, and it may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0339] The following will specifically introduce each component of the smart phone in conjunction with Figure 15 :

[0340] The RF circuit 610 can be used for receiving and transmitting signals during information reception or call processes. Specifically, after receiving the downlink information from the base station, it is given to the processor 680 for processing; in addition, the uplink data designed is sent to the base station. Generally, the RF circuit 610 includes but is not limited to antennas, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LNA), a duplexer, etc. In addition, the RF circuit 610 can also communicate with the network and other devices through wireless communication. The above wireless communication can use any communication standard or protocol, including but not limited to the Global System of Mobile Communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.

[0341] The memory 620 can be used to store software programs and modules. The processor 680 executes various functional applications and data processing of the smart phone by running the software programs and modules stored in the memory 620. The memory 620 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the smart phone (such as audio data, a phone book, etc.). In addition, the memory 620 can include high-speed random access memory and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices.

[0342] The input unit 630 can be used to receive input numerical or character information and generate key signal inputs related to the user settings and function controls of the smart phone. Specifically, the input unit 630 can include a touch panel 631 and other input devices 632. The touch panel 631, also known as a touch screen, can collect touch operations of the user thereon or nearby (such as operations of the user using any suitable object or accessory such as a finger, a stylus, etc. on or near the touch panel 631), and drive corresponding connection devices according to a preset program. Optionally, the touch panel 631 can include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch position of the user, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 680, and can also receive and execute commands sent by the processor 680. In addition, various types such as resistive, capacitive, infrared, and surface acoustic wave can be used to implement the touch panel 631. In addition to the touch panel 631, the input unit 630 can also include other input devices 632. Specifically, the other input devices 632 can include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, a joystick, etc.

[0343] The display unit 640 can be used to display information input by the user or information provided to the user and various menus of the smart phone. The display unit 640 can include a display panel 641. Optionally, the display panel 641 can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. Further, the touch panel 631 can cover the display panel 641. When the touch panel 631 detects a touch operation thereon or nearby, it is transmitted to the processor 680 to determine the type of touch event. Subsequently, the processor 680 provides corresponding visual output on the display panel 641 according to the type of touch event. Although in Figure 15 it, the touch panel 631 and the display panel 641 are implemented as two independent components to realize the input and input functions of the smart phone, but in some embodiments, the touch panel 631 and the display panel 641 can be integrated to realize the input and output functions of the smart phone.

[0344] The smart phone may also include at least one sensor 650, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display panel 641 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 641 and / or the backlight when the smart phone is moved to the ear. As a kind of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in all directions (generally three axes), and can detect the magnitude and direction of gravity when stationary, and can be used in applications for identifying the posture of the smart phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors such as gyroscopes, barometers, hygrometers, thermometers, infrared sensors that the smart phone can also be configured with, they will not be elaborated here.

[0345] The audio circuit 660, the speaker 661, and the microphone 662 can provide an audio interface between the user and the smart phone. The audio circuit 660 can transmit the electrical signal converted from the received audio data to the speaker 661, and the speaker 661 converts it into a sound signal for output; on the other hand, the microphone 662 converts the collected sound signal into an electrical signal, which is received by the audio circuit 660 and then converted into audio data. After the audio data is output to the processor 680 for processing, it is sent through the RF circuit 610 to, for example, another smart phone, or the audio data is output to the memory 620 for further processing.

[0346] WiFi belongs to short - range wireless transmission technology. The smart phone can help users send and receive emails, browse the web, and access streaming media through the WiFi module 670, which provides users with wireless broadband Internet access. Although Figure 15 the WiFi module 670 is shown, it can be understood that it does not belong to the essential components of the smart phone and can be omitted completely within the scope of not changing the essence of the invention according to needs.

[0347] The processor 680 is the control center of the smart phone, connecting various parts of the entire smart phone through various interfaces and lines. By running or executing software programs and / or modules stored in the memory 620, and by calling the data stored in the memory 620, it executes various functions of the smart phone and processes data. Optionally, the processor 680 may include one or more processing units; optionally, the processor 680 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above - mentioned modem processor may not be integrated into the processor 680 either.

[0348] The smart phone further includes a power supply 690 (such as a battery) for powering each component. Optionally, the power supply can be logically connected to the processor 680 through a power management system, so as to manage functions such as charging, discharging, and power consumption management through the power management system.

[0349] Although not shown, the smart phone may further include a camera, a Bluetooth module, etc., which will not be elaborated herein.

[0350] In the embodiment of the present application, the processor 680 included in the terminal device further has the following functions:

[0351] Obtain a pair of samples to be trained, where the pair of samples to be trained includes a text to be trained and an audio to be trained with a corresponding relationship, the text to be trained corresponds to a labeled phoneme sequence, and the audio to be trained corresponds to a real Mel spectrogram;

[0352] Based on the text to be trained, obtain a first Mel spectrogram through a speech synthesis model;

[0353] Based on the first Mel spectrogram, obtain a first phoneme sequence through a speech recognition model;

[0354] Update the model parameters of the speech synthesis model according to the loss value between the first Mel spectrogram and the real Mel spectrogram, and the loss value between the first phoneme sequence and the labeled phoneme sequence.

[0355] In the embodiment of the present application, the processor 680 included in the terminal device further has the following function: obtain a text to be synthesized;

[0356] Based on the text to be synthesized, obtain a target Mel spectrogram through a speech synthesis model;

[0357] Generate a target synthesized speech according to the target Mel spectrogram.

[0358] The steps performed by the terminal device in the above embodiments can be based on the Figure 15 shown terminal device structure.

[0359] In the embodiment of the present application, a computer-readable storage medium is further provided. A computer program is stored in the computer-readable storage medium. When it runs on a computer, it causes the computer to execute the methods described in the foregoing various embodiments.

[0360] In the embodiment of the present application, a computer program product including a program is further provided. When it runs on a computer, it causes the computer to execute the methods described in the foregoing various embodiments.

[0361] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0362] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0363] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0364] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0365] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0366] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A training method for a speech synthesis model, characterized in that, Including: Obtain pairs of samples to be trained, where each pair of samples to be trained includes a text to be trained and an audio to be trained with a corresponding relationship. The text to be trained corresponds to an annotated phoneme sequence, and the audio to be trained corresponds to a true Mel spectrogram; Based on the text to be trained, obtain a first Mel spectrogram through a speech synthesis model; Based on the first Mel spectrogram, obtain a first phoneme sequence through a speech recognition model; Determine a mean squared error loss value according to the first Mel spectrogram and the true Mel spectrogram; Determine a first cross-entropy loss value according to the first phoneme sequence and the annotated phoneme sequence; Determine a first target loss value according to the mean squared error loss value and the first cross-entropy loss value; Update the model parameters of the speech synthesis model according to the first target loss value.

2. The training method according to claim 1, wherein The audio to be trained is from a first object, and the first object corresponds to a first identity identifier; The step of obtaining a first Mel spectrogram through a speech synthesis model based on the text to be trained includes: Based on the text to be trained and the first identity identifier, obtain the first Mel spectrogram through the speech synthesis model.

3. The training method according to claim 1, wherein The step of determining a mean squared error loss value according to the first Mel spectrogram and the true Mel spectrogram includes: Obtain an M-frame predicted frequency amplitude vector corresponding to the first Mel spectrogram, where each frame of the predicted frequency amplitude vector in the M-frame predicted frequency amplitude vector corresponds to a frame of audio signal in the audio to be trained, and M is an integer greater than or equal to 1; Obtain an M-frame annotated frequency amplitude vector corresponding to the true Mel spectrogram, where each frame of the annotated frequency amplitude vector in the M-frame annotated frequency amplitude vector corresponds to a frame of audio signal in the audio to be trained; Determine an average predicted frequency amplitude according to the M-frame predicted frequency amplitude vector; Determine an average annotated frequency amplitude according to the M-frame annotated frequency amplitude vector; Determine an M-frame frequency amplitude difference according to the average predicted frequency amplitude and the average annotated frequency amplitude; Perform an averaging process on the M-frame frequency amplitude difference to obtain the mean squared error loss value.

4. The training method according to claim 1, wherein The step of determining a first cross-entropy loss value according to the first phoneme sequence and the annotated phoneme sequence includes: Obtain an M-frame predicted phoneme vector corresponding to the first phoneme sequence, where each frame of the predicted phoneme vector in the M-frame predicted phoneme vector corresponds to a frame of audio signal in the audio to be trained, and M is an integer greater than or equal to 1; Obtain an M-frame annotated phoneme vector corresponding to the annotated phoneme sequence, where each frame of the annotated phoneme vector in the M-frame annotated phoneme vector corresponds to a frame of audio signal in the audio to be trained; Determine a cross-entropy loss value of M-frame phonemes according to the M-frame predicted phoneme vector and the M-frame annotated phoneme vector; Perform an averaging process on the cross-entropy loss value of the M-frame phonemes to obtain the first cross-entropy loss value.

5. The training method according to any one of claims 1 to 4, characterized in that, After updating the model parameters of the speech synthesis model according to the first target loss value, the method further includes: Obtain the text to be tested and the second identity identifier corresponding to the text to be tested, where the second identity identifier corresponds to a second object; Based on the text to be tested, obtain a second Mel spectrum through the speech synthesis model; Based on the second Mel spectrum, obtain a predicted identity identifier through an object recognition model; Based on the second Mel spectrum, obtain a second phoneme sequence through the speech recognition model; Based on the text to be tested, obtain a weight matrix through the speech synthesis model; Determine a target phoneme sequence according to the weight matrix; Update the model parameters of the speech synthesis model according to the loss value between the second identity identifier and the predicted identity identifier, and the loss value between the second phoneme sequence and the target phoneme sequence.

6. The training method according to claim 5, wherein The updating the model parameters of the speech synthesis model according to the loss value between the second identity identifier and the predicted identity identifier, and the loss value between the second phoneme sequence and the target phoneme sequence includes: Determine a second cross-entropy loss value according to the second identity identifier and the predicted identity identifier; Determine a third cross-entropy loss value according to the second phoneme sequence and the target phoneme sequence; Determine a second target loss value according to the second cross-entropy loss value and the third cross-entropy loss value; Update the model parameters of the speech synthesis model according to the second target loss value.

7. The training method according to claim 6, characterized in that, The determining a second cross-entropy loss value according to the second identity identifier and the predicted identity identifier includes: Obtain the labeled identity vector corresponding to the second identity identifier; Obtain the predicted identity vector corresponding to the predicted identity identifier; Determine a second cross-entropy loss value according to the labeled identity vector and the predicted identity vector.

8. The training method according to claim 6, wherein The determining a third cross-entropy loss value according to the second phoneme sequence and the target phoneme sequence includes: Obtain N-frame predicted phoneme vectors corresponding to the second phoneme sequence, where each predicted phoneme vector in the N-frame predicted phoneme vectors corresponds to a frame of audio signal, and N is an integer greater than or equal to 1; Obtain N-frame phoneme vectors corresponding to the target phoneme sequence, where each phoneme vector in the N-frame phoneme vectors corresponds to a frame of audio signal; Determine the cross-entropy loss value of N-frame phonemes according to the N-frame predicted phoneme vectors and the N-frame phoneme vectors; Perform an averaging process on the cross-entropy loss value of the N-frame phonemes to obtain the third cross-entropy loss value.

9. A method for speech synthesis, characterized in that, Includes: Obtain the text to be synthesized; Based on the text to be synthesized, obtain a target Mel spectrum through a speech synthesis model, where the speech synthesis model is trained according to the training method described in any one of claims 1 to 8; Generate a target synthesized speech according to the target Mel spectrum.

10. The method according to claim 9, characterized in that The method further includes: Obtain a target identity identifier; The obtaining a target Mel spectrum based on the text to be synthesized through a speech synthesis model includes: Based on the text to be synthesized and the target identity identifier, obtain the target Mel spectrum through the speech synthesis model.

11. A voice synthesis model training device, characterized in that, Includes: An acquisition module, configured to acquire pairs of samples to be trained, where each pair of samples to be trained includes a text to be trained and an audio to be trained with a corresponding relationship, the text to be trained corresponds to an annotated phoneme sequence, and the audio to be trained corresponds to a true Mel spectrogram; The acquisition module is further configured to obtain a first Mel spectrogram based on the text to be trained through a speech synthesis model; The acquisition module is further configured to obtain a first phoneme sequence based on the first Mel spectrogram through a speech recognition model; A training module, configured to determine a mean square error loss value according to the first Mel spectrogram and the true Mel spectrogram; determine a first cross-entropy loss value according to the first phoneme sequence and the annotated phoneme sequence; determine a first target loss value according to the mean square error loss value and the first cross-entropy loss value; and update the model parameters of the speech synthesis model according to the first target loss value.

12. The training device according to claim 11, wherein The audio to be trained is from a first object, and the first object corresponds to a first identity identifier; Specifically, the acquisition module is configured to obtain the first Mel spectrogram based on the text to be trained and the first identity identifier through the speech synthesis model.

13. The training device according to claim 11, characterized in that, Specifically, the training module is configured to: Obtain an M-frame predicted frequency amplitude vector corresponding to the first Mel spectrogram, where each frame of the predicted frequency amplitude vector in the M-frame predicted frequency amplitude vector corresponds to a frame of audio signal in the audio to be trained, and M is an integer greater than or equal to 1; Obtain an M-frame annotated frequency amplitude vector corresponding to the true Mel spectrogram, where each frame of the annotated frequency amplitude vector in the M-frame annotated frequency amplitude vector corresponds to a frame of audio signal in the audio to be trained; Determine a predicted frequency amplitude average value according to the M-frame predicted frequency amplitude vector; Determine an annotated frequency amplitude average value according to the M-frame annotated frequency amplitude vector; Determine an M-frame frequency amplitude difference according to the predicted frequency amplitude average value and the annotated frequency amplitude average value; Perform an averaging process on the M-frame frequency amplitude difference to obtain the mean square error loss value.

14. The training device according to claim 11, wherein Specifically, the training module is configured to: Obtain an M-frame predicted phoneme vector corresponding to the first phoneme sequence, where each frame of the predicted phoneme vector in the M-frame predicted phoneme vector corresponds to a frame of audio signal in the audio to be trained, and M is an integer greater than or equal to 1; Obtain an M-frame annotated phoneme vector corresponding to the annotated phoneme sequence, where each frame of the annotated phoneme vector in the M-frame annotated phoneme vector corresponds to a frame of audio signal in the audio to be trained; Determine a cross-entropy loss value of M-frame phonemes according to the M-frame predicted phoneme vector and the M-frame annotated phoneme vector; Perform an averaging process on the cross-entropy loss value of the M-frame phonemes to obtain the first cross-entropy loss value.

15. The training device according to any one of claims 11 to 14, characterized in that The apparatus further includes: a determination module; After the training module updates the model parameters of the speech synthesis model according to the first target loss value, the acquisition module is further configured to acquire a text to be tested and a second identity identifier corresponding to the text to be tested, where the second identity identifier corresponds to a second object; The obtaining module is further configured to obtain a second Mel spectrogram based on the text to be tested through the speech synthesis model; The obtaining module is further configured to obtain a predicted identity identifier based on the second Mel spectrogram through an object recognition model; The obtaining module is further configured to obtain a second phoneme sequence based on the second Mel spectrogram through the speech recognition model; The obtaining module is further configured to obtain a weight matrix based on the text to be tested through the speech synthesis model; The determining module is configured to determine a target phoneme sequence according to the weight matrix; The training module is further configured to update the model parameters of the speech synthesis model according to the loss value between the second identity identifier and the predicted identity identifier, and the loss value between the second phoneme sequence and the target phoneme sequence.

16. The training device according to claim 15, characterized in that, Specifically, the training module is configured to: Determine a second cross-entropy loss value according to the second identity identifier and the predicted identity identifier; Determine a third cross-entropy loss value according to the second phoneme sequence and the target phoneme sequence; Determine a second target loss value according to the second cross-entropy loss value and the third cross-entropy loss value; Update the model parameters of the speech synthesis model according to the second target loss value.

17. The training device according to claim 16, characterized in that, Specifically, the training module is configured to: Obtain a labeled identity vector corresponding to the second identity identifier; Obtain a predicted identity vector corresponding to the predicted identity identifier; Determine a second cross-entropy loss value according to the labeled identity vector and the predicted identity vector.

18. The training device according to claim 16, characterized in that Specifically, the training module is configured to: Obtain N-frame predicted phoneme vectors corresponding to the second phoneme sequence, where each frame of the N-frame predicted phoneme vectors corresponds to a frame of audio signal, and N is an integer greater than or equal to 1; Obtain N-frame phoneme vectors corresponding to the target phoneme sequence, where each frame of the N-frame phoneme vectors corresponds to a frame of audio signal; Determine the cross-entropy loss value of the N-frame phonemes according to the N-frame predicted phoneme vectors and the N-frame phoneme vectors; Perform an averaging process on the cross-entropy loss value of the N-frame phonemes to obtain the third cross-entropy loss value.

19. A voice synthesis device, characterized in that, It includes: An obtaining module, configured to obtain a text to be synthesized; The obtaining module is further configured to obtain a target Mel spectrogram based on the text to be synthesized through a speech synthesis model, where the speech synthesis model is trained according to the training method described in any one of claims 1 to 8; A generating module, configured to generate a target synthesized speech according to the target Mel spectrogram.

20. The apparatus according to claim 19, wherein The obtaining module is further configured to obtain a target identity identifier; Specifically, the obtaining module is configured to obtain the target Mel spectrogram based on the text to be synthesized and the target identity identifier through the speech synthesis model.

21. A computer device, characterized in that, It includes: A memory, a processor, and a bus system; Wherein, the memory is used to store programs; The processor is configured to execute the program in the memory, and the processor is configured to execute the training method according to any one of claims 1 to 8, or execute the method according to claim 9 or 10, based on the instructions in the program code. The bus system is configured to connect the memory and the processor to enable communication between the memory and the processor.

22. A computer-readable storage medium, characterized in that, It includes instructions that, when running on a computer, cause the computer to execute the training method according to any one of claims 1 to 8, or execute the method according to claim 9 or 10.

23. A computer program product, characterized in that, The computer program product includes computer instructions stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to execute the training method according to any one of claims 1 to 8, or execute the method according to claim 9 or 10.

Citation Information

Patent Citations

  • Acoustic model training method and device, speech recognition method and device, equipment and medium

    CN107680582A

  • Voice synthesis model training method, and voice synthesis method and device

    CN110288972A