Method for generating speech synthesis model based on prosodic boundary information and VAE structure

Through the training data set, the speech synthesis network, the pronunciation coding network and the pronunciation prediction network are optimized, and the pronunciation synthesis model based on pronunciation boundary information and VAE structure is generated, which solves the problem of insufficient synthesis of speech accuracy and intonation changes in traditional pronunciation synthesis technology, and achieves higher accuracy speech synthesis.

CN119107931BActive Publication Date: 2025-07-22CHINA SOUTHERN POWER GRID ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411037244.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-31
Publication Date
2025-07-22
Estimated Expiration
2044-07-31

AI Technical Summary

Technical Problem

Traditional pronunciation synthesis technology has shortcomings in the accuracy of synthetic speech and the changes in sentence intonation, and cannot reflect accurate pronunciation rhythm, resulting in low accuracy of synthetic speech.

Method used

By obtaining the training data set, including sample speech, pronunciation boundary information, text information and phoneme information, the speech synthesis network, pronunciation coding network and pronunciation prediction network are trained, and the speech synthesis model based on pronunciation boundary information and VAE structure is generated, and the performance of the speech synthesis network, pronunciation coding network and pronunciation prediction network is optimized to improve the accuracy of synthesized speech.

Benefits of technology

It improves the accuracy of synthetic pronunciations, generates synthetic pronunciations with accurate rhythms, and enhances the effect of speech synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119107931B_ABST
    Figure CN119107931B_ABST
Patent Text Reader

Abstract

The present application relates to a method for generating a speech synthesis model based on prosodic boundary information and a VAE structure. The method includes: obtaining a training data set; the training data set includes sample speech and corresponding prosodic boundary information, text information, and phoneme information; training a speech synthesis network according to the sample speech and the corresponding phoneme information to obtain a trained speech synthesis network; training a prosodic encoding network according to the sample speech and the corresponding prosodic boundary information to obtain a trained prosodic encoding network; training a prosodic prediction network according to the sample speech and the corresponding text information, phoneme information, and prosodic boundary information to obtain a trained prosodic prediction network; generating a speech synthesis model based on prosodic boundary information and a VAE structure according to the trained speech synthesis network, the trained prosodic encoding network, and the trained prosodic prediction network. Using this method can generate synthetic speech with accurate prosody and improve the accuracy of the synthetic speech.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technologies, and particularly to a method, apparatus, computer device, computer-readable storage medium, and computer program product for generating a speech synthesis model based on prosodic boundary information and a VAE structure. Background Art

[0002] With the development of information technology and deep learning technology, human-computer interaction has become an important technology for enhancing the user experience in fields such as education, entertainment, and security. Among the numerous human-computer interaction methods, voice, with its intuitive and efficient information interaction advantages, has become the most commonly used human-computer interaction method in the current human-computer interaction field. Voice interaction generally includes voice understanding and voice output, and voice output is realized by speech synthesis technology. Speech synthesis, that is, the technology of text-to-speech (TTS), can convert the text information input into the speech synthesis system into human-intelligible and natural and fluent speech information for output.

[0003] When traditional technologies perform speech synthesis, they mainly use parametric synthesis and waveform splicing. By pre-building a speech database, combining the input text and corresponding rules to select corresponding speech data from the speech database, and splicing the speech data, speech synthesis is thus achieved. However, the accuracy of the synthesized speech obtained by traditional technologies depends on the speech synthesis rules, and the intonation of the synthesized speech sentences is flat and lacks variation, unable to reflect accurate speech prosody. Therefore, traditional technologies have the problem of low accuracy of synthesized speech. Summary of the Invention

[0004] Based on this, in view of the above technical problems, it is necessary to provide a method, apparatus, computer device, computer-readable storage medium, and computer program product for generating a speech synthesis model based on prosodic boundary information and a VAE structure, which can improve the accuracy of synthesized speech.

[0005] In a first aspect, this application provides a method for generating a speech synthesis model based on prosodic boundary information and a VAE structure, including:

[0006] Obtaining a training data set; the training data set includes sample speech and the prosodic boundary information, text information, and phoneme information corresponding to the sample speech;

[0007] Training a speech synthesis network according to the sample speech and the phoneme information corresponding to the sample speech to obtain a trained speech synthesis network;

[0008] Training a prosody encoding network according to the sample speech and the prosodic boundary information corresponding to the sample speech to obtain a trained prosody encoding network;

[0009] Train a prosody prediction network based on the sample speech and the corresponding text information, phoneme information, and prosody boundary information of the sample speech to obtain a trained prosody prediction network;

[0010] Generate the speech synthesis model based on prosody boundary information and VAE structure according to the trained speech synthesis network, the trained prosody encoding network, and the trained prosody prediction network.

[0011] In one embodiment, the training of the prosody prediction network according to the sample speech and the corresponding text information, phoneme information, and prosody boundary information of the sample speech includes:

[0012] Input the phoneme information corresponding to the sample speech into the trained speech synthesis network to obtain the phoneme feature information corresponding to the phoneme information;

[0013] Input the phoneme feature information and the text information corresponding to the sample speech into the prosody prediction network to obtain a prosody prediction result;

[0014] Input the sample speech and the prosody boundary information corresponding to the sample speech into the trained prosody encoding network to obtain prosody feature information;

[0015] Train the prosody prediction network according to the prosody prediction result and the prosody feature information.

[0016] In one embodiment, the training of the prosody prediction network according to the prosody prediction result and the prosody feature information includes:

[0017] Screen out candidate prosody feature information from the prosody feature information;

[0018] Train the prosody prediction network according to the prosody prediction result and the candidate prosody feature information.

[0019] In one embodiment, the training of the speech synthesis network according to the sample speech and the phoneme information corresponding to the sample speech includes:

[0020] Input the sample speech and the phoneme information corresponding to the sample speech into the speech synthesis network to obtain the latent variable feature information corresponding to the sample speech;

[0021] Generate a speech synthesis result corresponding to the sample speech through the speech synthesis network according to the latent variable feature information;

[0022] Train the speech synthesis network according to the speech synthesis result and the sample speech.

[0023] In one embodiment, training the prosody encoding network according to the sample speech and the prosody boundary information corresponding to the sample speech includes:

[0024] Input the phoneme information corresponding to the sample speech into the speech synthesis network to obtain the speech synthesis result corresponding to the sample speech;

[0025] Input the speech synthesis result into the prosody encoding network to obtain the prosody feature information corresponding to the speech synthesis result;

[0026] Train the prosody encoding network according to the prosody feature information corresponding to the speech synthesis result and the prosody boundary information corresponding to the sample speech.

[0027] In a second aspect, the present application also provides a speech generation method for a speech synthesis model based on prosody boundary information and a VAE structure, including:

[0028] Input the text information corresponding to the speech to be generated into the prosody prediction network in the pre-trained speech synthesis model based on prosody boundary information and a VAE structure to obtain the prosody prediction result corresponding to the speech to be generated;

[0029] Input the phoneme information corresponding to the speech to be generated and the prosody prediction result corresponding to the speech to be generated into the speech synthesis network in the pre-trained speech synthesis model based on prosody boundary information and a VAE structure to generate the speech to be generated.

[0030] In a third aspect, the present application also provides a speech synthesis model generation device based on prosody boundary information and a VAE structure, including:

[0031] A training data acquisition module for acquiring a training data set; the training data set includes a sample speech and the prosody boundary information, text information, and phoneme information corresponding to the sample speech;

[0032] A synthesis network training module for training a speech synthesis network according to the sample speech and the phoneme information corresponding to the sample speech to obtain a trained speech synthesis network;

[0033] An encoding network training module for training a prosody encoding network according to the sample speech and the prosody boundary information corresponding to the sample speech to obtain a trained prosody encoding network;

[0034] A prediction network training module for training a prosody prediction network according to the sample speech and the text information, phoneme information, and prosody boundary information corresponding to the sample speech to obtain a trained prosody prediction network;

[0035] A model generation module, configured to generate the speech synthesis model based on prosody boundary information and a VAE structure according to the trained speech synthesis network, the trained prosody encoding network, and the trained prosody prediction network.

[0036] In a fourth aspect, the present application further provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the computer program is executed by the processor, the steps of the above method are implemented.

[0037] In a fifth aspect, the present application further provides a computer-readable storage medium. On the computer-readable storage medium, a computer program is stored, and when the computer program is executed by the processor, the steps of the above method are implemented.

[0038] In a sixth aspect, the present application further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by the processor, the steps of the above method are implemented.

[0039] The above method, device, computer device, computer-readable storage medium, and computer program product for generating a speech synthesis model based on prosodic boundary information and VAE structure obtain a training data set; the training data set includes sample speech and the corresponding prosodic boundary information, text information, and phoneme information of the sample speech, and train a speech synthesis network based on the sample speech and the corresponding phoneme information of the sample speech to obtain a trained speech synthesis network, so that an accurate speech synthesis network can be trained using the sample speech in the training data set and the corresponding phoneme information of the sample speech, optimizing the performance of the speech synthesis network; train a prosody encoding network based on the sample speech and the corresponding prosodic boundary information of the sample speech to obtain a trained prosody encoding network, so that an accurate prosody encoding network can be trained using the sample speech in the training data set and the corresponding prosodic boundary information of the sample speech, optimizing the performance of the prosody encoding network, improving the accuracy of the prosody information generated by the prosody encoding network, and facilitating the speech synthesis network to generate prosodic synthetic speech by combining the prosody information generated by the prosody encoding network; train a prosody prediction network based on the sample speech and the corresponding text information, phoneme information, and prosodic boundary information of the sample speech to obtain a trained prosody prediction network, so that an accurate prosody prediction network can be trained using the sample speech in the training data set and the corresponding text information, phoneme information, and prosodic boundary information of the sample speech, optimizing the performance of the prosody prediction network, improving the accuracy of the prosody prediction results generated by the prosody prediction network, and facilitating the prosody encoding network to generate accurate prosody information by combining the prosody prediction results generated by the prosody prediction network; generate a speech synthesis model based on prosodic boundary information and VAE structure according to the trained speech synthesis network, trained prosody encoding network, and trained prosody prediction network, so that the prosody encoding network can generate accurate prosody information by combining the prediction results of the prosody prediction network, and the speech synthesis network can generate prosodic synthetic speech by combining the prosody information generated by the prosody encoding network, thereby improving the accuracy of the synthetic speech. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0041] Figure 1 It is an application environment diagram of a method for generating a speech synthesis model based on prosodic boundary information and VAE structure in an embodiment;

[0042] Figure 2 Schematic flow chart of a method for generating a speech synthesis model based on prosodic boundary information and VAE structure in an embodiment;

[0043] Figure 3 Schematic structural diagram of a speech synthesis network in an embodiment;

[0044] Figure 4 Schematic structural diagram of a feedforward neural network in an embodiment;

[0045] Figure 5 Schematic diagram of training a speech synthesis network in an embodiment;

[0046] Figure 6 Schematic structural diagram of a speech synthesis model based on prosodic boundary information and VAE structure in an embodiment;

[0047] Figure 7 Schematic structural diagram of a prosody encoding network in an embodiment;

[0048] Figure 8 Schematic structural diagram of a prosody prediction network in an embodiment;

[0049] Figure 9 Schematic flow chart of a speech synthesis method of a speech synthesis model based on prosodic boundary information and VAE structure in an embodiment;

[0050] Figure 10 Schematic block diagram of a device for generating a speech synthesis model based on prosodic boundary information and VAE structure in an embodiment;

[0051] Figure 11 Internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0052] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0053] The method for generating a speech synthesis model based on prosodic boundary information and VAE structure provided by the embodiments of the present application can be applied to, for example Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or placed on the cloud or other network servers. The server 104 obtains a training data set; the training data set includes sample voices and the corresponding prosody boundary information, text information, and phoneme information of the sample voices; the server 104 trains a speech synthesis network according to the sample voices and the corresponding phoneme information of the sample voices to obtain a trained speech synthesis network; the server 104 trains a prosody encoding network according to the sample voices and the corresponding prosody boundary information of the sample voices to obtain a trained prosody encoding network; the server 104 trains a prosody prediction network according to the sample voices and the corresponding text information, phoneme information, and prosody boundary information of the sample voices to obtain a trained prosody prediction network; the server 104 generates a speech synthesis model based on prosody boundary information and VAE structure according to the trained speech synthesis network, the trained prosody encoding network, and the trained prosody prediction network. Among them, the terminal 102 can be but is not limited to various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The head-mounted devices can be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc. The server 104 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0054] In an exemplary embodiment, as Figure 2 shown, a method for generating a speech synthesis model based on prosody boundary information and VAE structure is provided. Taking this method applied to a server as an example, it includes the following steps S202 to step S210. Among them:

[0055] Step S202, obtain a training data set; the training data set includes sample voices and the corresponding prosody boundary information, text information, and phoneme information of the sample voices.

[0056] Among them, the training data set can refer to the set composed of the data for training the speech synthesis network, prosody encoding network, and prosody prediction network in the speech synthesis model based on prosody boundary information and VAE structure. In practical applications, the training data set can include, but is not limited to, sample speech, prosody boundary information corresponding to the sample speech, text information corresponding to the sample speech, and phoneme information corresponding to the sample speech. Among them, the prosody boundary information can refer to the information describing the prosody of the sample speech. In specific implementations, the prosody boundary information can include prosody boundary label information. The text information can refer to the information describing the content of the sample speech, and the text information can include a text sequence. The phoneme information can refer to the information describing the phonemes corresponding to the sample speech, and the phoneme information can include a phoneme sequence.

[0057] As an example, in order to train the speech synthesis model based on prosody boundary information and VAE structure, the server can pre-build a sample database and, when it is necessary to train the speech synthesis model based on prosody boundary information and VAE structure, obtain a number of data from the sample database as the training data set. The training data set includes sample speech and the prosody boundary information, text information, and phoneme information corresponding to the sample speech for training the speech synthesis model.

[0058] Step S204: Train the speech synthesis network based on the sample speech and the phoneme information corresponding to the sample speech to obtain the trained speech synthesis network.

[0059] Among them, the speech synthesis network can refer to a model that generates corresponding speech using text and / or phonemes. In practical applications, the speech synthesis network can include a non-autoregressive end-to-end speech synthesis acoustic model (such as the FastSpeech model). The speech output by the speech synthesis network can be used as the synthesized speech.

[0060] As an example, the server inputs the sample speech and the phoneme information corresponding to the sample speech into the speech synthesis network to be trained. The speech synthesis network can generate corresponding speech as the output result using the phoneme information. The server can adjust the network parameters of the speech synthesis network based on the difference between the output result of the speech synthesis network and the sample speech until the difference between the output result of the speech synthesis network and the sample speech meets the preset speech requirements. The server takes the speech synthesis network at this time as the trained speech synthesis network.

[0061] Step S206: Train the prosody encoding network based on the sample speech and the prosody boundary information corresponding to the sample speech to obtain the trained prosody encoding network.

[0062] Among them, the prosody encoding network may refer to a model used to obtain the prosody boundary information of the output result of the speech synthesis network and / or the sample speech. In practical applications, the prosody boundary information output by the prosody encoding network can be used as prosody feature information.

[0063] As an example, the server inputs the sample speech and the prosody boundary information corresponding to the sample speech into the prosody encoding network to be trained. The prosody encoding network can analyze the prosody features of the sample speech and generate prosody feature information. The server can adjust the network parameters of the prosody encoding network based on the difference between the prosody feature information output by the prosody encoding network and the prosody boundary information corresponding to the sample speech until the difference between the prosody feature information output by the prosody encoding network and the prosody boundary information corresponding to the sample speech meets the preset prosody difference requirement. The server then uses the prosody encoding network at this time as the trained prosody encoding network.

[0064] Step S208: Train the prosody prediction network according to the sample speech, the text information, phoneme information, and prosody boundary information corresponding to the sample speech to obtain the trained prosody prediction network.

[0065] Among them, the prosody prediction network may refer to a model used to generate the prosody boundary information of the sample speech by using the text information and phoneme information corresponding to the sample speech. In practical applications, the prosody prediction result output by the prosody prediction network can be used as the prosody prediction result.

[0066] As an example, the server inputs the text information and phoneme information corresponding to the sample speech into the prosody prediction network to be trained. The prosody prediction network can analyze the text information and phoneme information corresponding to the sample speech and generate a prosody prediction result. The server can adjust the network parameters of the prosody prediction network according to the difference between the prosody prediction result and the prosody boundary information corresponding to the sample speech until the difference between the prosody prediction result and the prosody boundary information corresponding to the sample speech meets the preset prediction difference requirement. The server then uses the prosody prediction network at this time as the trained prosody prediction network. The server can also, after the prosody prediction network generates a prosody prediction result, adjust the network parameters of the prosody prediction network according to the prosody feature information output by the prosody encoding network, the prosody prediction result generated by the prosody prediction network, and the prosody boundary information corresponding to the sample speech until the difference between the prosody prediction result and the prosody boundary information corresponding to the sample speech meets the preset prediction difference requirement. The server then uses the prosody prediction network at this time as the trained prosody prediction network.

[0067] Step S210: Generate a speech synthesis model based on prosody boundary information and VAE structure according to the trained speech synthesis network, the trained prosody encoding network, and the trained prosody prediction network.

[0068] Among them, as an example, the server can construct a speech synthesis model based on prosody boundary information and VAE structure according to the trained speech synthesis network, the trained prosody encoding network, and the trained prosody prediction network. The server can also further train and validate the speech synthesis model based on prosody boundary information and VAE structure using the training dataset, so as to optimize the model performance of the speech synthesis model based on prosody boundary information and VAE structure and ensure that the synthesized speech has accurate prosody.

[0069] In the above method for generating a speech synthesis model based on prosody boundary information and VAE structure, by obtaining a training dataset; the training dataset includes sample speech and the corresponding prosody boundary information, text information, and phoneme information of the sample speech, and training a speech synthesis network according to the sample speech and the corresponding phoneme information of the sample speech to obtain a trained speech synthesis network, so that an accurate speech synthesis network can be trained using the sample speech in the training dataset and the corresponding phoneme information of the sample speech, and the performance of the speech synthesis network can be optimized; training a prosody encoding network according to the sample speech and the corresponding prosody boundary information of the sample speech to obtain a trained prosody encoding network, so that an accurate prosody encoding network can be trained using the sample speech in the training dataset and the corresponding prosody boundary information of the sample speech, and the performance of the prosody encoding network can be optimized, and the accuracy of the prosody information generated by the prosody encoding network can be improved, facilitating the speech synthesis network to generate synthesized speech with prosody by combining the prosody information generated by the prosody encoding network; training a prosody prediction network according to the sample speech and the corresponding text information, phoneme information, and prosody boundary information of the sample speech to obtain a trained prosody prediction network, so that an accurate prosody prediction network can be trained using the sample speech in the training dataset and the corresponding text information, phoneme information, and prosody boundary information of the sample speech, and the performance of the prosody prediction network can be optimized, and the accuracy of the prosody prediction results generated by the prosody prediction network can be improved, facilitating the prosody encoding network to generate accurate prosody information by combining the prosody prediction results generated by the prosody prediction network; generating a speech synthesis model based on prosody boundary information and VAE structure according to the trained speech synthesis network, the trained prosody encoding network, and the trained prosody prediction network, so as to train the speech synthesis network, the prosody encoding network, and the prosody prediction network respectively based on the training dataset, enabling the prosody encoding network to generate accurate prosody information by combining the prediction results of the prosody prediction network, and enabling the speech synthesis network to generate synthesized speech with accurate prosody by combining the prosody information generated by the prosody encoding network, thereby improving the accuracy of the synthesized speech.

[0070] In an exemplary embodiment, a prosody prediction network is trained based on a sample speech, as well as text information, phoneme information, and prosody boundary information corresponding to the sample speech, including: inputting the phoneme information corresponding to the sample speech into a trained speech synthesis network to obtain phoneme feature information corresponding to the phoneme information; inputting the phoneme feature information and the text information corresponding to the sample speech into the prosody prediction network to obtain a prosody prediction result; inputting the sample speech and the prosody boundary information corresponding to the sample speech into a trained prosody encoding network to obtain prosody feature information; and training the prosody prediction network according to the prosody prediction result and the prosody feature information.

[0071] Among them, the phoneme feature information may refer to the information obtained by encoding the phoneme information of the sample speech.

[0072] Among them, the prosody prediction result may refer to the information representing prosody obtained after the prosody prediction network analyzes the phoneme feature information of the phoneme information corresponding to the sample speech and the text information corresponding to the sample speech.

[0073] Among them, the prosody feature information may refer to the information representing prosody obtained after the prosody encoding network analyzes the sample speech and the prosody boundary information corresponding to the sample speech.

[0074] As an example, the server inputs the phoneme information corresponding to the sample speech into a trained speech synthesis network. The trained speech synthesis network can encode the phoneme information to obtain phoneme feature information corresponding to the phoneme information. Then, the server can input the phoneme feature information corresponding to the phoneme information and the text information corresponding to the sample speech into the prosody prediction network. The prosody prediction network can analyze the phoneme feature information and the text information to generate a prosody prediction result. The server can input the sample speech and the prosody boundary information corresponding to the sample speech into a trained prosody encoding network. The trained prosody encoding network can analyze the sample speech and the prosody boundary information corresponding to the sample speech to generate prosody feature information. The server can use the prosody feature information and / or the prosody boundary information corresponding to the sample speech as a reference to analyze the difference between the prosody feature information and / or the prosody boundary information corresponding to the sample speech and the prosody prediction result output by the prosody prediction network, and adjust the prosody prediction network until the difference between the prosody prediction result and the prosody feature information and / or the prosody boundary information corresponding to the sample speech meets the preset prediction difference requirement. The server uses the prosody prediction network at this time as the trained prosody prediction network.

[0075] In this embodiment, by inputting the phoneme information corresponding to the sample speech into the trained speech synthesis network, the phoneme feature information corresponding to the phoneme information is obtained; the phoneme feature information and the text information corresponding to the sample speech are input into the prosody prediction network to obtain a prosody prediction result; the sample speech and the prosody boundary information corresponding to the sample speech are input into the trained prosody encoding network to obtain prosody feature information; according to the prosody prediction result and the prosody feature information, the prosody prediction network is trained, and it is possible to generate a prosody prediction result based on the phoneme feature information generated by the trained speech synthesis network and the text information corresponding to the sample speech, and combine the prosody feature information output by the trained prosody encoding network to train the prosody prediction network, optimize the performance of the prosody prediction network, improve the accuracy of the prosody prediction result output by the prosody prediction network, provide a data basis for the speech synthesis network to generate a synthesized speech with accurate prosody by combining the prosody prediction result output by the prosody prediction network, and further improve the accuracy of the synthesized speech.

[0076] In some embodiments, training the prosody prediction network according to the prosody prediction result and the prosody feature information includes: screening out candidate prosody feature information from the prosody feature information; training the prosody prediction network according to the prosody prediction result and the candidate prosody feature information.

[0077] Among them, the candidate prosody feature information may refer to the information screened out from the prosody feature information according to a preset condition. In practical applications, the candidate prosody feature information may be the information screened out from the prosody feature information according to a preset ratio.

[0078] As an example, during the process of training the prosody prediction network, the server may not use all the prosody feature information to train the prosody prediction network, screen out candidate prosody feature information from the prosody feature information according to a preset ratio, and use the prosody prediction result and the candidate prosody feature information to train the prosody prediction network, so as to reduce the difference between the training process and the application process / inference process, and avoid the problem of mismatch between training and inference caused by the gap between the prosody feature information predicted in the inference stage and the true counterpart (such as prosody boundary information).

[0079] In this embodiment, by screening out candidate prosody feature information from the prosody feature information; training the prosody prediction network according to the prosody prediction result and the candidate prosody feature information, it is possible to avoid the prosody prediction network relying on all the prosody feature information for training, thereby reducing the difference between the training process and the inference process of the prosody prediction network, and further improving the accuracy of the prosody prediction result output by the prosody prediction network, providing a data basis for the speech synthesis network to generate a synthesized speech with accurate prosody by combining the prosody prediction result output by the prosody prediction network, and improving the accuracy of the synthesized speech.

[0080] In some embodiments, training a speech synthesis network based on a sample speech and the phoneme information corresponding to the sample speech includes: inputting the sample speech and the phoneme information corresponding to the sample speech into the speech synthesis network to obtain the latent variable feature information corresponding to the sample speech; generating, by the speech synthesis network, a speech synthesis result corresponding to the sample speech according to the latent variable feature information; and training the speech synthesis network based on the speech synthesis result and the sample speech.

[0081] Among them, the latent variable feature information may refer to information characterizing the features of the sample speech in a preset space (such as a latent space). In practical applications, the latent variable feature information may include latent variables.

[0082] Among them, the speech synthesis result may refer to the synthesized speech generated by the speech synthesis network based on the phoneme information corresponding to the sample speech.

[0083] As an example, the server inputs the sample speech and the phoneme information corresponding to the sample speech into the speech synthesis network. The speech synthesis network can encode the phoneme information corresponding to the sample speech to obtain phoneme feature information, and convert the phoneme feature information into a preset space to obtain the latent variable feature information of the phoneme information corresponding to the sample speech. The speech synthesis network can decode the latent variable feature information to obtain the speech synthesis result. The server can analyze the difference between the speech synthesis result and the sample speech, and adjust the network parameters of the speech synthesis network until the difference between the speech synthesis result and the sample speech meets the preset speech requirements. The server uses the speech synthesis network at this time as the trained speech synthesis network.

[0084] In this embodiment, by inputting the sample speech and the phoneme information corresponding to the sample speech into the speech synthesis network to obtain the latent variable feature information corresponding to the sample speech; generating, by the speech synthesis network, a speech synthesis result corresponding to the sample speech according to the latent variable feature information; and training the speech synthesis network based on the speech synthesis result and the sample speech, it is possible to convert the phoneme information corresponding to the sample speech into latent variable feature information, and train the speech synthesis network in combination with the difference between the speech synthesis result generated by the speech synthesis network based on the latent variable feature information and the sample speech, thereby improving the accuracy of the synthesized speech.

[0085] In some embodiments, training a prosody encoding network based on a sample speech and the prosody boundary information corresponding to the sample speech includes: inputting the phoneme information corresponding to the sample speech into the speech synthesis network to obtain a speech synthesis result corresponding to the sample speech; inputting the speech synthesis result into the prosody encoding network to obtain prosody feature information corresponding to the speech synthesis result; and training the prosody encoding network based on the prosody feature information corresponding to the speech synthesis result and the prosody boundary information corresponding to the sample speech.

[0086] As an example, the server inputs the phoneme information corresponding to the sample speech into the speech synthesis network. The speech synthesis network can encode the phoneme information corresponding to the sample speech, convert the encoding result into latent variable feature information, and then generate a speech synthesis result corresponding to the phoneme information of the sample speech using the latent variable feature information. The server inputs the speech synthesis result into the prosody encoding network. The prosody encoding network can analyze the speech synthesis result and generate prosody feature information corresponding to the speech synthesis result. Then, the server can adjust the network parameters of the prosody encoding network according to the difference between the prosody feature information corresponding to the speech synthesis result and the prosody boundary information corresponding to the sample speech until the difference between the prosody feature information corresponding to the speech synthesis result and the prosody boundary information corresponding to the sample speech meets the preset prosody difference requirement. The server uses the prosody encoding network at this time as the trained prosody encoding network.

[0087] In this embodiment, by inputting the phoneme information corresponding to the sample speech into the speech synthesis network, a speech synthesis result corresponding to the sample speech is obtained; by inputting the speech synthesis result into the prosody encoding network, prosody feature information corresponding to the speech synthesis result is obtained; and by training the prosody encoding network according to the prosody feature information corresponding to the speech synthesis result and the prosody boundary information corresponding to the sample speech, it is possible to utilize the difference between the prosody feature information obtained by analyzing the speech synthesis result using the prosody encoding network and the prosody boundary information corresponding to the sample speech to train the prosody encoding network, thereby improving the accuracy of the prosody feature information output by the prosody encoding network, providing a data basis for the subsequent speech synthesis network to generate a synthesized speech by combining the prosody feature information output by the prosody encoding network, and further improving the accuracy of the synthesized speech.

[0088] In some embodiments, a speech generation method for a speech synthesis model based on prosody boundary information and VAE structure is provided, including: inputting the text information corresponding to the speech to be generated into the prosody prediction network in the pre-trained speech synthesis model based on prosody boundary information and VAE structure to obtain a prosody prediction result corresponding to the speech to be generated; and inputting the phoneme information corresponding to the speech to be generated and the prosody prediction result corresponding to the speech to be generated into the speech synthesis network in the pre-trained speech synthesis model based on prosody boundary information and VAE structure to generate the speech to be generated.

[0089] As an example, the server inputs the phoneme information corresponding to the speech to be generated into the speech synthesis network in the pre-trained speech synthesis model based on prosodic boundary information and VAE structure. The speech synthesis network can encode the phoneme information corresponding to the speech to be generated to obtain the phoneme feature information of the speech to be generated. The server inputs the text information corresponding to the speech to be generated and the phoneme feature information of the speech to be generated into the prosody prediction network in the pre-trained speech synthesis model based on prosodic boundary information and VAE structure. The prosody prediction network analyzes the text information corresponding to the speech to be generated and the phoneme feature information of the speech to be generated, and generates a prosody prediction result corresponding to the speech to be generated. The speech synthesis network can use the phoneme feature information of the speech to be generated and the prosody prediction result corresponding to the speech to be generated to generate the latent variable feature information corresponding to the speech to be generated. The speech synthesis network can decode the latent variable feature information corresponding to the speech to be generated to obtain the speech to be generated. Further, the prosody encoding network in the pre-trained speech synthesis model based on prosodic boundary information and VAE structure can analyze the speech to be generated and generate the prosody feature information corresponding to the speech to be generated. The pre-trained speech synthesis model based on prosodic boundary information and VAE structure can perform splicing and other processing on the prosody feature information corresponding to the speech to be generated and the prosody prediction result corresponding to the speech to be generated to obtain the prosody embedding feature information. Then, the speech synthesis network uses the phoneme feature information of the speech to be generated and the prosody embedding feature information to generate the new latent variable feature information corresponding to the speech to be generated. The speech synthesis network can decode the new latent variable feature information to obtain the new speech to be generated.

[0090] In this embodiment, by inputting the text information corresponding to the speech to be generated into the prosody prediction network in the pre-trained speech synthesis model based on prosodic boundary information and VAE structure, a prosody prediction result corresponding to the speech to be generated is obtained; by inputting the phoneme information corresponding to the speech to be generated and the prosody prediction result corresponding to the speech to be generated into the speech synthesis network in the pre-trained speech synthesis model based on prosodic boundary information and VAE structure, the speech to be generated can be generated. The pre-trained speech synthesis model based on prosodic boundary information and VAE structure can combine the text information and phoneme information corresponding to the speech to be generated to generate a synthetic speech with accurate prosody, thereby improving the accuracy of the synthetic speech.

[0091] In some embodiments, the speech synthesis network may include a non-autoregressive acoustic model. The non-autoregressive structure enables the speech synthesis network to have a significant improvement in inference speed compared to traditional acoustic models. Moreover, since the non-autoregressive method of using phoneme-level feature expansion to complete the mapping from the phoneme sequence to the Mel spectrogram sequence length not only accelerates the inference speed of the model, but also reduces the problems of mispronunciation and repeated pronunciation caused by the autoregressive model relying on the attention mechanism to learn the alignment between phonemes and spectrograms, and improves the robustness of the acoustic model. For example, Figure 3As shown, a schematic structural diagram of a voice synthesis network is provided. The voice synthesis network can be a sequence-to-sequence (Seq2Seq) model, adopting an encoder-decoder structure. The basic module of the voice synthesis network is a feed-forward neural module (Feed-Forward Transformer, FFT). The encoder and decoder of the voice synthesis network are composed of several stacked feed-forward neural modules. The voice synthesis network also includes a length regulator and a duration predictor to model information such as phoneme duration. The linear layer in the voice synthesis network is used to output the generated Mel spectrum features, and then generate the synthesized speech based on the Mel spectrum features. The phoneme sequence is processed by the phoneme embedding layer of the voice synthesis network. The output result of the phoneme embedding layer is combined with the position encoding, and after being processed by the feed-forward neural network of the encoder and the length regulator, it enters the decoder. The feed-forward neural network of the decoder combines the position encoding to process the output result of the encoder, and finally generates the Mel spectrum features through the linear output layer, and then generates the synthesized speech.

[0092] In some embodiments, such as Figure 4As shown, a schematic structural diagram of a feedforward neural network is provided. The feedforward neural network is a feedforward structure composed of a multi-head attention network and a one-dimensional convolution. In a speech synthesis network, multiple FFT blocks are stacked to form (amplified to the frame level) an encoder and a decoder. The encoder encodes the input phoneme sequence, and the decoder can decode the output of the encoder, thereby completing the conversion from text information to Mel spectrogram. Among them, the encoder on the phoneme side has N repeated FFT blocks, and the decoder on the Mel spectrogram side also has N FFT blocks. The length regulator is located between the encoder and the decoder to complete the mapping alignment of the length of the phoneme sequence to the Mel spectrogram sequence, that is, to achieve the amplification of phoneme-level features, so that the one-to-many generation mapping can be completed without using an autoregressive method. The duration predictor in the length regulator realizes the prediction of the duration of each input phoneme, providing duration information for the length regulator. The self-attention network extracts input feature information through multi-head attention. The FFT module in the speech synthesis network also includes two one-dimensional convolutional networks with ReLU activation functions outside the attention layer. This convolutional structure can better utilize the neighborhood information of speech. At the same time, residual connections, layer normalization, and dropout layers are added after the self-attention network and the one-dimensional convolutional network. The length regulator can solve the problem of inconsistent lengths between the phoneme sequence and the Mel spectrogram in speech synthesis. In human speech, each phoneme has a corresponding phoneme duration, which is generally between several frames and more than a dozen frames. If the phoneme duration of a phoneme is n, it means that this phoneme corresponds to n frames of Mel spectrogram. Therefore, the speech synthesis network uses the length regulator to directly copy the phoneme-level features n times to the frame level according to the duration of the phoneme. In this way, frame-level features of the same length as the Mel spectrogram can be obtained by performing the same copying operation on each phoneme. At the same time, the rhythm performance of the final synthesized speech can be changed by directly modifying the duration of pauses and silences in the synthesized speech, thereby achieving the purpose of changing the interruption between words, phrases, and sentences.

[0093] In some embodiments, such as Figure 5As shown, a schematic diagram of training a voice synthesis network is provided. The voice synthesis network can have a model structure of a variational autoencoder, which is a probabilistic model based on variational inference. The variational autoencoder and the autoencoder have a similar encoding-decoding structure and belong to an unsupervised generative model. The difference is that the autoencoder is simply for compression encoding, while in variational inference, in addition to known data (such as observed data, training data, etc.), there is also a latent variable z (latent variable feature information / latent variable). The variational autoencoder assumes that the distribution of the dataset x (manifest variable) of the input data is completely controlled by a set of latent variables z, and these latent variables are independent of each other and follow a Gaussian distribution. The variational autoencoder can let the encoder learn the latent variable model of the input data, that is, learn the parameters mean and variance of the Gaussian probability distribution of this set of latent variables, and the latent variable z can be sampled from the normal distribution of this set of distribution parameters, and then the latent variable z is decoded by the decoder to reconstruct the input. Essentially, it realizes a continuous and smooth latent space representation. These latent variables (latent variables) contain useful information about the type of output that the model needs to generate. The probability distribution of the latent variable z is represented by p(z). The variational autoencoder selects the Gaussian distribution as the prior to learn the distribution p(z) so as to sample new samples more easily in the inference stage. The desired goal of the variational autoencoder is to define a generative model containing the latent variable z, and maximize the likelihood of the distribution p(x) of the X variable under the parameters θ obtained during training. At this time, p(x) can be expressed as p θ (x), p θ (x) can be expressed as:

[0094] .

[0095] However, since it is impossible to traverse the latent variable z, p θ (x) is difficult to solve, and the variational autoencoder introduces a new probability distribution q φ (z|x), and tries to obtain the lower bound of the log-likelihood function, and then maximize the lower bound, which is equivalent to approximately maximizing the log-likelihood function. At this time, there is:

[0096] .

[0097] The above formula finally consists of three terms. The first two terms can be calculated, and the third term cannot be calculated. However, according to the properties of the KL divergence, it can be known that the third term must be greater than or equal to 0 (and by fitting with the encoder neural network, it can be ensured as much as possible that this term is 0, so that the first two terms are as close as possible to the target likelihood function). Therefore, it can be obtained:

[0098] .

[0099] The variational autoencoder refers to the right side of the above inequality as a variational lower bound (ELBO). Therefore, the goal of the variational autoencoder is to maximize the variational lower bound, that is, to use the variational lower bound as the loss function of the model.

[0100] As a neural network model, the variational autoencoder uses the backpropagation algorithm to train the network. However, since the sampling operation is non-differentiable, the variational autoencoder introduces the reparameterization trick to solve the problem of parameter conduction in model training. For the distribution N(μ, σ) of the latent variable z, first sample from the standard normal distribution N(0, 1), and then consider the mean and variance parameters learned by the encoder, and let z = μ + σ * τ. In this way, the normal propagation of the gradient can be completed, and thus the training of the variational autoencoder model can be completed.

[0101] In some embodiments, such as Figure 6As shown in the figure, a schematic structural diagram of a speech synthesis model based on prosodic boundary information and VAE structure is provided. The speech synthesis model based on prosodic boundary information and VAE structure is a non-autoregressive model that takes phoneme sequences and text sequences as inputs to predict Mel spectrogram features. The speech synthesis model based on prosodic boundary information and VAE structure includes a speech synthesis network, a prosody encoding network, and a prosody prediction network. Based on the encoder-decoder structure and variance adaptation module of the speech synthesis network, with prosodic boundaries as the time resolution, prosody modeling networks at the prosodic word level and prosodic phrase level can be introduced, as well as a prosody prediction network trained considering the lack of reference speech in the inference stage. Among them, the variance adaptation module conducts supervised training, and the training data includes variance information such as phoneme-level fundamental frequency, phoneme-level pitch, and duration, so as to assist in completing the one-to-many mapping between text and speech and provide the missing speech prosody variation information of the text. During the training process, the phoneme-level encoder converts the phoneme sequence into its hidden representation (such as latent variables), the word-level encoder completes the encoding of the input Chinese character sequence, the variance adaptation module predicts variance information (including phoneme-level fundamental frequency, phoneme-level pitch, duration, etc.), the prosody encoding network provides prosody modeling encoding related to the prosodic boundary hierarchy, and finally the decoder inputs frame-level encoded latent variables with variance information, prosody information (such as multi-layer prosody embedding vectors) and phoneme information to complete the prediction of the frame-level Mel spectrogram. In the inference stage, the prosody encoding network is predicted by the prosody prediction module using phoneme-level encoding information and word-level encoding information as inputs. Prosody modeling is not just for prosody modeling at the prosody word level and prosody phrase level alone, but hierarchical dependence modeling referring to the prosodic boundary structure. In the encoding stage, the model first extracts prosody word-level Mel spectrogram features from the Mel spectrogram, then uses the prosody word-level prosody encoder to obtain the posterior distribution of the data, and obtains the prosody word-level prosody embedding vector (hereinafter referred to as the prosody word embedding vector) through random sampling. Then, referring to the characteristic that the higher-level prosody features in the prosodic boundary structure are established on the basis of the lower level, the prosody phrase-level prosody encoder adds the prosody word embedding vector as additional information input in addition to the Mel spectrogram features at the prosody phrase scale.

[0102] In some embodiments, the prosody modeling network has two main parts: the prosody encoding network and the prosody prediction network, as Figure 7As shown in the figure, a schematic structural diagram of a prosodic encoding network is provided. The (multi-level) prosodic encoding network can be composed of a prosodic word-level prosodic encoder and a prosodic phrase-level prosodic encoder. The prosodic encoder is composed of two one-dimensional convolutional layers and a bidirectional gated recurrent unit (GRU) layer. The mel spectrogram of the speech is used to divide the prosodic word boundaries and prosodic phrase boundaries. The mel spectrograms at the prosodic word level / prosodic phrase level after division pass through the corresponding prosodic encoders respectively. Finally, the amplified and spliced results are output to form the final multi-level prosodic embedding vector at the phoneme level, which is fused with the phoneme encoding vector. The prosodic word and prosodic phrase boundary information is obtained by combining the results of the front-end prosodic prediction with the forced alignment phoneme boundary information. The (multi-level) prosodic prediction network includes a prosodic word-level prosodic predictor and a prosodic phrase-level prosodic predictor, as Figure 8As shown, a schematic structural diagram of a prosody prediction network is provided. In the inference stage, since the model cannot obtain the reference speech corresponding to the required text, the prosody encoder cannot be used to obtain multi-level prosody embeddings. Therefore, it is necessary to use the prosody predictor to complete the prediction of multi-level prosody embeddings from the text information. It is composed of a GRU layer and a one-dimensional convolutional layer. At the same time, referring to the idea of BERT-Tacotron2, a BERT encoder is additionally introduced to provide character-level feature information outside the phoneme sequence, so as to improve the prediction ability of this part. For the prosody word-level prosody predictor, its prediction is based on the phoneme sequence information and the text sequence information. The phoneme-level information input uses the phoneme-level encoder of the main body of the acoustic model, and the structure is the same as that of FastSpeech2, while the text sequence is character-level feature information. Here, a pre-trained BERT model is used as the character-level encoder. After obtaining these two types of features, the scale conversion from the phoneme level and the character level to the prosody word level is completed through average pooling. Then, these two embeddings are concatenated and sent to the predictor. The text information input of the prosody phrase-level encoder is the same as that of the prosody word level, but at the same time, it is consistent with the encoding module. Referring to the characteristic that the high-level prosody features in the prosody boundary structure are based on the low-level ones, an additional prosody word-level prosody embedding vector is added as an additional conditional input. Among them, the prosody embedding vector is sampled from the posterior distribution encoded by the multi-level prosody encoder. At this time, the parameters to be trained include the main body of the acoustic model and the multi-level prosody encoder. For each sample in the training stage, its mel spectrogram, phoneme sequence, and word-level sequence are respectively sent to the encoder modules of the prosody encoder and the synthesizer. Usually, the model obtains the mel spectrogram cut components at the prosody word and prosody phrase time resolutions under the supervision of the phoneme boundary alignment information obtained by Montreal-Forced-Aligner (MFA) and the alignment relationship between the prosody word, prosody phrase, and phoneme. Then, the corresponding posterior distributions are obtained by using the corresponding encoders at these time scales, and then the prosody embedding vectors corresponding to the prosody word and prosody phrase are obtained through random sampling. Then, using the alignment relationship, the alignment is reversely expanded and jointly composes the input of the decoder with the output of the text encoder, and then the decoder is used to complete the frame-level decoding of the mel spectrogram. The training criteria of the model are similar to those of FastSpeech, including the mel spectrogram loss and the phoneme time prediction loss. In addition, the kl divergence loss between the VAE training criterion and the standard normal distribution is added to train the multi-level prosody encoder.

[0103] In some embodiments, such as Figure 9As shown, a flow diagram of a speech synthesis method for a speech synthesis model based on prosodic boundary information and a VAE structure is provided. The server can establish a speech synthesis dataset containing prosodic tags, text and phoneme information, and corresponding speech data. The server can combine a dictionary and use the speech synthesis dataset to train a speech synthesis (acoustic) model based on prosodic boundary information and a VAE structure. Specifically, the server can use the speech synthesis data in the speech synthesis dataset to model phoneme information, text information, and prosodic information related to prosodic boundaries. Then, the speech synthesis network and the (multi-level) prosodic encoding network are first trained. After the speech synthesis network and the (multi-level) prosodic encoding network are trained for a certain stage, as the kl loss of the prosodic encoding network converges, the training of the (multi-level) prosodic prediction network is added, and the result of the prosodic encoding network is used to train the (multi-level) prosodic prediction network. During the training of the prosodic prediction network, to solve the problem of the mismatch between training and inference caused by the gap between the prosodic encoding information predicted in the inference stage and the true counterpart, not all the features encoded by the encoder are completely provided to the decoder during training. Instead, the scheduled sampling mechanism is used to replace the prosodic representations of the encoder with the prosodic representations at the prosodic word level and prosodic phrase level predicted by the prosodic prediction network at a certain ratio to reduce the training and inference gap. The server can combine the scheduled sampling mechanism to make the speech synthesis network converge under the guidance of the multi-level prosodic prediction network, and then obtain a trained speech synthesis model based on prosodic boundary information and a VAE structure. During the training of the speech synthesis (acoustic) model based on prosodic boundary information and a VAE structure, the server can also, by judging the overall model loss and the prosodic listening test of the synthesized speech under the guidance of the prosodic encoder, after the prosodic encoding has been initially stabilized, add the prosodic predictor to the training process and use the L1 loss of the mean and variance of the prediction vector and the mean and variance of the encoded posterior distribution to train the prosodic predictor and stop the gradient flow to ensure that the prosodic prediction error does not affect the parameter learning of the prosodic encoder. For each training speech, the text sequence and phoneme sequence of the training speech are used to predict the mean and variance of the mixture Gaussian distribution of all prosodic words and prosodic phrases after this process, and the prosodic embedding vector is obtained by sampling.When using a speech synthesis (acoustic) model based on prosodic boundary information and VAE structure to generate speech, the server can input text, prosodic boundary information, and phoneme sequence into the speech synthesis (acoustic) model based on prosodic boundary information and VAE structure. The speech synthesis (acoustic) model based on prosodic boundary information and VAE structure can generate corresponding Mel spectrograms, and then generate synthetic speech based on the Mel spectrograms. For example: The server segments the text by character and converts it into the corresponding digital serial number of the text according to the dictionary, segments the phonemes by character, and converts them into the corresponding digital serial number of the phonemes according to the dictionary, processes the prosodic boundary information (prosodic boundary labels) into the matrix format required by the model. The server inputs linguistic features such as text information, phoneme information, and prosodic boundary information into the speech synthesis model based on prosodic boundary information and VAE structure. The speech synthesis model based on prosodic boundary information and VAE structure can output the reconstructed speech Mel spectrogram, and then generate the Mel spectrogram of the speech corresponding to the text, and finally generate the synthetic speech based on the Mel spectrogram.

[0104] In this embodiment, by combining a variational autoencoder and the prior information of prosodic boundaries in the text on the basis of a speech synthesis network to model and generate multi-level prosodic features that conform to the Chinese prosodic hierarchical structure, and by introducing a pre-trained BERT encoder to introduce word-level semantic information outside the phoneme sequence, the ability to predict speech prosody in the model can be improved, thereby generating synthetic speech with accurate prosody and improving the accuracy of the synthetic speech.

[0105] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.

[0106] Based on the same inventive concept, an embodiment of the present application further provides a speech synthesis model generation device based on prosodic boundary information and VAE structure for implementing the speech synthesis model generation method based on prosodic boundary information and VAE structure involved above. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the speech synthesis model generation device based on prosodic boundary information and VAE structure provided below can refer to the limitations on the speech synthesis model generation method based on prosodic boundary information and VAE structure in the above text, and will not be repeated here.

[0107] In an exemplary embodiment, as Figure 10 shown, a speech synthesis model generation device based on prosodic boundary information and VAE structure is provided, including: a training data acquisition module 1002, a synthesis network training module 1004, an encoding network training module 1006, a prediction network training module 1008, and a model generation module 1010, where:

[0108] The training data acquisition module 1002 is configured to acquire a training data set; the training data set includes sample speech and the prosodic boundary information, text information, and phoneme information corresponding to the sample speech.

[0109] The synthesis network training module 1004 is configured to train a speech synthesis network according to the sample speech and the phoneme information corresponding to the sample speech to obtain a trained speech synthesis network.

[0110] The encoding network training module 1006 is configured to train a prosodic encoding network according to the sample speech and the prosodic boundary information corresponding to the sample speech to obtain a trained prosodic encoding network.

[0111] The prediction network training module 1008 is configured to train a prosodic prediction network according to the sample speech and the text information, phoneme information, and prosodic boundary information corresponding to the sample speech to obtain a trained prosodic prediction network.

[0112] The model generation module 1010 is configured to generate the speech synthesis model based on prosodic boundary information and VAE structure according to the trained speech synthesis network, the trained prosodic encoding network, and the trained prosodic prediction network.

[0113] In one exemplary embodiment, the above-mentioned prediction network training module 1008 is further specifically configured to input the phoneme information corresponding to the sample speech into the trained speech synthesis network to obtain the phoneme feature information corresponding to the phoneme information; input the phoneme feature information and the text information corresponding to the sample speech into the prosody prediction network to obtain a prosody prediction result; input the sample speech and the prosody boundary information corresponding to the sample speech into the trained prosody encoding network to obtain prosody feature information; and train the prosody prediction network according to the prosody prediction result and the prosody feature information.

[0114] In one exemplary embodiment, the above-mentioned prediction network training module 1008 is further specifically configured to screen out candidate prosody feature information from the prosody feature information; and train the prosody prediction network according to the prosody prediction result and the candidate prosody feature information. In one exemplary embodiment, the above-mentioned synthesis network training module 1004 is further specifically configured to input the sample speech and the phoneme information corresponding to the sample speech into the speech synthesis network to obtain the latent variable feature information corresponding to the sample speech; generate a speech synthesis result corresponding to the sample speech through the speech synthesis network according to the latent variable feature information; and train the speech synthesis network according to the speech synthesis result and the sample speech.

[0115] In one exemplary embodiment, the above-mentioned encoding network training module 1006 is further specifically configured to input the phoneme information corresponding to the sample speech into the speech synthesis network to obtain a speech synthesis result corresponding to the sample speech; input the speech synthesis result into the prosody encoding network to obtain the prosody feature information corresponding to the speech synthesis result; and train the prosody encoding network according to the prosody feature information corresponding to the speech synthesis result and the prosody boundary information corresponding to the sample speech.

[0116] Each module in the above-mentioned speech synthesis model generation device based on prosody boundary information and VAE structure can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory in the computer device in the form of software, so as to facilitate the processor to call and execute the operations corresponding to each of the above modules.

[0117] In an exemplary embodiment, a speech generation device for a speech synthesis model based on prosody boundary information and VAE structure is provided, including: an input module and a generation module, where:

[0118] An input module, configured to input text information corresponding to the speech to be generated into a prosody prediction network in a pre-trained speech synthesis model based on prosody boundary information and a VAE structure, so as to obtain a prosody prediction result corresponding to the speech to be generated.

[0119] A generation module, configured to input phoneme information corresponding to the speech to be generated and the prosody prediction result corresponding to the speech to be generated into a speech synthesis network in the pre-trained speech synthesis model based on prosody boundary information and a VAE structure, and generate the speech to be generated.

[0120] Each module in the above speech generation device of the speech synthesis model based on prosody boundary information and a VAE structure can be implemented in whole or in part by software, hardware, or a combination thereof. The above-mentioned modules can be embedded in the processor in the computer device in the form of hardware or be independent of the processor, or can be stored in the memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above-mentioned modules.

[0121] In an exemplary embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 11 shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used for exchanging information between the processor and external devices. The communication interface of the computer device is used for communicating with external terminals in a wired or wireless manner. The wireless manner can be achieved through WIFI, a mobile cellular network, near field communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a method for generating a speech synthesis model based on prosody boundary information and a VAE structure. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, a touchpad, or a mouse, etc.

[0122] Those skilled in the art can understand, Figure 11The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0123] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0124] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0125] In one embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0126] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0127] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.

[0128] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope recorded in the present application.

[0129] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A method for generating a speech synthesis model based on prosodic boundary information and VAE structure, characterized in that The method includes: Obtaining a training data set; the training data set includes sample voices and prosody boundary information, text information, and phoneme information corresponding to the sample voices; Training a speech synthesis network according to the sample voice and the phoneme information corresponding to the sample voice; Training a prosody encoding network according to the sample voice and the prosody boundary information corresponding to the sample voice; the prosody encoding network is composed of a prosody word-level prosody encoder and a prosody phrase-level prosody encoder; the output of the prosody word-level prosody encoder is a prosody word-level prosody embedding vector; the output of the prosody phrase-level prosody encoder is a prosody phrase-level prosody embedding vector; the prosody word-level prosody embedding vector is obtained by inputting the prosody word-level Mel spectrum features extracted from the Mel spectrum into the prosody word-level prosody encoder, obtaining the posterior distribution by using the prosody word-level prosody encoder, and then performing random sampling; the prosody phrase-level prosody embedding vector is obtained by inputting the prosody word phrase-level Mel spectrum features extracted from the Mel spectrum and the prosody word-level prosody embedding vector into the prosody phrase-level prosody encoder, obtaining the posterior distribution by using the prosody phrase-level prosody encoder, and then performing random sampling; After a preset stage of training the speech synthesis network and the prosody encoding network, as the KL loss of the prosody encoding network converges, training a prosody prediction network according to the sample voice and the text information, phoneme information, and prosody boundary information corresponding to the sample voice; the prosody prediction network includes a prosody word-level prosody predictor and a prosody phrase-level prosody predictor; the prosody word-level prosody predictor is established on phoneme sequence information and text sequence information; the phoneme sequence information is encoded based on a phoneme-level encoder; the text sequence information is encoded based on a character-level encoder; Generating a speech synthesis model based on prosody boundary information and a VAE structure according to the trained speech synthesis network, the trained prosody encoding network, and the trained prosody prediction network; Wherein, during the training of the prosody prediction network, a scheduled sampling mechanism is adopted to enable the speech synthesis network to converge under the guidance of the prosody prediction network, so as to obtain the trained speech synthesis model based on prosody boundary information and a VAE structure.

2. The method according to claim 1, wherein The training of the prosody prediction network according to the sample voice and the text information, phoneme information, and prosody boundary information corresponding to the sample voice includes: Inputting the phoneme information corresponding to the sample voice into the trained speech synthesis network to obtain phoneme feature information corresponding to the phoneme information; Inputting the phoneme feature information and the text information corresponding to the sample voice into the prosody prediction network to obtain a prosody prediction result; Inputting the sample voice and the prosody boundary information corresponding to the sample voice into the trained prosody encoding network to obtain prosody feature information; Training the prosody prediction network according to the prosody prediction result and the prosody feature information.

3. The method according to claim 2, characterized in that, The training of the prosody prediction network according to the prosody prediction result and the prosody feature information includes: Screen out candidate prosodic feature information from the prosodic feature information; Train the prosody prediction network according to the prosody prediction result and the candidate prosodic feature information.

4. The method according to claim 1, wherein The training of the speech synthesis network according to the sample speech and the phoneme information corresponding to the sample speech includes: Input the sample speech and the phoneme information corresponding to the sample speech into the speech synthesis network to obtain the latent variable feature information corresponding to the sample speech; Generate the speech synthesis result corresponding to the sample speech through the speech synthesis network according to the latent variable feature information; Train the speech synthesis network according to the speech synthesis result and the sample speech.

5. The method according to claim 1, characterized in that, The training of the prosody encoding network according to the sample speech and the prosody boundary information corresponding to the sample speech includes: Input the phoneme information corresponding to the sample speech into the speech synthesis network to obtain the speech synthesis result corresponding to the sample speech; Input the speech synthesis result into the prosody encoding network to obtain the prosodic feature information corresponding to the speech synthesis result; Train the prosody encoding network according to the prosodic feature information corresponding to the speech synthesis result and the prosody boundary information corresponding to the sample speech.

6. A method for generating speech of a speech synthesis model based on prosodic boundary information and VAE structure, characterized in that, The method includes: Input the text information corresponding to the speech to be generated into the prosody prediction network in the pre-trained speech synthesis model based on prosody boundary information and VAE structure to obtain the prosody prediction result corresponding to the speech to be generated; Input the phoneme information corresponding to the speech to be generated and the prosody prediction result corresponding to the speech to be generated into the speech synthesis network in the pre-trained speech synthesis model based on prosody boundary information and VAE structure to generate the speech to be generated; Wherein, the pre-trained speech synthesis model based on prosody boundary information and VAE structure includes a pre-trained speech synthesis network, a pre-trained prosody encoding network and a pre-trained prosody prediction network; The pre-trained prosody encoding network is composed of a pre-trained prosody word-level prosody encoder and a pre-trained prosody phrase-level prosody encoder; the output of the prosody word-level prosody encoder is a prosody word-level prosody embedding vector; the output of the prosody phrase-level prosody encoder is a prosody phrase-level prosody embedding vector; the prosody word-level prosody embedding vector is obtained by inputting the prosody word-level Mel spectrogram features extracted from the Mel spectrogram into the prosody word-level prosody encoder, obtaining the posterior distribution by using the prosody word-level prosody encoder, and then performing random sampling; the prosody phrase-level prosody embedding vector is obtained by inputting the prosody word phrase-level Mel spectrogram features extracted from the Mel spectrogram and the prosody word-level prosody embedding vector into the prosody phrase-level prosody encoder, obtaining the posterior distribution by using the prosody phrase-level prosody encoder, and then performing random sampling; The pre-trained prosody prediction network is obtained by training as the KL loss of the prosody encoding network converges after a preset stage of training the speech synthesis network and the prosody encoding network; the pre-trained prosody prediction network includes a pre-trained prosody word-level prosody predictor and a pre-trained prosody phrase-level prosody predictor; the pre-trained prosody word-level prosody predictor is built on phoneme sequence information and text sequence information; the phoneme sequence information is encoded based on a phoneme-level encoder; the text sequence information is encoded based on a character-level encoder; Among them, during the training of the prosody prediction network, a scheduled sampling mechanism is adopted to make the speech synthesis network in the speech synthesis model converge under the guidance of the prosody prediction network, so as to obtain the trained speech synthesis model based on prosody boundary information and VAE structure.

7. A voice synthesis model generation device based on prosodic boundary information and VAE structure, characterized in that, The device includes: A training data acquisition module, configured to acquire a training data set; the training data set includes sample speech and prosody boundary information, text information, and phoneme information corresponding to the sample speech; A synthesis network training module, configured to train a speech synthesis network according to the sample speech and phoneme information corresponding to the sample speech; An encoding network training module, configured to train a prosody encoding network according to the sample speech and prosody boundary information corresponding to the sample speech; the prosody encoding network is composed of a prosody word-level prosody encoder and a prosody phrase-level prosody encoder; the output of the prosody word-level prosody encoder is a prosody word-level prosody embedding vector; the output of the prosody phrase-level prosody encoder is a prosody phrase-level prosody embedding vector; the prosody word-level prosody embedding vector is obtained by inputting the prosody word-level Mel spectrum features extracted from the Mel spectrum into the prosody word-level prosody encoder, obtaining the posterior distribution using the prosody word-level prosody encoder, and then performing random sampling; the prosody phrase-level prosody embedding vector is obtained by inputting the prosody word phrase-level Mel spectrum features extracted from the Mel spectrum and the prosody word-level prosody embedding vector into the prosody phrase-level prosody encoder, obtaining the posterior distribution using the prosody phrase-level prosody encoder, and then performing random sampling; A prediction network training module, configured to train a prosody prediction network according to the sample speech and text information, phoneme information, and prosody boundary information corresponding to the sample speech as the KL loss of the prosody encoding network converges after a preset stage of training the speech synthesis network and the prosody encoding network; the prosody prediction network includes a prosody word-level prosody predictor and a prosody phrase-level prosody predictor; the prosody word-level prosody predictor is built on phoneme sequence information and text sequence information; the phoneme sequence information is encoded based on a phoneme-level encoder; the text sequence information is encoded based on a character-level encoder; A model generation module, configured to generate a speech synthesis model based on prosody boundary information and a VAE structure according to a trained speech synthesis network, a trained prosody encoding network, and a trained prosody prediction network; wherein, during the training of the prosody prediction network, a scheduling sampling mechanism is adopted to enable the speech synthesis network to converge under the guidance of the prosody prediction network, so as to obtain the trained speech synthesis model based on prosody boundary information and a VAE structure.

8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Speech synthesis method and device based on rhythm, equipment and medium

    CN115273805A

  • Speech synthesis model training method, speech synthesis method and related equipment

    CN116129856A