Speech Synthesis Method, Apparatus, Device, and Storage Medium

Through the combination of graph encoder and attention mechanism, the problem of rhythm control error accumulation in existing speech synthesis technology is solved, fully automated rhythm adjustment is achieved, and the accuracy of speech synthesis is improved.

CN112349269BActive Publication Date: 2025-07-18PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011446751.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-11
Publication Date
2025-07-18
Estimated Expiration
2040-12-11

AI Technical Summary

Technical Problem

During the pronunciation control process, manual selection of reference speech may lead to the accumulation of model errors and the accuracy of synthetic speech is low.

Method used

The graph encoder and attention mechanism are used to convert the text to be synthesized into the graph embed vector information, generate prosthetic vector information, and use the multi-head attention mechanism to generate the Mel's spectral information, and finally output the speech synthesis information, achieving fully automated prosthetic adjustment.

Benefits of technology

The accuracy of speech synthesis is improved, making the rhythm adjustment process a fully automated process, and improving the quality of speech synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112349269B_ABST
    Figure CN112349269B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of semantic synthesis technology, and discloses a speech synthesis method, device, computer device and computer-readable storage medium. The method includes: obtaining the text to be synthesized, converting the text to be synthesized into graph embedding vector information through a speech synthesis model, encoding the graph embedding vector information according to a graph encoder to generate corresponding first intermediate vector information, generating corresponding Mel spectrogram information according to the first intermediate vector information, and outputting speech synthesis information corresponding to the Mel spectrogram information, so as to realize mapping to different speech prosody rhythms by analyzing the specific semantic information of text information through a graph-assisted encoder, making the prosody adjustment process a fully automated process and improving the accuracy of speech synthesis. At the same time, the present invention also relates to blockchain technology, and the present invention can be applied to fields such as smart government affairs, smart education, and smart healthcare, thereby further promoting the construction of a smart city.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of semantic synthesis, and particularly to a speech synthesis method, apparatus, computer device, and computer-readable storage medium. Background Art

[0002] The TTS (Text To Speech) speech synthesis system is an indispensable part of the intelligent dialogue system. The academic and industrial communities have tried to use limited resources and time to achieve the synthesis of human-like speech. In recent years, after the release of Google's Tacotron and Wavenet, neural network methods have become the mainstream solutions in the field of speech synthesis.

[0003] Currently, neural network-based TTS models have shown good synthesis effects. However, in the process of speech synthesis, prosody embedding is still a challenging task. The prosody vector is first attempted to be extracted from the Mel spectrogram and then input into the attention mechanism together with the output of the encoder at the attention mechanism of the end-to-end model. However, this method is sensitive to the sentence length and the synthesis effect has poor robustness. For this reason, multi-head global style tokens are proposed to represent different speaking styles of speech. These methods control the global style of the synthesized speech, but local speaking prosodies such as pauses, stresses, and intonations are still crucial for the naturalness of the synthesized speech. Therefore, scholars have proposed using a time structure to control the speaking style of the synthesized speech, or using a variational autoencoder to learn the hidden state vector of the speaking style, so that the end-to-end model can be more easily used for local style control. Although it solves the local prosody control in the speech synthesis process to a certain extent, in the process of prosody control, the process of manually selecting reference speech may cause the accumulation of model errors, and the accuracy of the synthesized speech is relatively low. Summary of the Invention

[0004] The main purpose of this application is to provide a speech synthesis method, apparatus, computer device, and computer-readable storage medium, aiming to solve the technical problem that in the existing prosody control process, the process of manually selecting reference speech may cause the accumulation of model errors, and the accuracy of the synthesized speech is relatively low.

[0005] In a first aspect, this application provides a speech synthesis method, and the speech synthesis method includes the following steps:

[0006] Obtain the text to be synthesized and input the text to be synthesized into a speech synthesis model, where the speech synthesis model includes an application layer, an output layer, a graph encoder, and an attention mechanism;

[0007] Based on the application layer, convert the text to be synthesized into graph embedding vector information;

[0008] Encode the graph embedding vector information according to the graph encoder to generate corresponding first prosody vector information, and use the first prosody vector information as the first intermediate vector information;

[0009] Based on the attention mechanism, generate corresponding Mel spectrogram information according to the first intermediate vector information;

[0010] Output the speech synthesis information corresponding to the Mel spectrogram information through the output layer.

[0011] In a second aspect, the present application also provides a speech synthesis device, which includes:

[0012] A first acquisition module, configured to acquire the text to be synthesized and input the text to be synthesized into a speech synthesis model, where the speech synthesis model includes an application layer, an output layer, a graph encoder, and an attention mechanism;

[0013] A conversion model, configured to convert the text to be synthesized into graph embedding vector information based on the application layer;

[0014] A first generation module, configured to encode the graph embedding vector information according to the graph encoder to generate corresponding first prosody vector information, and use the first prosody vector information as the first intermediate vector information;

[0015] A second generation module, configured to generate corresponding Mel spectrogram information based on the attention mechanism according to the first intermediate vector information;

[0016] A second acquisition module, configured to output the speech synthesis information corresponding to the Mel spectrogram information through the output layer.

[0017] In a third aspect, the present application also provides a computer device, which includes a processor, a memory, and a computer program stored on the memory and executable by the processor. When the computer program is executed by the processor, the steps of the speech synthesis method as described above are implemented.

[0018] In a fourth aspect, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the speech synthesis method as described above are implemented.

[0019] The present application provides a speech synthesis method, apparatus, computer device, and computer-readable storage medium. By obtaining the text to be synthesized and inputting the text to be synthesized into a speech synthesis model, where the speech synthesis model includes an application layer, an output layer, a graph encoder, and an attention mechanism; converting the text to be synthesized into graph embedding vector information based on the application layer; encoding the graph embedding vector information according to the graph encoder to generate corresponding first prosody vector information, and using the first prosody vector information as first intermediate vector information; generating corresponding Mel spectrogram information based on the first intermediate vector information according to the attention mechanism; and outputting speech synthesis information corresponding to the Mel spectrogram information through the output layer, it is realized that the specific semantic information of the text information is analyzed by a graph-assisted encoder and mapped to different speech prosody rhythms, making the prosody adjustment process a fully automated process and improving the accuracy of speech synthesis. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0021] Figure 1 It is a schematic flowchart of a speech synthesis method provided by an embodiment of the present application;

[0022] Figure 2 For Figure 1 it is a schematic flowchart of sub-steps of the speech synthesis method in

[0023] Figure 3 For Figure 1 it is a schematic flowchart of sub-steps of the speech synthesis method in

[0024] Figure 4 It is a schematic flowchart of another speech synthesis method provided by an embodiment of the present application;

[0025] Figure 5 It is a schematic block diagram of a speech synthesis apparatus provided by an embodiment of the present application;

[0026] Figure 6 It is a schematic block diagram of the structure of a computer device related to an embodiment of the present application.

[0027] The realization, functional features, and advantages of the objectives of the present application will be further described in conjunction with the embodiments and with reference to the drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0028] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0029] The flowchart shown in the accompanying drawings is only an example illustration, and does not necessarily include all the content and operations / steps, nor does it necessarily need to be executed in the described order. For example, some operations / steps can also be decomposed, combined, or partially merged, so the actual execution order may change according to the actual situation.

[0030] The embodiments of the present application provide a voice synthesis method, device, computer device, and computer-readable storage medium. Among them, the voice synthesis method can be applied to a computer device, and the computer device can be an electronic device such as a laptop computer or a desktop computer.

[0031] Next, some embodiments of the present application will be described in detail in conjunction with the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0032] Please refer to Figure 1 , Figure 1 , which is a schematic flowchart of a voice synthesis method provided for the embodiments of the present application.

[0033] As Figure 1 shown, the voice synthesis method includes steps S101 to S105.

[0034] Step S101, obtain the text to be synthesized, and input the text to be synthesized into a voice synthesis model.

[0035] Exemplarily, obtain the text to be synthesized, and the text to be synthesized includes short sentences and short texts, etc. The obtaining method includes obtaining the text input by the user, or obtaining the text stored in the preset storage path, etc., where the preset storage path includes a blockchain. When the text to be synthesized is obtained, the text to be synthesized is input into the semantic synthesis model, and the voice synthesis model can be stored in the preset blockchain. The voice synthesis model includes an application layer, an output layer, a graph encoder, and an attention mechanism, etc.

[0036] In one embodiment, before obtaining the text to be synthesized, the method further includes: obtaining a speech text to be trained, where the speech text to be trained includes text information and speech information corresponding to the text information; training a preset speech sequence model with the text information and the speech information to obtain graph embedding vector information corresponding to the text information and Mel spectrogram information corresponding to the speech information; obtaining a corresponding loss function based on the graph embedding vector information and the Mel spectrogram information, and updating model parameters of the preset speech sequence model with the loss function to generate a corresponding speech synthesis model.

[0037] Exemplarily, obtain a speech text to be trained, where the speech text to be trained includes text information and speech information corresponding to the text information. Train a preset speech sequence model with the text information and the speech information, obtain corresponding graph embedding vector information through the text information and a graph encoder in the preset speech sequence model, obtain a corresponding Mel spectrogram through the speech information and an attention mechanism in the preset speech sequence model, obtain a corresponding loss function based on the graph embedding vector information and the Mel spectrogram, optimize model parameters of the preset speech sequence model with the loss function, and when it is determined that the model parameters of the preset speech sequence model are optimized and the preset speech sequence model is in a converged state, generate a corresponding speech synthesis model with the preset speech sequence model.

[0038] Step S102: Convert the text to be synthesized into graph embedding vector information based on the application layer.

[0039] Exemplarily, when inputting the text to be synthesized into a speech synthesis model, the speech synthesis model includes an application layer. When the application layer detects the text to be synthesized, it converts the text to be synthesized into graph embedding vector information. Graph embedding is a process of mapping a high-dimensional dense matrix of graph data into a low-dimensional dense vector. By representing a graph as a set of low-dimensional vectors, there are different types of graphs, such as isomorphic graphs, heterogeneous graphs, attributed graphs, etc. The graph embedding vector information includes node vector information and edge vector information. Vector information of each word is obtained through the node vector information, and the prosody relationship between each word is obtained through the edge vector information. Among them, the edge vector information includes directed edge vector information, reverse edge vector information, and sequential edge vector information.

[0040] In one embodiment, specifically, referring to Figure 2 , step S102 includes: sub-step S1021 to sub-step S1022.

[0041] Sub-step S1021: Split the text to be synthesized into each word through the application layer, and obtain the sequential relationship between each word.

[0042] Exemplarily, when the application layer detects the synthetic text, it splits the text to be synthesized into individual words and obtains the sequential relationship between the individual words. For example, if the text to be synthesized is "I love China", the "I love China" is split into "I", "love", "China", and "country". And the sequence of "I", "love", "China", and "country" is obtained as "I" → "love" → "China" → "country".

[0043] Sub-step S1022: Perform mapping transformation on each word and the sequential relationship between each of the words to obtain graph embedding vector information corresponding to the text to be synthesized.

[0044] Exemplarily, when the individual words of the text to be synthesized and the sequential relationship between the individual words are obtained, each word and the sequential relationship of each word are mapped to obtain word vector information of each word and sequential vector information between each word, that is, edge vector information. The obtained word vector information and edge vector information are combined to obtain corresponding graph embedding vector information, where the weight in the edge vector information is 0.

[0045] Step S103: Encode the graph embedding vector information according to the graph encoder to generate corresponding first prosody vector information, and use the first prosody vector information as the first intermediate vector information.

[0046] Exemplarily, when the graph embedding vector information of the text to be synthesized is obtained, the graph embedding vector information is encoded by the graph encoder in the speech synthesis model to generate corresponding first prosody vector information. For example, the graph encoder includes a mapping function, and the graph embedding vector information is mapped and encoded through the mapping function to obtain the first prosody vector information corresponding to the graph embedding vector information. When the first prosody vector information is obtained, the first prosody vector information is used as the first intermediate vector information.

[0047] In a real-time example, the graph embedding vector information includes multiple node vectors and multiple edge vectors; encoding the graph embedding vector information according to the graph encoder to generate corresponding first prosody vector information includes: obtaining edge vectors between each of the node vectors through the graph encoder, and encoding the edge vectors to obtain the first prosody vector information corresponding to the graph embedding vector information, where the edge vector represents the prosody relationship between two corresponding node vectors.

[0048] Exemplarily, when the graph embedding vector information of the text to be synthesized is obtained, the edge vectors between each of the node vectors are encoded by the graph encoder to obtain prosody vector information between each of the node vectors, and corresponding first prosody vector information is obtained through the sequential relationship between each of the nodes and the prosody vector information between each of the node vectors.

[0049] Step S104: Based on the attention mechanism, generate corresponding Mel spectrogram information according to the first intermediate vector information.

[0050] Exemplarily, when the first intermediate vector information is obtained, the first intermediate vector information is input into the attention mechanism, and the attention mechanism performs context learning on the first intermediate vector information to generate corresponding Mel spectrogram information. For example, when the attention mechanism is a multi-head attention mechanism, information that should not be known during sequence generation (i.e., illegal information) is masked through the multi-head attention. Among them, multi-head attention is mainly for consistency during training and inference. For example, during training, when trying to predict the pronunciation of "w", in fact, the entire prosody vector will enter the network when it enters the network. The sequences after "w" in this prosody vector should be masked from the network to prevent the network from seeing information that needs to be predicted in the future, because this information cannot be seen during inference.

[0051] It should be noted that multi-head attention consists of several self-attention mechanisms. For example, 4-head attention is essentially performing self-attention on the sequence 4 times.

[0052] In one embodiment, specifically, referring to Figure 3 , step S104 includes: sub-step S1041 to sub-step S1042.

[0053] Sub-step S1041: Input the first intermediate vector information into the attention mechanism, and obtain the context prosody information of each node in the first intermediate vector information through the weight matrix in the attention mechanism.

[0054] Exemplarily, when the first prosody vector information is obtained, the first prosody vector information is used as the intermediate vector information and input into the attention mechanism, and the context prosody information of each node in the first intermediate vector is obtained through the weight matrix in the attention mechanism.

[0055] Sub-step S1042: Generate corresponding Mel spectrogram information by decoding the context prosody information of each node in the first intermediate vector information.

[0056] Exemplarily, when the context prosody information of each node in the first intermediate vector is obtained, the context prosody information of each node in the first intermediate vector is decoded through the pre-set decoder in the attention mechanism to obtain the corresponding Mel spectrogram information.

[0057] Step S105: Output the speech synthesis information corresponding to the Mel spectrogram information through the output layer.

[0058] Exemplarily, when the Mel spectrogram information is obtained, the speech synthesis information corresponding to the Mel spectrogram information is output through the output layer. For example, the output layer includes a vocoder that obtains the speech frequency domain feature information in the Mel spectrogram information and generates the corresponding speech synthesis information by synthesizing the speech frequency domain feature information.

[0059] In one embodiment, outputting the speech synthesis information corresponding to the Mel spectrogram information through the output layer includes: extracting the speech frequency domain features in the Mel spectrogram information through the output layer; and mapping the speech frequency domain features to obtain the speech synthesis information output by the output layer.

[0060] Exemplarily, when the Mel spectrogram information is obtained, the speech frequency domain features in the Mel spectrogram information are extracted through the output layer. When the speech frequency domain features in the Mel spectrogram information are extracted, the speech frequency domain features are mapped to obtain the speech synthesis information output by the output layer. For example, the output layer includes an extraction layer and a mapping layer. The speech frequency domain features in the Mel spectrogram information are extracted through the extraction layer, and the speech frequency domain features are activation-mapped through the activation function in the mapping layer to obtain the corresponding speech synthesis information.

[0061] In an embodiment of the present invention, a text to be synthesized is obtained and input into a speech synthesis model. The application layer converts the text to be synthesized into graph embedding vector information. The graph encoder encodes the graph embedding vector information to generate the corresponding first prosody vector information. The attention mechanism generates the corresponding Mel spectrogram information according to the first intermediate vector information. The output layer outputs the speech synthesis information corresponding to the Mel spectrogram information, realizing mapping to different speech prosody rhythms by analyzing the specific semantic information of the text information through the graph-assisted encoder, making the prosody adjustment process a fully automated process and improving the accuracy of speech synthesis.

[0062] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of another speech synthesis method provided by the embodiment of the present application.

[0063] As Figure 4 shown, the speech synthesis method includes steps S201 to S208.

[0064] Step S201: Obtain a text to be synthesized and input the text to be synthesized into a speech synthesis model.

[0065] Exemplarily, obtain the text to be synthesized, which includes short sentences and short texts, etc. The obtaining method includes obtaining the text input by the user, or obtaining the text stored in the preset storage path, etc., where the preset storage path includes a blockchain. When the text to be synthesized is obtained, input the text to be synthesized into the semantic synthesis model, and the speech synthesis model can be stored in the preset blockchain. The speech synthesis model includes an application layer, an output layer, a graph encoder, and an attention mechanism, etc.

[0066] In one embodiment, before obtaining the text to be synthesized, it further includes: obtaining the speech text to be trained, where the speech text to be trained includes text information and the speech information corresponding to the text information; training the preset speech sequence model through the text information and the speech information to obtain the graph embedding vector information corresponding to the text information and the Mel spectrogram information corresponding to the speech information; obtaining the corresponding loss function through the graph embedding vector information and the Mel spectrogram information, and updating the model parameters of the preset speech sequence model through the loss function to generate the corresponding speech synthesis model.

[0067] Exemplarily, obtain the speech text to be trained, which includes text information and the speech information corresponding to the text information. Train the preset speech sequence model through the text information and the speech information. Through the text information and the graph encoder in the preset speech sequence model, obtain the corresponding graph embedding vector information. Through the speech information and the attention mechanism in the preset speech sequence model, obtain the corresponding Mel spectrogram. Obtain the corresponding loss function through the graph embedding vector information and the Mel spectrogram, and optimize the model parameters of the preset speech sequence model through the loss function. When it is determined that the model parameters of the preset speech sequence model are optimized and the preset speech sequence model is in a converged state, generate the corresponding speech synthesis model from the preset speech sequence model.

[0068] Step S202: Convert the text to be synthesized into graph embedding vector information based on the application layer.

[0069] Exemplarily, when the text to be synthesized is input into the speech synthesis model, the speech synthesis model includes an application layer. When the application layer detects the text to be synthesized, it converts the text to be synthesized into graph embedding vector information. Graph embedding is a process of mapping a high-dimensional dense matrix of graph data into a low-dimensional dense vector. By representing the graph as a set of low-dimensional vectors, there are different types of graphs, such as isomorphic graphs, heterogeneous graphs, attribute graphs, etc. The graph embedding vector information includes node vector information and edge vector information. Through the node vector information, the vector information of each word is obtained, and through the edge vector information, the prosody relationship between each word is obtained. Among them, the edge vector information includes directed edge vector information, reverse edge vector information, and sequential edge vector information.

[0070] Step S203: Encode the graph embedding vector information according to the graph encoder to generate corresponding first prosody vector information, and use the first prosody vector information as the first intermediate vector information.

[0071] Exemplarily, when the graph embedding vector information of the text to be synthesized is obtained, the graph encoder in the speech synthesis model encodes the graph embedding vector information to generate corresponding first prosody vector information. For example, the graph encoder includes a mapping function, and the graph embedding vector information is mapped and encoded through the mapping function to obtain the first prosody vector information corresponding to the graph embedding vector information. When the first prosody vector information is obtained, the first prosody vector information is used as the first intermediate vector information.

[0072] Step S204: Convert the text to be synthesized into text vector information based on the application layer.

[0073] Exemplarily, when the text to be synthesized is input into the speech synthesis model, the application layer in the speech synthesis model converts the text to be synthesized into text vector information. The application layer extracts the positions of each word and the corresponding pinyin in the text to be synthesized, obtains the numbers or letters corresponding to the positions of each word and the corresponding pinyin in the preset coding rules, and converts the text to be synthesized into text vector information through the numbers or letters.

[0074] Step S205: Encode the text vector information according to the encoder to generate corresponding hidden vector information.

[0075] Exemplarily, when the text vector information is obtained, the encoder encodes the text vector information to generate corresponding hidden vector information. For example, the encoder includes an encoding rule, and the text vector information is encoded through the encoding rule to obtain the corresponding hidden vector information.

[0076] Step S206: Concatenate the hidden vector information and the first intermediate vector information to generate corresponding second intermediate vector information.

[0077] Exemplarily, when the hidden vector information and the first intermediate vector information are obtained, the dimension information of the hidden vector information and the dimension information of the first intermediate vector information are respectively obtained. Through the dimension information, it is determined that the hidden vector information and the first intermediate vector information are in the same dimension, and the hidden vector information and the first intermediate vector information are concatenated in the same dimension to generate corresponding second intermediate vector information.

[0078] Step S207: Generate corresponding Mel spectrogram information based on the attention mechanism and the second intermediate vector information.

[0079] Exemplarily, when the second intermediate vector information is obtained, it is input into the attention mechanism, and the context prosody information of each node in the second intermediate vector is obtained through the weight matrix in the attention mechanism. When the context prosody information of each node in the second intermediate vector is obtained, the context prosody information of each node in the second intermediate vector is decoded through a preset decoder to obtain the corresponding Mel spectrogram information.

[0080] Step S208: Output the speech synthesis information corresponding to the Mel spectrogram information through the output layer.

[0081] Exemplarily, when the Mel spectrogram information is obtained, the speech synthesis information corresponding to the Mel spectrogram information is output through the output layer. For example, the output layer includes a vocoder, and the vocoder obtains the speech frequency domain feature information in the Mel spectrogram information and generates the corresponding speech synthesis information by synthesizing the speech frequency domain feature information.

[0082] In the embodiment of the present invention, by inputting the text to be synthesized into the speech synthesis model, the text vector information and the graph embedding vector information corresponding to the text to be synthesized are obtained. The graph embedding vector information is encoded by the graph encoder and the text vector information is encoded by the encoder, and the Mel spectrogram information output by the attention mechanism is obtained, and the speech synthesis information corresponding to the Mel spectrogram information output by the output layer is obtained. The semantic structure information is embedded into the speech synthesis model, and the graph auxiliary encoder analyzes the prosody information from the text side, and the encoder analyzes the word position information from the text side, so as to realize mapping the specific semantic information of the text information analyzed by the graph auxiliary encoder to different speech prosody rhythms, making the prosody adjustment process a fully automated process and improving the accuracy of speech synthesis.

[0083] Please refer to Figure 5 , Figure 5 which is a schematic block diagram of a speech synthesis device provided by an embodiment of the present application.

[0084] As Figure 5 shown, the speech synthesis device 400 includes: a first acquisition module 401, a conversion model 402, a first generation module 403, a second generation module 404, and a second acquisition module 405.

[0085] The first acquisition module 401 is configured to acquire the text to be synthesized and input the text to be synthesized into the speech synthesis model, where the speech synthesis model includes an application layer, an output layer, a graph encoder, and an attention mechanism;

[0086] The conversion model 402 is configured to convert the text to be synthesized into graph embedding vector information based on the application layer;

[0087] The first generation module 403 is configured to encode the graph embedding vector information according to the graph encoder, generate corresponding first prosody vector information, and use the first prosody vector information as first intermediate vector information;

[0088] The second generation module 404 is configured to generate corresponding Mel spectrogram information based on the attention mechanism and according to the first intermediate vector information;

[0089] The second acquisition module 405 is configured to output speech synthesis information corresponding to the Mel spectrogram information through the output layer.

[0090] Wherein, the speech synthesis device is further specifically configured to:

[0091] Convert the text to be synthesized into text vector information based on the application layer;

[0092] Encode the text vector information according to the encoder to generate corresponding hidden vector information;

[0093] Concatenate the hidden vector information and the first intermediate vector information to generate corresponding second intermediate vector information;

[0094] Generate corresponding Mel spectrogram information based on the attention mechanism and the second intermediate vector information.

[0095] Wherein, the first generation module 403 is further specifically configured to:

[0096] Obtain edge vectors between each of the node vectors through the graph encoder, and encode the edge vectors to obtain first prosody vector information corresponding to the graph embedding vector information, wherein the edge vector represents the prosody relationship between two corresponding node vectors.

[0097] Wherein, the second generation module 404 is further specifically configured to:

[0098] Input the first intermediate vector information into the attention mechanism, and obtain context prosody information of each node in the first intermediate vector information through a weight matrix in the attention mechanism;

[0099] Decode the context prosody information of each node in the first intermediate vector information to generate corresponding Mel spectrum information.

[0100] Wherein, the second acquisition module 405 is further specifically configured to:

[0101] Extract speech frequency domain features in the Mel spectrum information through the output layer;

[0102] Map the speech frequency domain features to obtain speech synthesis information output by the output layer.

[0103] Among them, the conversion model 402 is specifically further configured to:

[0104] Split the text to be synthesized into individual words and phrases through the application layer, and obtain the sequential relationship between the individual words and phrases;

[0105] Perform mapping conversion on each word and phrase and the sequential relationship between each of the words and phrases to obtain the graph embedding vector information corresponding to the text to be synthesized.

[0106] Among them, the speech synthesis device is further configured to:

[0107] Obtain the speech text to be trained, where the speech text to be trained includes text information and the speech information corresponding to the text information;

[0108] Train a preset speech sequence model through the text information and the speech information to obtain the graph embedding vector information corresponding to the text information and the Mel spectrogram information corresponding to the speech information;

[0109] Obtain a corresponding loss function through the graph embedding vector information and the Mel spectrogram information, and update the model parameters of the preset speech sequence model through the loss function to generate a corresponding speech synthesis model.

[0110] It should be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described device and each module and unit can refer to the corresponding processes in the foregoing speech synthesis method embodiments, and will not be elaborated herein.

[0111] The device provided in the above embodiment can be implemented in the form of a computer program, and the computer program can run on a computer device as shown in Figure 6 shown.

[0112] Please refer to Figure 6 , Figure 6 , which is a schematic block diagram of the structure of a computer device provided in an embodiment of the present application. The computer device can be a terminal.

[0113] As shown in Figure 6 shown, the computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the memory can include a non-volatile storage medium and an internal memory.

[0114] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions, and when the program instructions are executed, the processor can execute any one of the speech synthesis methods.

[0115] The processor is used to provide computing and control capabilities to support the operation of the entire computer device.

[0116] The internal memory provides an environment for the operation of a computer program in a non-volatile storage medium. When the computer program is executed by the processor, the processor can execute any voice synthesis method.

[0117] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art can understand that Figure 6 The structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0118] It should be understood that the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0119] Wherein, in one embodiment, the processor is used to run a computer program stored in the memory to implement the following steps:

[0120] Obtain the text to be synthesized and input the text to be synthesized into a voice synthesis model, where the voice synthesis model includes an application layer, an output layer, a graph encoder, and an attention mechanism;

[0121] Convert the text to be synthesized into graph embedding vector information based on the application layer;

[0122] Encode the graph embedding vector information according to the graph encoder to generate corresponding first prosody vector information, and use the first prosody vector information as the first intermediate vector information;

[0123] Generate corresponding Mel spectrogram information based on the first intermediate vector information based on the attention mechanism;

[0124] Output voice synthesis information corresponding to the Mel spectrogram information through the output layer.

[0125] In one embodiment, the speech synthesis model of the processor further includes an encoder; when the method is implemented, it is used to implement:

[0126] Convert the text to be synthesized into text vector information based on the application layer;

[0127] Encode the text vector information according to the encoder to generate corresponding hidden vector information;

[0128] Concatenate the hidden vector information and the first intermediate vector information to generate corresponding second intermediate vector information;

[0129] Generate corresponding Mel spectrogram information based on the attention mechanism and the second intermediate vector information.

[0130] In one embodiment, the graph embedding vector information of the processor includes a plurality of node vectors and a plurality of edge vectors;

[0131] When implementing encoding the graph embedding vector information according to the graph encoder to generate corresponding first prosody vector information, it is used to implement:

[0132] Obtain the edge vectors between each of the node vectors through the graph encoder, and encode the edge vectors to obtain the first prosody vector information corresponding to the graph embedding vector information, where the edge vector represents the prosody relationship between the corresponding two node vectors.

[0133] In one embodiment, when the processor generates corresponding Mel spectrogram information based on the attention mechanism according to the first intermediate vector information, it is used to implement:

[0134] Input the first intermediate vector information into the attention mechanism, and obtain the context prosody information of each node in the first intermediate vector information through the weight matrix in the attention mechanism;

[0135] Decode the context prosody information of each node in the first intermediate vector information to generate corresponding Mel spectrogram information.

[0136] In one embodiment, when the processor outputs the speech synthesis information corresponding to the Mel spectrogram information through the output layer, it is used to implement:

[0137] Extract the speech frequency domain features in the Mel spectrogram information through the output layer;

[0138] And map the speech frequency domain features to obtain the speech synthesis information output by the output layer.

[0139] In one embodiment, when the processor converts the text to be synthesized into graph embedding vector information based on the application layer, it is used to implement:

[0140] Split the text to be synthesized into individual words and phrases through the application layer, and obtain the sequential relationship between the individual words and phrases;

[0141] Perform mapping transformation on each word and phrase and the sequential relationship between the words and phrases to obtain the graph embedding vector information corresponding to the text to be synthesized.

[0142] In one embodiment, when the processor obtains the text to be synthesized, it is used to implement:

[0143] Obtain the speech text to be trained, where the speech text to be trained includes text information and the speech information corresponding to the text information;

[0144] Train a preset speech sequence model through the text information and the speech information to obtain the graph embedding vector information corresponding to the text information and the Mel spectrogram information corresponding to the speech information;

[0145] Obtain a corresponding loss function through the graph embedding vector information and the Mel spectrogram information, and update the model parameters of the preset speech sequence model through the loss function to generate a corresponding speech synthesis model.

[0146] The embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. The computer program includes program instructions, and the method implemented when the program instructions are executed can refer to the various embodiments of the speech synthesis method of the present application.

[0147] Among them, the computer-readable storage medium may be an internal storage unit of the computer device described in the foregoing embodiment, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk equipped on the computer device, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc.

[0148] Further, the computer-readable storage medium mainly includes a storage program area and a storage data area. Among them, the storage program area can store an operating system, application programs required for at least one function, etc.; the storage data area can store data created according to the use of blockchain nodes, etc.

[0149] The blockchain referred to in the present invention is a new application mode of computer technologies such as the storage of speech synthesis models, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Blockchain, in essence, is a decentralized database, a series of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block. The blockchain can include the blockchain underlying platform, the platform product service layer, and the application service layer, etc.

[0150] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent in such a process, method, article or system. Without further limitations, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or system including the element.

[0151] The serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments. The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art in the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A voice synthesis method, characterized in that, Including: Obtain the text to be synthesized, and input the text to be synthesized into a speech synthesis model, where the speech synthesis model includes an application layer, an output layer, a graph encoder, and an attention mechanism; Based on the application layer, convert the text to be synthesized into graph embedding vector information, where the graph embedding vector information includes a plurality of node vectors and a plurality of edge vectors; Encode the graph embedding vector information according to the graph encoder to generate corresponding first prosody vector information, and use the first prosody vector information as first intermediate vector information; Based on the attention mechanism, generate corresponding Mel spectrogram information according to the first intermediate vector information; Output speech synthesis information corresponding to the Mel spectrogram information through the output layer; The encoding the graph embedding vector information according to the graph encoder to generate corresponding first prosody vector information includes: Obtain edge vectors between each of the node vectors through the graph encoder, and encode the edge vectors to obtain first prosody vector information corresponding to the graph embedding vector information, where the edge vector represents the prosody relationship between two corresponding node vectors.

2. The speech synthesis method according to claim 1, characterized in that, The speech synthesis model further includes an encoder; after using the first prosody vector information as first intermediate vector information and before outputting speech synthesis information corresponding to the Mel spectrogram information through the output layer, it further includes: Convert the text to be synthesized into text vector information based on the application layer; Encode the text vector information according to the encoder to generate corresponding hidden vector information; Concatenate the hidden vector information and the first intermediate vector information to generate corresponding second intermediate vector information; Generate corresponding Mel spectrogram information based on the attention mechanism and the second intermediate vector information.

3. The speech synthesis method according to claim 1, characterized in that, The generating corresponding Mel spectrogram information according to the first intermediate vector information based on the attention mechanism includes: Input the first intermediate vector information into the attention mechanism, and obtain context prosody information of each node in the first intermediate vector information through a weight matrix in the attention mechanism; Decode the context prosody information of each node in the first intermediate vector information to generate corresponding Mel spectrogram information.

4. The speech synthesis method according to claim 1, wherein The outputting speech synthesis information corresponding to the Mel spectrogram information through the output layer includes: Extract speech frequency domain features in the Mel spectrogram information through the output layer; Map the speech frequency domain features to obtain speech synthesis information output by the output layer.

5. The speech synthesis method according to claim 1, wherein The converting the text to be synthesized into graph embedding vector information based on the application layer includes: Split the text to be synthesized into each word through the application layer, and obtain the sequential relationship between each word; Perform mapping transformation on each word and the sequential relationship between each word to obtain graph embedding vector information corresponding to the text to be synthesized.

6. The speech synthesis method according to claim 1, wherein Before obtaining the text to be synthesized, it further includes: Obtain a speech text to be trained, where the speech text to be trained includes text information and speech information corresponding to the text information; Train a preset voice sequence model through the text information and the voice information to obtain graph embedding vector information corresponding to the text information and Mel spectrogram information corresponding to the voice information; Obtain a corresponding loss function through the graph embedding vector information and the Mel spectrogram information, and update the model parameters of the preset voice sequence model through the loss function to generate a corresponding voice synthesis model.

7. A voice synthesis device, characterized in that, It includes: A first acquisition module, configured to acquire a text to be synthesized and input the text to be synthesized into a voice synthesis model, where the voice synthesis model includes an application layer, an output layer, a graph encoder, and an attention mechanism; A conversion model, configured to convert the text to be synthesized into graph embedding vector information based on the application layer, where the graph embedding vector information includes a plurality of node vectors and a plurality of edge vectors; A first generation module, configured to encode the graph embedding vector information according to the graph encoder to generate corresponding first prosody vector information, and use the first prosody vector information as first intermediate vector information; A second generation module, configured to generate corresponding Mel spectrogram information based on the first intermediate vector information according to the attention mechanism; A second acquisition module, configured to output voice synthesis information corresponding to the Mel spectrogram information through the output layer; The encoding the graph embedding vector information according to the graph encoder to generate corresponding first prosody vector information includes: Obtain edge vectors between the respective node vectors through the graph encoder, and encode the edge vectors to obtain first prosody vector information corresponding to the graph embedding vector information, where the edge vectors represent the prosody relationship between two corresponding node vectors.

8. A computer device, characterized in that, The computer device includes a processor, a memory, and a computer program stored on the memory and executable by the processor, where when the computer program is executed by the processor, the steps of the voice synthesis method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, where when the computer program is executed by a processor, the steps of the voice synthesis method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Speech synthesis method, speech synthesis system, terminal equipment and readable storage medium

    CN110335587A

  • Voice processing method and device, electronic equipment and storage medium

    CN111326136A