Heterogeneous relation graph-based end-to-end speech synthesis method, apparatus and device, and medium

Through the end-to-end speech synthesis method based on heterogeneous relationship diagrams, the problem of lack of a representation framework integrating speech and syntactic features in the prior art is solved, and a more natural and smooth speech synthesis effect is achieved.

CN120148469APending Publication Date: 2025-06-13PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510221509.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art lacks a general representation framework and it is difficult to integrate all forms of structured speech, acoustic and syntactic features.

Method used

The end-to-end speech synthesis method based on heterogeneous relationship graph is adopted. By receiving a given text, linguistic information is extracted, and a heterogeneous relationship graph is encoded. The node features are processed using a graph convolution network, and the end-to-end TTS model is input to generate a speech waveform.

Benefits of technology

It realizes effective characterization of phonemes and syntactic information, improves the performance of the TTS model through structured representation, and significantly improves the naturalness and fluency of speech synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148469A_ABST
    Figure CN120148469A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and discloses an end-to-end speech synthesis method and device based on a heterogeneous relational graph, equipment and a medium, and the method comprises the steps: receiving a given text, and extracting linguistic information in the given text; encoding the linguistic information to obtain a corresponding heterogeneous relation graph; performing initialization, feature aggregation and normalization processing on each node in the heterogeneous relation graph by adopting a graph convolutional network to obtain a node feature of each node; and inputting the node features into an end-to-end TTS model to generate a corresponding voice waveform. According to the method, the deterministic linguistic information is input into the TTS model in the form of the heterogeneous relational graph, and phoneme and syntactic information can be effectively represented. In the process, structured representation is learned by means of a graph convolutional network, and the structured representation is applied to an encoder of a TTS model. Therefore, a text-to-speech synthesis scheme based on the heterogeneous relational graph is realized, and the method has the advantage of remarkably improving the performance by virtue of syntactic advantages.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an end-to-end speech synthesis method, apparatus, device and medium based on a heterogeneous relationship graph. Background Art

[0002] In the speech processing technology in the fields of finance and healthcare, a neural TTS system can be used to take free text generated in the fields of finance and healthcare as input and treat it as a character sequence. However, this input method ignores key speech information, such as phoneme information, stress patterns, syllable structures, and lexical structures. In contrast, past statistical parametric synthesis models have demonstrated the advantage of incorporating language structures into the models, which explicitly encode the local neighborhood structure around each speech or syntactic unit, such as the number of phonemes in a word. Inspired by this, researchers have begun to explore neural speech models and representation methods that can recover language information, which can capture the continuous subunit representations of syllables and phonemes. Although there has been a large amount of research on structural encoding and inferring potential speech relationships, there has been no research exploring a general representation framework that can encapsulate all forms of structured speech, acoustic, and syntactic features. Summary of the Invention

[0003] The present invention provides an end-to-end speech synthesis method, apparatus, computer device and medium based on a heterogeneous relationship graph to solve the problem that the prior art lacks a general representation framework for integrating all speech and syntactic features.

[0004] In a first aspect, an end-to-end speech synthesis method based on a heterogeneous relationship graph is provided, including:

[0005] Receiving a given text and extracting linguistic information in the given text;

[0006] Encoding the linguistic information to obtain a corresponding heterogeneous relationship graph;

[0007] Using a graph convolutional network to perform initialization, feature aggregation, and normalization processing on each node in the heterogeneous relationship graph to obtain node features of each node;

[0008] Inputting the node features into an end-to-end TTS model to generate a corresponding speech waveform.

[0009] In a second aspect, an end-to-end speech synthesis apparatus based on a heterogeneous relationship graph is provided, including:

[0010] An information extraction module, configured to receive a given text and extract linguistic information in the given text;

[0011] An encoding module, configured to encode the linguistic information to obtain a corresponding heterogeneous relationship graph;

[0012] A feature extraction module, which is used to initialize, aggregate features and normalize each node in the heterogeneous relationship graph by using a graph convolutional network, so as to obtain the node features of each node.

[0013] A speech synthesis module, which is used to input the node features into an end-to-end TTS model to generate corresponding speech waveforms.

[0014] In a third aspect, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned end-to-end speech synthesis method based on a heterogeneous relationship graph are implemented.

[0015] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned end-to-end speech synthesis method based on a heterogeneous relationship graph are implemented.

[0016] In the solution implemented by the above-mentioned end-to-end speech synthesis method, device, computer device and storage medium based on a heterogeneous relationship graph, by receiving a given text and extracting the linguistic information in the given text; encoding the linguistic information to obtain a corresponding heterogeneous relationship graph; using a graph convolutional network to initialize, aggregate features and normalize each node in the heterogeneous relationship graph, so as to obtain the node features of each node; inputting the node features into an end-to-end TTS model to generate corresponding speech waveforms. The present invention inputs deterministic linguistic information into the TTS model in the form of a heterogeneous relationship graph, which can effectively represent phoneme and syntactic information. In this process, a graph convolutional network is used to learn a structured representation and apply it to the encoder of the TTS model. Thus, an end-to-end text-to-speech synthesis solution based on a heterogeneous relationship graph is realized, which has the advantage of significantly improving performance by virtue of syntactic advantages. Description of the Drawings

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required to be used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 is a schematic diagram of an application environment of an end-to-end speech synthesis method based on a heterogeneous relationship graph in an embodiment of the present invention;

[0019] Figure 2 is a schematic flowchart of an end-to-end speech synthesis method based on a heterogeneous relationship graph in an embodiment of the present invention;

[0020] Figure 3 is Figure 2 A schematic flowchart of a specific implementation manner of step S202 in

[0021] Figure 4 is Figure 2 A schematic flowchart of a specific implementation manner of step S203 in

[0022] Figure 5 is Figure 2 A schematic flowchart of a specific implementation manner of step S204 in

[0023] Figure 6 A schematic structural diagram of an end-to-end speech synthesis device based on a heterogeneous relationship graph in an embodiment of the present invention;

[0024] Figure 7 A schematic structural diagram of a computer device in an embodiment of the present invention;

[0025] Figure 8 A schematic structural diagram of another computer device in an embodiment of the present invention;

[0026] Figure 9 A schematic diagram of a hierarchical tree structure in an embodiment of the present invention. Specific implementation manner

[0027] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.

[0028] The end-to-end speech synthesis method based on a heterogeneous relationship graph provided by the embodiments of the present invention can be applied, for example, in Figure 1In the application environment, the client communicates with the server through the network. The server can receive a given text through the client and extract the linguistic information in the given text; encode the linguistic information to obtain a corresponding heterogeneous relationship graph; use a graph convolutional network to initialize, aggregate features, and normalize each node in the heterogeneous relationship graph to obtain the node features of each node; input the node features into an end-to-end TTS model to generate corresponding speech waveforms. A text-to-speech synthesis scheme based on a heterogeneous relationship graph is implemented, which has the advantage of significantly improving performance by virtue of syntactic advantages. Among them, the client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail below through specific embodiments.

[0029] Please refer to Figure 2 as shown in Figure 2 FIG. is a schematic flowchart of an end-to-end speech synthesis method based on a heterogeneous relationship graph provided by an embodiment of the present invention, including the following steps.

[0030] S201. Receive a given text and extract the linguistic information in the given text;

[0031] S202. Encode the linguistic information to obtain a corresponding heterogeneous relationship graph;

[0032] In steps S201-S202, the heterogeneous relationship graph, as a self-contained architecture, is used to structurally represent the heterogeneous relationship between words, syllables, and phonemes in the pronunciation of a given text, facilitating more accurate language processing and analysis in the subsequent process.

[0033] S203. Use a graph convolutional network to initialize, aggregate features, and normalize each node in the heterogeneous relationship graph to obtain the node features of each node;

[0034] In this step, through initialization, feature aggregation, and normalization processing, each node can learn the structural information and feature information of itself and its neighbor nodes.

[0035] S204. Input the node features into an end-to-end TTS model to generate corresponding speech waveforms.

[0036] In this step, the TTS model can adopt a variational autoencoder - temporal generation network (VITS). This model combines the latent variable modeling ability of the variational autoencoder and the efficient speech synthesis ability of the temporal generation network, and can generate natural and fluent speech waveforms.

[0037] In this embodiment, inputting deterministic linguistic information into the TTS model in the form of a heterogeneous relational graph can effectively represent phoneme and syntactic information. In this process, a graph convolutional network is used to learn the structural representation and apply it to the encoder of the TTS model. Thus, a text-to-speech synthesis scheme based on a heterogeneous relational graph is realized, which has the advantage of achieving significant performance improvement by virtue of syntactic advantages. That is, the present invention reexamines the practicality of the heterogeneous relational graph as a general framework, represents the phoneme / syntactic information of the pronunciation of the input text in the TTS model through the heterogeneous relational graph, and provides some theoretical basis for the observed improvement by analyzing the ability of the heterogeneous relational graph in predicting the duration (acoustic information) of each phoneme in the input.

[0038] In a specific scenario in the financial field, when a user hopes to convert their financial report into a voice report so that they can also listen to it while driving or doing other activities. Using the solution of the present invention, first, the text in the financial report is converted into a heterogeneous relational graph, which contains the complex relationships between words, syllables and phonemes, as well as syntactic structure information. Then, these information are deeply learned through a graph convolutional network to extract key features. Next, these features are input into the VITS model to generate a natural and smooth speech waveform. Finally, the user can listen to the converted voice report through a smart device, which is both convenient and efficient. This process fully reflects the innovation and practicality of the present invention in the field of text-to-speech synthesis.

[0039] In a specific scenario in the medical field, when a doctor needs to convert a diagnosis report or drug usage instructions into voice for better patient understanding. Similarly, using the solution of the present invention, the doctor can convert the text of the diagnosis report or drug usage instructions into a heterogeneous relational graph, which contains the complex relationships between words, phrases and sentences, as well as syntactic structure information. Through the deep learning of the graph convolutional network, the key medical terms and key information in the text are extracted. Then, these features are input into the VITS model to generate clear and accurate voice. Finally, the patient can listen to the converted voice information through a mobile phone or other smart devices, which is of great significance for improving the patient's medical experience and treatment effect. This process once again proves the wide applicability and practical application value of the present invention in the field of text-to-speech synthesis.

[0040] It can be understood that the present invention can provide a more accurate and efficient solution for text-to-speech synthesis in different scenarios. Through the effective representation of the heterogeneous relational graph and the learning of the graph convolutional network, the TTS model of the present invention can better capture and express the phoneme and syntactic information in the text, thus generating a more natural and smooth speech waveform.

[0041] In one embodiment, step S201 includes: analyzing and processing the given text using predefined deterministic rules in a system with speech synthesis function to extract the linguistic information in the given text.

[0042] In this embodiment, the system with speech synthesis function can be the Festvox system. The Festvox system is an open-source software framework for speech synthesis, which contains a series of tools and predefined rules to help analyze text and convert text into speech. The linguistic information in the given text refers to language features such as words, syllables, phonemes, pronunciation, grammar, semantics, etc. These information are very important for the speech synthesis system and determine the accuracy and naturalness of the synthesized speech.

[0043] In this embodiment, there are predefined deterministic rules that guide how the text should be correctly pronounced and understood. These deterministic rules can be set based on linguistic principles, such as grammar rules, pronunciation rules, etc. By analyzing the given text through these rules, the system can identify the linguistic features in the text, such as sentence structure, lexical meaning, speech rhythm, etc., and then provide the necessary information for synthesizing speech.

[0044] In this embodiment, since the information from HMM alignment is not used, there is no need for speech signals to obtain the linguistic structure of the pronunciation of the given text. Limiting this embodiment to only using pre-aligned features has beneficial effects in two aspects: i) allowing the use of heterogeneous relational graphs during the inference process (when there is no speech waveform); ii) given a character sequence, the corresponding heterogeneous relational graph can be derived in a deterministic manner, and the additional processing cost can be ignored.

[0045] In one embodiment, please refer to Figure 3 as shown, step S202 includes:

[0046] S301. Encode the linguistic information into a graph G=(V, E), where V contains a set of nodes of words, syllables and phonemes, and E contains the relationships between the nodes;

[0047] S302. Based on the associated time, F0 parameters and cepstrum of the phonemes, perform time information encoding on the graph G, and classify the graph G=(V, E) into graph G1=(V, E1) and graph G2=(V, E2) through the heterogeneity of the relationships, where E1 and E2 represent two types of relationships between the nodes.

[0048] In this embodiment, the heterogeneous relationship graph, as a self - contained architecture, is used to structurally represent the heterogeneous relationships among syllables, words, and phonemes in the pronunciation of a given text. This form allows linguistic information to be encoded as a graph G=(V, E), where V contains a set of phonemes, syllables, words, and even phrase boundaries. The edge set E represents the relationships between nodes (i.e., allows interesting relationships to be represented). For example, a hierarchical tree decomposes a word into phonemes (as shown in Figure 9 ). On the other hand, a multi - linear list allows associations to be established between a series of intonation sounds and the corresponding syllables. In addition to features derived solely from the vocabulary, the heterogeneous relationship graph can also contain features obtained after Hidden Markov Model (HMM) alignment between phonemes and speech frames. Specifically, the heterogeneous relationship graph can encode time information: the time (in milliseconds) associated with each phoneme, as well as the F0 parameter and the cepstrum (stored as a multi - linear list). Finally, the heterogeneity of the relationships is exploited by having the same set of nodes V be part of different views / graphs G1=(V, E1), G2=(V, E2). For example, a syllable node is part of both the hierarchical structure (E1) between words and phonemes and the metric structure (E2) that explains the degree of stress on the syllable subtree.

[0049] Specifically, the fundamental frequency parameter refers to the lowest frequency component in a sound signal and is a key factor determining the pitch of a sound; the cepstrum is a representation method extracted from the spectrum (the frequency distribution of a sound), which reflects the characteristic structure of the sound and is commonly used in sound signal processing.

[0050] In one embodiment, as shown in Figure 4 , step S203 includes:

[0051] S401. Randomly sample each node in the heterogeneous relationship graph from a normal distribution to obtain the initial node features of each node;

[0052] S402. In each layer of the graph convolutional network, aggregate the initial node features of each node with the features of its neighbor nodes to obtain new features for each node;

[0053] S403. Normalize the new features of each node to balance the influence of the degrees of different nodes on the aggregation process.

[0054] In this embodiment, through the use of graph convolutional network technology, a series of processing steps are performed on each node in the aforementioned heterogeneous relationship graph. First, each node is initialized, which means assigning an initial node feature vector to each node. These vectors can be predefined or randomly generated. Next, a feature aggregation operation is carried out. This step involves integrating the feature information of the neighbor nodes of each node into the current node. Finally, to ensure the stability and consistency of the features, we perform normalization processing on the features of each node. Through this series of operations, the final node features of each node can be effectively obtained. These features not only contain the initial information of the node itself but also incorporate the information of its neighbor nodes, thus providing rich feature representations for subsequent graph analysis and tasks.

[0055] Specifically, the process of steps S401 - S403 is as follows:

[0056] Use a graph convolutional network (GCN) to learn rich node representations from the heterogeneous relationship graph. GCN can be used to learn node - level or graph - level representations (features) and is applicable to tasks such as node classification and graph classification. Our model architecture consists of L layers of GCN. For a graph G=(V, E), the initial node feature of each node v∈V is randomly initialized (seed embedding) from a normal distribution N(0, 0.3). Each layer refines the node features by aggregating the information of its neighbor nodes:

[0057]

[0058] where σ represents the activation function (such as ReLU), represents the feature representation of node v at the l - th layer, N(v) is the set of neighbor nodes of node v, c u,u represents the normalization constant, usually the reciprocal of the degree of node v, and W (l) is the weight matrix of the l - th layer. Among them, the normalization formula is:

[0059] In one embodiment, please refer to Figure 5 shown, step S204 includes:

[0060] S501. Input the node features into the decoder module of the TTS model to generate a voice spectrogram;

[0061] S502. Input the voice spectrogram into a preset vocoder module to convert it into voice waveform data;

[0062] S503. Post - process the voice waveform data to eliminate noise and adjust waveform parameters to obtain the final voice waveform output.

[0063] In this embodiment, a noise cancellation algorithm can be used to process the initial speech waveform data to obtain denoised speech waveform data; according to a preset waveform parameter adjustment rule, the denoised speech waveform data is optimized in parameters to generate optimized speech waveform data; if the signal-to-noise ratio of the optimized speech waveform data is lower than a preset threshold, the waveform parameters are readjusted; the optimized speech waveform data is input into a waveform output module to generate a final speech waveform output; through the joint optimization of a spectrum generation module and a waveform conversion module, the conversion accuracy between the speech spectrogram and the speech waveform data is ensured.

[0064] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0065] In one embodiment, an end-to-end speech synthesis device based on a heterogeneous relationship graph is provided, and the end-to-end speech synthesis device based on the heterogeneous relationship graph corresponds one-to-one to the end-to-end speech synthesis method based on the heterogeneous relationship graph in the above embodiment. As Figure 6 shown, the end-to-end speech synthesis device based on the heterogeneous relationship graph includes an information extraction module 601, an encoding module 602, a feature extraction module 603, and a speech synthesis module 604. The detailed description of each functional module is as follows:

[0066] The information extraction module 601 is configured to receive a given text and extract linguistic information in the given text;

[0067] The encoding module 602 is configured to encode the linguistic information to obtain a corresponding heterogeneous relationship graph;

[0068] The feature extraction module 603 is configured to perform initialization, feature aggregation, and normalization processing on each node in the heterogeneous relationship graph by using a graph convolutional network to obtain the node feature of each node;

[0069] The speech synthesis module 604 is configured to input the node feature into an end-to-end TTS model to generate a corresponding speech waveform.

[0070] In one embodiment, the information extraction module 601 is specifically configured to:

[0071] Analyze and process the given text by using a deterministic rule predefined in the Festvox system to extract the linguistic information in the given text.

[0072] In one embodiment, the encoding module 602 is specifically configured to:

[0073] Encode the linguistic information into a graph G=(V, E), where V includes a node set of words, syllables, and phonemes, and E includes the relationships between the nodes;

[0074] Encode the time information of graph G based on the associated time, F0 parameters, and cepstrum of phonemes, and classify the graph G=(V, E) into graph G1=(V, E1) and graph G2=(V, E2) through the heterogeneity of the relationships, where E1 and E2 represent two relationships between nodes and nodes.

[0075] In one embodiment, the feature extraction module 603 is further configured to:

[0076] Enable each node in the heterogeneous relationship graph to perform random sampling from a normal distribution to obtain the initial node features of each node;

[0077] In each layer of the graph convolutional network, aggregate the initial node features of each node with the features of its neighbor nodes to obtain the new features of each node;

[0078] Normalize the new features of each node to balance the influence of the degrees of different nodes on the aggregation process.

[0079] In one embodiment, the speech synthesis module 604 is specifically configured to:

[0080] Input the node features into the decoder module of the TTS model to generate a speech spectrogram;

[0081] Input the speech spectrogram into a preset vocoder module to convert it into speech waveform data;

[0082] Perform post-processing on the speech waveform data to eliminate noise and adjust waveform parameters to obtain the final speech waveform output.

[0083] The present invention provides an end-to-end speech synthesis device based on a heterogeneous relationship graph. Inputting deterministic linguistic information in the form of a heterogeneous relationship graph into the TTS model can effectively represent phoneme and syntactic information. In this process, learn the structured representation with the help of a graph convolutional network and apply it to the encoder of the TTS model. Thus, an end-to-end text-to-speech synthesis scheme based on a heterogeneous relationship graph is realized, which has the advantage of achieving significant performance improvement by virtue of syntactic advantages.

[0084] For the specific limitations of the end-to-end speech synthesis device based on a heterogeneous relationship graph, reference can be made to the limitations of the end-to-end speech synthesis method based on a heterogeneous relationship graph in the above text, which will not be elaborated here. Each module in the above end-to-end speech synthesis device based on a heterogeneous relationship graph can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0085] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in Figure 7 Figure 1. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of an end-to-end speech synthesis method based on a heterogeneous relationship graph.

[0086] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as shown in Figure 8 Figure 2. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the client side of an end-to-end speech synthesis method based on a heterogeneous relationship graph.

[0087] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:

[0088] Receive the given text and extract the linguistic information in the given text;

[0089] Encode the linguistic information to obtain a corresponding heterogeneous relationship graph;

[0090] Use a graph convolutional network to perform initialization, feature aggregation, and normalization processing on each node in the heterogeneous relationship graph to obtain the node features of each node;

[0091] Input the node features into an end-to-end TTS model to generate a corresponding speech waveform.

[0092] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0093] Receive a given text and extract the linguistic information in the given text;

[0094] Encode the linguistic information to obtain a corresponding heterogeneous relationship graph;

[0095] Use a graph convolutional network to perform initialization, feature aggregation, and normalization processing on each node in the heterogeneous relationship graph to obtain the node features of each node;

[0096] Input the node features into an end-to-end TTS model to generate corresponding speech waveforms.

[0097] It should be noted that for the functions or steps that can be realized by the above-mentioned computer-readable storage medium or computer device, reference can be made to the relevant descriptions on the server side and the client side in the foregoing method embodiments. To avoid repetition, they will not be described in detail here.

[0098] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0099] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.

[0100] The above-described embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included within the protection scope of the present invention. The non-company software tools or components that appear in the embodiments of this application are only introduced by way of example and do not represent actual use.

Claims

1. An end-to-end speech synthesis method based on a heterogeneous relationship graph, characterized in that: include: Receiving a given text and extracting linguistic information from the given text; Encoding the linguistic information to obtain a corresponding heterogeneous relationship graph; Using a graph convolutional network to initialize, aggregate and normalize each node in the heterogeneous relationship graph to obtain node features of each node; The node features are input into an end-to-end TTS model to generate corresponding speech waveforms.

2. The end-to-end speech synthesis method based on heterogeneous relationship graph according to claim 1, characterized in that: The receiving a given text and extracting linguistic information from the given text comprises: The given text is analyzed and processed using deterministic rules predefined in a system with a speech synthesis function to extract linguistic information from the given text.

3. The end-to-end speech synthesis method based on heterogeneous relationship graph according to claim 1, characterized in that: The encoding of the linguistic information to obtain a corresponding heterogeneous relationship graph includes: Encode the linguistic information into a graph G = (V, E), where V contains a node set of words, syllables and phonemes, and E contains nodes and relationships between nodes; Based on the associated time, F0 parameters and inverse spectrum of the phonemes, the time information of the graph G is encoded, and the graph G = (V, E) is classified into graph G1 = (V, E1) and graph G2 = (V, E2) through the heterogeneity of the relationship, where E1 and E2 represent nodes and two types of relationships between nodes.

4. The end-to-end speech synthesis method based on heterogeneous relationship graph according to claim 1, characterized in that: The graph convolutional network is used to initialize, aggregate features and normalize each node in the heterogeneous relationship graph to obtain node features of each node, including: Randomly sampling each node in the heterogeneous relationship graph from a normal distribution to obtain an initial node feature of each node; In each layer of the graph convolutional network, the initial node features of each node are aggregated with the features of its neighboring nodes to obtain new features of each node; The new features of each node are normalized to balance the impact of the degrees of different nodes on the aggregation process.

5. The end-to-end speech synthesis method based on heterogeneous relationship graph according to claim 4, characterized in that: In each layer of the graph convolutional network, the initial node features of each node are aggregated with the features of its neighboring nodes to obtain new features of each node, including: The aggregation process is performed according to the following formula: Among them, σ represents the activation function, represents the feature representation of node v at layer l, N(v) is the set of neighbor nodes of node v, c u,u represents a normalization constant, usually the inverse of the degree of node v, W (l) is the weight matrix of the lth layer.

6. The end-to-end speech synthesis method based on heterogeneous relationship graph according to claim 5, characterized in that: The normalization formula is:

7. The end-to-end speech synthesis method based on heterogeneous relationship graph according to claim 1, characterized in that: The step of inputting the node features into an end-to-end TTS model to generate a corresponding speech waveform includes: Inputting the node features into a decoder module of a TTS model to generate a speech spectrogram; Inputting the speech spectrogram into a preset vocoder module to convert it into speech waveform data; The speech waveform data is post-processed to remove noise and adjust waveform parameters to obtain the final speech waveform output.

8. An end-to-end speech synthesis device based on a heterogeneous relationship graph, characterized in that: include: An information extraction module, configured to receive a given text and extract linguistic information from the given text; An encoding module, used for encoding the linguistic information to obtain a corresponding heterogeneous relationship graph; A feature extraction module, used to use a graph convolutional network to initialize, aggregate features and normalize each node in the heterogeneous relationship graph to obtain node features of each node; The speech synthesis module is used to input the node features into an end-to-end TTS model to generate a corresponding speech waveform.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the end-to-end speech synthesis method based on a heterogeneous relationship graph as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the end-to-end speech synthesis method based on a heterogeneous relationship graph as described in any one of claims 1 to 7 are implemented.