Speech synthesis method, device, equipment and storage medium

By splitting text, semantic analysis and template speech library matching, the target speech is generated, and the problem of invisible and realistic speech synthesis in the existing technology is solved, and higher quality speech synthesis is achieved.

CN119600990BActive Publication Date: 2025-08-26GUANGZHOU QUWAN NETWORK TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411784364.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-06
Publication Date
2025-08-26
Estimated Expiration
2044-12-06

AI Technical Summary

Technical Problem

When synthesising speech, the existing method of speech synthesis cannot accurately express text information, resulting in the speech being not vivid and realistic enough, and sounding stiff and mechanical.

Method used

By splitting the target text into subtext, semantic analysis is performed, and the pre-trained speech element prediction model and template speech library are used to match the target template speech and generate the target speech.

Benefits of technology

It improves the authenticity and vividness of speech synthesis, so that the generated speech can express text information more accurately and be close to the voice effects issued by real people.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119600990B_ABST
    Figure CN119600990B_ABST
Patent Text Reader

Abstract

The present application discloses a speech synthesis method, apparatus, device and storage medium. The method comprises the following steps: obtaining a target text, splitting it into sub-texts, performing semantic analysis on each sub-text to determine the target semantic vectors of each sub-text, inputting the target semantic vectors of each sub-text into a pre-trained speech element prediction model to obtain the target speech element vectors corresponding to each sub-text, matching the target template speech in a pre-established template speech library with the target semantic vectors and target speech element vectors of the sub-text for each sub-text, and synthesizing the target speech based on the sub-texts, target semantic vectors, target speech element vectors and the target template speech corresponding to each sub-text. The present application determines a plurality of vectors that can enrich speech features and content, and then matches the template speech to obtain a realistic and vivid target speech.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of speech synthesis technology, and in particular to a speech synthesis method, apparatus, device and storage medium. Background Art

[0002] Nowadays, the application of speech synthesis technology is becoming more and more extensive. This technology generates corresponding speech data based on text information. Currently, there are a large number of speech synthesis needs in many fields, and even many studies have been conducted. With the advancement of technology, the requirements for speech synthesis are gradually increasing. Many fields have begun to focus on some realistic features of synthesized speech, such as emotional expression and rhythm, which has led to the emergence of some speech synthesis methods.

[0003] However, existing speech synthesis methods usually only focus on features such as the pronunciation of phonemes and changes in tone when improving realism, but rarely pay attention to the semantic content of the text information itself or the phonetic elements of the speech to be synthesized. This results in the synthesized speech not accurately expressing the text information, lacking in liveliness and authenticity, and sounding stiff and mechanical. Summary of the Invention

[0004] In view of this, the present application provides a speech synthesis method, apparatus, device and storage medium for solving the problem that the speech synthesized by the existing speech synthesis method cannot accurately express the text information, is not vivid and realistic, and sounds stiff and mechanical.

[0005] To achieve the above objectives, the following solutions are proposed:

[0006] In a first aspect, a speech synthesis method comprises:

[0007] Obtaining a target text and splitting the target text into sub-texts;

[0008] Performing semantic analysis on each of the sub-texts to determine target semantic vectors for each of the sub-texts;

[0009] Inputting each target semantic vector of each subtext into a pre-trained speech element prediction model to obtain each target speech element vector corresponding to each subtext; the speech element prediction model is trained using the semantic vectors of the text sample set as training samples and the speech elements of each text sample in the text sample set as sample labels;

[0010] For each of the subtexts, using each target semantic vector and each target speech element vector of the subtext to match a corresponding target template speech in a pre-established template speech library;

[0011] The target speech is synthesized based on each of the sub-texts, the target semantic vector, the target speech element vector and the target template speech corresponding to each of the sub-texts.

[0012] Preferably, the step of using each target semantic vector and each target speech element vector of the subtext to match a corresponding target template speech in a pre-established template speech library includes:

[0013] Using each target semantic vector of the subtext, selecting corresponding template voices from the template voice library as first template voices;

[0014] Determining the voice element labels of each of the first template voices;

[0015] Matching each of the target speech element vectors with each of the speech element labels of each of the first template speech, and calculating a matching degree;

[0016] Determining whether the highest matching degree is a first preset threshold;

[0017] If so, taking the first template speech whose matching degree is the first preset threshold as the target template speech;

[0018] If not, determining whether the number of first template voices corresponding to the highest matching degree is not less than 2;

[0019] If the number of first template voices corresponding to the highest matching degree is not less than 2, the first template voice corresponding to the highest matching degree is used as the second template voice;

[0020] A target template speech is determined from each of the second template speech.

[0021] Preferably, determining the target template speech from each of the second template speech comprises:

[0022] Counting the number of labels of each voice element label of each second template voice, and determining the number of vectors of the target voice pixel vector of the subtext;

[0023] Determine one or more second template voices, wherein the number of labels is smaller than the number of vectors, as respective third template voices;

[0024] The third template speech with the largest number of labels is used as the target template speech.

[0025] Preferably, the step of using each target semantic vector of the subtext to select corresponding template voices from the template voice library as first template voices includes:

[0026] Determining the semantic type of each target semantic vector of the subtext;

[0027] Filtering target semantic vectors whose semantic types are preset scene types and preset emotion types as first semantic vectors;

[0028] Determining each template semantic vector of each template speech;

[0029] taking each template speech containing each first semantic vector in the template semantic vector as each template speech to be determined;

[0030] Target semantic vectors of semantic types other than the scene type and the emotion type are used to screen the template voices to be determined to determine the first template voices.

[0031] Preferably, synthesizing the target speech based on each of the subtexts, the target semantic vector, the target speech element vector, and the target template speech corresponding to each of the subtexts includes:

[0032] Get the timbre to be synthesized;

[0033] determining a text order of each of the subtexts in the target text;

[0034] For each of the subtexts, extracting various acoustic features in the target template speech corresponding to the subtext;

[0035] Generate a sub-speech corresponding to the sub-text by using the timbre to be synthesized and each target semantic vector, target speech element vector, and acoustic feature corresponding to the sub-text;

[0036] The sub-speech corresponding to each of the sub-texts is combined according to the text order to form a target speech.

[0037] Preferably, performing semantic analysis on each of the subtexts to determine target semantic vectors for each of the subtexts includes:

[0038] For each of the sub-texts, determining each word in the sub-text;

[0039] Determine the basic semantics of each word in the subtext;

[0040] Counting the connection relationship between each word and the previous word and / or the next word in the subtext;

[0041] Optimizing the basic semantics of each of the words according to the connection relationship to obtain semantic information of each of the words;

[0042] determining the semantic role of each of the words;

[0043] Based on the semantic information and semantic role of each word, each target semantic vector is generated.

[0044] Preferably, the process of establishing the template speech library includes:

[0045] Set a variety of template sounds;

[0046] Collect multiple real voices with different acoustic elements, different language elements, different background noise elements, and different emotional elements;

[0047] For each of the template timbre, multiple real voices are configured according to the template timbre to obtain multiple template voices;

[0048] A plurality of template voices corresponding to each of the template timbre are aggregated to establish a template voice library.

[0049] In a second aspect, a speech synthesis device includes:

[0050] A splitting module is used to obtain a target text and split the target text into sub-texts;

[0051] A semantic analysis module, configured to perform semantic analysis on each of the sub-texts to determine target semantic vectors of each of the sub-texts;

[0052] a target speech element vector determination module, configured to input each target semantic vector of each subtext into a pre-trained speech element prediction model to obtain each target speech element vector corresponding to each subtext; the speech element prediction model is trained using the semantic vectors of the text sample set as training samples and using each speech element of each text sample in the text sample set as sample labels;

[0053] A matching module is used to match each subtext with a corresponding target template speech in a pre-established template speech library using each target semantic vector and each target speech element vector of the subtext;

[0054] The synthesis module is used to synthesize the target speech based on each of the sub-texts, the target semantic vector, the target speech element vector and the target template speech corresponding to each of the sub-texts.

[0055] In a third aspect, a speech synthesis device includes a memory and a processor;

[0056] The memory is used to store programs;

[0057] The processor is used to execute the program to implement each step of the speech synthesis method as described in any one of the first aspects.

[0058] In a fourth aspect, a storage medium stores a computer program thereon, which, when executed by a processor, implements the various steps of the speech synthesis method as described in any one of the first aspects.

[0059] It can be seen from the above technical solution that the present application obtains the target text and splits the target text to obtain sub-texts; performs semantic analysis on each sub-text to determine the target semantic vectors of each sub-text; inputs the target semantic vectors of each sub-text into a pre-trained speech element prediction model to obtain the target speech element vectors corresponding to each sub-text; the speech element prediction model is trained with the semantic vectors of the text sample set as training samples and the speech elements of each text sample in the text sample set as sample labels; for each sub-text, uses the target semantic vectors and target speech element vectors of the sub-text to match the corresponding target template speech in a pre-established template speech library; synthesizes the target speech based on the sub-texts, target semantic vectors, target speech element vectors and the target template speech corresponding to each sub-text. This application first splits the target text to be synthesized into various sub-texts, and then generates speech for each sub-text separately, which can improve the quality of the final target speech. Then, the sub-text is semantically analyzed to determine the target semantic vector. The target semantic vector can represent the grammatical and structural information of the sub-text. In addition, this application also pre-trains a speech element prediction model. This model can process the target semantic vector, thereby outputting the various target speech elements corresponding to the sub-text, thereby determining multiple vectors that can enrich the speech features and content, and then matching the pre-established template speech to obtain more acoustic features, that is, a real and vivid target speech can be obtained, and its authenticity is completely comparable to the effect of a real person directly speaking. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0061] Figure 1 An optional flow chart of a speech synthesis method provided in an embodiment of the present application;

[0062] Figure 2 A schematic diagram of the structure of a speech synthesis device provided in an embodiment of the present application;

[0063] Figure 3 A schematic diagram of the structure of a speech synthesis device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0064] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0065] The present invention can be used in a variety of general-purpose or special-purpose computing device environments or configurations, such as personal computers, server computers, handheld or portable devices, tablet devices, multi-processor devices, and distributed computing environments including any of the above.

[0066] The embodiment of the present invention provides a speech synthesis method, which can be applied to various computer terminals or smart terminals. The execution subject can be a processor or server of the computer terminal or smart terminal. The method flow chart of the method is as follows: Figure 1 As shown, specifically including:

[0067] S1: Obtain a target text, and split the target text into sub-texts.

[0068] The target text in this application is used to generate the target speech, that is, the text content corresponding to the target speech finally generated is the content of the target text.

[0069] The target text can be any text consisting of sentences or paragraphs, such as news, novels, product manuals, papers, conversations, instructions, and any sentence structure. It can be a combination of declarative sentences, interrogative sentences, exclamatory sentences, and other sentence structures, such as: "Have you had dinner? If not, let's go have dinner together!" It can also be text with formatting or markup, etc., which is not limited in this embodiment.

[0070] After obtaining the target text, the target text is split. The splitting methods include but are not limited to the following:

[0071] 1) Split by sentence, for example, based on spaces or punctuation. Using the example above, you can split the text based on punctuation into "Have you had dinner?", "If not?", and "Let's go have dinner together." This splitting method is particularly convenient for target text with punctuation or spaces. Furthermore, you can further split the text based on vocabulary within each sentence.

[0072] 2) Split by paragraph or chapter structure. In target texts with relatively standardized document formats, such as novels and essays, there are usually obvious paragraph markers, such as blank lines or indentations, to divide paragraphs. In this case, you can split according to this format.

[0073] 3) Split the target text into semantic units, perform semantic understanding on the target text, and extract key elements in the target text, such as time, place, people, actions, etc., to determine the boundaries of the split.

[0074] After splitting, each sub-text is obtained, and then the corresponding speech is generated in the form of sub-text. Compared with the method of directly generating speech from the entire text, this split form of generating multiple speech can highlight each sub-text itself and will not ignore certain small speech features like when generating speech from the entire text, thereby improving the accuracy and precision of the final target speech.

[0075] S2: Perform semantic analysis on each of the sub-texts to determine target semantic vectors for each of the sub-texts.

[0076] Semantic analysis of subtexts can help us understand their true meaning at multiple levels. At the lexical level, semantic analysis can clarify the semantic information of each word in the subtext, analyze the semantic relationships between words, and identify collocational conventions. At the sentence level, it can identify the core events, meaning, and grammatical structure of the subtext. For longer subtexts, it can clarify the subtext's thematic information and main idea. By combining the analysis results at these levels, we can determine the target semantic vector for each subtext, which can then be used to determine the target speech elements.

[0077] Each sub-text can be processed by calling a large language model, which will output the target semantic vector for each sub-text. Alternatively, a separate model can be trained to perform semantic analysis and predict the target semantic vector, such as a language model based on the Transformer architecture. This model can learn from a large amount of text, predict semantic relationships, and generate semantic vectors.

[0078] S3: Input each target semantic vector of each sub-text into a pre-trained speech element prediction model to obtain each target speech element vector corresponding to each sub-text; the speech element prediction model is trained using the semantic vectors of the text sample set as training samples and the speech elements of each text sample in the text sample set as sample labels.

[0079] The text sample set is the basic data set used to train the speech element model. These text samples can be collected from multiple sources and channels, such as a large number of news articles, novels, papers, conversations, etc., so that they can cover text content of various types, styles, and themes. Before training, it is necessary to extract multiple semantic vectors corresponding to each text sample in the text sample set, and use these semantic vectors to train the model, so that the trained model can handle diverse semantic vectors when it is actually applied and can adapt to more application scenarios.

[0080] The target speech element vector involved in this application refers to various features and attributes related to speech, which may include acoustic features, such as pitch, loudness, duration, timbre, etc.; it may also include prosodic features, such as intonation, stress, rhythm, etc.; it may also include features or content of the semantics and emotions to be expressed, so that the final generated target speech can be more realistic, more vivid, and not rigid.

[0081] An autoregressive model based on the Transformer structure can be used. Transformer-structured models have powerful representational learning capabilities when processing natural language tasks and can effectively capture the complex mapping relationship between semantic vectors and speech element vectors. An autoregressive model is a model that generates sequences based on probability distributions. For example, it constructs a subtext sequence based on the order of each subtext in the target text. When predicting each speech element vector corresponding to the semantic vector in the sequence, it is based on previously predicted speech elements. Combining these two models will result in a more powerful speech element prediction model with better ability to predict the target speech element vector.

[0082] Many existing models that can predict speech elements are trained directly with text or speech, which increases the training intensity and is not as accurate as the method of first splitting the text, then extracting semantic vectors, and then predicting speech element vectors. Moreover, since semantic vectors can express the true meaning of the text and the desired emotion more deeply, it is more effective to use them to predict speech element vectors.

[0083] S4: For each of the subtexts, the target semantic vectors and target speech element vectors of the subtext are used to match the corresponding target template speech in a pre-established template speech library.

[0084] In the above steps, the target semantic vector and target speech elements are determined respectively. However, in order to make the target speech more realistic and accurate, the present application also obtains other features or elements required for synthesizing speech, namely acoustic features, through other methods. This is more comprehensive. Therefore, a template speech library is established in advance. The template speech library contains a large number of template speech. Each template speech has or is marked with scenes, speech elements, semantic vectors, acoustic features, language features, background noise, emotions, rhythm, emotions, etc. It is very comprehensive and realistic. These template speech can be real readings, real conversations, etc. According to the target semantic vector and target speech element vector of the sub-text, the template speech that best matches it can be screened to fill the gaps and obtain more comprehensive features, elements and content.

[0085] S5: synthesizing a target speech based on each of the sub-texts, the target semantic vector, the target speech element vector, and the target template speech corresponding to each of the sub-texts.

[0086] In the above steps, various features, elements, and content that can accurately generate the target speech have been obtained. They are expressed in the form of subtext, target semantic vector, target speech element vector, and target template speech. Then, we can use zero-shot speech synthesis capabilities to fuse these elements to generate the target speech.

[0087] In this step, you can use existing large speech models such as NaturalSpeech, SeedTTS, and MaskGCT.

[0088] It can be seen from the above technical solution that the present application obtains the target text and splits the target text to obtain sub-texts; performs semantic analysis on each sub-text to determine the target semantic vectors of each sub-text; inputs the target semantic vectors of each sub-text into a pre-trained speech element prediction model to obtain the target speech element vectors corresponding to each sub-text; the speech element prediction model is trained with the semantic vectors of the text sample set as training samples and the speech elements of each text sample in the text sample set as sample labels; for each sub-text, uses the target semantic vectors and target speech element vectors of the sub-text to match the corresponding target template speech in a pre-established template speech library; synthesizes the target speech based on the sub-texts, target semantic vectors, target speech element vectors and the target template speech corresponding to each sub-text. This application first splits the target text to be synthesized into various sub-texts, and then generates speech for each sub-text separately, which can improve the quality of the final target speech. Then, the sub-text is semantically analyzed to determine the target semantic vector. The target semantic vector can represent the grammatical and structural information of the sub-text. In addition, this application also pre-trains a speech element prediction model. This model can process the target semantic vector, thereby outputting the various target speech elements corresponding to the sub-text, thereby determining multiple vectors that can enrich the speech features and content, and then matching the pre-established template speech to obtain more acoustic features, that is, a real and vivid target speech can be obtained, and its authenticity is completely comparable to the effect of a real person directly speaking.

[0089] In addition, with the development of the self-media industry and the availability of artificial intelligence content generation tools, more and more industries and individuals need to perform speech synthesis, such as anchors, TV stations, hosts, etc. The speech synthesis method provided in this application can achieve highly natural speech synthesis and be applied to speech cloning and synthesis services to meet the speech synthesis needs of anchors, TV stations, hosts and other users in live broadcasts, news, speeches, audio novels and other fields, thereby helping them improve work efficiency, such as shortening the program production cycle and saving resources.

[0090] In the method provided in the embodiment of the present invention, the process of performing semantic analysis on each subtext to determine each target semantic vector of each subtext is specifically described as follows:

[0091] For each of the sub-texts, determining each word in the sub-text;

[0092] Determine the basic semantics of each word in the subtext;

[0093] Counting the connection relationship between each word and the previous word and / or the next word in the subtext;

[0094] Optimizing the basic semantics of each of the words according to the connection relationship to obtain semantic information of each of the words;

[0095] determining the semantic role of each of the words;

[0096] Based on the semantic information and semantic role of each word, each target semantic vector is generated.

[0097] Specifically, when conducting semantic analysis, on the one hand, we first need to determine the various words in the sub-text. Each word has its own meaning, that is, basic semantics. In addition, the position / order of each word in the sub-text is different. Except for the first word and the last word, other words in the middle position have a previous word and a next word, and there is a connection relationship between the previous word and the next word. This connection relationship refers to the semantic relationship. At the same time, considering that a word may have multiple basic semantics, the basic semantics expressed by the same word in different scenarios or emotions are different. Therefore, the basic semantics of each word are optimized according to the connection relationship. Some words themselves only have one basic semantics, so the optimized basic semantics are the same as before optimization, so that the semantic information of each word in the sub-text can be obtained, and this semantic information is consistent with the sub-text.

[0098] On the other hand, in addition to semantic information, words also have the characteristic of semantic roles. Semantic roles refer to lexical components such as subject, predicate, object, and adverbial. For example, a sub-text is "cat eats fish", "cat" is the subject, "eat" is the predicate, and "fish" is the object. More specifically, for the subject "cat", the predicate "eat" clearly expresses the ongoing behavior of "cat", so the predicate is closely connected with the subject and is the action performed by the subject. Through this action, the subject and the object establish a connection. Therefore, in a sub-text, two or more words can establish a connection through their own semantic roles. Therefore, based on the semantic information and semantic roles of the words, the target semantic vector can be generated. The generation method can be splicing, or the semantic information and semantic roles can be encoded separately, and then weighted summed after encoding and expressed in the form of a vector.

[0099] Representing the semantic content of the subtext in the form of a target semantic vector can reduce the amount of data and make the subsequent speech synthesis process more concise and efficient.

[0100] In order to make the final synthesized target speech more realistic and vivid, this application not only utilizes the target semantic vector and target speech element vector in the generation process, but also pre-builds a template speech library. The purpose is to match a corresponding target template speech in the template speech library, thereby extracting some features and content of the target template speech to enrich the target speech, improve the freedom and flexibility of speech synthesis, and enhance the user experience. The establishment process includes:

[0101] Set a variety of template sounds;

[0102] Collect multiple real voices with different acoustic elements, different language elements, different background noise elements, and different emotional elements;

[0103] For each of the template timbre, multiple real voices are configured according to the template timbre to obtain multiple template voices;

[0104] A plurality of template voices corresponding to each of the template timbre are aggregated to establish a template voice library.

[0105] Specifically, the template timbre is set to unify the template voices in the template voice library, making it easier to manage the template voice library and facilitate the matching process. When collecting voices, no restrictions are placed on any elements. Multiple real voices with different acoustic elements, different language elements, different background noise elements, and different emotional elements are collected. This allows us to cover a variety of scenarios, and these real voices are actually human voices.

[0106] It should be noted that since there are no restrictions on any factors when collecting real voices, after the collection is completed, each template voice must be configured separately to obtain multiple template voices. The purpose of setting up multiple template voices and configuring them separately is to enrich the template voices and to take into account that people of different age groups and genders have different timbre characteristics. Configuring each template voice separately can improve authenticity. For example, for men, four template voices are set according to age groups: teenagers, young people, middle-aged people, and old people. At the same time, for women, four template voices are set according to age groups: teenagers, young people, middle-aged people, and old people. Thus, eight template voices are obtained. After summarizing them, the library is established, and a template voice library with rich scenes, rich elements, and rich emotions is obtained.

[0107] In addition, the authenticity of speech is inseparable from the addition of background noise. Therefore, when establishing the template speech library, real speech with different background noise elements is specially collected. Among them, background noise includes many types, such as traffic noise, machine noise, human voice noise, etc.

[0108] Next, the process of matching the corresponding target template speech in the pre-established template speech library using each target semantic vector and each target speech element vector of the subtext in the present application is described in detail.

[0109] The process of matching the target template speech in this application can be roughly divided into two steps. First, multiple template speech are matched according to the target semantic vector. Then, the most suitable and matching template speech is accurately found as the target template speech according to the target speech element vector. Since each subtext needs to be matched with the corresponding target template speech, the following description is based on a subtext:

[0110] 1) First, using each target semantic vector of the subtext, the corresponding template speech is screened out from the template speech library as each first template speech. The steps include:

[0111] Determining the semantic type of each target semantic vector of the subtext;

[0112] Filtering target semantic vectors whose semantic types are preset scene types and preset emotion types as first semantic vectors;

[0113] Determining each template semantic vector of each template speech;

[0114] taking each template speech containing each first semantic vector in the template semantic vector as each template speech to be determined;

[0115] Target semantic vectors of semantic types other than the scene type and the emotion type are used to screen the template voices to be determined to determine the first template voices.

[0116] Specifically, it is understood that semantic vectors (i.e., target semantic vectors and template semantic vectors) contain multiple semantic types, and different semantic types represent different semantic meanings, such as scenario type, emotion type, intent type, style type, etc., which are not limited in this embodiment. Scenario type refers to the scenario in which the subtext is used, such as a product purchase scenario, a learning scenario, a teaching scenario, or a play scenario; emotion type refers to the mood of the corresponding user in the subtext, such as happiness, frustration, excitement, etc.; intent type refers to the purpose or intention expressed by the subtext, such as inquiry intent, purchase intent, etc.; style type refers to the language style expressed by the subtext, such as formality, casualness, etc.

[0117] Taking into account that scene type and emotion type are very important for speech, during screening, screening is first performed according to these two semantic types. Since template speech also has semantic vectors, this application calls them template semantic vectors. Therefore, the template speech with template semantic vectors of scene type and emotion type is screened out as the template speech to be determined, and then the target semantic vectors of other semantic types are used for screening, that is, the target semantic vectors of other semantic types except scene type and emotion type are used to screen in each template speech to be determined. In order to improve the screening efficiency, the target semantic vectors of other semantic types may not have corresponding template semantic types. Therefore, when screening the template speech to be determined, the similarity between the target semantic vectors of other semantic types and the target semantic vectors of other semantic types can be calculated. If the similarity is greater than the preset threshold, it is considered to be corresponding.

[0118] 2) After determining each first template speech, further determining each speech element label marked by each first template speech using the target speech element vector;

[0119] Matching each of the target speech element vectors with each of the speech element labels of each of the first template speech, and calculating a matching degree;

[0120] Determining whether the highest matching degree is a first preset threshold;

[0121] If so, taking the first template speech whose matching degree is the first preset threshold as the target template speech;

[0122] If not, determining whether the number of first template voices corresponding to the highest matching degree is not less than 2;

[0123] If the number of first template voices corresponding to the highest matching degree is not less than 2, the first template voice corresponding to the highest matching degree is used as the second template voice;

[0124] A target template speech is determined from each of the second template speech.

[0125] Specifically, in addition to the template semantic vector, the template speech is also marked with a speech element label. The speech element label of the template speech and the target speech element vector of the sub-text are two different forms of expression for the speech element, so there will be a certain matching relationship between the two. Therefore, each target speech element is matched with each speech element label of each first template speech, and the matching degree is calculated. The template speech that is more suitable for the sub-text is determined as the target template speech based on the size of the matching degree. The present application sets the first preset threshold value to 100%, that is, there may be a completely matched template speech, so first determine whether the highest matching degree is 100%. If so, it means a complete match, and the corresponding first template speech can be directly determined as the target template speech; if the highest matching degree is not 100%, that is, less than 100%, it means a partial match, but it should be noted that since a sub-text may have multiple target speech element vectors, the template speech may also have multiple speech element labels, and the number may be different, so there will be two situations, namely, the number of target speech element vectors of a sub-text is greater than the number of speech element labels of a template speech, or, the number of target speech element vectors of a sub-text is less than the number of speech element labels of a template speech. For the template speech, the more speech element labels there are, the more specific and detailed the template speech is.

[0126] In one example, a subtext has 4 target speech element vectors, one first template speech a has 5 speech element labels, and another first template speech b has 3 speech element labels. These two first template speech vectors have the same degree of match with the subtext. If b is selected as the target template speech, there may be a deviation from the subtext. Selecting a as the target template speech is optimal. Therefore, if the highest degree of match is not 100%, further confirmation is required, including:

[0127] Counting the number of labels of each voice element label of each second template voice, and determining the number of vectors of the target voice pixel vector of the subtext;

[0128] Determine one or more second template voices, wherein the number of labels is smaller than the number of vectors, as respective third template voices;

[0129] The third template speech with the largest number of labels is used as the target template speech.

[0130] Another case is that the highest match degree is not 100%, but there is only one corresponding first template speech. In this case, to ensure accuracy, the number of labels and the number of vectors are determined. If the number of labels is greater than the number of vectors, the first template speech is removed and the highest match degree among the remaining first template speech is recalculated.

[0131] Furthermore, the process of synthesizing the target speech based on each of the subtexts, the target semantic vector, the target speech element vector, and the target template speech corresponding to each of the subtexts includes the following specific steps:

[0132] Get the timbre to be synthesized;

[0133] determining a text order of each of the subtexts in the target text;

[0134] For each of the subtexts, extracting various acoustic features in the target template speech corresponding to the subtext;

[0135] Generate a sub-speech corresponding to the sub-text by using the timbre to be synthesized and each target semantic vector, target speech element vector, and acoustic feature corresponding to the sub-text;

[0136] The sub-speech corresponding to each of the sub-texts is combined according to the text order to form a target speech.

[0137] Specifically, the timbre to be synthesized can be freely set, such as the timbre of a user using the speech synthesis method provided in this application, or the timbre can be extracted from any prompt voice. Prompt voice is an existing voice content that exists as a guide or prompt information in voice-related tasks or application scenarios; the timbre of the target template voice can also be used. The text order of the subtext in the target text also refers to the text position. For example, if a target text is "The weather is nice today, let's go for a walk together", its corresponding subtexts include: "The weather is nice today" and "Let's go for a walk together". Then the text order corresponding to "The weather is nice today" is 1, and the text order corresponding to "Let's go for a walk together" is 2. In this way, the context order of each subtext in the target text can be clearly defined.

[0138] Next, based on the timbre to be synthesized and the target semantic vectors, target speech element vectors and acoustic features corresponding to the sub-text, a sub-speech corresponding to the sub-text is generated, and then the sub-texts are spliced, combined and summarized according to the text order of the sub-texts in the target text to form the target speech.

[0139] and Figure 1 Corresponding to the method described above, the embodiment of the present invention further provides a speech synthesis device for Figure 1 The specific implementation of the method in the embodiment of the present invention is that the speech synthesis device provided by the embodiment of the present invention can be used in a computer terminal or various mobile devices, combined with Figure 2 , introduce the speech synthesis device, such as Figure 2 As shown, the device may include:

[0140] A splitting module 10 is used to obtain a target text and split the target text into sub-texts;

[0141] A semantic analysis module 20, configured to perform semantic analysis on each of the sub-texts to determine target semantic vectors of each of the sub-texts;

[0142] The target speech element vector determination module 30 is configured to input the target semantic vectors of each subtext into a pre-trained speech element prediction model to obtain target speech element vectors corresponding to each subtext; the speech element prediction model is trained using the semantic vectors of the text sample set as training samples and the speech elements of each text sample in the text sample set as sample labels;

[0143] A matching module 40 is configured to match each subtext with a corresponding target template speech in a pre-established template speech library using each target semantic vector and each target speech element vector of the subtext;

[0144] The synthesis module 50 is configured to synthesize a target speech based on each of the sub-texts, the target semantic vector, the target speech element vector, and the target template speech corresponding to each of the sub-texts.

[0145] It can be seen from the above technical solution that the present application obtains the target text and splits the target text to obtain sub-texts; performs semantic analysis on each sub-text to determine the target semantic vectors of each sub-text; inputs the target semantic vectors of each sub-text into a pre-trained speech element prediction model to obtain the target speech element vectors corresponding to each sub-text; the speech element prediction model is trained with the semantic vectors of the text sample set as training samples and the speech elements of each text sample in the text sample set as sample labels; for each sub-text, uses the target semantic vectors and target speech element vectors of the sub-text to match the corresponding target template speech in a pre-established template speech library; synthesizes the target speech based on the sub-texts, target semantic vectors, target speech element vectors and the target template speech corresponding to each sub-text. This application first splits the target text to be synthesized into various sub-texts, and then generates speech for each sub-text separately, which can improve the quality of the final target speech. Then, the sub-text is semantically analyzed to determine the target semantic vector. The target semantic vector can represent the grammatical and structural information of the sub-text. In addition, this application also pre-trains a speech element prediction model. This model can process the target semantic vector, thereby outputting the various target speech elements corresponding to the sub-text, thereby determining multiple vectors that can enrich the speech features and content, and then matching the pre-established template speech to obtain more acoustic features, that is, a real and vivid target speech can be obtained, and its authenticity is completely comparable to the effect of a real person directly speaking.

[0146] Furthermore, the embodiment of the present application provides a speech synthesis device. Optionally, Figure 3 The hardware structure diagram of the speech synthesis device is shown in FIG. Figure 3 The hardware structure of the speech synthesis device may include: at least one processor 01, at least one communication interface 02, at least one memory 03 and at least one communication bus 04.

[0147] In the embodiment of the present application, the number of the processor 01 , the communication interface 02 , the memory 03 , and the communication bus 04 is at least one, and the processor 01 , the communication interface 02 , and the memory 03 communicate with each other through the communication bus 04 .

[0148] The processor 01 may be a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention.

[0149] The memory 03 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), for example, at least one disk memory.

[0150] The memory stores a program, and the processor can call the program stored in the memory, and the program is used to execute the following speech synthesis method, including:

[0151] Obtaining a target text and splitting the target text into sub-texts;

[0152] Performing semantic analysis on each of the sub-texts to determine target semantic vectors for each of the sub-texts;

[0153] Inputting each target semantic vector of each subtext into a pre-trained speech element prediction model to obtain each target speech element vector corresponding to each subtext; the speech element prediction model is trained using the semantic vectors of the text sample set as training samples and the speech elements of each text sample in the text sample set as sample labels;

[0154] For each of the subtexts, using each target semantic vector and each target speech element vector of the subtext to match a corresponding target template speech in a pre-established template speech library;

[0155] The target speech is synthesized based on each of the sub-texts, the target semantic vector, the target speech element vector and the target template speech corresponding to each of the sub-texts.

[0156] Optionally, the detailed functions and extended functions of the program may refer to the description of the speech synthesis method in the method embodiment.

[0157] The present application also provides a storage medium that can store a program suitable for execution by a processor. When the program is executed, the device where the storage medium is located is controlled to perform the following speech synthesis method, including:

[0158] Obtaining a target text and splitting the target text into sub-texts;

[0159] Performing semantic analysis on each of the sub-texts to determine target semantic vectors for each of the sub-texts;

[0160] Inputting each target semantic vector of each subtext into a pre-trained speech element prediction model to obtain each target speech element vector corresponding to each subtext; the speech element prediction model is trained using the semantic vectors of the text sample set as training samples and the speech elements of each text sample in the text sample set as sample labels;

[0161] For each of the subtexts, using each target semantic vector and each target speech element vector of the subtext to match a corresponding target template speech in a pre-established template speech library;

[0162] The target speech is synthesized based on each of the sub-texts, the target semantic vector, the target speech element vector and the target template speech corresponding to each of the sub-texts.

[0163] Specifically, the storage medium may be a computer-readable storage medium, and the computer-readable storage medium may be an electronic memory such as a flash memory, an EEPROM (Electrically Erasable Programmable Read-Only Memory), an EPROM, a hard disk or a ROM.

[0164] Optionally, the detailed functions and extended functions of the program may refer to the description of the speech synthesis method in the method embodiment.

[0165] In addition, the functional modules in the various embodiments of the present disclosure can be integrated together to form an independent part, or each module can exist separately, or two or more modules can be integrated to form an independent part. If the function is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for a computer device (which can be a personal computer, a live broadcast device, or a network device, etc.) to perform all or part of the steps of the methods of the various embodiments of the present disclosure.

[0166] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.

[0167] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0168] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A speech synthesis method, characterized in that: include: Obtaining a target text and splitting the target text into sub-texts; Performing semantic analysis on each of the sub-texts to determine target semantic vectors for each of the sub-texts; Inputting each target semantic vector of each subtext into a pre-trained speech element prediction model to obtain each target speech element vector corresponding to each subtext; the speech element prediction model is trained using the semantic vectors of the text sample set as training samples and the speech elements of each text sample in the text sample set as sample labels; For each of the sub-texts, the target semantic vectors and target speech element vectors of the sub-text are used to match the corresponding target template speech in a pre-established template speech library; including: using the target semantic vectors of the sub-text to screen out the corresponding template speech in the template speech library as each first template speech; respectively determining the speech element labels marked by each of the first template speech; matching each of the target speech element vectors with each of the speech element labels of the first template speech, and calculating the matching degree; determining whether the highest matching degree is a first preset threshold; if so, taking the first template speech with the matching degree of the first preset threshold as the target template speech; if not, judging whether the number of the first template speech corresponding to the highest matching degree is not less than 2; if the number of the first template speech corresponding to the highest matching degree is not less than 2, taking the first template speech corresponding to the highest matching degree as the second template speech; determining the target template speech from each of the second template speech; The target speech is synthesized based on each of the sub-texts, the target semantic vector, the target speech element vector and the target template speech corresponding to each of the sub-texts.

2. The method according to claim 1, characterized in that The determining of a target template speech from each of the second template speech comprises: Counting the number of labels of each voice element label of each second template voice, and determining the number of vectors of the target voice pixel vector of the subtext; Determine one or more second template voices, wherein the number of labels is smaller than the number of vectors, as respective third template voices; The third template speech with the largest number of labels is used as the target template speech.

3. The method according to any one of claims 1 or 2, characterized in that The method of selecting corresponding template voices from the template voice library using the target semantic vectors of the subtext as first template voices includes: Determining the semantic type of each target semantic vector of the subtext; Filtering target semantic vectors whose semantic types are preset scene types and preset emotion types as first semantic vectors; Determining each template semantic vector of each template speech; taking each template speech containing each first semantic vector in the template semantic vector as each template speech to be determined; Target semantic vectors of semantic types other than the scene type and the emotion type are used to screen the template voices to be determined to determine the first template voices.

4. The method according to any one of claims 1 to 2, characterized in that The synthesizing the target speech based on each of the subtexts, the target semantic vector, the target speech element vector, and the target template speech corresponding to each of the subtexts includes: Get the timbre to be synthesized; determining a text order of each of the subtexts in the target text; For each of the subtexts, extracting various acoustic features in the target template speech corresponding to the subtext; Generate a sub-speech corresponding to the sub-text by using the timbre to be synthesized and each target semantic vector, target speech element vector, and acoustic feature corresponding to the sub-text; The sub-speech corresponding to each of the sub-texts is combined according to the text order to form a target speech.

5. The method according to any one of claims 1 to 2, characterized in that The performing semantic analysis on each of the sub-texts to determine target semantic vectors of each of the sub-texts includes: For each of the sub-texts, determining each word in the sub-text; Determine the basic semantics of each word in the subtext; Counting the connection relationship between each word and the previous word and / or the next word in the subtext; Optimizing the basic semantics of each of the words according to the connection relationship to obtain semantic information of each of the words; determining the semantic role of each of the words; Based on the semantic information and semantic role of each word, each target semantic vector is generated.

6. The method according to any one of claims 1 to 2, characterized in that: The process of establishing the template speech library includes: Set a variety of template sounds; Collect multiple real voices with different acoustic elements, different language elements, different background noise elements, and different emotional elements; For each of the template timbre, multiple real voices are configured according to the template timbre to obtain multiple template voices; A plurality of template voices corresponding to each of the template timbre are aggregated to establish a template voice library.

7. A speech synthesis device, characterized in that: include: A splitting module is used to obtain a target text and split the target text into sub-texts; A semantic analysis module, configured to perform semantic analysis on each of the sub-texts to determine target semantic vectors of each of the sub-texts; a target speech element vector determination module, configured to input each target semantic vector of each subtext into a pre-trained speech element prediction model to obtain each target speech element vector corresponding to each subtext; the speech element prediction model is trained using the semantic vectors of the text sample set as training samples and using each speech element of each text sample in the text sample set as sample labels; A matching module is used to match the corresponding target template speech in a pre-established template speech library using the target semantic vectors and target speech element vectors of the subtext for each subtext; including: using the target semantic vectors of the subtext to screen out the corresponding template speech in the template speech library as each first template speech; respectively determining the speech element labels marked by each first template speech; matching each target speech element vector with each speech element label of each first template speech, and calculating the matching degree; determining whether the highest matching degree is a first preset threshold; if so, taking the first template speech with the matching degree of the first preset threshold as the target template speech; if not, determining whether the number of the first template speech corresponding to the highest matching degree is not less than 2; if the number of the first template speech corresponding to the highest matching degree is not less than 2, taking the first template speech corresponding to the highest matching degree as the second template speech; determining the target template speech from each second template speech; The synthesis module is used to synthesize the target speech based on each of the sub-texts, the target semantic vector, the target speech element vector and the target template speech corresponding to each of the sub-texts.

8. A speech synthesis device, characterized in that: including memory and processor; The memory is used to store programs; The processor is used to execute the program to implement the various steps of the speech synthesis method according to any one of claims 1 to 6.

9. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, each step of the speech synthesis method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Voice dialogue processing method and system

    CN111862977A

  • Voice synthesis method and device, server and storage medium

    CN113096634A