Speech synthesis method and related apparatus, electronic device and storage medium

By performing colloquialization conversion and phoneme sequence extraction on the text to be synthesized, and combining it with a speech synthesis model to generate colloquial speech, the problem of speech synthesis in existing technologies not conforming to everyday spoken language is solved, thus improving the user interaction experience.

CN114299911BActive Publication Date: 2025-12-19UNIV OF SCI & TECH OF CHINA +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111630204.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-28
Publication Date
2025-12-19
Estimated Expiration
2041-12-28

AI Technical Summary

Technical Problem

Existing speech synthesis technology struggles to generate speech that conforms to everyday spoken communication, resulting in a poor user experience.

Method used

The text to be synthesized is converted into spoken language, phoneme sequences are extracted and spoken language control tags are predicted, and spoken language speech is generated by combining the speech synthesis model. The model includes a spoken language conversion module, a phoneme extraction module and a tag prediction module, and the pronunciation state is controlled by referring to at least one conversion mode and spoken language control tags.

Benefits of technology

It achieves dual optimization at both the text and acoustic levels, generating speech that is more in line with colloquial expressions and improving the user interaction experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114299911B_ABST
    Figure CN114299911B_ABST
Patent Text Reader

Abstract

The application discloses a speech synthesis method and related device, electronic equipment and storage medium, wherein the speech synthesis method comprises: converting a text to be synthesized into a colloquial text; wherein the conversion refers to at least one conversion mode; extracting a phoneme sequence of the colloquial text and predicting a colloquial control label of the colloquial text; wherein the colloquial control label is used to control a pronunciation state; and synthesizing a colloquial speech of the text to be synthesized based on the phoneme sequence and the colloquial control label. The above scheme can realize colloquial speech synthesis, thereby improving user interaction experience.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of speech synthesis, in particular to a speech synthesis method and related device, electronic equipment and storage medium. BACKGROUND

[0002] In recent years, speech synthesis technology is widely used in intelligent customer service, voice assistants, novel reading, virtual characters and other directions. For example, a voice assistant can use speech recognition technology to recognize the user's voice, then use natural language processing technology to generate an answer text, and finally generate fluent speech for communication with the user through speech synthesis technology according to the text.

[0003] At present, the speech synthesized by the speech synthesis technology is formal fluent language. However, in the process of daily communication, there are usually various informal colloquial phenomena. Therefore, compared with the daily colloquial communication between people, the formal fluent language synthesized by the speech synthesis is difficult to give users a good interactive experience because it does not conform to the colloquial expression. Therefore, how to realize colloquial speech synthesis to improve the user interactive experience has become a problem to be solved. SUMMARY

[0004] The technical problem solved by the present application is to provide a speech synthesis method and related device, electronic equipment and storage medium, which can realize colloquial speech synthesis to improve the user interactive experience.

[0005] In order to solve the above technical problem, the first aspect of the present application provides a speech synthesis method, comprising: converting a to-be-synthesized text into colloquial language to obtain a colloquial text; wherein the colloquial conversion refers to at least one conversion mode; extracting a phoneme sequence of the colloquial text and predicting a colloquial control label of the colloquial text; wherein the colloquial control label is used to control the pronunciation state; and synthesizing the colloquial speech of the to-be-synthesized text based on the phoneme sequence and the colloquial control label.

[0006] In order to solve the above technical problem, the second aspect of the present application provides a speech synthesis device, comprising: a colloquial conversion module, a phoneme extraction module, a label prediction module and a sound synthesis module, the colloquial conversion module is used to convert a to-be-synthesized text into colloquial language to obtain a colloquial text; wherein the colloquial conversion refers to at least one conversion mode; the phoneme extraction module is used to extract a phoneme sequence of the colloquial text; the label prediction module is used to predict a colloquial control label of the colloquial text; wherein the colloquial control label is used to control the pronunciation state; and the sound synthesis module is used to synthesize the colloquial speech of the to-be-synthesized text based on the phoneme sequence and the colloquial control label.

[0007] To solve the above technical problems, the third aspect of the present application provides an electronic device, comprising a memory and a processor coupled with each other, the memory stores program instructions, and the processor is configured to execute the program instructions to implement the speech synthesis method in the first aspect.

[0008] To solve the above technical problems, the fourth aspect of the present application provides a computer readable storage medium, which stores program instructions capable of being executed by a processor, and the program instructions are used to implement the speech synthesis method in the first aspect.

[0009] The above scheme converts the text to be synthesized into a colloquial text, and the colloquial text refers to at least one conversion mode, and extracts the phoneme sequence of the colloquial text, and predicts the colloquial control label of the colloquial text, and the colloquial control label is used to control the pronunciation state, and on this basis, the colloquial speech of the text to be checked is synthesized based on the phoneme sequence and the colloquial control label. On the one hand, the text to be synthesized is converted into a colloquial text by referring to at least one conversion mode, which is beneficial to make the colloquial text as much as possible to meet the colloquial expression, and on the other hand, the colloquial control label of the colloquial text is predicted on this basis, which can further provide reference for colloquial speech synthesis from the acoustic level on the basis of the foregoing text level. Therefore, the colloquial speech synthesis can be realized from two different levels of text level and acoustic level, so as to improve the user interaction experience. BRIEF DESCRIPTION OF DRAWINGS

[0010] Figure 1 is a flowchart of an embodiment of the speech synthesis method of the present application;

[0011] Figure 2 is a process diagram of an embodiment of obtaining sample data;

[0012] Figure 3 is a framework diagram of an embodiment of the colloquial prediction network;

[0013] Figure 4 is a framework diagram of another embodiment of the colloquial prediction network;

[0014] Figure 5 is a process diagram of an embodiment of obtaining the first label;

[0015] Figure 6 is a process diagram of an embodiment of obtaining the phoneme category;

[0016] Figure 7 is a process diagram of an embodiment of obtaining the phoneme label;

[0017] Figure 8 is a process diagram of an embodiment of obtaining the emotional feature representation;

[0018] Figure 9 is a framework schematic diagram of an embodiment of the voice emotion network;

[0019] Figure 10 is a process schematic diagram of an embodiment of the voice synthesis method of the present application;

[0020] Figure 11 is a framework schematic diagram of an embodiment of the voice synthesis device of the present application;

[0021] Figure 12 is a framework schematic diagram of an embodiment of the electronic device of the present application;

[0022] Figure 13 is a framework schematic diagram of an embodiment of the computer readable storage medium of the present application. DETAILED DESCRIPTION

[0023] The scheme of the embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0024] In the following description, specific details are set forth in order to provide a thorough understanding of the present application. The present application may, however, be practiced without some or all of these details. Otherwise, well known structures have not been described in detail in order not to obscure the understanding of this description.

[0025] The terms "system" and "network" are often used interchangeably herein. The term "and / or" herein merely describes an associated relationship, which means that there can be three relationships, for example, A and / or B, which means that A exists alone, A and B exist together, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects. In addition, "multiple" herein means two or more than two.

[0026] Please refer to Figure 1 , Figure 1 is a flow schematic diagram of an embodiment of the voice synthesis method of the present application. Specifically, it can include the following steps:

[0027] Step S11: converting the text to be synthesized into colloquial, obtaining colloquial text.

[0028] In the embodiments of the present disclosure, the oralization conversion refers to at least one conversion mode. For example, the conversion mode can include, but is not limited to, adding a modal particle, adding a catchphrase, adding a repeated sentence, correcting content, and reversing the order of a sentence. For example, a modal particle "um" can be added at the beginning of the text to be synthesized; and / or a catchphrase "that" can be added at the beginning of the text to be synthesized; and / or a part of the text to be synthesized can be repeated, such as the word "understand" (i.e., repeated as "understand understand") and the word "good" (i.e., repeated as "good good"); and / or the relevant expression in the text to be synthesized can be corrected, such as the incorrect expression (e.g., "I can go there tomorrow, no, the day after tomorrow") in the text to be synthesized; and / or the order of the expression (e.g., "You are mistaking the subject and the object") can be reversed. Other cases can be similarly deduced, and will not be listed one by one here.

[0029] In one implementation scenario, in order to improve the efficiency of speech synthesis, a speech synthesis model can be pre-trained, and the speech synthesis model can include an oralization prediction network, so that the oralization prediction network can be used to predict the text to be synthesized to obtain an oralization text. In addition, the oralization model can be trained by using a plurality of sample texts, and the second sample text is obtained by oralization recording, and the first sample text is obtained by written conversion based on the second sample text. In the above manner, the oralization text is obtained based on the oralization prediction network, the oralization model is trained by using a plurality of sample texts, the sample text pair includes the first sample text and the second sample text, the second sample text is obtained by oralization recording, and the first sample text is obtained by written conversion based on the second sample text. The oralization prediction network can learn the oralization expression features, which is beneficial to improving the accuracy and efficiency of oralization conversion.

[0030] In a specific implementation scenario, the above-mentioned sample data can be recorded based on a dialogue topic and a dialogue outline related to a preset oralization scenario. For example, for a chat scenario, a dialogue topic (e.g., a travel topic) and a dialogue outline (e.g., places visited, the most impressive place, interesting things happened during the trip, and places planned to visit) related to the chat can be designed; or for a debate scenario, a dialogue topic and a dialogue outline related to the debate can be designed, and other scenarios can be similarly deduced, and will not be listed one by one here. In the above manner, the speech synthesis model is trained by using sample data, and the sample data is recorded based on a dialogue topic and a dialogue outline related to a preset oralization scenario, which can ensure the quality of the training sample at the level of audio library recording, and is beneficial to improving the accuracy of the speech synthesis model.

[0031] In one specific implementation scenario, the sample data can include, but is not limited to, a recorded sample colloquial speech, a second sample text transcribed from the sample colloquial speech, and a first sample text converted from the second sample text. Please refer to Figure 2 , Figure 2 is a process diagram of an embodiment of obtaining sample data. As shown in Figure 2 , the conversation can be conducted according to a designed conversation topic and a conversation outline, and before that, the conversants can be agreed to freely improvise the conversation topic according to their own colloquial habits. In addition, the two conversants can also conduct the conversation by calling each other in different recording rooms, and record the sample colloquial speech of the two conversants respectively, so as to reduce the sound overlap as much as possible during recording, and improve the sample quality. Further, as shown in Figure 2 , during the recording process, the emotions of the conversants can be obtained in real time by using artificial monitoring or algorithm monitoring, and the emotions of the conversants can be called in time to make the conversants express emotions in line with the context. For example, in a chatting scenario, the emotions of the conversants can be mobilized to be in a relatively relaxed state; or in a debating scenario, the emotions of the conversants can be mobilized to be in a relatively excited state, and other scenarios can be similarly deduced, which will not be exemplified one by one. After obtaining the sample colloquial speech, the sample colloquial speech can be converted by using an ASR (Automatic Speech Recognition) tool to obtain a second sample text, and the second sample text can be converted to a written form to obtain a first sample text. Exemplarily, in the conversation process of the foregoing tourism topic, the sample speech "I remember it was very cold at that time" is recorded, the second sample text "I remember it was very cold at that time" is recognized, it is found that there is a colloquial phenomenon of inverted word order, so the written conversion is performed to correct the inverted word order, and the first sample text "I remember it was very cold at that time" is obtained. Other cases can be similarly deduced, which will not be exemplified one by one.

[0032] In one specific implementation scenario, the colloquialization prediction network can include, but is not limited to, an end-to-end network such as an Encoder-Decoder structure. In this case, please refer to Figure 3 , Figure 3 is a framework diagram of an embodiment of the colloquialization prediction network. As shown in Figure 3As shown, the first sample text can be input to an encoder (Encoder) of the spokenization prediction network for encoding to obtain sample semantic feature representations of each sample word in the first sample text, and then the sample semantic feature representations of each sample word can be input to a decoder (Decoder) of the spokenization prediction network for decoding to obtain a plurality of predicted words. On this basis, the prediction loss of the spokenization prediction network can be obtained based on the difference between the ith predicted word and the ith sample word in the second sample text, and the network parameters of the spokenization prediction network can be adjusted based on the prediction loss. Specifically, the decoder can obtain the prediction probability values of each preset word in the preset word table during each decoding, so that the preset word corresponding to the maximum prediction probability value can be taken as the predicted word output by this decoding. In addition, the prediction loss can be calculated based on the difference between the ith predicted word and the ith sample word in the second sample text, and the prediction probability value of the ith predicted word, through a loss function such as cross-entropy. Further, the network parameters of the spokenization prediction network can be adjusted based on the prediction loss through an optimization method such as gradient descent. It should be noted that the specific calculation process of the prediction loss can refer to the technical details of the loss function such as cross-entropy, and the specific adjustment process of the network parameters can refer to the technical details of the optimization method such as gradient descent, which will not be described here. Based on this, the spokenization prediction network can be used to predict the to-be-synthesized text to obtain the corresponding spokenization text. Specifically, in the case of an Encoder-Decoder structure of the spokenization prediction network, the encoder of the spokenization prediction network can encode each word in the to-be-synthesized text to obtain semantic feature representations of each word, and input the semantic feature representations of each word to the decoder of the spokenization prediction network for decoding to obtain a plurality of predicted words, so that the combination of the predicted words obtained by sequential decoding can be taken as the spokenization text. As mentioned earlier, the conversion mode can include but is not limited to: adding mood words, adding catchphrases, adding repeated sentences, correcting content, and reversing the word order. Please refer to Table 1, which is a schematic table of an embodiment of spokenization conversion. As shown in Table 1, the underlined words represent the spokenization conversion made on the basis of the to-be-synthesized text. For example, for the conversion mode of adding mood words, the to-be-synthesized text is "I also think it's very good", and the spokenization text after spokenization is "Hmm, I also think it's very good". The other conversion modes in Table 1 can be similarly extended, and will not be exemplified one by one here. It should be noted that the examples shown in Table 1 are only several typical examples of spokenization conversion, and other spokenization conversion modes are not excluded.

[0033] Table 1 Schematic table of an embodiment of spokenization conversion

[0034]

[0035] It should be noted that the encoder can include but is not limited to a multi-layer long short-term memory network, a Transformer structure, etc., which is not limited herein. The encoder can further include an input layer, a hidden layer and an output layer, and the output result of the output layer is input into the decoder through the Attention mechanism. In addition, the decoder can include but is not limited to a multi-layer long short-term memory network, a Transformer structure, etc., which is not limited herein. The input of the decoder is the output obtained by the Attention mechanism, and the decoder outputs the corresponding spoken text in a word-by-word autoregressive manner. As shown in FIG. 8, in the case of inputting the text to be synthesized "I also think it's good", the spoken text "Hmm, I also think it's good" can be output. Figure 3

[0036] In a specific implementation scenario, the spoken language prediction network can also not be an end-to-end network. Illustratively, the spoken language prediction network can be trained based on the edit distance between the first sample text and the second sample text. Specifically, the first sample text and the second sample text can be aligned based on the edit distance between the first sample text and the second sample text, and the sample edit label of each sample word in the first sample text can be obtained based on the alignment result between the first sample text and the second sample text, and the sample edit label includes a sample edit type and a sample edit text. Please refer to Table 2, which is a sample edit label table from the first sample text to the second sample text. As shown in Table 2, the first sample text "I understand I will go immediately" and the second sample text "I understand I understand I will go immediately" can be aligned based on the minimum edit distance algorithm. The alignment result shows that a sample word "understand" is inserted in the first sample text "understand", that is, the second sample text is obtained, so the sample edit label of each sample word in the first sample text shown in Table 2 can be obtained, such as the sample edit label "Keep|understand" of the sample word "understand", wherein the sample edit type is "Keep", and the sample edit text is "understand", which means that the sample edit text "understand" is added based on the "Keep" (i.e. keep) of the sample word "understand" in the first sample text. Similarly, the sample edit label of the sample word "I" is "Keep", wherein the sample edit type is "Keep", and the sample edit text is empty, which means that the sample word "I" in the first sample text is "Keep" (kept) and no other word needs to be added. The sample edit labels of other sample words are similar and will not be repeated here. In addition, it can also include but is not limited to editing types such as insertion (Insert), deletion (Delete), etc., which are not limited herein.

[0037] Table 2 is a sample edit label table from the first sample text to the second sample text

[0038]

[0039]

[0040] On this basis, the spoken language prediction network can be used to predict the prediction editing label of each sample word in the first sample text, and the prediction editing label includes a prediction editing type and a prediction editing text. Thus, based on the difference between the sample editing label and the prediction editing label, the network parameters of the spoken language prediction network can be adjusted. Specifically, please refer to Figure 4 , Figure 4 is a framework schematic diagram of another embodiment of the spoken language prediction network. As Figure 4 indicated, the spoken language prediction network can include a semantic extraction subnetwork and a label prediction subnetwork. The semantic extraction subnetwork can include, but is not limited to, BERT (Bidirectional Encoder Representation from Transformers, i.e., a bidirectional encoder representation based on Transformers), etc., for extracting a semantic representation sequence of the input written text. The label prediction subnetwork can include, but is not limited to, a bidirectional long short-term memory network (Bi-LSTM), a feedforward neural network, etc., for predicting an editing label based on the semantic representation sequence. For the specific process of adjusting the parameters according to the difference, please refer to the foregoing loss calculation and gradient descent, etc., which will not be described here. The above method aligns the first sample text and the second sample text based on the editing distance between the first sample text and the second sample text, and obtains the sample editing label of each sample word in the first sample text based on the alignment result between the first sample text and the second sample text. The sample editing label includes a sample editing type and a sample editing text. Then, the spoken language prediction network is used to predict the prediction editing label of each sample word in the first sample text, and the prediction editing label includes a prediction editing type and a prediction editing text. Thus, based on the difference between the sample editing label and the prediction editing label, the network parameters of the spoken language prediction network are adjusted. Therefore, the spoken language prediction network can learn the editing difference between the written text and the spoken language text from the editing perspective, which is conducive to improving the accuracy and interpretability of the spoken language conversion.

[0041] After the oralization prediction network is trained, the editing labels of each word in the text to be synthesized can be predicted based on the oralization prediction network, and the editing labels can include an editing type and an editing text. On this basis, the editing operation corresponding to the editing type of each word can be performed based on the editing text of the word, and the oralized text is obtained. For example, taking the text to be synthesized "I also think it is very good" as an example, as shown in Table 3, the editing labels of each word in the text to be synthesized can be predicted. The specific meaning of the editing labels can be referred to Table 2 and the related description, which will not be repeated here.

[0042] Table 3: Editing labels of each word in the text to be synthesized

[0043] Number 1 2 3 4 5 Text to be synthesized I also feel very good Editorial label Keep | Hmm Keep Keep Keep Keep

[0044] On this basis, for the word "I", the editing operation corresponding to the editing type (i.e. Keep) of the word can be performed based on the editing text "um", that is, under the premise of "Keep" (retain) the word "I", the editing text "um" is added; similarly, the same can be applied to other words, and finally the oralized text "um I also think it is very good" is obtained. Other texts to be synthesized can be similarly applied, and will not be repeated here. The above method predicts the editing labels of each word in the text to be synthesized based on the oralization prediction network, and the editing labels include the editing type and the editing text. On this basis, the editing operation corresponding to the editing type of each word is performed based on the editing text of the word, and the oralized text is obtained, which enables the oralization prediction network to learn the editing difference between the written text and the oralized text from the editing perspective, and is beneficial to improve the accuracy and interpretability of oralization conversion.

[0045] Step S12: Extracting the phoneme sequence of the oralized text and predicting the oralization control label of the oralized text.

[0046] In the embodiments of the present disclosure, the oralization control label is used to control the pronunciation state. It should be noted that the pronunciation state can represent the pronunciation urgency, pronunciation emotion, pronunciation pause, etc. of each phoneme in the phoneme sequence in the finally synthesized oralized speech, which is not limited here.

[0047] In an embodiment scenario, a phoneme represents the smallest unit of speech divided according to the natural properties of speech. Illustratively, for the spoken text "en#uo#ie#j#ve#d#e#h#en#b#u#c#uo", its phoneme sequence can be represented as "en#uo#ie#j#ve#d#e#h#en#b#u#c#uo". Other texts can be similarly represented, which are not listed one by one herein. In addition, in order to improve the convenience of extracting the phoneme sequence, a front-end text analysis tool can be used to analyze the spoken text to obtain the phoneme sequence. The front-end text analysis tool can include, but is not limited to, phonemizer, etc., which are not limited herein.

[0048] In an embodiment scenario, the spoken tag can include a first tag, and the first tag can represent whether the word to which each phoneme in the phoneme sequence belongs is a discourse particle. In this way, by setting the spoken tag to include the first tag, and the first tag representing whether the word to which each phoneme in the phoneme sequence belongs is a discourse particle, the pronunciation of the discourse particle in the finally synthesized spoken speech can be controlled, and the finally synthesized spoken speech can be more consistent with spoken expression in the acoustic level.

[0049] In a specific embodiment scenario, please refer to Figure 5 , Figure 5 is a process diagram of an embodiment of obtaining the first tag. As Figure 5 shown, the words in the spoken text that are located in the discourse particle word table can be used as candidate words, and whether the candidate words belong to the discourse particle can be determined based on the word position of the candidate words in the spoken text, and the first tag can be obtained based on whether the word to which each phoneme in the phoneme sequence belongs is a discourse particle. Specifically, a discourse particle word table (for example, it can include discourse particles such as "en", "ah", "hi", "oh", etc.) can be maintained in advance, and then a front-end text analysis tool can be used to segment the spoken text to obtain each word in the spoken text. For example, the spoken text "en#uo#ie#j#ve#d#e#h#en#b#u#c#uo" can be segmented to obtain the following words: en, I, also, think, very, not bad, and other texts can be similarly segmented, which are not listed one by one herein. In addition, as Figure 5As shown, the candidate word can be further determined according to the word position of the candidate word in the oralized text, whether the candidate word is located at the phrase boundary (L3 boundary) and the starting position, and if so, it can be determined that the candidate word belongs to the discourse word. For example, in the aforementioned oralized text "hmm I also think very good", the word "hmm" is located in the discourse word table, so it can be determined that it is a candidate word, and the candidate word is also at the phrase boundary and the starting position, so it can be determined that the candidate word is a discourse word. In addition, for example, in order to improve the convenience of marking the label, if the word to which the phoneme in the phoneme sequence belongs is a discourse word, it can be marked as 1, otherwise, if the word to which the phoneme in the phoneme sequence belongs is not a discourse word, it can be marked as 0. In this case, the first label of the aforementioned oralized text "hmm I also think very good" can be expressed in the form of a vector as [1000000000000]. Other texts can be similarly deduced and will not be repeated here. The above method, by taking the words in the oralized text located in the discourse word table as candidate words, and determining whether the candidate word belongs to the discourse word based on the word position of the candidate word in the oralized text, and then obtaining the first label based on whether the word to which each phoneme in the phoneme sequence belongs is a discourse word, can improve the accuracy of the first label.

[0050] In one implementation scenario, the oralization label can include a second label, and the second label represents the duration of each phoneme in the phoneme sequence. The above method further configures the oralization label to include a second label, and the second label represents the duration of each phoneme in the phoneme sequence, which can be beneficial to control the pronunciation duration of each phoneme in the finally synthesized oralized speech, so that the finally synthesized oralized speech is more consistent with oral expression in the acoustic level.

[0051] In a specific implementation scenario, the semantic feature representation of the oralized text can be extracted, and the prosodic boundary information of the oralized text can be extracted. On this basis, the phoneme category of each phoneme in the phoneme sequence can be obtained based on the semantic feature representation and the prosodic boundary information, and the phoneme category includes any one of a diphone phoneme and a normal phoneme. On this basis, the duration of each phoneme can be obtained by performing duration prediction on the diphone phoneme and the normal phoneme respectively, and the second label can be obtained based on the duration of each phoneme in the phoneme sequence. It should be noted that the diphone phoneme represents that there is a diphone phenomenon when pronouncing on this phoneme, and the normal phoneme represents that there is no diphone phenomenon when pronouncing on this phoneme. In oral expression, there is usually a diphone phenomenon. For example, for "that" in the oralized text "that I forgot", it belongs to a verbal tic expression, and there is usually a diphone phenomenon when pronouncing the word "ge". Other cases can be similarly deduced and will not be repeated here. In addition, the prosodic boundary information represents the phrase boundary information, which can represent which words are located at the boundary, such as the aforementioned "that" which is located at the phrase boundary. Please refer to Figure 6 ,Figure 6 is a process schematic diagram of an embodiment of obtaining a phoneme class. As shown in Figure 6 order to improve the efficiency of the prediction of the trailing phoneme, a trailing phoneme prediction network can be pre-trained, and the trailing phoneme prediction network can include a semantic extraction sub-network such as BERT, used to extract a semantic feature representation of the oralized text, and the trailing phoneme prediction network can also include a front-end text analysis tool, used to analyze the oralized text to obtain prosodic boundary information. The trailing phoneme prediction network can also include a label prediction sub-network (such as Bi-LSTM, feedforward neural network, etc.), and on this basis, the semantic feature representation and the prosodic boundary information can be input into the label prediction sub-network to obtain the phoneme type of each phoneme. In the above manner, by extracting the semantic feature representation of the oralized text and extracting the prosodic boundary information of the oralized text, on this basis, the trailing phoneme is predicted based on the semantic feature representation and the prosodic boundary information, and the phoneme class of each phoneme in the phoneme sequence is obtained, and the phoneme class is any one of the trailing phoneme and the normal phoneme, so that the duration of each phoneme is predicted respectively for the trailing phoneme and the normal phoneme, and then the second label is obtained based on the duration of each phoneme in the phoneme sequence, which can be beneficial to improve the accuracy of the second label.

[0052] In a specific implementation scenario, as described above, the phoneme class is obtained based on the trailing phoneme prediction network, and the trailing phoneme prediction network is trained based on a plurality of sample texts, and the sample phoneme sequence of the sample text is annotated with the phoneme class of each sample phoneme, and the phoneme class of the sample phoneme is obtained based on the duration difference between the actual duration and the predicted duration of the sample phoneme. For example, if the duration difference between the actual duration and the predicted duration of the sample phoneme is greater than a duration threshold, it can be considered that the phoneme class of the sample phoneme is a trailing phoneme, and otherwise, it can be considered that the phoneme class of the sample phoneme is a normal phoneme. In addition, the sample text can be obtained by sample speech recognition, the actual duration of the sample phoneme is obtained by the sample speech, and the predicted duration of the sample phoneme is predicted by a pre-trained duration prediction network. For example, the above sample text can be a second sample text in the sample data, i.e., a sample text transcribed from a sample speech recorded in an oralized recording process. Please refer to Figure 7 , Figure 7 is a process schematic diagram of an embodiment of obtaining a phoneme annotation. As shown in Figure 7As shown, for the sample phoneme sequence of the sample text, the pre-trained duration prediction network can be used to predict each sample phoneme to obtain the predicted duration, and the difference between the predicted duration of the sample phoneme and the actual duration of the sample phoneme is obtained to obtain the duration difference. On this basis, the duration difference and the duration threshold can be compared to determine whether the sample phoneme is labeled as a dragged phoneme or a normal phoneme. It should be noted that the duration threshold can be obtained according to the actual duration of each sample phoneme in the sample phoneme sequence of the sample text (for example, the average value of the actual duration can be taken). In addition, the duration prediction network can be pre-trained using sample phonemes labeled with actual phoneme duration. The duration prediction network can include, but is not limited to, a long short-term memory network, a feedforward neural network, etc., which is not limited herein. In the above manner, the phoneme category is obtained based on the dragged phoneme prediction network, the dragged phoneme prediction network is trained using a plurality of sample texts, the sample phoneme sequence of the sample text is labeled with the phoneme category of each sample phoneme, and the phoneme category of the sample phoneme is obtained based on the duration difference between the actual duration and the predicted duration of the sample phoneme. The sample text is obtained by sample speech recognition, the actual duration of the sample phoneme is obtained by the sample speech, and the predicted duration of the sample phoneme is obtained by the pre-trained duration prediction network. The duration prediction network can be used to predict the duration of the sample phoneme sequence, and manual labeling of the phoneme category of each sample phoneme in the sample phoneme sequence is avoided, which is beneficial to improve the training efficiency.

[0053] In a specific implementation scenario, in order to improve the efficiency and accuracy of duration prediction of normal phonemes and dragged phonemes, a hybrid duration prediction network can be pre-trained, and the hybrid duration prediction network includes a first prediction network for predicting the duration of a normal phoneme and a second prediction network for predicting the duration of a dragged phoneme. On this basis, the first prediction network or the second prediction network can be selected for duration prediction according to the phoneme category of each phoneme in the phoneme sequence, which is beneficial to improve the accuracy of duration prediction. In addition, considering that the dragged phenomenon is less in the display scenario, speech recognition data can be used to expand the training data size, which is beneficial to improve the network stability.

[0054] In an implementation scenario, the spokenization label can include an emotional feature representation for representing the emotional classification of the spokenization text, such as calm, relaxed, excited, etc., which is not limited herein. It should be noted that the emotional feature representation can be expressed in the form of a vector. Specifically, a plurality of reference texts of the spokenization text can be obtained, and the plurality of reference texts include interactive texts before and / or after the spokenization text. On this basis, the emotional feature representation of the spokenization text is obtained based on the semantic feature representation of the spokenization text and the semantic feature representation of each reference text. Please refer to Figure 8 , Figure 8 is a process schematic diagram of an embodiment of obtaining an emotional feature representation. As shown,Figure 8 As shown, for the convenience of description, the colloquial text can be referred to as the current text, the text generated before the interaction can be referred to as the historical text, and the text generated after the interaction can be referred to as the future text. In order to improve the efficiency of obtaining the sentiment feature representation, the sentiment prediction network can be pre-trained, and the sentiment prediction network can include a semantic extraction sub-network, which can include but is not limited to BERT and the like, for extracting semantic feature representation, and in addition, the sentiment prediction network can include a representation prediction sub-network, which can include but is not limited to Bi-LSTM, feedforward neural network and the like, for predicting the sentiment feature representation. In the above manner, the colloquial control label includes the sentiment feature representation, and a plurality of reference texts of the colloquial text are obtained, and the plurality of reference texts include the interaction texts before and / or after the colloquial text, and on this basis, the sentiment feature representation of the colloquial text is obtained based on the semantic feature representation of the colloquial text and the semantic feature representation of each reference text, that is, the sentiment feature representation of the colloquial text is obtained in combination with the reference texts, which is beneficial to improve the accuracy of the sentiment feature representation.

[0055] In one specific implementation scenario, as described above, the sentiment feature representation is obtained based on the sentiment prediction network, and the sentiment prediction network is trained using a plurality of sample texts, the sample texts are labeled with sample sentiment feature representations, the sample texts are obtained by sample speech recognition, and the sample sentiment feature representations are predicted by pre-training the voice emotion network on the sample speech. Please refer to Figure 9 Figure 9 is a framework schematic diagram of an embodiment of the voice emotion network. As shown in Figure 9 ​As shown, the speech emotion network can include an encoder, a decoder and an emotion recognition sub-network, the sample speech is labeled with a sample emotion category, the encoder of the speech emotion network can be used to encode the sample emotion category to obtain a predicted emotion feature representation, on this basis, the emotion recognition sub-network can be used to predict the emotion of the predicted emotion feature representation to obtain a predicted emotion category, on this basis, the network parameters of the speech emotion network can be adjusted based on the difference between the predicted emotion category and the sample emotion category. It should be noted that the encoder can be implemented by VAE (Variational AutoEncoder), GST (Global Style Token) and the like. In addition, the decoder can further decode the predicted emotion feature representation to obtain a predicted speech, and in the training process, the emotion expression contained in the sample speech in the predicted speech is as much as possible to be improved, and the substantial content contained in the sample speech in the predicted speech is as much as possible to be suppressed, so as to coordinate the recognition task of the emotion category, as much as possible to improve the accuracy of the predicted emotion feature representation, so as to as much as possible contain the feature information related to emotion and as much as possible contain the feature information related to the substantial content. After the speech emotion network training converges, the sample text corresponding to the sample speech can be encoded to obtain the predicted emotion feature representation of the sample speech, and the sample emotion feature representation of the sample text can be labeled. On this basis, the semantic feature representation of the sample text can be processed by using the emotion prediction network to obtain the predicted emotion feature representation of the sample text, and the network parameters of the emotion prediction network can be adjusted based on the difference between the predicted emotion feature representation of the sample text and the sample emotion feature representation until convergence, that is, the emotion prediction network can be used to predict the emotion feature representation of the oral text. The above-mentioned manner, the emotion feature representation is obtained based on the emotion prediction network, the emotion prediction network is trained based on a plurality of sample texts, the sample texts are labeled with sample emotion feature representations, and the sample texts are recognized by sample speech, the sample emotion feature representations are predicted by the pre-trained speech emotion network based on the sample speech, so that the sample emotion feature representations can be labeled by the pre-trained speech emotion network, and on this basis, the emotion prediction network is trained, which is beneficial to improve the accuracy of sample labeling.

[0056] Step S13: based on the phoneme sequence and the oralization control label, the oralization speech of the to-be-synthesized text is synthesized.

[0057] Specifically, the colloquial speech is synthesized based on a speech synthesis model, and the speech synthesis model is trained based on sample data recorded based on a dialogue topic and a dialogue outline related to a preset colloquial scenario. For example, the sample data can include a first sample text, a second sample text, and sample speech, and the related meanings can refer to the foregoing related description, which will not be repeated here. In addition, the speech synthesis model can include, but is not limited to, the foregoing colloquial prediction network, the diphone prediction network, the duration prediction network, the hybrid duration prediction network, the emotion prediction network, the speech emotion network, and the like, which will not be limited here. Please refer to Figure 10 , Figure 10 is a process schematic diagram of an embodiment of the speech synthesis method of the present application. As shown in Figure 10 , the text to be synthesized is first obtained by a colloquial prediction network to obtain a colloquial text. The colloquial text can be combined with the foregoing emotion prediction network to obtain an emotion feature representation. At the same time, the colloquial text can be analyzed by a front-end text analysis tool to obtain a phoneme sequence, prosodic boundary information, and phrase boundary. The phrase boundary can be determined as a first label in combination with the colloquial text, and the prosodic boundary information can be combined with the colloquial text and the foregoing diphone prediction network to obtain a phoneme category (i.e., a normal phoneme or a diphone phoneme). Further, the speech synthesis model can further include an acoustic model, and the foregoing hybrid duration prediction network can be included in the acoustic model to predict the phoneme duration of the normal phoneme and the phoneme duration of the diphone phoneme (i.e., a second label) by the hybrid duration prediction network, and combine the foregoing phoneme sequence, first label, and emotion feature representation to obtain a plurality of acoustic parameters. Specifically, the phoneme sequence feature can be analyzed by the front-end text analysis tool, and on this basis, the phoneme sequence feature and the first label (i.e., the tone word label) and the emotion feature representation can be spliced to obtain spliced features. Further, the phoneme type can be predicted (i.e., the diphone label prediction, that is, the phoneme is predicted to be a normal phoneme or a diphone phoneme), and then the phoneme duration is predicted according to the phoneme type (i.e., the first prediction network or the second prediction network in the foregoing hybrid duration prediction network is selected). Finally, the spliced features can be expanded (copied) to the frame level according to the phoneme duration, and the frame-level acoustic parameters can be obtained by inputting the acoustic model. In addition, the speech synthesis model can also include a vocoder for generating a speech waveform based on the acoustic parameters to obtain colloquial speech.

[0058] The scheme converts the to-be-synthesized text into a colloquial text, and the colloquial text refers to at least one conversion mode, and the phoneme sequence of the colloquial text is extracted, and the colloquial control label of the colloquial text is predicted, and the colloquial control label is used to control the pronunciation state. On this basis, the colloquial speech of the to-be-verified text is synthesized based on the phoneme sequence and the colloquial control label. On the one hand, the to-be-synthesized text is converted into a colloquial text by referring to at least one conversion mode, which is beneficial to make the colloquial text as much as possible to conform to the colloquial expression. On the other hand, the colloquial control label of the colloquial text is predicted on this basis, which can further provide a reference for colloquial speech synthesis from the acoustic level on the basis of the foregoing text level. Therefore, the colloquial speech synthesis can be realized from two different levels of text level and acoustic level, so as to improve the user interaction experience.

[0059] Please refer to Figure 11 , Figure 11 is a frame schematic diagram of an embodiment of the speech synthesis device 110. The speech synthesis device 110 comprises a colloquial conversion module 111, a phoneme extraction module 112, a label prediction module 113, and a sound synthesis module 114. The colloquial conversion module 111 is configured to convert the to-be-synthesized text into a colloquial text. The colloquial conversion refers to at least one conversion mode. The phoneme extraction module 112 is configured to extract the phoneme sequence of the colloquial text. The label prediction module 113 is configured to predict the colloquial control label of the colloquial text. The colloquial control label is used to control the pronunciation state. The sound synthesis module 114 is configured to synthesize the colloquial speech of the to-be-synthesized text based on the phoneme sequence and the colloquial control label.

[0060] The scheme converts the to-be-synthesized text into a colloquial text by referring to at least one conversion mode, which is beneficial to make the colloquial text as much as possible to conform to the colloquial expression. On the other hand, the colloquial control label of the colloquial text is predicted on this basis, which can further provide a reference for colloquial speech synthesis from the acoustic level on the basis of the foregoing text level. Therefore, the colloquial speech synthesis can be realized from two different levels of text level and acoustic level, so as to improve the user interaction experience.

[0061] In some disclosed embodiments, the colloquial text is obtained based on a colloquial prediction network, the colloquial prediction network is trained based on a plurality of sample text pairs, each sample text pair comprises a first sample text and a second sample text, the second sample text is obtained by colloquial recording, and the first sample text is obtained by written conversion based on the second sample text.

[0062] Therefore, the oral text is obtained based on the oral prediction network, and the oral model is trained based on a plurality of sample text pairs, each sample text pair including a first sample text and a second sample text, and the second sample text is obtained through oral recording, and the first sample text is obtained through written conversion based on the second sample text, so that the oral prediction network can learn oral expression features, and the accuracy and efficiency of oral conversion can be improved.

[0063] In some disclosed embodiments, the speech synthesis device 110 includes a text alignment module configured to align the first sample text and the second sample text based on an edit distance between the first sample text and the second sample text; the text alignment module includes a label annotation module configured to obtain a sample edit label of each sample word in the first sample text based on an alignment result between the first sample text and the second sample text; the sample edit label includes a sample edit type and a sample edit text; the speech synthesis device 110 includes a label prediction module configured to predict a prediction edit label of each sample word in the first sample text based on the oral prediction network; the prediction edit label includes a prediction edit type and a prediction edit text; and the speech synthesis device 110 includes a parameter adjustment module configured to adjust network parameters of the oral prediction network based on a difference between the sample edit label and the prediction edit label.

[0064] Therefore, the first sample text and the second sample text are aligned based on an edit distance between the first sample text and the second sample text, and a sample edit label of each sample word in the first sample text is obtained based on an alignment result between the first sample text and the second sample text, and the sample edit label includes a sample edit type and a sample edit text, and a prediction edit label of each sample word in the first sample text is predicted based on the oral prediction network, and the prediction edit label includes a prediction edit type and a prediction edit text, so that network parameters of the oral prediction network can be adjusted based on a difference between the sample edit label and the prediction edit label, and the oral prediction network can learn an edit difference between written text and oral text from an edit perspective, and the accuracy and interpretability of oral conversion can be improved.

[0065] In some disclosed embodiments, the oral conversion module 111 includes a label prediction submodule configured to predict an edit label of each word in the text to be synthesized based on the oral prediction network; the edit label includes an edit type and an edit text; and the oral conversion module 111 includes a text editing submodule configured to, for each word, perform an edit operation corresponding to the edit type of the word based on the edit text of the word to obtain the oral text.

[0066] Therefore, the spoken language prediction network predicts the editing label of each word in the to-be-composed text, and the editing label includes an editing type and an editing text, and on this basis, the editing operation corresponding to the editing type of each word is performed based on the editing text of the word, to obtain the spoken language text, which can enable the spoken language prediction network to learn the editing difference between the written text and the spoken language text from the editing perspective, and is beneficial to improving the accuracy and interpretability of the spoken language conversion.

[0067] In some disclosed embodiments, the spoken language control label includes a first label, and the first label represents whether the word to which each phoneme in the phoneme sequence belongs belongs to an interjection.

[0068] Therefore, by setting the spoken language label to include the first label, and the first label representing whether the word to which each phoneme in the phoneme sequence belongs belongs to an interjection, the pronunciation of the interjection in the finally synthesized spoken language speech can be controlled, and the finally synthesized spoken language speech is more in line with spoken language expression in the acoustic level.

[0069] In some disclosed embodiments, the label prediction module 113 includes a candidate word extraction submodule for extracting, as a candidate word, the word in the spoken language text that is located in the interjection word table; the label prediction module 113 includes an interjection determination submodule for determining, based on the word position of the candidate word in the spoken language text, whether the candidate word belongs to an interjection; and the label prediction module 113 includes a first label acquisition submodule for obtaining the first label based on whether the word to which each phoneme in the phoneme sequence belongs belongs to an interjection.

[0070] Therefore, by extracting, as a candidate word, the word in the spoken language text that is located in the interjection word table, and determining, based on the word position of the candidate word in the spoken language text, whether the candidate word belongs to an interjection, and then obtaining the first label based on whether the word to which each phoneme in the phoneme sequence belongs belongs to an interjection, the accuracy of the first label can be improved.

[0071] In some disclosed embodiments, the spoken language control label includes a second label, and the second label represents the duration of each phoneme in the phoneme sequence.

[0072] Therefore, the spoken language label is further set to include the second label, and the second label represents the duration of each phoneme in the phoneme sequence, which can be beneficial to controlling the pronunciation duration of each phoneme in the finally synthesized spoken language speech, and the finally synthesized spoken language speech is more in line with spoken language expression in the acoustic level.

[0073] In some disclosed embodiments, the label prediction module 113 comprises an information extraction sub-module configured to extract a semantic feature representation of the spoken text and extract prosodic boundary information of the spoken text; the label prediction module 113 comprises a phoneme class prediction sub-module configured to perform diphone prediction based on the semantic feature representation and the prosodic boundary information to obtain a phoneme class of each phoneme in the phoneme sequence; wherein the phoneme class is any one of a diphone phoneme and a normal phoneme; the label prediction module 113 comprises a phoneme duration prediction sub-module configured to perform duration prediction on the diphone phoneme and the normal phoneme respectively to obtain a duration of each phoneme; and the label prediction module 113 comprises a second label acquisition sub-module configured to obtain a second label based on the duration of each phoneme in the phoneme sequence.

[0074] Therefore, by extracting a semantic feature representation of the spoken text and extracting prosodic boundary information of the spoken text, and on this basis, performing diphone prediction based on the semantic feature representation and the prosodic boundary information to obtain a phoneme class of each phoneme in the phoneme sequence, and the phoneme class being any one of a diphone phoneme and a normal phoneme, duration prediction is performed on the diphone phoneme and the normal phoneme respectively to obtain a duration of each phoneme, and then a second label is obtained based on the duration of each phoneme in the phoneme sequence, which can help improve the accuracy of the second label.

[0075] In some disclosed embodiments, the phoneme class is obtained based on a diphone prediction network, the diphone prediction network is trained using a plurality of sample texts, the sample phoneme sequence of the sample text is annotated with a phoneme class of each sample phoneme, and the phoneme class of the sample phoneme is obtained based on a duration difference between an actual duration and a predicted duration of the sample phoneme; wherein the sample text is obtained by sample speech recognition, the actual duration of the sample phoneme is obtained by the sample speech, and the predicted duration of the sample phoneme is predicted by a pre-trained duration prediction network.

[0076] Therefore, the phoneme class is obtained based on a diphone prediction network, the diphone prediction network is trained using a plurality of sample texts, the sample phoneme sequence of the sample text is annotated with a phoneme class of each sample phoneme, and the phoneme class of the sample phoneme is obtained based on a duration difference between an actual duration and a predicted duration of the sample phoneme, and the sample text is obtained by sample speech recognition, the actual duration of the sample phoneme is obtained by the sample speech, and the predicted duration of the sample phoneme is predicted by a pre-trained duration prediction network, which can perform duration prediction on the sample phoneme sequence through the duration prediction network, and avoid manual annotation of the phoneme class of each sample phoneme in the sample phoneme sequence, which helps improve training efficiency.

[0077] In some disclosed embodiments, the colloquialization control label comprises an emotional feature representation, the label prediction module 113 comprises a reference text acquisition sub-module configured to acquire a plurality of reference texts of the colloquialization text; wherein the plurality of reference texts comprise interactive texts before and / or after the colloquialization text; and the label prediction module 113 comprises a feature representation prediction sub-module configured to obtain the emotional feature representation of the colloquialization text based on the semantic feature representation of the colloquialization text and the semantic feature representation of each reference text.

[0078] Therefore, the colloquialization control label comprises an emotional feature representation, and a plurality of reference texts of the colloquialization text are acquired, and the plurality of reference texts comprise interactive texts before and / or after the colloquialization text. On this basis, the emotional feature representation of the colloquialization text is obtained based on the semantic feature representation of the colloquialization text and the semantic feature representation of each reference text, that is, the emotional feature representation of the colloquialization text is obtained in combination with the reference texts, which is beneficial to improving the accuracy of the emotional feature representation.

[0079] In some disclosed embodiments, the emotional feature representation is obtained based on an emotion prediction network, the emotion prediction network is trained based on a plurality of sample texts, the sample texts are labeled with sample emotional feature representations, and the sample texts are recognized from sample speeches, and the sample emotional feature representations are predicted from the sample speeches by a pre-trained speech emotion network.

[0080] Therefore, the emotional feature representation is obtained based on an emotion prediction network, the emotion prediction network is trained based on a plurality of sample texts, the sample texts are labeled with sample emotional feature representations, and the sample texts are recognized from sample speeches, and the sample emotional feature representations are predicted from the sample speeches by a pre-trained speech emotion network. Therefore, the sample emotional feature representations can be labeled by the pre-trained speech emotion network, and on this basis, the emotion prediction network is trained, which is beneficial to improving the accuracy of sample labeling.

[0081] In some disclosed embodiments, the colloquialization speech is synthesized based on a speech synthesis model, and the speech synthesis model is trained based on a plurality of sample data, and the sample data is recorded based on a dialogue theme and a dialogue outline related to a preset colloquialization scenario.

[0082] Therefore, the speech synthesis model is trained using sample data, and the sample data is recorded based on a dialogue theme and a dialogue outline related to a preset colloquialization scenario, which can ensure the quality of training samples at the level of audio library recording, and is beneficial to improving the accuracy of the speech synthesis model.

[0083] Please refer to Figure 12 , Figure 12is a frame diagram of an embodiment of the electronic device 120 of the present application. The electronic device 120 comprises a memory 121 and a processor 122 coupled with each other, the memory 121 stores program instructions, and the processor 122 is configured to execute the program instructions to implement the steps in any of the above voice synthesis method embodiments. Specifically, the electronic device 120 can include but is not limited to a desktop computer, a notebook computer, a server, a mobile phone, a tablet computer, and the like, which are not limited herein.

[0084] Specifically, the processor 122 is configured to control itself and the memory 121 to implement the steps in any of the above voice synthesis method embodiments. The processor 122 can also be referred to as a CPU (Central Processing Unit). The processor 122 can be an integrated circuit chip having a processing capability of signals. The processor 122 can also be a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor or the like. In addition, the processor 122 can be implemented by an integrated circuit chip together.

[0085] The above scheme, on the one hand, first converts the text to be synthesized into a colloquial text according to at least one conversion mode, which is conducive to making the colloquial text as much as possible to conform to the colloquial expression, and on the other hand, predicts the colloquial control tags of the colloquial text on this basis, which can further provide a reference for colloquial speech synthesis from the acoustic level on the basis of the aforementioned text level, so as to realize colloquial speech synthesis from two different levels of text level and acoustic level at the same time, so as to improve the user interaction experience.

[0086] Please refer to Figure 13 , Figure 13 is a frame diagram of an embodiment of the computer readable storage medium 130 of the present application. The computer readable storage medium 130 stores program instructions 131 capable of being executed by a processor, and the program instructions 131 are used to implement the steps in any of the above voice synthesis method embodiments.

[0087] The above scheme, on the one hand, first refers to at least one conversion mode to perform oralization conversion on the to-be-synthesized text to obtain an oralized text, which is beneficial to making the oralized text as much as possible to conform to oralized expression, and on the other hand, predicts oralized control tags of the oralized text on this basis, which can further provide a reference for oralized speech synthesis from the acoustic level on the basis of the foregoing text level, and thus can realize oralized speech synthesis from two different levels of the text level and the acoustic level, so as to improve the user interaction experience.

[0088] In some embodiments, the apparatus provided by the embodiments of the present disclosure has functions or includes modules that can be used to perform the methods described in the above method embodiments, and the specific implementation can refer to the description of the above method embodiments. For brevity, details are not repeated here.

[0089] The above description of various embodiments tends to emphasize the differences between various embodiments, and the same or similar parts can be mutually referred to. For brevity, details are not repeated here.

[0090] In several embodiments provided in the present application, it should be understood that the disclosed methods and apparatuses can be implemented in other ways. For example, the above-described apparatus implementation is only schematic, for example, the division of modules or units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the shown or discussed mutual elements can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms.

[0091] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the present embodiment scheme.

[0092] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The above integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0093] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application, essentially or in other words, the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to perform all or part of the steps of the methods in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

Claims

1. A speech synthesis method characterized by, The method comprises: Converting the text to be synthesized into a colloquial text; wherein the conversion refers to at least one conversion mode; Extracting the phoneme sequence of the colloquial text and predicting the colloquial control label of the colloquial text; wherein the colloquial control label is used to control the pronunciation state; Based on the phoneme sequence and the colloquial control label, the colloquial speech of the text to be synthesized is synthesized; Wherein the colloquial control label includes a second label, the second label represents the duration of each phoneme in the phoneme sequence, and the second label is obtained by: Extracting the semantic feature representation of the colloquial text and extracting the prosodic boundary information of the colloquial text; Based on the semantic feature representation and the prosodic boundary information, the phoneme category of each phoneme in the phoneme sequence is obtained by predicting the diphthong; wherein the phoneme category is any one of the diphthong phoneme and the ordinary phoneme, the phoneme category is obtained based on the diphthong prediction network, the diphthong prediction network is trained based on a plurality of sample texts, the sample phoneme sequence of the sample text is labeled with the phoneme category of each sample phoneme, and the phoneme category of the sample phoneme is obtained based on the time difference between the actual duration and the predicted duration of the sample phoneme, in response to the time difference being greater than the duration threshold, the phoneme category of the sample phoneme is the diphthong phoneme, and in response to the time difference being not greater than the duration threshold, the phoneme category of the sample phoneme is the ordinary phoneme, the sample text is obtained by sample speech recognition, the actual duration of the sample phoneme is obtained by the sample speech, the duration threshold is obtained based on the actual duration of each sample phoneme in the sample phoneme sequence of the sample text, and the predicted duration of the sample phoneme is obtained by the pre-trained duration prediction network; The duration of each phoneme is obtained by predicting the duration of the diphthong phoneme and the ordinary phoneme respectively; wherein the duration of each phoneme is predicted based on a hybrid duration prediction network, the hybrid duration prediction network includes a first prediction network for predicting the duration of the ordinary phoneme, and a second prediction network for predicting the duration of the diphthong phoneme, so as to select the first prediction network or the second prediction network for duration prediction according to the phoneme category of each phoneme in the phoneme sequence; Based on the duration of each phoneme in the phoneme sequence, the second label is obtained.

2. The method of claim 1, wherein, The colloquial text is obtained based on a colloquial prediction network, the colloquial prediction network is trained based on a plurality of sample text pairs, the sample text pair includes a first sample text and a second sample text, and the second sample text is obtained by colloquial recording, and the first sample text is obtained by written conversion based on the second sample text.

3. The method of claim 2, wherein, The training steps of the colloquial prediction network include: Aligning the first sample text and the second sample text based on the edit distance between them; obtaining a sample editing label of each sample word in the first sample text based on the alignment result between the first sample text and the second sample text, wherein the sample editing label comprises a sample editing type and a sample editing text; predicting a predicted editing label of each sample word in the first sample text based on the spoken language prediction network, wherein the predicted editing label comprises a predicted editing type and a predicted editing text; adjusting network parameters of the spoken language prediction network based on a difference between the sample editing label and the predicted editing label.

4. The method of claim 3, wherein, The spoken language conversion of the to-be-synthesized text comprises: predicting an editing label of each word in the to-be-synthesized text based on the spoken language prediction network, wherein the editing label comprises an editing type and an editing text; performing an editing operation corresponding to the editing type of each word based on the editing text of the word to obtain the spoken language text.

5. The method of claim 1, wherein, The spoken language control label comprises a first label, and the first label represents whether a word to which each phoneme in the phoneme sequence belongs is a mood word.

6. The method of claim 5, wherein, The first label comprises: words in the spoken language text that are located in a mood word table are taken as candidate words; whether the candidate word belongs to a mood word is determined based on a word position of the candidate word in the spoken language text; the first label is obtained based on whether a word to which each phoneme in the phoneme sequence belongs is a mood word.

7. The method of claim 1, wherein, The spoken language control label comprises an emotional feature representation, and the emotional feature representation comprises: a plurality of reference texts of the spoken language text are obtained; wherein the plurality of reference texts comprise interactive texts before and / or after the spoken language text; an emotional feature representation of the spoken language text is obtained based on a semantic feature representation of the spoken language text and a semantic feature representation of each reference text.

8. The method of claim 7, wherein, The emotional feature representation is obtained based on an emotional prediction network, the emotional prediction network is trained based on a plurality of sample texts, the sample texts are labeled with sample emotional feature representations, and the sample texts are obtained by sample speech recognition, and the sample emotional feature representations are predicted by a pre-trained voice emotion network based on the sample speech.

9. The method of claim 1, wherein, The spoken language voice is synthesized based on a voice synthesis model, and the voice synthesis model is trained based on a plurality of sample data, and the sample data is recorded based on a dialogue theme and a dialogue outline related to a preset spoken language scene.

10. A speech synthesis apparatus characterized by comprising: The spoken language conversion module is configured to convert a to-be-synthesized text into a spoken language text, wherein the spoken language conversion refers to at least one conversion mode; The phoneme extraction module is configured to extract a phoneme sequence of the spoken language text. The label prediction module is configured to predict a spoken language control label of the spoken language text, wherein the spoken language control label is used to control a pronunciation state. ​ a voice synthesis module, configured to synthesize a spoken voice of the text to be synthesized based on the phoneme sequence and the spokenization control label; wherein the spokenization control label comprises a second label, the second label representing a duration of each phoneme in the phoneme sequence; an extraction sub-module, configured to extract a semantic feature representation of the spoken text and extract prosodic boundary information of the spoken text; a phoneme class prediction sub-module, configured to perform diphthong prediction based on the semantic feature representation and the prosodic boundary information to obtain a phoneme class of each phoneme in the phoneme sequence; wherein the phoneme class is any one of a diphthong phoneme and a normal phoneme, the phoneme class is obtained based on a diphthong prediction network, the diphthong prediction network is trained based on a plurality of sample texts, a sample phoneme sequence of the sample text is annotated with a phoneme class of each sample phoneme, and the phoneme class of the sample phoneme is obtained based on a duration difference between an actual duration and a predicted duration of the sample phoneme, in response to the duration difference being greater than a duration threshold, the phoneme class of the sample phoneme is the diphthong phoneme, in response to the duration difference being not greater than the duration threshold, the phoneme class of the sample phoneme is the normal phoneme, the sample text is obtained by sample speech recognition, the actual duration of the sample phoneme is obtained by the sample speech, the duration threshold is obtained based on a statistical result of the actual duration of each sample phoneme in the sample phoneme sequence of the sample text, and the predicted duration of the sample phoneme is predicted by a pre-trained duration prediction network; a phoneme duration prediction sub-module, configured to perform duration prediction on the diphthong phoneme and the normal phoneme respectively to obtain a duration of each phoneme; wherein the duration of each phoneme is predicted based on a hybrid duration prediction network, the hybrid duration prediction network comprises a first prediction network for predicting the duration of the normal phoneme and a second prediction network for predicting the duration of the diphthong phoneme, and the first prediction network or the second prediction network is selected for duration prediction according to the phoneme class of each phoneme in the phoneme sequence; a second label acquisition sub-module, configured to obtain the second label based on the duration of each phoneme in the phoneme sequence.

11. An electronic device, comprising: A memory and a processor coupled to each other, the memory storing program instructions, and the processor being configured to execute the program instructions to implement the speech synthesis method of any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that, A memory storing program instructions executable by a processor, the program instructions being configured to implement the speech synthesis method of any one of claims 1 to 9.

Citation Information

Patent Citations

  • Audio synthesis method and device

    CN109599092A

  • Method and device for converting text into voice, storage medium and equipment

    CN113192483A