Speech synthesis method and related device, electronic device, storage medium

By predicting and fusion of intonation categories and pronunciation sequences of clause texts, the problem of poor pronunciation synthesis is solved, and the intonation performance and control of pronunciation synthesis is improved.

CN114664284BActive Publication Date: 2025-07-11IFLYTEK CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210289041.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-22
Publication Date
2025-07-11
Estimated Expiration
2042-03-22

AI Technical Summary

Technical Problem

The existing speech synthesis technology has poor synthesis effects, resulting in a large deviation from the real person's pronunciation and affecting the user experience.

Method used

By obtaining the text to be synthesized and reference data, predict the intonation category of the text in the clause, and combine the pronunciation sequence and intonation category to obtain the sequence to be synthesized, and finally synthesize the pronunciation sequence and intonation category to improve the sound synthesis effect.

Benefits of technology

It improves the tone performance and control of pronunciation synthesis and improves the overall effect of pronunciation synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114664284B_ABST
    Figure CN114664284B_ABST
Patent Text Reader

Abstract

The present application discloses a speech synthesis method and related devices, electronic devices, and storage media. Among them, the speech synthesis method includes: obtaining the text to be synthesized and reference data of the text to be synthesized; wherein, the text to be synthesized includes several clause texts ending with clause punctuation marks, and the reference data at least includes a pronunciation sequence used to characterize the pronunciation of the text to be synthesized; predicting the intonation category of each clause text based on the text to be synthesized; fusing to obtain a sequence to be synthesized based on the reference data and the intonation category; and synthesizing a synthesized voice corresponding to the text to be synthesized based on the sequence to be synthesized. The above solution helps to improve the intonation performance of the synthesized voice, and improves the control degree of intonation during the speech synthesis process, and finally can improve the speech synthesis effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech synthesis technology, and particularly relates to a speech synthesis method, related devices, electronic devices, and storage media. Background Art

[0002] With the rapid development of electronic information technology, speech synthesis has been widely applied in many scenarios such as mobile phone assistants and telephone customer service.

[0003] Currently, in traditional speech synthesis technology, the synthesized speech finally obtained still has problems such as poor synthesis effect, which will more or less cause some impacts in the actual application process. For example, due to the large deviation from the pronunciation of real people, the experience is greatly reduced. In view of this, how to improve the speech synthesis effect has become an urgent problem to be solved. Summary of the Invention

[0004] The main technical problem to be solved by this application is to provide a speech synthesis method, related devices, electronic devices, and storage media, which can improve the speech synthesis effect.

[0005] To solve the above technical problem, in the first aspect of this application, a speech synthesis method is provided, including: obtaining the text to be synthesized and reference data of the text to be synthesized; wherein, the text to be synthesized includes several clause texts ending with clause punctuation marks, and the reference data at least includes a pronunciation sequence for characterizing the pronunciation of the text to be synthesized; predicting the intonation category of each clause text based on the text to be synthesized; fusing to obtain a sequence to be synthesized based on the reference data and the intonation category; and synthesizing the synthesized speech corresponding to the text to be synthesized based on the sequence to be synthesized.

[0006] To solve the above technical problem, in the second aspect of this application, a speech synthesis device is provided, including: an obtaining module, a predicting module, a fusing module, and a synthesizing module. The obtaining module is used to obtain the text to be synthesized and reference data of the text to be synthesized; wherein, the text to be synthesized includes several clause texts ending with clause punctuation marks, and the reference data at least includes a pronunciation sequence for characterizing the pronunciation of the text to be synthesized; the predicting module is used to predict the intonation category of each clause text based on the text to be synthesized; the fusing module is used to fuse to obtain a sequence to be synthesized based on the reference data and the intonation category; and the synthesizing module is used to synthesize the synthesized speech corresponding to the text to be synthesized based on the sequence to be synthesized.

[0007] To solve the above technical problem, in the third aspect of this application, an electronic device is provided, including a memory and a processor coupled to each other. Program instructions are stored in the memory, and the processor is used to execute the program instructions to implement the speech synthesis method in the first aspect above.

[0008] To solve the above technical problems, a fourth aspect of the present application provides a computer-readable storage medium storing program instructions that can be run by a processor, and the program instructions are used to implement the speech synthesis method in the first aspect above.

[0009] In the above solution, the text to be synthesized and reference data of the text to be synthesized are obtained, and the text to be synthesized includes several clause texts ending with clause punctuation marks. The reference data includes at least a pronunciation sequence used to characterize the pronunciation of the text to be synthesized. On this basis, based on the text to be synthesized, the intonation category of each clause text is predicted, and based on the reference data and the intonation category, a sequence to be synthesized is fused. Thus, based on the sequence to be synthesized, the synthesized speech corresponding to the text to be synthesized is synthesized. Furthermore, in the process of speech synthesis, not only the pronunciation sequence is concerned, but also the intonation category of each clause text is further predicted, and the synthesized speech is finally synthesized by combining parameter data including at least the pronunciation sequence and the intonation category, which helps to improve the intonation performance of the synthesized speech and improve the control degree of intonation in the speech synthesis process. Therefore, the speech synthesis effect can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 is a schematic flowchart of an embodiment of the speech synthesis method of the present application;

[0011] Figure 2 is a schematic process diagram of an embodiment of the speech synthesis method of the present application;

[0012] Figure 3 is a schematic process diagram of another embodiment of the speech synthesis method of the present application;

[0013] Figure 4 is a schematic framework diagram of an embodiment of the speech synthesis device of the present application;

[0014] Figure 5 is a schematic framework diagram of an embodiment of the electronic device of the present application;

[0015] Figure 6 is a schematic framework diagram of an embodiment of the computer-readable storage medium of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0016] The solutions of the embodiments of the present application will be described in detail below with reference to the accompanying drawings of the specification.

[0017] In the following description, specific details such as specific system architectures, interfaces, and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the present application.

[0018] The terms "system" and "network" are often used interchangeably herein. The term "and / or" in this document is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this document generally represents an "or" relationship between the related objects before and after. In addition, "multiple" in this document means two or more than two.

[0019] Please refer to Figure 1 , Figure 1 is a schematic flowchart of an embodiment of the speech synthesis method of this application.

[0020] Specifically, it may include the following steps:

[0021] Step S11: Obtain the text to be synthesized and the reference data of the text to be synthesized.

[0022] In the embodiments of the present disclosure, the text to be synthesized includes several clause texts ending with clause punctuation, and the reference data includes at least a pronunciation sequence for characterizing the pronunciation of the text to be synthesized.

[0023] In an implementation scenario, the language to which the text to be synthesized belongs can be set according to the actual application scenario, which is not limited herein. Exemplarily, the text to be synthesized may only involve one language. For example, the text to be synthesized "Do you have a meeting at six o'clock tomorrow?" only involves Chinese; or, the text to be synthesized "what is the weather like today?" only involves English. Of course, the text to be synthesized may also involve multiple languages (such as, involving two languages, involving three languages, etc.). For example, the text to be synthesized "Please make sure to submit this manuscript before the agreed deadline." involves Chinese and English. Other situations can be inferred by analogy and will not be exemplified one by one here.

[0024] In an implementation scenario, the text to be synthesized can be obtained after text normalization processing of the user's input text. It should be noted that the text normalization processing may include but is not limited to: string preprocessing, non-word character normalization processing, word segmentation, etc., which are not limited herein.

[0025] In a specific implementation scenario, the string preprocessing may specifically include but is not limited to: correcting typos, unifying character encoding, etc., which are not limited herein. Please refer to Figure 2 , Figure 2 is a schematic diagram of the process of an embodiment of the speech synthesis method of this application. As Figure 2As shown, for the input text "Do you have a meeting at 18:0O tomorrow?", an obvious typo can be identified, that is, the second zero in "18:0O" is miswritten as the English letter "O", so it can be corrected to obtain the preprocessed input text "Do you have a meeting at 18:00 tomorrow?". Other cases can be inferred by analogy, and no more examples will be given here.

[0026] In a specific implementation scenario, through non-word character normalization, non-word characters in the preprocessed input text can be normalized into words. Please continue to refer to Figure 2 , still taking the aforementioned preprocessed input text "Do you have a meeting at 18:00 tomorrow?" as an example, after non-word normalization, the non-word character "18:00" in it can be processed into the word "six o'clock in the evening". Therefore, after the preprocessed input text "Do you have a meeting at 18:00 tomorrow?" is further processed by non-word normalization, the processed input text "Do you have a meeting at six o'clock in the evening tomorrow?" can be obtained. Other cases can be inferred by analogy, and no more examples will be given here.

[0027] In a specific implementation scenario, through word segmentation, the input text after the aforementioned preprocessing and non-word character normalization can be further divided into the smallest units by words. Please continue to refer to Figure 2 , still taking the aforementioned input text "Do you have a meeting at six o'clock in the evening tomorrow?" after preprocessing and non-word normalization as an example, after word segmentation, it can be obtained as "you / tomorrow / evening / six o'clock / have / meeting / right?", and this is used as the text to be synthesized, where the " / " symbol represents the word segmentation symbol, that is, the symbol " / " is used to separate different word segments. Other cases can be inferred by analogy, and no more examples will be given here. It should be noted that the above word segmentation process can be achieved through a word segmentation tool. The word segmentation tool can include but is not limited to Jieba Segmentation, etc., and no limitation is made here. For the specific process of word segmentation, the technical details of word segmentation tools such as Jieba Segmentation can be referred to, and no more details will be elaborated here.

[0028] In an implementation scenario, as mentioned above, the text to be synthesized may include several clause texts ending with clause punctuation marks. Further, the clause punctuation marks may include, but are not limited to: commas, colons, semicolons, periods, exclamation marks, question marks, etc., which are not limited herein. Still taking the aforementioned text to be synthesized "Will you have a meeting at six o'clock tomorrow evening?" as an example, since it only contains one clause punctuation mark, it can be considered that the text to be synthesized only includes one clause text, and this clause text is the text to be synthesized itself; or, taking the text to be synthesized "The weather is really nice today. Let's go for an outing!" as an example, since it contains two clause punctuation marks, it can be considered that the text to be synthesized includes two clause texts. The first clause text is "The weather is really nice today," ending with the clause punctuation mark "," (i.e., comma), and the second clause text is "Let's go for an outing!" ending with the clause punctuation mark "!" (i.e., exclamation mark). Other cases can be inferred by analogy and will not be exemplified one by one here. In addition, it should be noted that the embodiments of the present disclosure do not limit the specific number of clause texts included in the text to be synthesized. For example, the text to be synthesized may only include one clause text, may include two clause texts, may include three clause texts, or may also include more than three clause texts, which are not limited herein.

[0029] In an implementation scenario, as mentioned above, the reference data of the text to be synthesized may include a pronunciation sequence representing the pronunciation of the text to be synthesized. Specifically, the pronunciation sequence may include phonemes and tones. Please refer to Table 1, which is a schematic table of an embodiment of phonemes and tones. As shown in Table 1, for the text to be synthesized "Have you eaten?" its phonemes can be expressed as: nichi f an l e, and the tones corresponding to the above phonemes can be expressed as: / 3 / 1 / 4 / 0, that is, it means that the phoneme "n" has no tone, the first phoneme "i" is the third tone, the phoneme "ch" has no tone, the second phoneme "i" is the first tone, and so on. Similarly, for the text to be synthesized "Have you eaten?" its phonemes can be expressed as: n ichi f an l e, and the tones corresponding to the above phonemes can be expressed as: / 3 / 1 / 4 / 0. For the specific meaning, please refer to the above description and will not be elaborated here.

[0030] Table 1 Schematic table of an embodiment of phonemes and tones

[0031] Text to be synthesized Have you eaten? Did you eat? Phoneme n ichi f an l e n ichi f an l e Tone (0 - 4, / indicates none) / 3 / 1 / 4 / 0 / 3 / 1 / 4 / 0

[0032] In a specific implementation scenario, please continue to refer to Figure 2For each word segment in the text to be synthesized, it can be determined whether the word segment is included in the preset dictionary (such as the Chinese dictionary, etc.). If the word segment has been included in the preset dictionary, the preset dictionary can be used to obtain the phonemes and tones corresponding to the word segment. On the contrary, if the word segment is not included in the preset dictionary, pronunciation prediction can be performed to obtain the phonemes and tones corresponding to the word segment (for example, if "六点" is not included in the preset dictionary, the corresponding phonemes and tones "liu4 dian3" can be predicted). It should be noted that pronunciation prediction can be implemented based on LTS (LetterToSound). The specific process can refer to the technical details of LTS, which will not be repeated here. Furthermore, for the segmented words that have been included in the preset dictionary, it can be further determined whether they are polyphonic characters. If so, the phonemes and tones corresponding to the segmented words can be determined in combination with the context information of the text to be synthesized (e.g., "会" has two pronunciations, "kuai4" and "hui4", and the corresponding phonemes and tones should be "hui4" according to the context of the text to be synthesized "你明天晚六点吗"). On the contrary, if it is determined that it is not a polyphonic character, the phonemes and tones corresponding to the segmented words can be directly obtained from the preset dictionary (e.g., "你" is not a polyphonic character, and its corresponding phonemes and tones "ni3" can be directly determined according to the preset dictionary). Other situations can be deduced by analogy, and no examples are given here one by one.

[0033] In a specific implementation scenario, please continue to refer to Figure 2 After obtaining the phonemes and tones corresponding to each word in the text to be synthesized, the phonemes and tones corresponding to each word can be combined in order according to the position of each word in the text to be synthesized, so as to obtain a pronunciation sequence. Still taking the text to be synthesized "Do you have a meeting at six o'clock tomorrow night?" as an example, by combining the phonemes and tones corresponding to each word, the final pronunciation sequence "ni3 / ming2tian1 / wan3shang4 / liu4dian3 / you3 / hui4 / me0" can be obtained. Other cases can be deduced by analogy, and no examples are given here one by one.

[0034] In one implementation scenario, the reference data of the text to be synthesized may include the prosodic boundaries of words in the text to be synthesized. Figure 2, the prosodic boundary can be obtained through prosodic prediction, such as "|", "||" can be used to represent different categories of prosodic boundaries. Exemplarily, in actual application, different prosodic boundaries may include but are not limited to: phoneme prosodic level (i.e. L0), word prosodic level (i.e. L1), breath break prosodic level (i.e. L3), sentence prosodic level (i.e. L4), etc., which are not limited here. For the specific meaning of the prosodic boundary, please refer to the technical details related to prosody, which will not be repeated here. Please refer to Table 2, which is a schematic table of an embodiment of a prosodic boundary. As shown in Table 2, for the text to be synthesized "You've eaten.", the prosodic boundaries of each phoneme are "HHM MMM TT", where "H" represents head, "M" represents middle, and "T" represents tail. In addition, for the phoneme that becomes a prosodic boundary alone, it can be represented by "S" (i.e. single), which is not limited here. Similarly, for the text to be synthesized “Have you eaten?”, the prosodic boundaries of each phoneme can also be “HHM MMM TT” respectively. Other situations can be deduced by analogy, and no examples are given here one by one.

[0035] Table 2 Schematic diagram of an embodiment of rhythmic boundary

[0036]

[0037] Step S12: Based on the text to be synthesized, predict the intonation category of each sentence text.

[0038] In one implementation scenario, the text to be synthesized can be predicted based on the intonation prediction model to obtain the intonation category of each sentence text. It should be noted that the intonation category may include but is not limited to: rising tone, falling tone, and others, which are not limited here. Of course, the intonation category may also include: rising tone, falling tone, flat tone, etc., which are not limited here. The specific types of intonation categories can be set according to the actual application scenario and are not limited here. In the above method, the text to be synthesized is predicted based on the intonation prediction model to obtain the intonation category of each sentence text, thereby avoiding manual recognition, which is conducive to greatly reducing labor costs, and because the intonation prediction is performed through the intonation prediction model, it can also avoid problems such as intonation recognition errors caused by inconsistent manual cognition as much as possible, which can improve the efficiency and accuracy of intonation prediction.

[0039] In a specific implementation scenario, the intonation prediction model can be trained based on sample texts. Similar to the text to be synthesized, the sample texts can also include several sample clause texts, and each sample clause text in the sample texts is labeled with a sample intonation category. It should be noted that the sample intonation category represents the intonation category when the sample clause text is pronounced (e.g., rising tone, falling tone, flat tone, etc.). In addition, for the specific meaning of the sample clause text and the basis for defining different sample clause texts, reference can be made to the aforementioned relevant descriptions about the clause text, which will not be elaborated here. In the above manner, since the intonation prediction model is trained based on sample texts, the sample texts include several sample clause texts, and the sample clause texts are labeled with sample intonation categories, the intonation prediction model can learn the intonation features when different clause texts are pronounced through supervised training during the training process, which is beneficial to improving the accuracy of the intonation prediction model. In addition, compared with unsupervised training, since supervised training can better supervise the model training through annotation information, it can help reduce the requirements for the number and duration of speakers in the voice library, thereby improving the adaptability to the multi-language single voice library scenario.

[0040] In a specific implementation scenario, the intonation prediction model can include, but is not limited to, Bi-LSTM (Bi-directional Long Short Term Memory), GRU (Gated Recurrent Unit), CRF (Conditional Random Field), etc. The network structure of the intonation prediction model is not limited here.

[0041] In a specific implementation scenario, when marking the sample intonation category of the sample sentence text, the sample intonation category of each sample sentence text can be obtained based on the preset intonation rules and the sentence punctuation of each sample sentence text. Specifically, the preset intonation rules can define the intonation categories of each language at different punctuation marks. For example, in French, commas, colons, semicolons, and punctuation marks such as semicolons that indicate pauses within sentences generally have a slightly rising tone, while periods and exclamation marks generally have a falling tone, and question marks generally have a rising tone. Sample intonation categories such as "other" (or, "other"), "falling tone" (or "fall"), and "rising tone" (or, "rise") can be marked based on the preset intonation rules. Other situations can be deduced by analogy, and examples are not given one by one here. In addition, in order to further improve the accuracy of intonation annotation, after obtaining the sample intonation category of each sample sentence text based on the preset intonation rules and the sentence punctuation of each sample sentence text, the user's correction instruction for the sample intonation category can also be accepted to adapt to the application scenario of intonation annotation for some languages ​​with more complex intonation rules. For example, Italian has a rising tone, a falling tone, a falling tone first and then a rising tone at the question mark. Therefore, after obtaining the sample intonation category of the sample sentence text through the preset intonation rules, the user's correction instruction can be accepted to correct it to the correct sample intonation category. The above method, based on the preset intonation rules and the sentence punctuation of each sample sentence text, obtains the sample intonation category of each sample sentence text, which can reduce the manual participation in the intonation annotation process as much as possible, which is conducive to greatly reducing the annotation cost and annotation difficulty.

[0042] In a specific implementation scenario, different from the foregoing method of annotating the sample intonation category through a preset intonation rule, the sample intonation category of each sample clause text can also be obtained based on the fundamental frequency change of the fundamental frequency data extracted from the sample audio corresponding to the sample text at the reference position, and the reference position is the position corresponding to the fundamental frequency data at the end of the sample clause text. It should be noted that during the training process, the sample audio can be obtained in advance, and the sample text can be obtained by transcribing the sample audio. The transcribing can be done manually or by a pre-trained speech recognition model, which is not limited here. In addition, for the specific meaning of the fundamental frequency and the specific extraction process of the fundamental frequency data, the relevant technical details of the fundamental frequency can be referred to, which will not be elaborated here. The above method obtains the sample intonation category of each sample clause text based on the fundamental frequency change of the fundamental frequency data extracted from the sample audio corresponding to the sample text at the reference position, which can fully refer to the sample audio corresponding to the sample text during the process of annotating the sample intonation category, and is beneficial to further reducing the participation of humans in the intonation annotation process as much as possible, and further greatly reducing the annotation cost and the annotation difficulty. Exemplarily, after obtaining the fundamental frequency change at the position corresponding to the fundamental frequency data at the end of the sample clause text, if the fundamental frequency change includes a fundamental frequency increase, the sample intonation category can be determined to be a rising tone, and if the fundamental frequency change includes a fundamental frequency decrease, the sample intonation category can be determined to be a falling tone. It should be noted that in order to further improve the accuracy of the sample intonation category, the fundamental frequency change of a continuous preset number of frames at the reference position can be statistically analyzed. If the fundamental frequency change of the continuous preset number of frames is a fundamental frequency increase, the sample intonation category can be determined to be a rising tone, and vice versa, if the fundamental frequency change of the continuous preset number of frames is a fundamental frequency decrease, the sample intonation category can be determined to be a falling tone. The above method determines the sample intonation category to be a rising tone if the fundamental frequency change includes a fundamental frequency increase, and determines the sample intonation category to be a falling tone if the fundamental frequency change includes a fundamental frequency decrease. Therefore, during the sample intonation annotation process, it is only necessary to directly detect whether the fundamental frequency increases or decreases at the reference position to determine the sample intonation category, which is beneficial to further greatly reducing the annotation cost and the annotation difficulty.

[0043] In a specific implementation scenario, after the sample intonation categories of each sample clause text in the sample text are labeled, the intonation prediction model can be trained using the sample text. Specifically, the sample text can be predicted based on the intonation prediction model to obtain the predicted intonation categories of each sample clause text, and the network parameters of the intonation prediction model can be adjusted based on the differences between the predicted intonation categories of each sample clause text and the sample intonation categories. It should be noted that the intonation prediction model can include, but is not limited to, Bi-LSTM, GRU, CRF, etc., and the network structure of the intonation prediction model is not limited here. In the above manner, the sample text is predicted based on the intonation prediction model to obtain the predicted intonation categories of each sample clause text, and the network parameters of the intonation prediction model are adjusted based on the differences between the predicted intonation categories of each sample clause text and the sample intonation categories. Therefore, during the training process, the model training can be supervised by the labeled sample intonation categories, improving the prediction accuracy of the intonation prediction model.

[0044] In a specific implementation scenario, the intonation prediction model can specifically predict the predicted probability values that the sample clause text belongs to several preset intonation categories. Exemplarily, the several preset intonation categories can include, but are not limited to: rising tone (i.e., rise), falling tone (i.e., fall), other (i.e., other), etc., and are not limited here. On this basis, the preset intonation category corresponding to the maximum predicted probability value can be used as the predicted intonation category.

[0045] In a specific implementation scenario, as described above, the intonation prediction model can specifically predict the predicted probability values that the sample clause text belongs to several preset intonation categories. In this case, based on the sample intonation category, the cross-entropy loss function can be used to process the above predicted probability values to obtain the prediction loss of the intonation prediction model, and the network parameters of the intonation prediction model can be adjusted based on this prediction loss. It should be noted that for the measurement process of the prediction loss, the technical details of, such as, the cross-entropy loss function can be referred to, and for the adjustment process of the network parameters, the technical details of, such as, optimization methods like gradient descent can be referred to, which will not be elaborated here.

[0046] In a specific implementation scenario, after training the intonation prediction model, the trained and converged intonation prediction model is used to predict the to-be-synthesized text, and the probability values of the segmented text belonging to several preset intonation categories are obtained. On this basis, for each segmented text, the preset intonation category corresponding to the maximum value among the probability values of the segmented text belonging to several preset intonation categories can be used as the intonation category of the segmented text. Exemplarily, the probability values of the segmented text belonging to the rising tone (i.e., rise), falling tone (i.e., fall), and others (i.e., other) are 0.8, 0.15, and 0.05 respectively, then the intonation category of the segmented text can be determined as the rising tone (i.e., rise). Other cases can be deduced by analogy and will not be exemplified one by one here.

[0047] In an implementation scenario, different from the foregoing method of predicting the intonation category through the intonation prediction model, the intonation category of each segmented text can also be obtained based on the preset intonation rules and the segmented punctuation of each segmented text. It should be noted that the specific types of intonation categories can refer to the foregoing relevant descriptions and will not be elaborated here. The above method, obtaining the intonation category of each segmented text based on the preset intonation rules and the segmented punctuation of each segmented text, can minimize the participation of manual labor in the intonation prediction process, which is beneficial to greatly reducing the cost and difficulty of intonation prediction.

[0048] In a specific implementation scenario, as described above, the preset intonation rules can specifically define the intonation categories at different punctuation marks for each language, which can refer to the foregoing relevant descriptions and will not be elaborated here. On this basis, the target language to which the to-be-synthesized text belongs can be obtained, and the intonation category of the segmented text can be obtained based on the intonation category of the target language at the segmented punctuation in the preset intonation rules. Exemplarily, the to-be-synthesized text is in French, and the segmented punctuation of the segmented text is a question mark. Then, based on the foregoing relevant descriptions of the preset intonation rules, since the intonation category of French at the question mark is generally the rising tone, the intonation category of the segmented text can be determined as the rising tone. Other cases can be deduced by analogy and will not be exemplified one by one here. The above method, obtaining the target language to which the to-be-synthesized text belongs and obtaining the intonation category of the segmented text based on the intonation category of the target language at the segmented punctuation in the preset intonation rules, can determine the intonation category in combination with different languages, which is beneficial to improving the accuracy of intonation prediction.

[0049] In a specific implementation scenario, it should be noted that when using the preset intonation rules to determine the intonation category of the sample segmented text, categories such as "slightly rising tone" and "rising tone" in Italian can be not subdivided and uniformly determined as the intonation category "rising tone".

[0050] In an implementation scenario, to further improve the accuracy of intonation prediction, the intonation categories of each clause text can also be predicted by combining the above two methods. Exemplarily, the text to be synthesized can be predicted based on the intonation prediction model to obtain the intonation categories of each clause text. It should be noted that when using the intonation prediction model for prediction, the confidence levels of the intonation categories of each clause text can also be obtained. The greater the confidence level, the higher the probability that the predicted intonation category of the clause text is correct. Conversely, the smaller the confidence level, the lower the probability that the predicted intonation category of the clause text is correct. In this case, for the clause text with a confidence level of the predicted intonation category lower than the preset threshold, the intonation category of the clause text can be further obtained based on the preset intonation rules and the clause punctuation of the clause text. Of course, it is also possible to first obtain the intonation categories of each clause text based on the preset intonation rules and the clause punctuation of each clause text, and then further determine the final intonation categories of each clause text in combination with the intonation prediction model. This is not limited here. The above method combines the intonation prediction model and the preset intonation rules to jointly determine the intonation categories of each clause text, which is conducive to further improving the accuracy of intonation categories.

[0051] In an implementation scenario, please refer to Table 3, which is a schematic table of an embodiment of intonation categories. As shown in Table 3, for the text to be synthesized "Have you eaten?", the intonation categories of each phoneme are respectively " / / / / / / F F", while for the text to be synthesized "Have you eaten?", the intonation categories of each phoneme are respectively " / / / / / / R R", where " / " represents none, "F" represents a falling tone, and "R" represents a rising tone. Other cases can be deduced by analogy and will not be exemplified one by one here.

[0052] Table 3 Schematic Table of an Embodiment of Intonation Categories

[0053] Text to be synthesized Have you eaten? Did you eat? Whether the sentence ends with a question mark 0 (i.e., no) 1 (i.e., yes) Intonation category (Fall / Rise / none) / / / / / / F F / / / / / / RR

[0054] In an implementation scenario, please continue to refer to Figure 2 , as mentioned above, the text to be synthesized "Do you have a meeting at six o'clock tomorrow evening?" only contains one clause text (i.e., itself). In this case, intonation prediction can be performed based on the intonation prediction model and / or the preset intonation rules (as shown by the thick dashed box "Prosody Prediction" in Figure 2 ) to obtain the intonation category, that is, a rising tone (i.e., rise). Other cases can be deduced by analogy and will not be exemplified one by one here. It should be noted that different from only performing boundary prediction in the prosody prediction process (as shown by the thick dashed box "Prosody Prediction" in Figure 2 ), in the prosody prediction process, boundary prediction and intonation prediction can better assist speech synthesis, which is conducive to further improving the speech synthesis effect.

[0055] Step S13: Based on the reference data and intonation categories, fuse to obtain the sequence to be synthesized.

[0056] In an implementation scenario, based on the position of the end of the clause text in the text to be synthesized, the insertion position corresponding to the intonation category of the clause text in the pronunciation sequence can be obtained, and a character representing the intonation category is inserted at this insertion position to obtain the sequence to be synthesized. For example, if the end of the clause text is also the end of the text to be synthesized, it can be determined that the insertion position is also the end of the pronunciation sequence, and then the character representing the intonation category can be inserted at the end of the pronunciation sequence to obtain the sequence to be synthesized. It should be noted that the characters representing the intonation category can include but are not limited to: (rise), (fall), (other), which represent rising tone, falling tone, and others respectively. Here, the intonation category and the characters representing the intonation category are not limited.

[0057] In an implementation scenario, as mentioned above, the reference data can also include the prosodic boundaries of the words in the text to be synthesized. Then, based on the position of the word to which the prosodic boundary belongs in the text to be synthesized, the first position corresponding to the prosodic boundary can be obtained, and based on the position of the end of the clause text in the text to be synthesized, the second position corresponding to the intonation category of the clause text can be obtained. On this basis, a first character ensuring the prosodic boundary can be inserted at the first position, and a second character representing the intonation category can be inserted at the second position to obtain the sequence to be synthesized. Please refer to Figure 2 , still taking the text to be synthesized "Will you have a meeting at six o'clock tomorrow evening?" as an example. By performing the above character insertion steps in the pronunciation sequence, the sequence to be synthesized "ni3|ming2tian1 / wan3shang4 / liu4dian3|you3 / hui4 / me0||(rise)" can be obtained. Other situations can be deduced by analogy, and no further examples will be given here. In the above manner, based on the position of the word to which the prosodic boundary belongs in the text to be synthesized, the first position corresponding to the prosodic boundary is obtained, and based on the position of the end of the clause text in the text to be synthesized, the second position corresponding to the intonation category of the clause text is obtained. On this basis, a first character representing the prosodic boundary is inserted at the first position, and a second character representing the intonation category is inserted at the second position to obtain the sequence to be synthesized. Therefore, it is possible to jointly determine the sequence to be synthesized finally used for backend synthesis through simple steps such as character insertion by combining the pronunciation sequence, prosodic boundaries, and intonation categories, which is beneficial to improving the speech synthesis efficiency.

[0058] Step S14: Based on the sequence to be synthesized, synthesize the synthetic speech corresponding to the text to be synthesized.

[0059] Specifically, each character in the sequence to be synthesized can be subjected to feature transformation to obtain a feature sequence, and the feature sequence includes the feature representations of each character, and speech synthesis is performed based on the feature sequence to obtain synthesized speech. It should be noted that the feature representations after feature transformation of different characters are also different. In addition, the feature representation can be expressed in the form of a vector, and the feature dimension of the feature representation can be set according to the actual situation, such as 128 dimensions, 256 dimensions, etc., which are not limited herein. In the above manner, each character in the sequence to be synthesized is subjected to feature transformation to obtain a feature sequence, and the feature sequence includes the feature representations of each character, and then speech synthesis is performed based on the feature sequence to obtain synthesized speech, that is, after obtaining the sequence to be synthesized, synthesized speech is finally obtained through a series of steps such as feature transformation, which can help improve the accuracy of speech synthesis.

[0060] In one implementation scenario, since the total number of the above three types of characters, namely the characters in the pronunciation sequence, the characters representing intonation categories, and the characters representing prosodic boundaries, is limited, the feature representations of different characters can be pre-constructed. On this basis, after obtaining the sequence to be synthesized, as long as the feature sequence can be obtained based on each character in the sequence to be synthesized and the pre-constructed feature representations. Exemplarily, taking the total number of the above three types of characters as N and the sequence to be synthesized containing M characters in total as an example, one-hot encoding can be used to pre-construct feature representations with a dimension of N for each character, so that a feature sequence of M*N can be obtained by combining each character in the sequence to be synthesized. For the specific construction process of the feature representation, the technical details of coding methods such as one-hot encoding can be referred to, which will not be elaborated herein.

[0061] In one implementation scenario, please refer to Figure 3 , Figure 3 which is a schematic diagram of the process of another embodiment of the speech synthesis method of the present application. As Figure 3 shown, the feature sequence can be processed by an acoustics model obtained through pre-training to obtain acoustic parameters. On this basis, the acoustic parameters can be processed by a vocoder obtained through pre-training to obtain synthesized speech.

[0062] In a specific implementation scenario, the acoustics model can include but is not limited to: Tacotron, Tacotron2, Fast2Speech, etc., and the network structure of the acoustics model is not limited herein.

[0063] In a specific implementation scenario, the vocoder can include but is not limited to: LPCNet, WaveRNN, etc., and the network structure of the vocoder is not limited herein.

[0064] In a specific implementation scenario, during the process of training an acoustic model, the sample audio can be first converted into a spectrogram, and then the spectrogram can be sampled to obtain the sample acoustic parameters corresponding to the sample audio. At the same time, the sample feature sequence of the sample text can be extracted. Specifically, the text obtained by manually transcribing or speech recognition of the sample audio can be used as the sample text. In addition, for the extraction method of the sample feature sequence, reference can be made to the specific process of extracting the feature sequence for the text to be synthesized as described above, which will not be elaborated here. On this basis, the sample feature sequence can be processed by the acoustic model to obtain the predicted acoustic parameters, and based on the difference between the sample acoustic parameters and the predicted acoustic parameters, the network parameters of the acoustic model can be adjusted. It should be noted that taking the acoustic model using Tacotron2 as an example, in contrast, for the unsupervised learning solution, a Prosody Encoder needs to be added on the basis of Tacotron2. However, since the sample feature sequence contains prosody-related feature information such as prosody boundaries and intonation categories, there is no need to add a new Prosody Encoder, and the existing encoder of the acoustic model can be directly used. In addition, for the specific process of training the acoustic model, reference can be made to the technical details of acoustic models such as Fast2Speech, which will not be elaborated here.

[0065] In a specific implementation scenario, during the process of training a vocoder, the sample acoustic parameters can be processed by the vocoder to obtain the sample synthesized speech, and based on the difference between the sample audio and the sample synthesized speech, the network parameters of the vocoder can be adjusted. In addition, for the specific process of training the vocoder, reference can be made to the technical details of vocoders such as Tacotron2, which will not be elaborated here.

[0066] In a specific implementation scenario, for the data distribution situation where the proportions of exclamatory sentences, interrogative sentences, and complex sentences in a single voice library are relatively too small or too large compared to ordinary sentence patterns, during the training process, data augmentation can be achieved by simply replicating the data with a relatively small proportion, so as to improve the tendency of end-to-end speech synthesis and facilitate optimizing the synthesis effect of intonation.

[0067] In the above solution, the text to be synthesized and the reference data of the text to be synthesized are obtained. The text to be synthesized includes several clause texts ending with clause punctuation marks, and the reference data includes at least a pronunciation sequence for characterizing the pronunciation of the text to be synthesized. On this basis, based on the text to be synthesized, the intonation category of each clause text is predicted, and based on the reference data and the intonation category, a synthesized sequence is fused. Then, based on the synthesized sequence, the synthesized speech corresponding to the text to be synthesized is synthesized. Furthermore, in the process of speech synthesis, not only the pronunciation sequence is concerned, but also the intonation category of each clause text is further predicted, and the parameter data including at least the pronunciation sequence and the intonation category are combined to finally synthesize the synthesized speech, which helps to improve the intonation performance of the synthesized speech and improve the control degree of intonation in the process of speech synthesis. Therefore, the speech synthesis effect can be improved.

[0068] Please refer to Figure 4 , Figure 4 FIG. is a schematic framework diagram of an embodiment of the speech synthesis device 40 of the present application. Specifically, the speech synthesis device 40 includes: an acquisition module 41, a prediction module 42, a fusion module 43, and a synthesis module 44. The acquisition module 41 is configured to acquire the text to be synthesized and the reference data of the text to be synthesized. Among them, the text to be synthesized includes several clause texts ending with clause punctuation marks, and the reference data includes at least a pronunciation sequence for characterizing the pronunciation of the text to be synthesized. The prediction module 42 is configured to predict the intonation category of each clause text based on the text to be synthesized. The fusion module 43 is configured to fuse a synthesized sequence based on the reference data and the intonation category. The synthesis module 44 is configured to synthesize the synthesized speech corresponding to the text to be synthesized based on the synthesized sequence.

[0069] In the above solution, the text to be synthesized and the reference data of the text to be synthesized are obtained. The text to be synthesized includes several clause texts ending with clause punctuation marks, and the reference data includes at least a pronunciation sequence for characterizing the pronunciation of the text to be synthesized. On this basis, based on the text to be synthesized, the intonation category of each clause text is predicted, and based on the reference data and the intonation category, a synthesized sequence is fused. Then, based on the synthesized sequence, the synthesized speech corresponding to the text to be synthesized is synthesized. Furthermore, in the process of speech synthesis, not only the pronunciation sequence is concerned, but also the intonation category of each clause text is further predicted, and the parameter data including at least the pronunciation sequence and the intonation category are combined to finally synthesize the synthesized speech, which helps to improve the intonation performance of the synthesized speech and improve the control degree of intonation in the process of speech synthesis. Therefore, the speech synthesis effect can be improved.

[0070] In some disclosed embodiments, the prediction module 42 includes a first prediction submodule, which is used to predict the text to be synthesized based on the intonation prediction model to obtain the intonation category of each sentence text; the prediction module 42 includes a second prediction submodule, which is used to obtain the intonation category of each sentence text based on preset intonation rules and sentence punctuation of each sentence text.

[0071] Therefore, the text to be synthesized is predicted based on the intonation prediction model to obtain the intonation category of each sentence text, thereby avoiding manual recognition, which is conducive to greatly reducing labor costs. In addition, since the intonation prediction is performed through the intonation prediction model, it can also avoid problems such as intonation recognition errors caused by inconsistent manual cognition as much as possible, which can improve the efficiency and accuracy of intonation prediction. Based on the preset intonation rules and the sentence punctuation of each sentence text, the intonation category of each sentence text is obtained, which can minimize the involvement of humans in the intonation prediction process as much as possible, which is conducive to greatly reducing the cost and difficulty of intonation prediction.

[0072] In some disclosed embodiments, the preset intonation rules define intonation categories of different languages ​​at different punctuation marks; the second prediction submodule includes a language acquisition unit, which is used to acquire the target language to which the text to be synthesized belongs; the second prediction submodule includes a category determination unit, which is used to obtain the intonation category of the sentence text based on the intonation category of the target language at the sentence punctuation marks in the preset intonation rules.

[0073] Therefore, the target language of the text to be synthesized is obtained, and the intonation category of the sentence text is obtained based on the intonation category of the target language at the sentence punctuation in the preset intonation rules. The intonation category can be determined in combination with different languages, which is conducive to improving the accuracy of intonation prediction.

[0074] In some disclosed embodiments, the intonation category is predicted based on an intonation prediction model, and the intonation prediction model is trained based on sample texts. The sample texts include a number of sample sentence texts, and each sample sentence text of the sample text is annotated with a sample intonation category.

[0075] Therefore, the intonation prediction model is trained based on sample texts, which contain several sample sentence texts, and the sample sentence texts are annotated with sample intonation categories. Therefore, supervised training can be used to enable the intonation prediction model to learn the intonation features of different sentence texts during the training process, which is conducive to improving the accuracy of the intonation prediction model. In addition, compared with unsupervised training, since supervised training can better supervise model training through annotation information, it can help reduce the requirements for the number of speakers and duration of the sound library, thereby improving the adaptability to multi-language single sound library scenarios.

[0076] In some disclosed embodiments, the speech synthesis device 40 includes a first determination module configured to obtain the sample intonation categories of respective sample clause texts based on a preset intonation rule and the clause punctuation of each sample clause text; the speech synthesis device 40 includes a second determination module configured to obtain the sample intonation categories of respective sample clause texts based on the fundamental frequency change condition of the fundamental frequency data extracted from the sample audio corresponding to the sample text at a reference position; wherein, the reference position is the position corresponding to the fundamental frequency data at the end of the sample clause text.

[0077] Therefore, obtaining the sample intonation categories of respective sample clause texts based on a preset intonation rule and the clause punctuation of each sample clause text can minimize the participation of humans in the intonation annotation process as much as possible, which is beneficial to greatly reducing the annotation cost and the annotation difficulty; while obtaining the sample intonation categories of respective sample clause texts based on the fundamental frequency change condition of the fundamental frequency data extracted from the sample audio corresponding to the sample text at a reference position can fully refer to the sample audio corresponding to the sample text during the process of annotating the sample intonation categories, which is beneficial to further minimizing the participation of humans in the intonation annotation process and further greatly reducing the annotation cost and the annotation difficulty.

[0078] In some disclosed embodiments, the second determination module includes a first response sub-module configured to determine that the sample intonation category is a rising tone in response to the fundamental frequency change condition including a rising fundamental frequency; the second determination module includes a second response sub-module configured to determine that the sample intonation category is a falling tone in response to the fundamental frequency change condition including a falling fundamental frequency.

[0079] Therefore, determining that the sample intonation category is a rising tone in response to the fundamental frequency change condition including a rising fundamental frequency, and determining that the sample intonation category is a falling tone in response to the fundamental frequency change condition including a falling fundamental frequency, so that during the process of annotating the sample intonation categories, it is possible to directly detect a rising or falling fundamental frequency at the reference position to determine the sample intonation category, which is beneficial to further greatly reducing the annotation cost and the annotation difficulty.

[0080] In some disclosed embodiments, the speech synthesis device 40 includes a sample prediction module configured to predict the sample text based on an intonation prediction model to obtain the predicted intonation categories of respective sample clause texts; the speech synthesis device 40 includes a parameter adjustment module configured to adjust the network parameters of the intonation prediction model based on the difference between the predicted intonation categories and the sample intonation categories of respective sample clause texts.

[0081] Therefore, predicting the sample text based on an intonation prediction model to obtain the predicted intonation categories of respective sample clause texts, and adjusting the network parameters of the intonation prediction model based on the difference between the predicted intonation categories and the sample intonation categories of respective sample clause texts, so that during the training process, the model training can be supervised by the annotated sample intonation categories, improving the prediction accuracy of the intonation prediction model.

[0082] In some disclosed embodiments, the reference data further includes the prosodic boundaries of the words in the text to be synthesized; the fusion module 43 includes a position determination sub-module, which is configured to obtain a first position corresponding to the prosodic boundary based on the position of the word to which the prosodic boundary belongs in the text to be synthesized, and obtain a second position corresponding to the intonation category of the clause text based on the position of the end of the clause text in the text to be synthesized; the fusion module 43 includes a character insertion sub-module, which is configured to insert a first character representing the prosodic boundary at the first position, and insert a second character representing the intonation category at the second position to obtain the sequence to be synthesized.

[0083] Therefore, a first position corresponding to the prosodic boundary is obtained based on the position of the word to which the prosodic boundary belongs in the text to be synthesized, and a second position corresponding to the intonation category of the clause text is obtained based on the position of the end of the clause text in the text to be synthesized. On this basis, a first character representing the prosodic boundary is inserted at the first position, and a second character representing the intonation category is inserted at the second position to obtain the sequence to be synthesized. Therefore, the sequence to be synthesized finally used for backend synthesis can be jointly determined by combining the pronunciation sequence, prosodic boundary and intonation category through simple steps such as character insertion, which is beneficial to improving the speech synthesis efficiency.

[0084] In some disclosed embodiments, the synthesis module 44 includes a feature transformation sub-module, which is configured to perform feature transformation on each character in the sequence to be synthesized to obtain a feature sequence; wherein, the feature sequence includes the feature representations of each character; the synthesis module 44 includes a speech synthesis sub-module, which is configured to perform speech synthesis based on the feature sequence to obtain synthesized speech.

[0085] Therefore, each character in the sequence to be synthesized is subjected to feature transformation to obtain a feature sequence, and the feature sequence includes the feature representations of each character, and then speech synthesis is performed based on the feature sequence to obtain synthesized speech, that is, after obtaining the sequence to be synthesized, synthesized speech is finally obtained through a series of steps such as feature transformation, which is beneficial to improving the speech synthesis accuracy.

[0086] Please refer to Figure 5 , Figure 5 which is a schematic framework diagram of an embodiment of the electronic device 50 of the present application. The electronic device 50 includes a memory 51 and a processor 52 that are coupled to each other. Program instructions are stored in the memory 51, and the processor 52 is configured to execute the program instructions to implement the steps in any of the above-mentioned speech synthesis method embodiments. Specifically, the electronic device 50 may include, but is not limited to: desktop computers, laptop computers, servers, mobile phones, tablet computers, etc., which are not limited herein.

[0087] Specifically, the processor 52 is used to control itself and the memory 51 to implement the steps in any of the above-described embodiments of the speech synthesis method. The processor 52 may also be referred to as a CPU (Central Processing Unit). The processor 52 may be an integrated circuit chip with signal processing capabilities. The processor 52 may also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 52 may be implemented jointly by integrated circuit chips.

[0088] In the above solution, during the speech synthesis process, not only the pronunciation sequence is concerned, but also the intonation categories of the texts of each clause are further predicted, and finally the synthesized speech is obtained by combining the parameter data at least including the pronunciation sequence and the intonation categories, which helps to improve the intonation performance of the synthesized speech and the control degree of the intonation during the speech synthesis process. Therefore, the speech synthesis effect can be improved.

[0089] Please refer to Figure 6 , Figure 6 which is a schematic framework diagram of an embodiment of the computer-readable storage medium 60 of the present application. The computer-readable storage medium 60 stores program instructions 61 that can be run by a processor, and the program instructions 61 are used to implement the steps in any of the above-described embodiments of the speech synthesis method.

[0090] In the above solution, during the speech synthesis process, not only the pronunciation sequence is concerned, but also the intonation categories of the texts of each clause are further predicted, and finally the synthesized speech is obtained by combining the parameter data at least including the pronunciation sequence and the intonation categories, which helps to improve the intonation performance of the synthesized speech and the control degree of the intonation during the speech synthesis process. Therefore, the speech synthesis effect can be improved.

[0091] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0092] The descriptions of the above embodiments tend to emphasize the differences between the embodiments. The similarities or similarities between them can be referred to each other. For the sake of brevity, they will not be repeated in this article.

[0093] In several embodiments provided by the present application, it should be understood that the disclosed methods and apparatuses can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the apparatuses or units can be in electrical, mechanical or other forms.

[0094] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0095] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0096] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in each embodiment of the present application. And the aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks or optical disks and other various media that can store program codes.

Claims

1. A voice synthesis method, characterized in that Including: Obtaining the text to be synthesized and reference data of the text to be synthesized; wherein, the text to be synthesized includes several clause texts ending with clause punctuation marks, and the reference data at least includes a pronunciation sequence for characterizing the pronunciation of the text to be synthesized and prosodic boundaries of words in the text to be synthesized; Based on the text to be synthesized, predicting the intonation categories of each of the clause texts; Based on the position of the word to which the prosodic boundary belongs in the text to be synthesized, obtaining a first position corresponding to the prosodic boundary, and based on the position of the end of the clause text in the text to be synthesized, obtaining a second position corresponding to the intonation category of the clause text; Inserting a first character representing the prosodic boundary at the first position, and inserting a second character representing the intonation category at the second position to obtain a sequence to be synthesized; Based on the sequence to be synthesized, synthesizing a synthesized voice corresponding to the text to be synthesized.

2. The method according to claim 1, wherein The predicting the intonation categories of each of the clause texts based on the text to be synthesized includes at least one of the following: Predicting the text to be synthesized based on an intonation prediction model to obtain the intonation categories of each of the clause texts; Based on a preset intonation rule and the clause punctuation marks of each of the clause texts, obtaining the intonation categories of each of the clause texts.

3. The method according to claim 2, wherein The preset intonation rule defines intonation categories at different punctuation marks for each language; the obtaining the intonation categories of each of the clause texts based on the preset intonation rule and the clause punctuation marks of each of the clause texts includes: Obtaining the target language to which the text to be synthesized belongs; Based on the intonation categories of the target language at the clause punctuation marks in the preset intonation rule, obtaining the intonation categories of the clause texts.

4. The method according to claim 1, characterized in that, The intonation categories are obtained by prediction based on an intonation prediction model, and the intonation prediction model is trained based on a sample text, the sample text includes several sample clause texts, and each of the sample clause texts in the sample text is labeled with a sample intonation category.

5. The method according to claim 4, characterized in that, The labeling steps of the sample intonation categories include at least one of the following: Based on a preset intonation rule and the clause punctuation marks of each of the sample clause texts, obtaining the sample intonation categories of each of the sample clause texts; Based on the fundamental frequency change situation of the fundamental frequency data extracted from the sample audio corresponding to the sample text at a reference position, obtaining the sample intonation categories of each of the sample clause texts; wherein, the reference position is the position of the end of the sample clause text corresponding to the fundamental frequency data.

6. The method according to claim 5, wherein The obtaining the sample intonation categories of each of the sample clause texts based on the fundamental frequency change situation of the fundamental frequency data extracted from the sample audio corresponding to the sample text at a reference position includes: In response to the fundamental frequency change situation including a rising fundamental frequency, determining that the sample intonation category is a rising tone; And / or, in response to the fundamental frequency change situation including a falling fundamental frequency, determining that the sample intonation category is a falling tone.

7. The method according to claim 4, characterized in that The training steps of the intonation prediction model include: Predicting the sample text based on the intonation prediction model to obtain the predicted intonation categories of each of the sample clause texts; Adjust the network parameters of the intonation prediction model based on the differences between the predicted intonation categories of each of the sample clause texts and the sample intonation categories.

8. The method according to claim 1, characterized in that, Synthesizing the synthesized speech corresponding to the text to be synthesized based on the sequence to be synthesized includes: Performing feature transformation on each character in the sequence to be synthesized to obtain a feature sequence; wherein the feature sequence includes feature representations of each of the characters; Performing speech synthesis based on the feature sequence to obtain the synthesized speech.

9. A voice synthesis device, characterized in that, Includes: An acquisition module for acquiring the text to be synthesized and reference data for the text to be synthesized; wherein the text to be synthesized includes several clause texts ending with clause punctuation marks, and the reference data includes at least a pronunciation sequence for characterizing the pronunciation of the text to be synthesized and prosodic boundaries of words in the text to be synthesized; A prediction module for predicting the intonation category of each of the clause texts based on the text to be synthesized; A fusion module for obtaining a first position corresponding to the prosodic boundary based on the position of the word to which the prosodic boundary belongs in the text to be synthesized, and obtaining a second position corresponding to the intonation category of the clause text based on the position of the end of the clause text in the text to be synthesized; inserting a first character representing the prosodic boundary at the first position, and inserting a second character representing the intonation category at the second position to obtain a sequence to be synthesized; A synthesis module for synthesizing the synthesized speech corresponding to the text to be synthesized based on the sequence to be synthesized.

10. An electronic device, characterized in that, Includes a memory and a processor coupled to each other, wherein program instructions are stored in the memory, and the processor is configured to execute the program instructions to implement the speech synthesis method according to any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that, Stores program instructions that can be run by a processor, and the program instructions are used to implement the speech synthesis method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Speech synthesis method and device, readable storage medium and electronic equipment

    CN114155829A