Speech synthesis method and apparatus, computer readable medium, and electronic device

By generating TOBI representation sequences and prosodic acoustic features at the phoneme level and combining them with a speech synthesis model, the problem of insufficient integration of prosodic features in existing technologies is solved, and the naturalness of synthesized audio is improved and the semantic expression is controlled.

CN114495902BActive Publication Date: 2025-10-17BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210179831.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-25
Publication Date
2025-10-17
Estimated Expiration
2042-02-25

AI Technical Summary

Technical Problem

Existing speech synthesis technology is difficult to effectively combine the rhythmic features of text, resulting in insufficient naturalness of the synthesized audio, uncontrollable intensity, and inability to accurately express semantic changes and emotions.

Method used

By generating TOBI representation sequences and prosodic acoustic features at the phoneme level and combining them with a speech synthesis model, acoustic feature information is generated to control the rhythm, emphasis, and intonation of the audio, thereby improving the naturalness of the prosody and semantic expression.

Benefits of technology

The naturalness of the rhythm of the synthesized audio is improved, the audio intensity is controllable, and different semantic changes can be reflected under the same rhythm, which is consistent with the semantic expression of the speaker's intention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114495902B_ABST
    Figure CN114495902B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a speech synthesis method, device, computer readable medium and electronic device. The method comprises: obtaining a phoneme sequence corresponding to a text to be synthesized; generating a TOBI representation sequence and prosody acoustic features corresponding to the text to be synthesized according to the phoneme sequence and the text to be synthesized, generating acoustic feature information corresponding to the text to be synthesized according to the TOBI representation sequence and the prosody acoustic features; and generating first audio information corresponding to the text to be synthesized according to the acoustic feature information. The TOBI representation sequence can give different sentences appropriate rhythm, emphasis and intonation characteristics, and the prosody acoustic features can explicitly represent the specific acoustic manifestations of the corresponding prosody events, thereby improving the prosody naturalness of the synthesized audio while controlling the audio intensity. Thus, under the same prosody language performance, different prosody acoustic features can represent different semantic changes, making the synthesized audio more natural, more rhythmic, and more consistent with the speaker's expressed meaning.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of speech synthesis, in particular, to a speech synthesis method, device, computer readable medium and electronic equipment. BACKGROUND

[0002] In linguistics, prosody refers to the components of non-independent phonemes (vowels and consonants) in the process of speaking, i.e., the properties of syllables or larger units. These properties form language functions such as intonation, tone, stress, and rhythm. Prosody can reflect various characteristics of the speaker or the speech: the emotional state of the speaker, the form of the speech (statement, question, or command), whether there is emphasis, contrast, focus, and other language elements that cannot be expressed by grammar and vocabulary. Different manifestations of the same prosodic event can convey rich semantics and emotional changes. In tasks such as speech synthesis, how to combine the prosodic features of the text to make the synthesized audio more natural and smooth becomes a research focus. SUMMARY

[0003] This section is provided to introduce the general concepts of the present application in a simplified form, which will be described in detail in the following detailed description section. This section does not intend to identify key or essential features of the claimed technology nor is it intended to be used to limit the scope of the claimed technology.

[0004] In a first aspect, the present disclosure provides a speech synthesis method, comprising:

[0005] obtaining a phoneme sequence corresponding to a text to be synthesized;

[0006] generating a TOBI representation sequence at the phoneme level and prosodic acoustic features corresponding to the text to be synthesized according to the phoneme sequence and the text to be synthesized, and generating acoustic feature information corresponding to the text to be synthesized according to the TOBI representation sequence and the prosodic acoustic features;

[0007] generating first audio information corresponding to the text to be synthesized according to the acoustic feature information.

[0008] In a second aspect, the present disclosure provides a speech synthesis device, comprising:

[0009] an obtaining module configured to obtain a phoneme sequence corresponding to a text to be synthesized;

[0010] a first generating module configured to generate a TOBI representation sequence at the phoneme level and prosodic acoustic features corresponding to the text to be synthesized according to the phoneme sequence and the text to be synthesized obtained by the obtaining module, and generate acoustic feature information corresponding to the text to be synthesized according to the TOBI representation sequence and the prosodic acoustic features;

[0011] a second generating module configured to generate first audio information corresponding to the text to be synthesized according to the acoustic feature information generated by the first generating module.

[0012] In a third aspect, the present disclosure provides a computer readable medium having stored thereon a computer program, which, when executed by a processing device, implements the steps of the method provided in the first aspect of the present disclosure.

[0013] In a fourth aspect, the present disclosure provides an electronic device comprising:

[0014] a storage device having stored thereon one or more computer programs;

[0015] one or more processing devices configured to execute the one or more computer programs in the storage device to implement the steps of the method provided in the first aspect of the present disclosure.

[0016] In the above technical solution, after obtaining the phoneme sequence corresponding to the text to be synthesized, the TOBI representation sequence and the prosody acoustic feature at the phoneme level corresponding to the text to be synthesized are generated according to the phoneme sequence and the text to be synthesized, and the acoustic feature information corresponding to the text to be synthesized is generated according to the TOBI representation sequence and the prosody acoustic feature. Finally, the first audio information corresponding to the text to be synthesized is generated according to the acoustic feature information. In the speech synthesis, the TOBI representation sequence and the prosody acoustic feature corresponding to the text to be synthesized are simultaneously referred to, that is, not only the prosody feature at the language level of the text to be synthesized is referred to, but also the prosody feature at the acoustic level of the text to be synthesized is referred to, and the performance of the prosody in different dimensions is considered. The TOBI representation sequence can give different sentences appropriate rhythm, emphasis and intonation characteristics, and the corresponding prosody acoustic feature can explicitly reflect the specific acoustic manifestation of the corresponding prosody event, so as to improve the prosody naturalness of the synthesized audio while controlling the intensity (i.e. amplitude) of the audio, such as assigning different intensities to different emphasis positions to realize different semantic expression emphasis, or adjusting the intensity to realize the intonation change of the interrogative sentence to convey different semantics (emotions). Thus, different prosody acoustic features can reflect different semantic changes under the same prosody language performance, so that the synthesized audio is more natural, has more rhythm and tone, and is more consistent with the semantic expression of the speaker.

[0017] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF DRAWINGS

[0018] The above and other features, advantages, and aspects of embodiments of the present disclosure will become more apparent by describing in detail exemplary embodiments thereof with reference to the attached drawings. The same or similar components have the same or similar reference labels. It should be understood that the drawings are schematic and elements in the drawings are not necessarily to scale. In the drawings:

[0019] Figure 1 is a flowchart of a speech synthesis method according to an exemplary embodiment.

[0020] Figure 2 is a structural schematic diagram of a speech synthesis model according to an exemplary embodiment.

[0021] Figure 3 is a structural schematic diagram of a prosody language feature prediction module according to an exemplary embodiment.

[0022] Figure 4 is a flowchart of a training method of a speech synthesis model according to an exemplary embodiment.

[0023] Figure 5 is a flowchart of a speech synthesis method according to another exemplary embodiment.

[0024] Figure 6 is a block diagram of a speech synthesis device according to an exemplary embodiment.

[0025] Figure 7 is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0026] As discussed in the background, in speech synthesis and other tasks, how to combine the prosody features of the text to make the synthesized audio more natural and smooth has become the focus of research. In order to improve the naturalness of the synthesized audio, the speech synthesis method at the present stage mainly uses the prosody features of the language level, i.e., the TOBI (Tones and Break Indices) data manually annotated to realize the prosody control of the synthesized audio, so as to improve the naturalness of the speech synthesis, but the intensity of the synthesized audio is uncontrollable.

[0027] In view of this, the present disclosure provides a speech synthesis method, device, computer readable medium, and electronic device.

[0028] Embodiments of the present disclosure will be described below in greater detail with reference to the accompanying drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be interpreted as being limited to the embodiments set forth herein; rather, these embodiments are provided so as to facilitate a more complete and thorough understanding of the present disclosure. It is understood that the drawings and embodiments of the present disclosure are for exemplary purposes only and are not intended to limit the scope of the present disclosure.

[0029] It should be understood that each step recited in the method embodiments of the present disclosure can be performed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0030] The term "comprising" and variations thereof as used herein are used inclusively, i.e., "comprising but not limited to". The term "based on" is "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Related definitions are given throughout the description.

[0031] It should be noted that the terms "first", "second", and the like in the present disclosure are merely used to distinguish different devices, modules or units, and do not imply the order or interdependence of the functions performed by these devices, modules or units.

[0032] It should be noted that the terms "one", "multiple" in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that "one" or "multiple" should be understood as "one or more" unless otherwise explicitly indicated in the context.

[0033] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are merely for illustrative purposes and are not intended to limit the scope of the messages or information.

[0034] Figure 1 is a flowchart of a speech synthesis method according to an exemplary embodiment. As shown in Figure 1 , the method includes S101-S103.

[0035] In S101, a phoneme sequence corresponding to the text to be synthesized is obtained.

[0036] In the present disclosure, the text to be synthesized described above can be Chinese, English, Japanese and the like. In addition, the phoneme sequence corresponding to the text to be synthesized can be obtained by a Grapheme-to-Phoneme (G2P) model.

[0037] For example, the G2P model can use a recurrent neural network (RNN) and a long short-term memory network (LSTM) to achieve the conversion from graphemes to phonemes.

[0038] In S102, based on the phoneme sequence and the text to be synthesized, a TOBI representation sequence and prosodic acoustic features at the phoneme level corresponding to the text to be synthesized are generated, and based on the TOBI representation sequence and prosodic acoustic features, acoustic feature information corresponding to the text to be synthesized is generated.

[0039] In the present disclosure, the TOBI representation sequence is used to reflect the prosodic features of the language level of the text to be synthesized, that is, the prosodic language features, which refer to the prosodic language phenomena defined by the ToBI system in original linguistics. They are discrete features and can specifically include tone, intonation, pitch stress and prosodic boundaries.

[0040] Tone refers to the changes in the pitch of a sound. For example, there are four tones in Chinese: yinping, yangping, shangsheng, and qusheng. English includes stressed, semi-stressed, and unstressed tones, and Japanese includes stressed and unstressed tones.

[0041] Intonation, or the tone of speech, is the arrangement and variation of speed, weight, and intensity within a sentence. Besides lexical meaning, a sentence also has intonation meaning. Intonation is the attitude or tone conveyed by the speaker's tone. The complete meaning of a sentence is the sum of its lexical meaning and intonation. The same sentence, with different intonations, can have a different meaning, sometimes dramatically different.

[0042] Pitch accent is used to describe the pitch change of stressed syllables and can control the emphasized information and the rhythm of stress-rhythm languages. Its scope of action is on the main stressed syllable, or on the main stressed syllable and the syllable after the main stressed syllable of the same word. In the present disclosure, pitch accent control is only performed on the main stressed syllable, ignoring redundant information on other syllables such as secondary stress and zero stress, so as to achieve the effect of information simplification. Accordingly, pitch accent information is used to indicate the syllable position of the specified stress phenomenon in the text to be synthesized, where the specified stress phenomenon can include high stress, low stress, rising stress, low rising stress, and high falling stress.

[0043] Specifically, for high stress, the pitch target is high, the fundamental frequency curve (f0) is high and flat, and the listening experience is Chinese yinping; for low stress, the pitch target is low, the fundamental frequency curve is low and flat, and the listening experience is the first half of Chinese shangsheng; for rising stress, the pitch target is high, the fundamental frequency curve is rising, and the listening experience is Chinese yangping; for low rising stress, the pitch target is low, and if it acts on a monosyllable, the fundamental frequency curve is declining, with a slight rise at the end; if it acts on a disyllabic, the fundamental frequency curve is declining on the main stress and rising on the syllable after the main stress, and the listening experience is Chinese shangsheng; for high falling stress, the pitch target is high, the fundamental frequency curve is declining, and the listening experience is Chinese qusheng.

[0044] Prosodic boundaries indicate where pauses should be placed in the text being synthesized. For example, prosodic boundaries are divided into four levels: "#1," "#2," "#3," and "#4," with increasing levels of pause intensity. English and Japanese lack distinct prosodic levels, so this field is left blank.

[0045] The prosodic acoustic features (i.e., the prosodic features at the acoustic level) are broadly defined as physical quantities that measure the acoustic characteristics of speech, such as timbre, resonance peaks, fundamental frequency, or resonance peak intensity. Among them, the acoustic features that are more closely related to the prosodic events defined by the ToBI system in linguistics are: duration, fundamental frequency, and energy. For example, the high rising tone of the prosodic language feature "intonation" can be specifically manifested as the corresponding fundamental frequency in a speech segment continuously climbing to the high point of the fundamental frequency in a sentence. Therefore, the prosodic acoustic features in the present disclosure include at least one of the fundamental frequency, energy, and pronunciation duration at the phoneme level corresponding to the text to be synthesized, which is a continuity feature.

[0046] The acoustic feature information may be, for example, a mel spectrum, a spectral envelope, or the like.

[0047] In S103, first audio information corresponding to the text to be synthesized is generated according to the acoustic feature information.

[0048] In the present disclosure, the first audio information corresponding to the text to be synthesized can be obtained by inputting the acoustic feature information into a vocoder, wherein the vocoder can be, for example, a Wavenet vocoder, a Griffin-Lim vocoder, etc.

[0049] In the above technical solution, after obtaining the phoneme sequence corresponding to the text to be synthesized, a phoneme-level TOBI representation sequence and prosodic acoustic features corresponding to the text to be synthesized are generated based on the phoneme sequence and the text to be synthesized. Acoustic feature information corresponding to the text to be synthesized is then generated based on the TOBI representation sequence and prosodic acoustic features. Finally, first audio information corresponding to the text to be synthesized is generated based on the acoustic feature information. During speech synthesis, both the TOBI representation sequence and prosodic acoustic features corresponding to the text to be synthesized are referenced. This refers not only to the prosodic features at the linguistic level of the text to be synthesized, but also to the prosodic features at the acoustic level of the text to be synthesized, taking into account the performance of prosody in different dimensions. The TOBI representation sequence can be used to assign appropriate rhythm, emphasis, and intonation characteristics to different sentences. The corresponding prosodic acoustic features can also explicitly reflect the specific acoustic manifestations of the corresponding prosodic events, thereby improving the rhythmic naturalness of the synthesized audio while controlling the intensity (i.e., amplitude) of the audio. For example, different intensities can be assigned to multiple stress positions to achieve different emphasis in semantic expression, or intensity adjustment can be used to achieve intonation changes in interrogative sentences to convey different semantics (emotions). As a result, different prosodic acoustic features can reflect different semantic changes under the same prosodic language expression, making the synthesized audio more natural, more rhythmic, and more consistent with the meaning expressed by the speaker.

[0050] The following describes in detail the specific implementation method of generating the TOBI representation sequence and prosodic acoustic features at the phoneme level corresponding to the text to be synthesized based on the phoneme sequence and the text to be synthesized, and generating the acoustic feature information corresponding to the text to be synthesized based on the TOBI representation sequence and prosodic acoustic features in the above S102.

[0051] Specifically, the phoneme sequence and the text to be synthesized can be input into a pre-trained speech synthesis model, so that the speech synthesis model can generate the phoneme-level TOBI representation sequence and prosodic acoustic features corresponding to the text to be synthesized based on the phoneme sequence and the text to be synthesized, and generate the acoustic feature information corresponding to the text to be synthesized based on the TOBI representation sequence and prosodic acoustic features.

[0052] like Figure 2 As shown, the above-mentioned speech synthesis model includes an encoding network, an attention network, a decoding network, a prosodic language feature prediction module, a prosodic acoustic feature prediction module, an embedding layer, a first splicing module, a second splicing module and a third splicing module, wherein the prosodic language feature prediction module, the first splicing module, the encoding network, the second splicing module, the prosodic acoustic feature prediction module, the third splicing module, the attention network and the decoding network are connected in sequence, and the first splicing module is also connected to the embedding layer, the second splicing module is also connected to the prosodic language feature prediction module, and the third splicing module is also connected to the encoding network.

[0053] Specifically, the prosody language feature prediction module is configured to generate a TOBI representation sequence at a phoneme level corresponding to the text to be synthesized according to the text to be synthesized.

[0054] The embedding layer is configured to generate a phoneme representation sequence corresponding to the text to be synthesized according to the phoneme sequence, wherein the phoneme representation sequence is formed by arranging word vectors corresponding to each phoneme in the text to be synthesized in the order of the corresponding phonemes in the text to be synthesized, and the word vectors corresponding to each phoneme in the text to be synthesized can be determined according to a pre-established correspondence between phonemes and word vectors.

[0055] The first splicing module is configured to splice the TOBI representation sequence at the phoneme level and the phoneme representation sequence to obtain a first spliced sequence.

[0056] The encoding network is configured to encode the first spliced sequence to generate an encoded sequence.

[0057] The second splicing module is configured to splice the encoded sequence and the TOBI representation sequence at the phoneme level to obtain a second spliced sequence.

[0058] The prosody acoustic feature prediction module is configured to generate prosody acoustic features corresponding to the text to be synthesized according to the second spliced sequence.

[0059] For example, the prosody acoustic feature prediction module can be a shallow network composed of a convolutional layer, a bidirectional LSTM layer, and a fully connected layer.

[0060] The third splicing module is configured to splice the encoded sequence and the prosody acoustic features to obtain a third spliced sequence.

[0061] The attention network is configured to generate semantic representations corresponding to the text to be synthesized according to the third spliced sequence. For example, the attention network can be a Locative Sensitive Attention or a Gaussian Mixture Model (GMM) based attention network, i.e., GMM attention.

[0062] The decoding network is configured to generate acoustic feature information corresponding to the text to be synthesized according to the semantic representations.

[0063] As shown in Figure 3 The prosody language feature prediction module includes a first sub-embedding layer, a prosody language feature prediction network, a second sub-embedding layer, and an expansion layer connected in sequence.

[0064] Specifically, the first sub-embedding layer is configured to extract deep representations at a word level corresponding to the text to be synthesized. For example, the first sub-embedding layer can be a TinyBert model based on distillation learning.

[0065] The prosodic language feature prediction network is configured to generate word-level TOBI labels based on the deep-level representation. The TOBI labels can include intonation, pitch accent, and prosodic boundaries.

[0066] For example, the prosodic language feature prediction network can be a shallow network composed of a convolutional layer, a bidirectional LSTM layer, and a fully connected layer.

[0067] The second sub-embedding layer is configured to generate a word-level TOBI representation sequence corresponding to the text to be synthesized based on the TOBI labels.

[0068] The expansion layer is configured to expand the word-level TOBI representation sequence to obtain a phoneme-level TOBI representation sequence corresponding to the text to be synthesized.

[0069] Specifically, for each word in the text to be synthesized, the word-level TOBI representation corresponding to the word is copied L-1 times to obtain a phoneme-level TOBI representation corresponding to the word, where L is the number of phonemes contained in the word.

[0070] For example, the text to be synthesized includes words A and B connected in sequence, where word A includes three phonemes, word B includes four phonemes, the word-level TOBI representation corresponding to word A is M, and the word-level TOBI representation corresponding to word B is N. The phoneme-level TOBI representation corresponding to word A is MMM, the TOBI representation corresponding to word B is NNNN, and the phoneme-level TOBI representation sequence corresponding to the text to be synthesized is MMMNNNN.

[0071] In addition, the speech synthesis model described above can be trained by S401-S403 shown in FIG. 4. Figure 4

[0072] In S401, a training text is obtained.

[0073] In S402, a training phoneme sequence, word-level training TOBI labels, training prosodic acoustic features, and training acoustic feature information corresponding to the training text are determined.

[0074] In the present disclosure, the training text can be a text extracted from a real existing speech. An annotator can first label the word-level TOBI (i.e., the word-level training TOBI labels) corresponding to the training text by listening to the speech corresponding to the training text.

[0075] The training phoneme sequence corresponding to the training text can be obtained in the same way as the phoneme sequence corresponding to the text to be synthesized in S101 described above.

[0076] ​In addition, the training prosodic acoustic features corresponding to the training text can be determined in the following manner: frame-level fundamental frequency and energy features can be extracted from real speech corresponding to the training text based on open source tools such as librosa or straight, etc., and then for each phoneme in the training text, the average of the fundamental frequencies of the multiple frames corresponding to the phoneme can be taken as the fundamental frequency of the phoneme, and the average of the energies of the multiple frames corresponding to the phoneme can be taken as the energy of the phoneme, that is, the phoneme-level fundamental frequency and the phoneme-level energy are obtained; at the same time, the pronunciation duration of each phoneme in the training text is obtained based on a forced alignment tool.

[0077] In addition, the training prosodic acoustic features corresponding to the training text can be determined in the following manner: frame-level fundamental frequency and energy features can be extracted from real speech corresponding to the training text based on open source tools such as librosa or straight, etc., and then for each phoneme in the training text, the average of the fundamental frequencies of the multiple frames corresponding to the phoneme can be taken as the fundamental frequency of the phoneme, and the average of the energies of the multiple frames corresponding to the phoneme can be taken as the energy of the phoneme, that is, the phoneme-level fundamental frequency and the phoneme-level energy are obtained; at the same time, the pronunciation duration of each phoneme in the training text is obtained based on a forced alignment tool.

[0078] In S403, the model is trained in the following manner: the training text is input into the first sub-embedding layer as input, the output of the first sub-embedding layer is input into the prosodic linguistic feature prediction network as input, the word-level training TOBI label is input into the prosodic linguistic feature prediction network as target output, the output of the prosodic linguistic feature prediction network is input into the second sub-embedding layer as input, the output of the second sub-embedding layer is input into the expansion layer as input, the training phoneme sequence is input into the embedding layer as input, the output of the expansion layer and the output of the embedding layer are input into the first splicing module as input, the output of the first splicing module is input into the encoding network as input, the output of the encoding network and the output of the expansion layer are input into the second splicing module as input, the output of the second splicing module is input into the prosodic acoustic feature prediction module as input, the training prosodic acoustic features are input into the prosodic acoustic feature prediction module as target output, the output of the prosodic acoustic feature prediction module and the output of the encoding network are input into the third splicing module as input, the output of the third splicing module is input into the attention network as input, the output of the attention network is input into the decoding network as input, and the training acoustic feature information is input into the decoding network as target output, so as to obtain the speech synthesis model.

[0079] In the present disclosure, the loss function during training of the speech synthesis model is the sum of the acoustic feature information loss and the prosodic feature loss. The acoustic feature information loss is the mean square error between the acoustic feature information predicted by the decoding network and the training acoustic feature information; the prosodic feature loss includes the prosodic linguistic feature prediction loss and the prosodic acoustic feature prediction loss, wherein the prosodic linguistic feature prediction loss is the cross-entropy loss between the word-level TOBI predicted by the prosodic linguistic feature prediction network and the word-level training TOBI label; the prosodic acoustic feature prediction loss is the mean square error between the acoustic feature information predicted by the prosodic acoustic feature prediction module and the training prosodic acoustic features.

[0080] In addition, in order to improve the user experience, after obtaining the first audio information corresponding to the text to be synthesized in step 103, background music can also be added to the first audio information, so that the user can more easily understand the corresponding text content according to the background music and the first audio information. Specifically, as shown in Figure 5 The method can further include the following S104.

[0081] In S104, the first audio information is synthesized with the target background music to obtain second audio information.

[0082] In an embodiment, the target background music can be preset music, that is, any music set by the user or default music.

[0083] In another embodiment, before synthesizing the first audio information with the target background music, the use scenario information corresponding to the text to be synthesized can be determined according to the text information of the text to be synthesized, wherein the use scenario information includes but is not limited to news broadcast, military weapon introduction, fairy tale, campus radio, etc. Then, according to the use scenario information, the target background music matching the use scenario information is determined.

[0084] In the present disclosure, the text information can be a keyword, at this time, the use scenario information of the text to be synthesized can be intelligently predicted according to the keyword by automatically recognizing the keyword of the text to be synthesized.

[0085] After determining the use scenario information corresponding to the text to be synthesized, the target background music matching the use scenario information can be determined according to the use scenario information and the corresponding relationship between the use scenario information and the background music stored in advance. For example, the use scenario information is military weapon introduction, and the corresponding background music can be stirring music; the use scenario information is a fairy tale, and the corresponding background music can be light and lively music.

[0086] Figure 6 is a block diagram of a speech synthesis device according to an example embodiment. As Figure 6 shown, the device 600 includes:

[0087] The acquisition module 601 is configured to acquire a phoneme sequence corresponding to a text to be synthesized.

[0088] The first generation module 602 is configured to generate a TOBI representation sequence and a prosody acoustic feature of a phoneme level corresponding to the text to be synthesized according to the phoneme sequence and the text to be synthesized acquired by the acquisition module 601, and generate acoustic feature information corresponding to the text to be synthesized according to the TOBI representation sequence and the prosody acoustic feature.

[0089] The second generation module 603 is configured to generate first audio information corresponding to the text to be synthesized according to the acoustic feature information generated by the first generation module 602.

[0090] In the above technical solution, after obtaining the phoneme sequence corresponding to the text to be synthesized, the phoneme-level TOBI representation sequence and the prosody acoustic feature corresponding to the text to be synthesized are generated according to the phoneme sequence and the text to be synthesized, and the acoustic feature information corresponding to the text to be synthesized is generated according to the TOBI representation sequence and the prosody acoustic feature. Finally, the first audio information corresponding to the text to be synthesized is generated according to the acoustic feature information. In the speech synthesis, the TOBI representation sequence and the prosody acoustic feature corresponding to the text to be synthesized are simultaneously referred to, that is, not only the prosody feature of the language level of the text to be synthesized is referred to, but also the prosody feature of the acoustic level of the text to be synthesized is referred to, and the performance of the prosody in different dimensions is considered. The TOBI representation sequence can give different sentences appropriate rhythm, emphasis and tone characteristics, and the corresponding prosody acoustic feature can explicitly reflect the specific acoustic manifestation of the corresponding prosody event, so as to improve the prosody naturalness of the synthesized audio while controlling the intensity (i.e. amplitude) of the audio, such as allocating different intensities at multiple emphasis positions to realize different emphasis of semantic expression, or adjusting the intensity to realize the tone change of the interrogative sentence to convey different semantics (emotions). Therefore, different prosody acoustic features can reflect different semantic changes under the same prosody language performance, so that the synthesized audio is more natural, has more rhythm and tone, and is more consistent with the semantic expression of the speaker.

[0091] Optionally, the first generation module 602 is configured to input the phoneme sequence and the text to be synthesized into a pre-trained speech synthesis model, so as to generate the phoneme-level TOBI representation sequence and the prosody acoustic feature corresponding to the text to be synthesized according to the phoneme sequence and the text to be synthesized through the speech synthesis model, and generate the acoustic feature information corresponding to the text to be synthesized according to the TOBI representation sequence and the prosody acoustic feature.

[0092] Optionally, the speech synthesis model comprises an encoding network, an attention network, a decoding network, a prosody language feature prediction module, a prosody acoustic feature prediction module, an embedding layer, a first splicing module, a second splicing module and a third splicing module.

[0093] The prosody language feature prediction module is configured to generate the phoneme-level TOBI representation sequence corresponding to the text to be synthesized according to the text to be synthesized.

[0094] The embedding layer is configured to generate the phoneme representation sequence corresponding to the text to be synthesized according to the phoneme sequence.

[0095] The first splicing module is configured to splice the phoneme-level TOBI representation sequence and the phoneme representation sequence to obtain a first spliced sequence.

[0096] The encoding network is configured to encode the first spliced sequence to generate an encoded sequence.

[0097] The second splicing module is configured to splice the encoded sequence and the phoneme-level TOBI representation sequence to obtain a second spliced sequence.

[0098] The prosody acoustic feature prediction module is configured to generate prosody acoustic features corresponding to the text to be synthesized according to the second spliced sequence.

[0099] The third splicing module is configured to splice the encoded sequence and the prosody acoustic features to obtain a third spliced sequence.

[0100] The attention network is configured to generate semantic representation corresponding to the text to be synthesized according to the third spliced sequence.

[0101] The decoding network is configured to generate acoustic feature information corresponding to the text to be synthesized according to the semantic representation.

[0102] Optionally, the prosody language feature prediction module comprises a first sub-embedding layer, a prosody language feature prediction network, a second sub-embedding layer and an expansion layer connected in sequence.

[0103] The first sub-embedding layer is configured to extract deep-level representation corresponding to the text to be synthesized.

[0104] The prosody language feature prediction network is configured to generate TOBI labels at the word level according to the deep-level representation.

[0105] The second sub-embedding layer is configured to generate TOBI representation sequence at the word level corresponding to the text to be synthesized according to the TOBI labels.

[0106] The expansion layer is configured to expand the TOBI representation sequence at the word level to obtain TOBI representation sequence at the phoneme level corresponding to the text to be synthesized.

[0107] Optionally, the speech synthesis model is obtained by training a model training device, wherein the model training device comprises:

[0108] A training text acquisition module is configured to acquire training text.

[0109] determining a training phoneme sequence corresponding to the training text, a word-level training TOBI label, training prosody acoustic features, and training acoustic feature information;

[0110] The training module is configured to train the speech synthesis model by taking the training text as an input of the first sub-embedding layer, taking an output of the first sub-embedding layer as an input of the prosody language feature prediction network, taking the word-level training TOBI label as a target output of the prosody language feature prediction network, taking an output of the prosody language feature prediction network as an input of the second sub-embedding layer, taking an output of the second sub-embedding layer as an input of the expansion layer, taking the training phoneme sequence as an input of the embedding layer, taking an output of the expansion layer and an output of the embedding layer as an input of the first splicing module, taking an output of the first splicing module as an input of the encoding network, taking an output of the encoding network and an output of the expansion layer as an input of the second splicing module, taking an output of the second splicing module as an input of the prosody acoustic feature prediction module, taking the training prosody acoustic features as a target output of the prosody acoustic feature prediction module, taking an output of the prosody acoustic feature prediction module and an output of the encoding network as an input of the third splicing module, taking an output of the third splicing module as an input of the attention network, taking an output of the attention network as an input of the decoding network, and taking the training acoustic feature information as a target output of the decoding network.

[0111] Optionally, the prosody acoustic features include at least one of a fundamental frequency, energy, and pronunciation duration of a phoneme corresponding to the text to be synthesized.

[0112] Optionally, the apparatus 600 further includes:

[0113] The synthesis module is configured to synthesize the first audio information and target background music to obtain second audio information.

[0114] It should be noted that the above model training apparatus can be integrated into the above speech synthesis apparatus 600 or can be independent of the above speech synthesis apparatus 600, and the present disclosure is not limited in this regard.

[0115] The present disclosure also provides a computer readable medium having a computer program stored thereon, the program being executed by a processing apparatus to implement the steps of the above speech synthesis method provided by the present disclosure.

[0116] Reference will be made to the following description Figure 7, which shows a schematic structural diagram of an electronic device (terminal device or server) 700 suitable for implementing the embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (such as in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 7 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.

[0117] like Figure 7 As shown, the electronic device 700 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. Various programs and data required for the operation of the electronic device 700 are also stored in the RAM 703. The processing device 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0118] Typically, the following devices may be connected to the I / O interface 705: an input device 706 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 707 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 708 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 709. The communication device 709 may allow the electronic device 700 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 7 The electronic device 700 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.

[0119] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 709, or installed from the storage device 708, or installed from the ROM 702. When the computer program is executed by the processing device 701, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.

[0120] It should be noted that the computer-readable medium described above can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any suitable combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in baseband or propagated as a carrier wave in a propagated data signal, in which the computer-readable program code is contained. Such a propagated data signal can take many forms, including but not limited to, an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium that can send, propagate or transfer the program for use by or in connection with the instruction execution system, apparatus or device. The program code contained in the computer-readable medium can be transmitted by any suitable medium, including but not limited to, wire, cable, RF (radio frequency), etc., or any suitable combination of the above.

[0121] In some embodiments, the client, server, or both can communicate using any current known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any current known or future developed networks.

[0122] The computer-readable medium described above can be included in the electronic device described above; or can exist separately from the electronic device, and is not assembled into the electronic device.

[0123] The computer readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, cause the electronic device to: acquire a phoneme sequence corresponding to a text to be synthesized; generate a TOBI representation sequence at a phoneme level and prosody acoustic features corresponding to the text to be synthesized according to the phoneme sequence and the text to be synthesized, and generate acoustic feature information corresponding to the text to be synthesized according to the TOBI representation sequence and the prosody acoustic features; and generate first audio information corresponding to the text to be synthesized according to the acoustic feature information.

[0124] Computer program code for carrying out operations of the present disclosure can be written in any of one or more programming languages or combinations of languages including object or visual programming languages such as Java, Smalltalk, C++ or conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0125] The flow and block diagrams in the drawings show architectural, functional, and operational representations of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow and block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may

[0126] The modules described in the embodiments of the present disclosure can be implemented in the form of software, or can be implemented in the form of hardware. In some cases, the name of the module does not constitute a limitation on the module itself, for example, the acquisition module can also be described as "a module for acquiring a phoneme sequence corresponding to a text to be synthesized".

[0127] The functionality described above in this document can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.

[0128] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0129] According to one or more embodiments of the present disclosure, example 1 provides a speech synthesis method, comprising: obtaining a phoneme sequence corresponding to a text to be synthesized; generating a TOBI representation sequence at a phoneme level and prosodic acoustic features corresponding to the text to be synthesized according to the phoneme sequence and the text to be synthesized, and generating acoustic feature information corresponding to the text to be synthesized according to the TOBI representation sequence and the prosodic acoustic features; and generating first audio information corresponding to the text to be synthesized according to the acoustic feature information.

[0130] According to one or more embodiments of the present disclosure, example 2 provides the method of example 1, wherein the generating a TOBI representation sequence at a phoneme level and prosodic acoustic features corresponding to the text to be synthesized according to the phoneme sequence and the text to be synthesized, and generating acoustic feature information corresponding to the text to be synthesized according to the TOBI representation sequence and the prosodic acoustic features comprises: inputting the phoneme sequence and the text to be synthesized into a pre-trained speech synthesis model, so as to generate the TOBI representation sequence at the phoneme level and the prosodic acoustic features corresponding to the text to be synthesized according to the phoneme sequence and the text to be synthesized by the speech synthesis model, and generate the acoustic feature information corresponding to the text to be synthesized according to the TOBI representation sequence and the prosodic acoustic features.

[0131] According to one or more embodiments of the present disclosure, example 3 provides the method of example 2, the speech synthesis model comprising an encoding network, an attention network, a decoding network, a prosody language feature prediction module, a prosody acoustic feature prediction module, an embedding layer, a first splicing module, a second splicing module, and a third splicing module; wherein the prosody language feature prediction module is configured to generate a phoneme-level TOBI representation sequence corresponding to the text to be synthesized according to the text to be synthesized; the embedding layer is configured to generate a phoneme representation sequence corresponding to the text to be synthesized according to the phoneme sequence; the first splicing module is configured to splice the phoneme-level TOBI representation sequence and the phoneme representation sequence to obtain a first splicing sequence; the encoding network is configured to encode the first splicing sequence to generate an encoding sequence; the second splicing module is configured to splice the encoding sequence and the phoneme-level TOBI representation sequence to obtain a second splicing sequence; the prosody acoustic feature prediction module is configured to generate prosody acoustic features corresponding to the text to be synthesized according to the second splicing sequence; the third splicing module is configured to splice the encoding sequence and the prosody acoustic features to obtain a third splicing sequence; the attention network is configured to generate semantic representations corresponding to the text to be synthesized according to the third splicing sequence; and the decoding network is configured to generate acoustic feature information corresponding to the text to be synthesized according to the semantic representations.

[0132] According to one or more embodiments of the present disclosure, example 4 provides the method of example 3, the prosody language feature prediction module comprising a first sub-embedding layer, a prosody language feature prediction network, a second sub-embedding layer, and an expansion layer connected in sequence; wherein the first sub-embedding layer is configured to extract deep-level representations of words corresponding to the text to be synthesized; the prosody language feature prediction network is configured to generate word-level TOBI labels according to the deep-level representations; the second sub-embedding layer is configured to generate word-level TOBI representation sequences corresponding to the text to be synthesized according to the TOBI labels; and the expansion layer is configured to expand the word-level TOBI representation sequences to obtain phoneme-level TOBI representation sequences corresponding to the text to be synthesized.

[0133] According to one or more embodiments of the present disclosure, example 5 provides the method of example 4, wherein the speech synthesis model is trained by: obtaining training text; determining training phoneme sequences, word-level training TOBI labels, training prosody acoustic features, and training acoustic feature information corresponding to the training text; training the speech synthesis model by inputting the training text into the first sub-embedding layer, inputting the output of the first sub-embedding layer into the prosody language feature prediction network, inputting the word-level training TOBI labels into the prosody language feature prediction network as target output, inputting the output of the prosody language feature prediction network into the second sub-embedding layer, inputting the output of the second sub-embedding layer into the expansion layer, inputting the training phoneme sequences into the embedding layer, inputting the output of the expansion layer and the output of the embedding layer into the first splicing module, inputting the output of the first splicing module into the encoding network, inputting the output of the encoding network and the output of the expansion layer into the second splicing module, inputting the output of the second splicing module into the prosody acoustic feature prediction module, inputting the training prosody acoustic features into the prosody acoustic feature prediction module as target output, inputting the output of the prosody acoustic feature prediction module and the output of the encoding network into the third splicing module, inputting the output of the third splicing module into the attention network, inputting the output of the attention network into the decoding network, and inputting the training acoustic feature information into the decoding network as target output.

[0134] According to one or more embodiments of the present disclosure, example 6 provides the method of any one of examples 1-5, wherein the prosody acoustic features comprise at least one of a fundamental frequency, an energy, and a pronunciation duration of a phoneme corresponding to the text to be synthesized.

[0135] According to one or more embodiments of the present disclosure, example 7 provides the method of any one of examples 1-5, further comprising: synthesizing the first audio information with target background music to obtain second audio information.

[0136] According to one or more embodiments of the present disclosure, example 8 provides a speech synthesis device, comprising: an obtaining module configured to obtain phoneme sequences corresponding to text to be synthesized; a first generating module configured to generate, according to the phoneme sequences and the text to be synthesized obtained by the obtaining module, TOBI representations of phonemes corresponding to the text to be synthesized and prosody acoustic features, and generate acoustic feature information corresponding to the text to be synthesized according to the TOBI representations of phonemes and the prosody acoustic features; and a second generating module configured to generate, according to the acoustic feature information generated by the first generating module, first audio information corresponding to the text to be synthesized.

[0137] According to one or more embodiments of the present disclosure, example 9 provides the apparatus of example 8, the first generation module is configured to input the phoneme sequence and the text to be synthesized into a pre-trained speech synthesis model, to generate, by the speech synthesis model, a TOBI representation sequence at a phoneme level and a prosody acoustic feature corresponding to the text to be synthesized according to the phoneme sequence and the text to be synthesized, and to generate acoustic feature information corresponding to the text to be synthesized according to the TOBI representation sequence and the prosody acoustic feature.

[0138] According to one or more embodiments of the present disclosure, example 10 provides the apparatus of example 9, the speech synthesis model comprises an encoding network, an attention network, a decoding network, a prosody language feature prediction module, a prosody acoustic feature prediction module, an embedding layer, a first splicing module, a second splicing module, and a third splicing module; wherein the prosody language feature prediction module is configured to generate a TOBI representation sequence at a phoneme level corresponding to the text to be synthesized according to the text to be synthesized; the embedding layer is configured to generate a phoneme representation sequence corresponding to the text to be synthesized according to the phoneme sequence; the first splicing module is configured to splice the TOBI representation sequence at the phoneme level and the phoneme representation sequence to obtain a first splicing sequence; the encoding network is configured to encode the first splicing sequence to generate an encoding sequence; the second splicing module is configured to splice the encoding sequence and the TOBI representation sequence at the phoneme level to obtain a second splicing sequence; the prosody acoustic feature prediction module is configured to generate a prosody acoustic feature corresponding to the text to be synthesized according to the second splicing sequence; the third splicing module is configured to splice the encoding sequence and the prosody acoustic feature to obtain a third splicing sequence; the attention network is configured to generate a semantic representation corresponding to the text to be synthesized according to the third splicing sequence; and the decoding network is configured to generate acoustic feature information corresponding to the text to be synthesized according to the semantic representation.

[0139] According to one or more embodiments of the present disclosure, example 11 provides the apparatus of example 10, the prosody language feature prediction module comprises a first sub-embedding layer, a prosody language feature prediction network, a second sub-embedding layer, and an expansion layer connected in sequence; wherein the first sub-embedding layer is configured to extract a deep representation at a word level corresponding to the text to be synthesized; the prosody language feature prediction network is configured to generate a TOBI label at a word level according to the deep representation; the second sub-embedding layer is configured to generate a TOBI representation sequence at a word level corresponding to the text to be synthesized according to the TOBI label; and the expansion layer is configured to expand the TOBI representation sequence at the word level to obtain the TOBI representation sequence at the phoneme level corresponding to the text to be synthesized.

[0140] According to one or more embodiments of the present disclosure, example 12 provides the apparatus of example 11, wherein the speech synthesis model is trained by a model training apparatus, and the model training apparatus comprises: a training text acquisition module configured to acquire training text; a determination module configured to determine a training phoneme sequence corresponding to the training text, a word-level training TOBI label, training prosody acoustic features, and training acoustic feature information; and a training module configured to train the speech synthesis model by taking the training text as an input of the first sub-embedding layer, taking an output of the first sub-embedding layer as an input of the prosody language feature prediction network, taking the word-level training TOBI label as a target output of the prosody language feature prediction network, taking an output of the prosody language feature prediction network as an input of the second sub-embedding layer, taking an output of the second sub-embedding layer as an input of the expansion layer, taking the training phoneme sequence as an input of the embedding layer, taking an output of the expansion layer and an output of the embedding layer as an input of the first splicing module, taking an output of the first splicing module as an input of the encoding network, taking an output of the encoding network and an output of the expansion layer as an input of the second splicing module, taking an output of the second splicing module as an input of the prosody acoustic feature prediction module, taking the training prosody acoustic features as a target output of the prosody acoustic feature prediction module, taking an output of the prosody acoustic feature prediction module and an output of the encoding network as an input of the third splicing module, taking an output of the third splicing module as an input of the attention network, taking an output of the attention network as an input of the decoding network, and taking the training acoustic feature information as a target output of the decoding network.

[0141] According to one or more embodiments of the present disclosure, example 13 provides the apparatus of any one of examples 8-12, wherein the prosody acoustic features comprise at least one of a fundamental frequency, an energy, and a pronunciation duration of a phoneme level corresponding to the text to be synthesized.

[0142] According to one or more embodiments of the present disclosure, example 14 provides the apparatus of any one of examples 8-12, wherein the apparatus further comprises: a synthesis module configured to synthesize the first audio information with target background music to obtain second audio information.

[0143] According to one or more embodiments of the present disclosure, example 15 provides a computer readable medium having stored thereon a computer program, which, when executed by a processing apparatus, causes the steps of the method of any one of examples 1-7.

[0144] According to one or more embodiments of the present disclosure, example 16 provides an electronic device comprising: a storage device having stored thereon one or more computer programs; and one or more processing devices configured to execute the one or more computer programs stored in the storage device to implement the steps of the method of any one of examples 1-7.

[0145] The above description is only preferred embodiments of the present disclosure and the explanation of the applied technical principles. It should be understood by those skilled in the art that the disclosed scope of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by the combinations of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the technical solutions formed by the mutual replacement of the above features and the technical features disclosed in the present disclosure (but not limited to) having similar functions.

[0146] In addition, although each operation is depicted in a particular order, this should not be understood as requiring the operations to be performed in the particular order shown or in a sequential order. In certain circumstances, multitasking and parallel processing can be advantageous. Similarly, although several implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be combined in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented separately or in any suitable subcombination.

[0147] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims. With respect to the devices in the above-described embodiments, the specific manner in which the various modules perform operations has been described in detail in the embodiments related to the method, and will not be described in detail here.

Claims

1. A speech synthesis method, characterized in that: include: Obtain the phoneme sequence corresponding to the text to be synthesized; Inputting the phoneme sequence and the text to be synthesized into a pre-trained speech synthesis model, so that the speech synthesis model generates a phoneme-level TOBI representation sequence and prosodic acoustic features corresponding to the text to be synthesized based on the phoneme sequence and the text to be synthesized, and generates acoustic feature information corresponding to the text to be synthesized based on the TOBI representation sequence and the prosodic acoustic features; Generating first audio information corresponding to the text to be synthesized according to the acoustic feature information; The speech synthesis model includes an encoding network, an attention network, a decoding network, a prosodic language feature prediction module, a prosodic acoustic feature prediction module, an embedding layer, a first splicing module, a second splicing module, and a third splicing module; The prosodic language feature prediction module is used to generate a TOBI representation sequence at the phoneme level corresponding to the text to be synthesized based on the text to be synthesized; The embedding layer is used to generate a phoneme representation sequence corresponding to the text to be synthesized according to the phoneme sequence; The first splicing module is configured to splice the phoneme-level TOBI representation sequence with the phoneme representation sequence to obtain a first splicing sequence; The encoding network is used to encode the first spliced ​​sequence to generate a coding sequence; The second splicing module is used to splice the coding sequence with the TOBI representation sequence at the phoneme level to obtain a second spliced ​​sequence; The prosodic acoustic feature prediction module is used to generate prosodic acoustic features corresponding to the text to be synthesized based on the second splicing sequence; The third splicing module is used to splice the coding sequence and the prosodic acoustic features to obtain a third spliced ​​sequence; The attention network is used to generate a semantic representation corresponding to the text to be synthesized according to the third splicing sequence; The decoding network is used to generate acoustic feature information corresponding to the text to be synthesized based on the semantic representation.

2. The method according to claim 1, characterized in that The prosodic language feature prediction module includes a first sub-embedding layer, a prosodic language feature prediction network, a second sub-embedding layer and an expansion layer connected in sequence; The first sub-embedding layer is used to extract the deep representation of the word level corresponding to the text to be synthesized; The prosodic language feature prediction network is used to generate word-level TOBI tags based on the deep representation; The second sub-embedding layer is used to generate a word-level TOBI representation sequence corresponding to the text to be synthesized according to the TOBI label; The expansion layer is used to expand the TOBI representation sequence at the word level to obtain the TOBI representation sequence at the phoneme level corresponding to the text to be synthesized.

3. The method according to claim 2, characterized in that The speech synthesis model is trained in the following way: Get training text; Determining a training phoneme sequence, word-level training TOBI tags, training prosodic acoustic features, and training acoustic feature information corresponding to the training text; The method uses the training text as the input of the first sub-embedding layer, the output of the first sub-embedding layer as the input of the prosodic language feature prediction network, the word-level training TOBI tag as the target output of the prosodic language feature prediction network, the output of the prosodic language feature prediction network as the input of the second sub-embedding layer, the output of the second sub-embedding layer as the input of the expansion layer, the training phoneme sequence as the input of the embedding layer, the output of the expansion layer and the output of the embedding layer as the input of the first splicing module, the output of the first splicing module as the input of the encoding network, and the encoding The output of the network and the output of the expansion layer are used as the input of the second splicing module, the output of the second splicing module is used as the input of the prosodic acoustic feature prediction module, the training prosodic acoustic feature is used as the target output of the prosodic acoustic feature prediction module, the output of the prosodic acoustic feature prediction module and the output of the encoding network are used as the input of the third splicing module, the output of the third splicing module is used as the input of the attention network, the output of the attention network is used as the input of the decoding network, and the training acoustic feature information is used as the target output of the decoding network to perform model training to obtain the speech synthesis model.

4. The method according to any one of claims 1 to 3, characterized in that The prosodic acoustic features include at least one of the fundamental frequency, energy, and pronunciation duration of the phoneme level corresponding to the text to be synthesized.

5. The method according to any one of claims 1 to 3, characterized in that The method further comprises: The first audio information is synthesized with target background music to obtain second audio information.

6. A speech synthesis device, characterized in that: include: An acquisition module is used to obtain the phoneme sequence corresponding to the text to be synthesized; A first generation module is configured to input the phoneme sequence and the text to be synthesized into a pre-trained speech synthesis model, so that the speech synthesis model generates a phoneme-level TOBI representation sequence and prosodic acoustic features corresponding to the text to be synthesized based on the phoneme sequence and the text to be synthesized acquired by the acquisition module, and generates acoustic feature information corresponding to the text to be synthesized based on the TOBI representation sequence and the prosodic acoustic features; A second generating module is configured to generate first audio information corresponding to the text to be synthesized based on the acoustic feature information generated by the first generating module; The speech synthesis model includes an encoding network, an attention network, a decoding network, a prosodic language feature prediction module, a prosodic acoustic feature prediction module, an embedding layer, a first splicing module, a second splicing module, and a third splicing module; The prosodic language feature prediction module is used to generate a TOBI representation sequence at the phoneme level corresponding to the text to be synthesized based on the text to be synthesized; The embedding layer is used to generate a phoneme representation sequence corresponding to the text to be synthesized according to the phoneme sequence; The first splicing module is configured to splice the phoneme-level TOBI representation sequence with the phoneme representation sequence to obtain a first splicing sequence; The encoding network is used to encode the first spliced ​​sequence to generate a coding sequence; The second splicing module is used to splice the coding sequence with the TOBI representation sequence at the phoneme level to obtain a second spliced ​​sequence; The prosodic acoustic feature prediction module is used to generate prosodic acoustic features corresponding to the text to be synthesized based on the second splicing sequence; The third splicing module is used to splice the coding sequence and the prosodic acoustic features to obtain a third spliced ​​sequence; The attention network is used to generate a semantic representation corresponding to the text to be synthesized according to the third splicing sequence; The decoding network is used to generate acoustic feature information corresponding to the text to be synthesized based on the semantic representation.

7. A computer-readable medium having a computer program stored thereon, characterized in that: When the program is executed by a processing device, the steps of the method according to any one of claims 1 to 5 are implemented.

8. An electronic device, characterized in that: include: a storage device having one or more computer programs stored thereon; One or more processing devices, configured to execute the one or more computer programs in the storage device to implement the steps of the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Voice synthesis method and device, synthesis model training method and device, medium and equipment

    CN112786006A