Speech synthesis methods, devices, media and electronic equipment
By acquiring phoneme sequences and target prosodic information, and combining them with speaker vectors to synthesize target speech, the problems of inaccurate pronunciation and poor naturalness in cross-language speech synthesis by monolingual models are solved, achieving highly accurate and natural cross-language speech synthesis.
Patent Information
- Application Number
- CN202210108033.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-28
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2042-01-28
AI Technical Summary
In existing technologies, monolingual models synthesize speech from non-speaker language texts or mixed language texts that have been authorized by the user but are not inaccurate and have poor naturalness. This results in low pronunciation accuracy and unnatural rhythm in sentences, which seriously affects the intelligibility and naturalness of the speech.
By acquiring the phoneme sequence and target prosodic information of the text to be synthesized, and combining the speech vector of the first speaker in the first language and the timbre vector of the second speaker, the target speech is synthesized. The synthesized speech references the rhythm information of the first speaker and the timbre information of the second speaker, thus solving the problem of cross-language speech synthesis.
It achieves accuracy and naturalness in cross-language speech synthesis, with accurate pronunciation and natural rhythm, improving the intelligibility and naturalness of the speech.
Smart Images

Figure CN114242035B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and more specifically, to a speech synthesis method, apparatus, medium, and electronic device. Background Technology
[0002] With the development of artificial intelligence technology, speech synthesis technology is receiving increasing attention. Speech synthesis can convert text into speech output. In related technologies, monolingual models are typically used to synthesize accurate and realistic speech from data in the language of a user-authorized speaker. For example, a Chinese speaker speech synthesis model synthesizes Chinese text, or an English speaker speech synthesis model synthesizes English text. However, for data in a language other than the user's authorized speaker, or mixed-language text, the synthesized speech is inaccurate and lacks naturalness. Summary of the Invention
[0003] This section is provided to briefly introduce the concepts, which will be described in detail in the Detailed Description section later. This section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0004] In a first aspect, this disclosure provides a speech synthesis method, including:
[0005] Obtain the phoneme sequence of the text to be synthesized and the target prosodic information corresponding to the phoneme sequence, wherein the target prosodic information is the prosodic information under the first language to which the text to be synthesized belongs;
[0006] Based on the phoneme sequence, the target prosodic information, the speech vector of the first speaker in the first language, and the timbre vector of the second speaker, the target speech is synthesized, wherein the target speech represents the speech of the second speaker speaking the text to be synthesized in the first language.
[0007] Secondly, this disclosure provides a speech synthesis apparatus, comprising:
[0008] The acquisition module is configured to acquire the phoneme sequence of the text to be synthesized and the target prosodic information corresponding to the phoneme sequence, wherein the target prosodic information is the prosodic information of the first language to which the text to be synthesized belongs;
[0009] The synthesis module is configured to synthesize target speech based on the phoneme sequence, the target prosodic information, the speech vector of the first speaker in the first language, and the timbre vector of the second speaker, wherein the target speech represents the speech of the second speaker speaking the text to be synthesized in the first language.
[0010] Thirdly, this disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the method described in the first aspect.
[0011] Fourthly, this disclosure provides an electronic device, comprising:
[0012] A storage device having at least one computer program stored thereon;
[0013] At least one processing device is configured to execute the at least one computer program in the storage device to implement the steps of the method described in the first aspect.
[0014] The above technical solution, by providing the first speaker's speech vector in the first language and the second speaker's pronunciation vector, allows the rhythmic information of the synthesized target speech to reference the rhythmic information of the first speaker's speech in the first language, and the timbre information to reference the timbre information of the second speaker. This enables the synthesis of speech spoken by the second speaker in the first language, even if the first language is not the second speaker's language, effectively solving the problem of cross-language speech synthesis. Furthermore, by synthesizing the target speech using phoneme sequences and target prosodic information, the synthesized target speech exhibits high pronunciation accuracy and natural prosody.
[0015] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description
[0016] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale. In the drawings:
[0017] Figure 1 This is a schematic diagram illustrating an implementation environment according to an exemplary embodiment of the present disclosure.
[0018] Figure 2 This is a flowchart illustrating a speech synthesis method according to an exemplary embodiment of the present disclosure.
[0019] Figure 3 This is a flowchart illustrating a method for synthesizing target speech according to an exemplary embodiment of the present disclosure.
[0020] Figure 4 This is a flowchart illustrating a method for training a speech synthesis model according to an exemplary embodiment of the present disclosure.
[0021] Figure 5 This is a flowchart illustrating a method for obtaining target prosodic information according to an exemplary embodiment of the present disclosure.
[0022] Figure 6 This is a flowchart illustrating a method for training a target prosodic prediction model according to an exemplary embodiment of the present disclosure.
[0023] Figure 7 This is a block diagram of a speech synthesis apparatus illustrated according to an exemplary embodiment of the present disclosure.
[0024] Figure 8 This is a schematic diagram of the structure of an electronic device according to an exemplary embodiment of the present disclosure. Detailed Implementation
[0025] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0026] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0027] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0028] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0029] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0030] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0031] In related technologies, monolingual models are typically used to synthesize accurate and realistic speech from text in the language of a user-authorized speaker. For example, a Chinese speaker speech synthesis model synthesizes Chinese text, or an English speaker speech synthesis model synthesizes English text. However, monolingual models produce inaccurate and unnatural speech from text in a language other than the user's authorized speaker (i.e., text in a user-authorized speaker's non-native language) or mixed language text. For example, a Chinese speaker speech synthesis model synthesizing English or mixed Chinese-English text, or an English speaker speech synthesis model synthesizing Chinese or mixed Chinese-English text, results in inaccurate pronunciation and unnatural performance.
[0032] Furthermore, inaccurate pronunciation and poor naturalness of synthesized speech can lead to the following defects: poor sentence pronunciation accuracy, severely affecting the intelligibility of the synthesized speech; unnatural sentence rhythm, with a significant difference from the pronunciation of a native speaker authorized by the user, mainly reflected in intonation, stress, and duration, as well as unnatural transitions in mixed-language texts. Therefore, it is clear that the relevant technology cannot effectively solve the problem of cross-language speech synthesis.
[0033] Figure 1 This is a schematic diagram illustrating an implementation environment according to an exemplary embodiment of this disclosure. Figure 1 As shown, the implementation environment may include a model training device 110 and a model usage device 120. In some embodiments, the model training device 110 may be a computer device such as a computer or server, used to train a speech synthesis model. The model training device 110 may use machine learning to train the speech synthesis model; the training process of the speech synthesis model can be found below. Figure 4 The details and related descriptions will not be repeated here.
[0034] The trained speech synthesis model can be deployed on the model-using device 120. The model-using device 120 can be a terminal device such as a mobile phone, tablet, personal computer, or multimedia playback device, or it can be a server. The model-using device 120 can synthesize target speech from the text to be synthesized using the speech synthesis model. Specific details regarding the synthesis of target speech can be found below. Figure 3 The details and related descriptions will not be repeated here.
[0035] Figure 2 This is a flowchart illustrating a speech synthesis method according to an exemplary embodiment of this disclosure. Figure 2 As shown, the method may include the following steps.
[0036] Step 210, obtain the phoneme sequence of the text to be synthesized and the target prosody information corresponding to the phoneme sequence, where the target prosody information is the prosody information in the first language to which the text to be synthesized belongs.
[0037] In some embodiments, the text to be synthesized may be the text for which speech needs to be synthesized, and the text to be synthesized includes the text in the first language. In some embodiments, the first language may be one or more. When the first language is one, the text to be synthesized is a monolingual text. For example, if the first language is Chinese or English, the text to be synthesized is a Chinese text or an English text. When the first language is multiple, the text to be synthesized is a multilingual text. For example, if the first language is Chinese and English, the text to be synthesized is a multilingual text including Chinese text and English text.
[0038] It should be noted that the types of the foregoing first languages are only illustrative examples and are not limited to Chinese, English, or Chinese-English. For example, it may also be German, French, or a mixed language of the two, etc. The present disclosure makes no restrictions on this.
[0039] In some embodiments, the phoneme sequence may be a sequence composed of the phonemes of the text in the first language included in the text to be synthesized, and a phoneme is the smallest speech unit divided according to the natural attributes of speech. For texts in different first languages, the method of dividing phonemes may be different. Exemplarily, taking the Chinese text "Hello" as an example of the text in the first language, the phoneme sequence is {N, I, H, A, O}; taking the English text "seattle" as an example of the text in the first language, the phoneme sequence is {S, IY, AE, T, AX, L}; taking the Chinese-English mixed text "Hello seattle" as an example of the text in the first language, the phoneme sequence is {N, I, H, A, O, S, IY, AE, T, AX, L}.
[0040] In some embodiments, the phoneme sequence of the text to be synthesized may be obtained by manually annotating the phonemes of the text to be synthesized according to statistical knowledge. In some embodiments, the phoneme sequence of the text to be synthesized may also be obtained by querying the phonemes of each character or word in the text to be synthesized in a preset dictionary, where the preset dictionary stores the phonemes of multiple characters or words in advance.
[0041] In some embodiments, the target prosody information is prosody information in the first language to which the text to be synthesized belongs, and the target prosody information is composed of the prosody information of the text in the first language included in the text to be synthesized. For texts in different first languages, the methods for determining their prosody information may be different. For example, if the text in the first language is a Chinese text, the prosody information may include prosodic words, prosodic phrases, intonation phrases, and focus stress. Among them, prosodic words, prosodic phrases, and intonation phrases can represent the level of semantic groups and reflect pauses in acoustics; focus stress can represent accent emphasis in acoustics. Another example is that if the text in the first language is an English text, the prosody information may include ToBI (Tones and Break Indices) features. ToBI features may include phrase stress, boundary tones, and pitch accents. The prosody information of the English text can be obtained by annotating the prosody of the English text according to the ToBI annotation system.
[0042] In some embodiments, the target prosody information corresponds to the phoneme sequence, that is, the target prosody information is obtained by expanding the prosody information of the text in the first language to the phoneme level. The prosody information of the text in the first language is at the text level. By expanding the prosody information at the text level to the phoneme level to obtain the target prosody information, the granularity alignment of the target prosody information and the phoneme sequence is achieved, so that the target prosody information and the phoneme sequence belong to the same level, that is, both belong to the phoneme level, which is convenient for subsequent processing. Exemplarily, taking the text in the first language as "一个" (a) as an example, assuming its prosody information is {04}, and its phoneme sequence is {YIGE}, where 0 represents the prosody of the character "一" (one), and 4 represents the prosody of the character "个" (a), then expanding the prosody information to the phoneme level can obtain the target prosody information {0044}. After expansion, "00" represents the prosody of the phonemes Y and I, and "44" represents the prosody of the phonemes G and E. It can be seen from this that the granularity of the target prosody information and the phoneme sequence is the same.
[0043] In some embodiments, the initial prosody information of the text to be synthesized can be obtained according to the trained target prosody prediction model; the initial prosody information of the text to be synthesized is expanded to the phoneme level to obtain the target prosody information. For the specific details of obtaining the target prosody information, reference can be made to the following Figure 5 and its related descriptions, which will not be elaborated here.
[0044] Step 220, synthesize the target speech according to the phoneme sequence, the target prosody information, the speech vector of the first speaker who has obtained user authorization to use in the first language, and the timbre vector of the second speaker who has obtained user authorization to use. The target speech represents the speech of the second speaker who has obtained user authorization to use speaking the text to be synthesized in the first language.
[0045] For specific details regarding the phoneme sequence and target prosodic information, please refer to step 210 above and its related description, which will not be repeated here. As mentioned earlier, the phoneme sequence is composed of the phonemes of the text to be synthesized. Since the target speech is the speech of the text to be synthesized, the phoneme sequence can reflect the pronunciation information of the target speech. The target prosodic information is the prosodic information of the text to be synthesized; therefore, the target prosodic information can reflect the prosodic information of the target speech.
[0046] In some embodiments, the speech vector of the first speaker authorized by the user in the first language can be an encoded vector of the speech data of the first speaker authorized by the user in the first language, which can be one or more speech data. If there are multiple speech data, each of the multiple speech data can be encoded to obtain multiple encoded vectors, and the average of these multiple encoded vectors can be used to obtain the speech vector of the first speaker authorized by the user in the first language.
[0047] In some embodiments, the speech vector of the first speaker authorized by the user in the first language can reflect the pronunciation duration of each phoneme of the text when the first speaker, authorized by the user, pronounces the text corresponding to the speech data in the first language, that is, reflect the pronunciation rhythm of the speech data. In some embodiments, the first language can be the language (or native language) of the first speaker authorized by the user, and the language (or native language) of a second speaker who is not authorized by the user.
[0048] In some embodiments, the duration of each phoneme in the phoneme sequence can be determined by the speech vector of a first speaker authorized by the user in the first language, thus obtaining the target duration sequence. In other words, the duration of each phoneme in the text to be synthesized can be predicted by the pronunciation duration of each phoneme in the text of the first language's speech data spoken by the first speaker authorized by the user. Therefore, the speech vector of the first speaker authorized by the user in the first language can indirectly reflect the pronunciation rhythm of the target speech (i.e., the pronunciation duration of each phoneme in the text to be synthesized), and also indirectly reflect the pronunciation rhythm of the speech spoken by the second speaker authorized by the user in the first language. Specific details regarding obtaining the target duration sequence can be found below. Figure 3 The details and related descriptions will not be repeated here.
[0049] In some embodiments, the timbre vector of the second speaker who has been authorized by the user can be the encoding vector of the speech data of the second speaker who has been authorized by the user. The method of obtaining the encoding vector of the speech data is the same as the method of obtaining the encoding vector of the speech data of the first speaker who has been authorized by the user in the first language. For details, please refer to the foregoing relevant description, which will not be repeated here.
[0050] In some embodiments, the timbre vector of a second speaker authorized by the user can reflect the timbre information of the second speaker when speaking; timbre is the characteristic of sound perceived by hearing. Since the target speech is the speech spoken by a second speaker authorized by the user, the timbre information of the target speech can be reflected through the timbre vector of the second speaker authorized by the user.
[0051] In some embodiments, the speech vectors of the first speaker authorized by the user in the first language and the pronunciation vectors of the second speaker authorized by the user can be obtained by processing the speech data of the first speaker authorized by the user in the first language and the speech data of the second speaker authorized by the user in the first language, respectively, using a pre-trained speaker coding model. The speech data of the second speaker authorized by the user can be language data of the language of the second speaker authorized by the user (e.g., a second language). The speaker coding model can be obtained through end-to-end training; specific training methods can be found in related technologies and will not be elaborated here.
[0052] In some embodiments, based on the identifiers of a first speaker and a second speaker who have been authorized by the user (e.g., the IDs of the first and second speakers), the speech vectors of the first speaker and the pronunciation vectors of the second speaker in a first language can be retrieved from a preset database. The preset database stores speech vectors of multiple authorized speakers in different languages and pronunciation vectors of multiple authorized speakers.
[0053] In this embodiment, the phoneme sequence reflects the pronunciation information of the target speech, the target prosody information reflects the prosodic information of the target speech, the speech vector of the first speaker authorized by the user in the first language reflects the rhythm information of the target speech, and the timbre vector of the second speaker authorized by the user reflects the timbre information of the target speech. Therefore, this disclosure decomposes speech from multiple dimensions to obtain decoupled label information representing pronunciation information, prosody information, rhythm information, and timbre information.
[0054] Therefore, by providing the speech vectors of a first speaker authorized by the user in the first language and the pronunciation vectors of a second speaker authorized by the user, the rhythmic information of the synthesized target speech references the rhythmic information of the first speaker authorized by the user in the first language, and the timbre information of the synthesized target speech references the timbre information of the second speaker authorized by the user. This allows for the synthesis of the text spoken by the second speaker authorized by the user in the first language, even if the first language is not the language of the second speaker authorized by the user (or the first language is not the second speaker's native language), effectively solving the problem of cross-language speech synthesis. Furthermore, even if the training data of the monolingual model lacks data in a language other than the speaker's language authorized by the user, the monolingual model can still synthesize text in a language other than the speaker's language authorized by the user. For example, a Chinese speaker speech synthesis model can synthesize English or mixed language text, or an English speaker speech synthesis model can synthesize Chinese or mixed language text.
[0055] Meanwhile, by synthesizing target speech through phoneme sequences and target prosodic information, speech with high pronunciation accuracy and natural prosody can be synthesized using pronunciation and prosodic information. The final synthesized target speech has high intelligibility and natural prosody.
[0056] Figure 3 This is a flowchart illustrating a method for synthesizing target speech according to an exemplary embodiment of this disclosure. Figure 3 As shown, the method may include the following steps.
[0057] Step 310: Based on the speech vector of the first speaker who has been authorized by the user in the first language, determine the duration of each phoneme in the phoneme sequence to obtain the target duration sequence.
[0058] Step 320: Obtain audio features based on the target duration sequence and the timbre vector of the second speaker who has been authorized by the user.
[0059] Step 330: Process the audio features according to the acoustic model to synthesize the target speech.
[0060] In some embodiments, the training of the speech synthesis model can be performed to determine the duration of each phoneme in the phoneme sequence based on the speech vector of the first speaker authorized by the user in the first language, thereby obtaining the target duration sequence and the audio features. That is, steps 310 and 320 above can be performed based on the training of the speech synthesis model. The training process of the speech synthesis model can be found below. Figure 4 The details and related descriptions will not be repeated here.
[0061] In some embodiments, the trained speech synthesis model may include an encoding model and a duration prediction model. Based on the speech vector of a first speaker authorized by the user in a first language, the duration of each phoneme in the phoneme sequence is determined to obtain a target duration sequence. This includes: fusing the phoneme sequence and target prosodic information to obtain a target phoneme sequence; encoding the target phoneme sequence according to the encoding model to obtain a first vector; and processing the speech vector and the first vector according to the duration prediction model to obtain the target duration sequence.
[0062] In some embodiments, the target prosodic information can be represented by a sequence. Fusing the phoneme sequence and the target prosodic information can refer to concatenating the phoneme sequence and the target prosodic information to obtain the target phoneme sequence. In some embodiments, the encoding model can be specifically determined according to the actual situation. For example, the encoding model can be a BERT model or a Transformer model, etc. This disclosure does not impose any restrictions on the specific type of encoding model.
[0063] In some embodiments, the phoneme length prediction model may consist of multiple convolutional layers and a linear layer. In some embodiments, the target duration sequence may be used to characterize the pronunciation duration of each phoneme in the phoneme sequence, and the pronunciation duration may be characterized by audio frames. For example, taking the aforementioned phoneme sequence {YIGE} as an example, the target duration sequence may be {4,12,10,8}, which may reflect that the pronunciation durations of phonemes Y, I, G, and E are 4 frames, 12 frames, 10 frames, and 8 frames, respectively.
[0064] In this embodiment, the target duration sequence is obtained by processing the vector of the target phoneme sequence and the speech vector of the first speaker authorized by the user in the first language. The target phoneme sequence integrates the target prosodic information and the phoneme sequence. As mentioned above, the target prosodic information and the phoneme sequence reflect the prosodic information and pronunciation information of the target speech, respectively, and the speech vector of the first speaker authorized by the user in the first language reflects the rhythmic information of the target speech. Therefore, the duration of each phoneme in the target duration sequence is obtained by referring to the prosodic information, pronunciation information, and rhythmic information, which improves the linguistic naturalness and authenticity of the target speech synthesized based on the target duration sequence (i.e., the speech spoken by the second speaker authorized by the user in the first language), as well as the prosodic naturalness of the target speech.
[0065] In some embodiments, the trained speech synthesis model further includes a decoding model to obtain audio features based on the target duration sequence and the timbre vector of a second speaker authorized by the user. This includes decoding the target duration sequence and the timbre vector using the decoding model to obtain the audio features. In some embodiments, the decoding model may be a recurrent neural network. The audio features may be Mel-spectral features.
[0066] In some embodiments, the acoustic model can be specifically determined according to the actual situation. For example, the acoustic model can be a Griffin-Lim model, a WaveRNN model, or an LPCNet model, etc. This disclosure does not impose any restrictions on the specific type of acoustic model. In some embodiments, the acoustic model can be used as a network layer in a speech synthesis model, or as a post-processing layer in a speech synthesis model.
[0067] Figure 4 This is a flowchart illustrating a method for training a speech synthesis model according to an exemplary embodiment of this disclosure. Figure 4 As shown, the method may include the following steps.
[0068] Step 410: Obtain multiple training samples corresponding to the first language, each training sample including the second training text and sample audio of the second training text.
[0069] For example, taking Chinese as the first language, multiple training samples correspond to Chinese. In this case, each training sample may include Chinese text and audio of the Chinese text. The Chinese text may be a second training text in that training sample, and the audio of the Chinese text may be sample audio of the second training text in that training sample. In some embodiments, the audio of the Chinese text may be the speech of the Chinese text spoken in Chinese by a Chinese speaker who has been authorized by the user. The Chinese speaker who has been authorized by the user may be a speaker whose native language is Chinese, or a speaker who can use Chinese as a language.
[0070] Taking Chinese and English as the first languages as an example, multiple training samples correspond to Chinese and English. Each training sample can include Chinese text and its audio, or English text and its audio. The Chinese and English texts can be second training texts in the training samples, and the audio of the Chinese and English texts can be sample audio of the second training texts in the training samples. In some embodiments, the audio of the English text can be the speech of an English speaker authorized by the user, speaking the English text in English. The authorized speaker can be a native English speaker or a speaker capable of using English as a language. Specific details regarding the audio of the Chinese text can be found in the foregoing description and will not be repeated here.
[0071] Step 420: For each training sample, obtain the sample phoneme sequence of the second training text and the sample target prosodic information corresponding to the sample phoneme sequence, as well as the sample audio features of the sample audio.
[0072] In some embodiments, the sample phoneme sequence may be a sequence of phonemes composed of the second training text. The sample phoneme sequence of the second training text is determined in the same way as the phoneme sequence of the text to be synthesized, as detailed in step 210 and its related description above.
[0073] In some embodiments, the sample target prosodic information is the prosodic information of the first language to which the second training text belongs. The sample target prosodic information corresponding to the sample phoneme sequence is determined in the same way as the target prosodic information corresponding to the phoneme sequence. For details, please refer to step 210 above and its related description, which will not be repeated here.
[0074] In some embodiments, for each training sample, the second training text can be manually annotated with phonemes and prosody based on the sample audio of the second training text in that training sample, combined with auditory perception and the waveform and spectrum of the sample audio, to obtain the sample phoneme sequence and sample target prosodic information. In some embodiments, the sample audio features can be the true features of the sample audio, or the sample audio features can be Mel-spectral features.
[0075] Step 430: Perform the following processing on each training sample according to the speech synthesis model to obtain the predicted audio features corresponding to the training sample. The processing includes: determining the pronunciation duration of each phoneme in the sample phoneme sequence based on the sample speech vector of the sample speaker who has been authorized by the user in the first language, to obtain the first sample duration sequence; and processing the first sample duration sequence and the sample speech vector to obtain the predicted audio features.
[0076] In some embodiments, the sample speech vector of the user-authorized speaker in the first language can be the encoding vector of the sample audio of the second training text in the training samples. For example, if the sample audio of the second training text is English text, then the sample speech vector of the user-authorized speaker in the English language is the encoding vector of the English text audio. In this case, the user-authorized speaker is an English speaker, and the speech synthesis model synthesizes the speech of the English text spoken by the user-authorized English speaker using the training samples. Thus, the user-authorized speaker provides both rhythm and timbre information for the synthesized speech. Therefore, in some embodiments, the sample speech vector of the user-authorized speaker in the first language can reflect the timbre information of the user-authorized speaker.
[0077] As previously mentioned, the speech synthesis model may include an encoding model, a duration prediction model, and a decoding model. In some embodiments, based on the sample speech vector of a sample speaker authorized by the user in a first language, the pronunciation duration of each phoneme in the sample phoneme sequence is determined to obtain a first sample duration sequence, including: fusing the sample phoneme sequence and sample target prosodic information to obtain a sample target phoneme sequence; encoding the sample target phoneme sequence according to the encoding model to obtain a sample first vector; and processing the sample speech vector and the sample first vector according to the duration prediction model to obtain the first sample duration sequence. The specific details of obtaining the first sample duration sequence are the same as those of obtaining the target duration sequence, and can be found in steps 310 and 320 above and their related descriptions, which will not be repeated here.
[0078] In some embodiments, processing the first sample duration sequence and sample speech vectors to obtain predicted audio features includes: decoding the first sample duration sequence and sample speech vectors according to a decoding model to obtain predicted audio features. In some embodiments, the predicted audio features may be predicted Mel spectrum features.
[0079] Step 440: Based on the difference between the first sample duration sequence and the second sample duration sequence, as well as the difference between the predicted audio features and the sample audio features, the second objective loss function value of the speech synthesis model is obtained; the second sample duration sequence is the pronunciation duration of each phoneme in the sample phoneme sequence in the sample audio.
[0080] In some embodiments, the second sample duration sequence can be the actual pronunciation duration of each phoneme in the sample phoneme sequence in the sample audio. In some embodiments, the sample audio and sample phoneme sequence can be processed by a forced alignment tool to obtain the pronunciation duration of each phoneme in the sample phoneme sequence in the sample audio, i.e., the second sample duration sequence. The forced alignment tool can be specifically determined according to the actual situation. For example, the forced alignment tool can be the speech recognition toolkit Kaldi, and this disclosure does not impose any limitations on it.
[0081] As mentioned earlier, the pronunciation duration is at the frame level, the phoneme sequence is at the phoneme level, the first sample duration sequence is at the frame level, and the sample phoneme sequence is at the phoneme level. By using a forced alignment tool to generate the pronunciation duration at the frame level for each phoneme in the sample phoneme sequence at the phoneme level, the alignment between the phoneme level and the frame level can be achieved.
[0082] In some embodiments, the corresponding loss function values can be determined based on the difference between the first sample duration sequence and the second sample duration sequence, and based on the difference between the predicted audio features and the sample audio features. For example, the cross-entropy loss function values can be determined separately, and the loss function values of the two can be fused to obtain the second target loss function value of the speech synthesis model. This fusion can refer to weighted averaging.
[0083] Step 450: Iteratively update the parameters of the speech synthesis model based on the value of the second objective loss function to reduce the value of the second objective loss function until a trained speech synthesis model is obtained.
[0084] During the training of a speech synthesis model, the parameters of the model can be continuously updated based on multiple training samples. For example, the parameters can be continuously adjusted to reduce the value of the second objective loss function corresponding to each training sample, ensuring that the second objective loss function value meets a preset condition. For instance, the loss function value converges, or the loss function value is less than a preset value. When the second objective loss function value meets the preset condition, the model training is complete, and a trained speech synthesis model is obtained. In some embodiments, the Adam optimizer can be used to optimize the training process of the speech synthesis model. For more information on the Adam optimizer, please refer to relevant technologies; details will not be elaborated here.
[0085] In the embodiments of the present disclosure, the speech synthesis model is trained using a plurality of training samples corresponding to the first language. For example, the speech synthesis model is trained using training samples composed of the audio of a Chinese speaker and the Chinese text of the audio for which user authorization has been obtained, as well as the audio of an English speaker and the English text of the audio for which user authorization has been obtained. Thus, it can be seen that the training data of the speech synthesis model does not include data in a language other than the language of the speaker for which user authorization has been obtained. In the application stage of the speech synthesis model, by decomposing the speech from multiple dimensions, decoupled information of tags representing pronunciation information, prosody information, rhythm information, and timbre information is obtained, and the problem of cross-language speech synthesis is effectively solved by this decoupled information of tags.
[0086] In some embodiments, there are multiple first languages to which the text to be synthesized belongs. In the case where there are multiple first languages, the target prosody information of the text to be synthesized can be obtained through a target prosody prediction model. Refer to Figure 5 , Figure 5 is a flowchart of a method for obtaining target prosody information shown according to an exemplary embodiment of the present disclosure. As Figure 5 shown, the method may include the following steps.
[0087] Step 510, extract the phonemes of the text in each first language of the text to be synthesized to obtain a phoneme sequence.
[0088] Exemplarily, taking the text to be synthesized as the aforementioned "Hello seattle" as an example, there are multiple first languages, and the texts in the first languages include the Chinese text "Hello" and the English text "seattle". Further, the phonemes of "Hello" and "seattle" can be extracted respectively, and the respective phonemes are spliced to obtain a phoneme sequence. The method for extracting the phonemes of the text corresponding to the first language is similar to the aforementioned step 210, and specific details can be referred to the aforementioned step 210 and its related descriptions, which will not be elaborated here.
[0089] Step 520, process the text to be synthesized according to the trained target prosody prediction model to obtain the prosody information of the text to be synthesized in each first language output by the prosody prediction model corresponding to each first language; the target prosody prediction model includes a prosody prediction model corresponding to each first language.
[0090] Exemplarily, still taking the aforementioned to-be-synthesized text "Hello Seattle" as an example, the target prosody prediction model may include a Chinese prosody prediction model and an English prosody prediction model. By processing the to-be-synthesized text "Hello Seattle" according to the target prosody prediction model, the prosody information in the Chinese language output by the Chinese prosody prediction model for this to-be-synthesized text, and the prosody information in the English language output by the English prosody prediction model for this to-be-synthesized text can be obtained. For the specific details of the training method of the target prosody prediction model, reference can be made to the following Figure 6 and its related descriptions, which will not be elaborated here.
[0091] Step 530: In the prosody information of the to-be-synthesized text in each first language, extract the prosody information of the text in each first language in the to-be-synthesized text to obtain the initial prosody information of the to-be-synthesized text.
[0092] Exemplarily, still taking the aforementioned example as an example, the texts in the first language in the to-be-synthesized text include the Chinese text "Hello" and the English text "Seattle". The prosody information of "Hello" can be extracted from the prosody information in the Chinese language output by the Chinese prosody prediction model for the to-be-synthesized text, and the prosody information of "Seattle" can be extracted from the prosody information in the English language output by the English prosody prediction model for this to-be-synthesized text, and they are spliced to obtain the initial prosody information of the to-be-synthesized text.
[0093] In some embodiments, the texts in the first language include Chinese texts and non-Chinese texts. After obtaining the initial prosody information of the to-be-synthesized text, the speech synthesis method further includes: when the Chinese text is after the non-Chinese text, determining the boundary tone corresponding to the non-Chinese text in the initial prosody information of the to-be-synthesized text as a falling tone. It can be understood that when the texts in the first language include Chinese texts and non-Chinese texts, the to-be-synthesized text is a mixed-language text.
[0094] Exemplarily, taking English as the non-Chinese language, when the Chinese text is after the English text in the to-be-synthesized text, the boundary tone corresponding to the English text in the initial prosody information is determined as a falling tone. When the English text is still followed by an English text in the to-be-synthesized text, the initial prosody information is not adjusted. By determining the boundary tone corresponding to the non-Chinese text in the initial prosody information of the to-be-synthesized text as a falling tone when the Chinese text is after the non-Chinese text, the prosody connection of different language texts in the mixed-language text is made more natural, and thus, when the to-be-synthesized text is a mixed-language text, the prosody of the synthesized speech of the to-be-synthesized text is made more natural.
[0095] Step 540: Based on the phoneme sequence, expand the initial prosody information to the corresponding phoneme level to obtain the target prosody information corresponding to the phoneme sequence.
[0096] In some embodiments, the initial prosody information is at the text level. Expanding the initial prosody information to the corresponding phoneme level based on the phoneme sequence may refer to performing granularity alignment between the initial prosody information and the phoneme sequence. In some embodiments, the initial prosody information of the same language text may be granularity-aligned with the phonemes corresponding to the language text in the phoneme sequence. For example, the initial prosody information corresponding to the Chinese text "Hello" is granularity-aligned with the phonemes corresponding to the Chinese text "Hello" in the phoneme sequence. For the specific details of expanding to the phoneme level, reference may be made to step 210 and its related descriptions above, which will not be elaborated here.
[0097] In some embodiments, when there is only one first language, the target prosody prediction model can process the text to be synthesized through the prosody prediction model corresponding to the first language, and output the initial prosody information of the text to be synthesized. For example, if the text to be synthesized is a Chinese text, the initial prosody information of the Chinese text can be obtained through the Chinese prosody prediction model in the target prosody prediction model.
[0098] Figure 6 is a flowchart of a method for training a target prosody prediction model according to an exemplary embodiment of the present disclosure. As Figure 6 shown, the method may include the following steps.
[0099] Step 610: Obtain each corresponding first training text in multiple first languages. The first training text includes a label for characterizing the prosody information of the first training text.
[0100] In some embodiments, the label may be used to characterize some true information of the first training text. In some embodiments, the label may be used to characterize the prosody information of the first training text. For the specific details of the prosody information, reference may be made to step 210 and its related descriptions above, which will not be elaborated here. By way of example, still taking the first languages as Chinese and English, the multiple first training texts may include Chinese texts and English texts, and the Chinese texts and English texts respectively include labels for characterizing the prosody information of their respective texts.
[0101] Step 620: For each first training text corresponding to a first language, process the vector of the first training text according to the prosody prediction model corresponding to the first language to obtain the predicted prosody information of the first training text; and obtain the loss function value of the prosody prediction model corresponding to the first language according to the difference between the predicted prosody information and the first label.
[0102] In some embodiments, the vector of the first training text can be obtained by encoding the first training text using a text encoding model, which may include a BERT model or a Transformer model, etc. In some embodiments, the prosody prediction model may be a convolutional neural network or a long short-term memory model, etc.
[0103] In some embodiments, the loss function value of the prosody prediction model corresponding to the first language can be specifically determined according to the actual situation. For example, it can be the cross-entropy loss function value obtained based on the predicted prosody information and the first label. For example, taking Chinese and English as the first languages, for Chinese text, the vector of the Chinese text can be processed according to the Chinese prosody prediction model to obtain the predicted prosody information of the Chinese text, and the loss function value of the Chinese prosody prediction model can be obtained based on the predicted prosody information and the first label of the Chinese text; for English text, the vector of the English text can be processed according to the English prosody prediction model to obtain the predicted prosody information of the English text, and the loss function value of the English prosody prediction model can be obtained based on the predicted prosody information and the first label of the English text.
[0104] Step 630: Determine the first target loss function value of the target prosodic prediction model based on the loss function value of the prosodic prediction model corresponding to each first language.
[0105] In some embodiments, the loss function values of the prosody prediction models corresponding to each first language can be fused to obtain the first target loss function value. Fusion can refer to weighted averaging. For example, taking the previous example, the loss function values of the Chinese prosody prediction model and the English prosody prediction model can be weighted and averaged to obtain the first target loss function value.
[0106] Step 640: Iteratively update the parameters of the target prosody prediction model based on the first target loss function value to reduce the first target loss function value until a well-trained target prosody prediction model is obtained.
[0107] During the training of the target prosodic prediction model, the parameters of the target prosodic prediction model can be continuously updated based on multiple first training texts. For example, the parameters of the target prosodic prediction model (i.e., the parameters of the prosodic prediction model corresponding to each first language) can be continuously adjusted to reduce the first target loss function value corresponding to each first training text, so that the first target loss function value meets a preset condition. For example, the loss function value converges, or the loss function value is less than a preset value. When the first target loss function value meets the preset condition, the model training is complete, and a trained target prosodic prediction model is obtained. In some embodiments, the training process of the target prosodic prediction model can be optimized using the Adam optimizer. For more information on the Adam optimizer, please refer to relevant technologies; further details will not be provided here.
[0108] In this embodiment of the disclosure, prosodic prediction models are constructed for different first languages, so that the final generated target prosodic model can predict the prosodicity of texts in different languages, ensuring the accuracy of the predicted prosodic information and further improving the prosodic naturalness of the synthesized speech.
[0109] Figure 7 This is a block diagram illustrating a speech synthesis apparatus according to an exemplary embodiment of the present disclosure. Figure 7 As shown, the device 700 includes:
[0110] The acquisition module 710 is configured to acquire the phoneme sequence of the text to be synthesized and the target prosodic information corresponding to the phoneme sequence, wherein the target prosodic information is the prosodic information of the first language to which the text to be synthesized belongs.
[0111] The synthesis module 720 is configured to synthesize target speech based on the phoneme sequence, the target prosodic information, the speech vector of a first speaker authorized by the user in the first language, and the timbre vector of a second speaker authorized by the user, wherein the target speech represents the speech of the second speaker authorized by the user speaking the text to be synthesized in the first language.
[0112] In some embodiments, the synthesis module 720 is further configured to:
[0113] Based on the speech vector of the first speaker who has been authorized by the user in the first language, the duration of each phoneme in the phoneme sequence is determined to obtain the target duration sequence;
[0114] Audio features are obtained based on the target duration sequence and the timbre vector of the second speaker who has been authorized by the user.
[0115] The audio features are processed according to the acoustic model to synthesize the target speech.
[0116] In some embodiments, the first language includes multiple languages, and the acquisition module 710 is further configured to:
[0117] Extract the phonemes of each text in the first language from the text to be synthesized to obtain the phoneme sequence;
[0118] The text to be synthesized is processed according to the trained target prosodic prediction model to obtain the prosodic information of the text to be synthesized in the first language, which is output by the prosodic prediction model corresponding to each first language; the target prosodic prediction model includes the prosodic prediction model corresponding to each first language.
[0119] In the prosodic information of the text to be synthesized in each of the first languages, the prosodic information of each of the first languages in the text to be synthesized is extracted to obtain the initial prosodic information of the text to be synthesized.
[0120] Based on the phoneme sequence, the initial prosodic information is extended to the corresponding phoneme level to obtain the target prosodic information corresponding to the phoneme sequence.
[0121] In some embodiments, the target prosodic prediction model is trained based on the following method:
[0122] Obtain a first training text corresponding to each of the multiple first languages, wherein the first training text includes a label for representing the prosodic information of the first training text;
[0123] For each first training text corresponding to the first language, the vector of the first training text is processed according to the prosody prediction model corresponding to the first language to obtain the predicted prosody information of the first training text; and the loss function value of the prosody prediction model corresponding to the first language is obtained according to the difference between the predicted prosody information and the first label.
[0124] Based on the loss function value of the prosody prediction model corresponding to each of the first languages, determine the first target loss function value of the target prosody prediction model;
[0125] The parameters of the target prosodic prediction model are iteratively updated based on the first target loss function value to reduce the first target loss function value until the trained target prosodic prediction model is obtained.
[0126] In some embodiments, the text in the first language includes Chinese text and non-Chinese text, and the synthesis module 720 is further configured to: when the non-Chinese text is followed by the Chinese text, determine the boundary tone corresponding to the non-Chinese text in the initial prosodic information of the text to be synthesized as a falling tone.
[0127] In some embodiments, the synthesis module 720 is further configured to: execute the steps of determining the duration of each phoneme in the phoneme sequence, obtaining a target duration sequence, and obtaining audio features based on the speech vector of the first speaker authorized by the user in the first language according to the trained speech synthesis model;
[0128] The trained speech synthesis model includes an encoding model and a duration prediction model, and the synthesis module 720 is further configured as follows:
[0129] The target phoneme sequence is obtained by fusing the phoneme sequence and the target prosodic information;
[0130] The target phoneme sequence is encoded according to the encoding model to obtain a first vector;
[0131] The speech vector and the first vector are processed according to the pitch length prediction model to obtain the target duration sequence.
[0132] In some embodiments, the trained speech synthesis model further includes a decoding model, and the synthesis module 720 is further configured to:
[0133] The target duration sequence and the timbre vector are decoded according to the decoding model to obtain the audio features.
[0134] In some embodiments, the trained speech synthesis model is obtained based on the following method:
[0135] Obtain multiple training samples corresponding to the first language, each training sample including a second training text and sample audio of the second training text;
[0136] For each training sample, obtain the sample phoneme sequence of the second training text and the sample target prosodic information corresponding to the sample phoneme sequence, as well as the sample audio features of the sample audio.
[0137] For each training sample, the speech synthesis model performs the following processing to obtain the predicted audio features corresponding to that training sample, the processing including:
[0138] Based on the sample speech vector of the sample speaker who has been authorized by the user in the first language, the pronunciation duration of each phoneme in the sample phoneme sequence is determined to obtain the first sample duration sequence.
[0139] The predicted audio features are obtained by processing the first sample duration sequence and the sample speech vector;
[0140] Based on the difference between the first sample duration sequence and the second sample duration sequence, and the difference between the predicted audio features and the sample audio features, the second objective loss function value of the speech synthesis model is obtained; the second sample duration sequence is the pronunciation duration of each phoneme in the sample phoneme sequence in the sample audio.
[0141] The parameters of the speech synthesis model are iteratively updated based on the value of the second objective loss function to reduce the value of the second objective loss function until a well-trained speech synthesis model is obtained.
[0142] The following is for reference. Figure 8 It illustrates an electronic device suitable for implementing embodiments of the present disclosure (e.g., Figure 1 The diagram below shows the structure of the terminal device or server 800. The terminal device in this embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and vehicle terminals (e.g., vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 8 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0143] like Figure 8 As shown, the electronic device 800 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 808 into a random access memory (RAM) 803. The RAM 803 also stores various programs and data required for the operation of the electronic device 800. The processing device 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0144] Typically, the following devices can be connected to I / O interface 805: input devices 806 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 807 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 808 including, for example, magnetic tapes, hard disks, etc.; and communication devices 809. Communication device 809 allows electronic device 800 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 8 An electronic device 800 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0145] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 809, or installed from a storage device 808, or installed from a ROM 802. When the computer program is executed by a processing device 801, it performs the functions defined in the methods of embodiments of this disclosure.
[0146] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0147] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0148] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0149] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: acquire a phoneme sequence of a text to be synthesized and target prosodic information corresponding to the phoneme sequence, wherein the target prosodic information is prosodic information in a first language to which the text to be synthesized belongs; and synthesize target speech based on the phoneme sequence, the target prosodic information, a speech vector of a first speaker authorized by the user in the first language, and a timbre vector of a second speaker authorized by the user, wherein the target speech represents the speech of the second speaker authorized by the user speaking the text to be synthesized in the first language.
[0150] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0151] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0152] The modules described in the embodiments of this disclosure can be implemented in software or hardware. The names of the modules are not, in some cases, intended to limit the functionality of the module itself.
[0153] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0154] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0155] According to one or more embodiments of this disclosure, Example 1 provides a speech synthesis method, including:
[0156] Obtain the phoneme sequence of the text to be synthesized and the target prosodic information corresponding to the phoneme sequence, wherein the target prosodic information is the prosodic information under the first language to which the text to be synthesized belongs;
[0157] Based on the phoneme sequence, the target prosodic information, the speech vector of the first speaker authorized by the user in the first language, and the timbre vector of the second speaker authorized by the user, a target speech is synthesized, wherein the target speech represents the speech of the second speaker authorized by the user speaking the text to be synthesized in the first language.
[0158] According to one or more embodiments of this disclosure, Example 2 provides a speech synthesis method of Example 1, wherein synthesizing target speech based on the phoneme sequence, the target prosodic information, the speech vector of a first speaker authorized by the user in the first language, and the timbre vector of a second speaker authorized by the user, includes:
[0159] Based on the speech vector of the first speaker who has been authorized by the user in the first language, the duration of each phoneme in the phoneme sequence is determined to obtain the target duration sequence;
[0160] Audio features are obtained based on the target duration sequence and the timbre vector of the second speaker who has been authorized by the user.
[0161] The audio features are processed according to the acoustic model to synthesize the target speech.
[0162] According to one or more embodiments of this disclosure, Example 3 provides a speech synthesis method of Example 1, wherein the first language includes multiple languages, and the step of obtaining the phoneme sequence of the text to be synthesized and the target prosodic information corresponding to the phoneme sequence includes:
[0163] Extract the phonemes of each text in the first language from the text to be synthesized to obtain the phoneme sequence;
[0164] The text to be synthesized is processed according to the trained target prosodic prediction model to obtain the prosodic information of the text to be synthesized in the first language, which is output by the prosodic prediction model corresponding to each first language; the target prosodic prediction model includes the prosodic prediction model corresponding to each first language.
[0165] In the prosodic information of the text to be synthesized in each of the first languages, the prosodic information of each of the first languages in the text to be synthesized is extracted to obtain the initial prosodic information of the text to be synthesized.
[0166] Based on the phoneme sequence, the initial prosodic information is extended to the corresponding phoneme level to obtain the target prosodic information corresponding to the phoneme sequence.
[0167] According to one or more embodiments of this disclosure, Example 4 provides a speech synthesis method of Example 3, wherein the target prosodic prediction model is trained based on the following:
[0168] Obtain a first training text corresponding to each of the multiple first languages, wherein the first training text includes a label for representing the prosodic information of the first training text;
[0169] For each first training text corresponding to the first language, the vector of the first training text is processed according to the prosody prediction model corresponding to the first language to obtain the predicted prosody information of the first training text; and the loss function value of the prosody prediction model corresponding to the first language is obtained according to the difference between the predicted prosody information and the first label.
[0170] Based on the loss function value of the prosody prediction model corresponding to each of the first languages, determine the first target loss function value of the target prosody prediction model;
[0171] The parameters of the target prosodic prediction model are iteratively updated based on the first target loss function value to reduce the first target loss function value until the trained target prosodic prediction model is obtained.
[0172] According to one or more embodiments of this disclosure, Example 5 provides the speech synthesis method of Example 3, wherein the text in the first language includes Chinese text and non-Chinese text, and after obtaining the initial prosodic information of the text to be synthesized, the method further includes:
[0173] When the non-Chinese text is followed by the Chinese text, the boundary tone corresponding to the non-Chinese text in the initial prosodic information of the text to be synthesized is determined to be a falling tone.
[0174] According to one or more embodiments of this disclosure, Example 6 provides a speech synthesis method of Example 2, which executes the steps of determining the duration of each phoneme in the phoneme sequence based on the speech vector of the first speaker who has been authorized by the user in the first language, obtaining a target duration sequence, and obtaining audio features based on a trained speech synthesis model;
[0175] The trained speech synthesis model includes an encoding model and a duration prediction model. The step of determining the duration of each phoneme in the phoneme sequence based on the speech vector of the first speaker (who has obtained user authorization) in the first language, to obtain a target duration sequence, includes:
[0176] The target phoneme sequence is obtained by fusing the phoneme sequence and the target prosodic information;
[0177] The target phoneme sequence is encoded according to the encoding model to obtain a first vector;
[0178] The speech vector and the first vector are processed according to the pitch length prediction model to obtain the target duration sequence.
[0179] According to one or more embodiments of this disclosure, Example 7 provides the speech synthesis method of Example 6, wherein the trained speech synthesis model further includes a decoding model, and the step of obtaining audio features based on the target duration sequence and the timbre vector of the second speaker who has obtained user authorization includes:
[0180] The target duration sequence and the timbre vector are decoded according to the decoding model to obtain the audio features.
[0181] According to one or more embodiments of this disclosure, Example 8 provides the speech synthesis method of Example 6, wherein the trained speech synthesis model is obtained based on the following method:
[0182] Obtain multiple training samples corresponding to the first language, each training sample including a second training text and sample audio of the second training text;
[0183] For each training sample, obtain the sample phoneme sequence of the second training text and the sample target prosodic information corresponding to the sample phoneme sequence, as well as the sample audio features of the sample audio.
[0184] For each training sample, the speech synthesis model performs the following processing to obtain the predicted audio features corresponding to that training sample, the processing including:
[0185] Based on the sample speech vector of the sample speaker who has been authorized by the user in the first language, the pronunciation duration of each phoneme in the sample phoneme sequence is determined to obtain the first sample duration sequence.
[0186] The predicted audio features are obtained by processing the first sample duration sequence and the sample speech vector;
[0187] Based on the difference between the first sample duration sequence and the second sample duration sequence, and the difference between the predicted audio features and the sample audio features, the second objective loss function value of the speech synthesis model is obtained; the second sample duration sequence is the pronunciation duration of each phoneme in the sample phoneme sequence in the sample audio.
[0188] The parameters of the speech synthesis model are iteratively updated based on the value of the second objective loss function to reduce the value of the second objective loss function until a well-trained speech synthesis model is obtained.
[0189] According to one or more embodiments of this disclosure, Example 9 provides a speech synthesis apparatus, including:
[0190] The acquisition module is configured to acquire the phoneme sequence of the text to be synthesized and the target prosodic information corresponding to the phoneme sequence, wherein the target prosodic information is the prosodic information of the first language to which the text to be synthesized belongs;
[0191] The synthesis module is configured to synthesize target speech based on the phoneme sequence, the target prosodic information, the speech vector of a first speaker authorized by the user in the first language, and the timbre vector of a second speaker authorized by the user, wherein the target speech represents the speech of the second speaker authorized by the user speaking the text to be synthesized in the first language.
[0192] According to one or more embodiments of this disclosure, Example 10 provides a speech synthesis apparatus of Example 9, wherein the synthesis module is further configured to:
[0193] Based on the speech vector of the first speaker who has been authorized by the user in the first language, the duration of each phoneme in the phoneme sequence is determined to obtain the target duration sequence;
[0194] Audio features are obtained based on the target duration sequence and the timbre vector of the second speaker who has been authorized by the user.
[0195] The audio features are processed according to the acoustic model to synthesize the target speech.
[0196] According to one or more embodiments of this disclosure, Example 11 provides a speech synthesis apparatus of Example 9, wherein the first language includes multiple languages, and the acquisition module is further configured to:
[0197] Extract the phonemes of each text in the first language from the text to be synthesized to obtain the phoneme sequence;
[0198] The text to be synthesized is processed according to the trained target prosodic prediction model to obtain the prosodic information of the text to be synthesized in the first language, which is output by the prosodic prediction model corresponding to each first language; the target prosodic prediction model includes the prosodic prediction model corresponding to each first language.
[0199] In the prosodic information of the text to be synthesized in each of the first languages, the prosodic information of each of the first languages in the text to be synthesized is extracted to obtain the initial prosodic information of the text to be synthesized.
[0200] Based on the phoneme sequence, the initial prosodic information is extended to the corresponding phoneme level to obtain the target prosodic information corresponding to the phoneme sequence.
[0201] According to one or more embodiments of this disclosure, Example 12 provides a speech synthesis apparatus of Example 11, wherein the target prosodic prediction model is trained based on the following:
[0202] Obtain a first training text corresponding to each of the multiple first languages, wherein the first training text includes a label for representing the prosodic information of the first training text;
[0203] For each first training text corresponding to the first language, the vector of the first training text is processed according to the prosody prediction model corresponding to the first language to obtain the predicted prosody information of the first training text; and the loss function value of the prosody prediction model corresponding to the first language is obtained according to the difference between the predicted prosody information and the first label.
[0204] Based on the loss function value of the prosody prediction model corresponding to each of the first languages, determine the first target loss function value of the target prosody prediction model;
[0205] The parameters of the target prosodic prediction model are iteratively updated based on the first target loss function value to reduce the first target loss function value until the trained target prosodic prediction model is obtained.
[0206] According to one or more embodiments of this disclosure, Example 13 provides a speech synthesis apparatus of Example 11, wherein the text of the first language includes Chinese text and non-Chinese text, and the synthesis module is further configured to: when the non-Chinese text is followed by the Chinese text, determine the boundary tone corresponding to the non-Chinese text in the initial prosodic information of the text to be synthesized as a falling tone.
[0207] According to one or more embodiments of this disclosure, Example 14 provides a speech synthesis apparatus of Example 10, wherein the synthesis module is further configured to: execute the steps of determining the duration of each phoneme in the phoneme sequence, obtaining a target duration sequence, and obtaining audio features based on the speech vector of the first speaker authorized by the user in the first language according to a trained speech synthesis model;
[0208] The trained speech synthesis model includes an encoding model and a duration prediction model, and the synthesis module is further configured as follows:
[0209] The target phoneme sequence is obtained by fusing the phoneme sequence and the target prosodic information;
[0210] The target phoneme sequence is encoded according to the encoding model to obtain a first vector;
[0211] The speech vector and the first vector are processed according to the pitch length prediction model to obtain the target duration sequence.
[0212] According to one or more embodiments of this disclosure, Example 15 provides a speech synthesis apparatus of Example 14, wherein the trained speech synthesis model further includes a decoding model, and the synthesis module is further configured to:
[0213] The target duration sequence and the timbre vector are decoded according to the decoding model to obtain the audio features.
[0214] According to one or more embodiments of this disclosure, Example 16 provides a speech synthesis apparatus of Example 14, wherein the trained speech synthesis model is obtained based on the following method:
[0215] Obtain multiple training samples corresponding to the first language, each training sample including a second training text and sample audio of the second training text;
[0216] For each training sample, obtain the sample phoneme sequence of the second training text and the sample target prosodic information corresponding to the sample phoneme sequence, as well as the sample audio features of the sample audio.
[0217] For each training sample, the speech synthesis model performs the following processing to obtain the predicted audio features corresponding to that training sample, the processing including:
[0218] Based on the sample speech vector of the sample speaker who has been authorized by the user in the first language, the pronunciation duration of each phoneme in the sample phoneme sequence is determined to obtain the first sample duration sequence.
[0219] The predicted audio features are obtained by processing the first sample duration sequence and the sample speech vector;
[0220] Based on the difference between the first sample duration sequence and the second sample duration sequence, and the difference between the predicted audio features and the sample audio features, the second objective loss function value of the speech synthesis model is obtained; the second sample duration sequence is the pronunciation duration of each phoneme in the sample phoneme sequence in the sample audio.
[0221] The parameters of the speech synthesis model are iteratively updated based on the value of the second objective loss function to reduce the value of the second objective loss function until a well-trained speech synthesis model is obtained.
[0222] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0223] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0224] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative forms of implementing the claims. Regarding the apparatus in the above embodiments, the specific manner in which the various modules perform their operations has been described in detail in the embodiments relating to the method, and will not be elaborated upon here.
Claims
1. A speech synthesis method, characterized in that, include: Obtain the phoneme sequence of the text to be synthesized and the target prosodic information corresponding to the phoneme sequence, wherein the target prosodic information is the prosodic information under the first language to which the text to be synthesized belongs; Based on the phoneme sequence, the target prosodic information, the speech vector of the first speaker in the first language, and the timbre vector of the second speaker, the target speech is synthesized, wherein the target speech represents the speech of the second speaker speaking the text to be synthesized in the first language; The first language includes multiple languages. The step of obtaining the phoneme sequence of the text to be synthesized and the target prosodic information corresponding to the phoneme sequence includes: extracting phonemes from each text in the first language of the text to be synthesized to obtain the phoneme sequence; processing the text to be synthesized according to a trained target prosodic prediction model to obtain the prosodic information of the text to be synthesized under that first language, output by the prosodic prediction model corresponding to each first language; the target prosodic prediction model includes a prosodic prediction model corresponding to each first language; extracting the prosodic information of each text in the first language of the text to be synthesized from the prosodic information of the text to be synthesized under each first language to obtain the initial prosodic information of the text to be synthesized; and expanding the initial prosodic information to the corresponding phoneme level based on the phoneme sequence to obtain the target prosodic information corresponding to the phoneme sequence.
2. The speech synthesis method according to claim 1, characterized in that, The process of synthesizing target speech based on the phoneme sequence, the target prosodic information, the speech vector of the first speaker in the first language, and the timbre vector of the second speaker includes: Based on the speech vector of the first speaker in the first language, the duration of each phoneme in the phoneme sequence is determined to obtain the target duration sequence; Based on the target duration sequence and the timbre vector of the second speaker, audio features are obtained; The audio features are processed according to the acoustic model to synthesize the target speech.
3. The speech synthesis method according to claim 1, characterized in that, The target prosodic prediction model is trained based on the following method: Obtain a first training text corresponding to each of the multiple first languages, wherein the first training text includes a label for representing the prosodic information of the first training text; For each first training text corresponding to the first language, the vector of the first training text is processed according to the prosody prediction model corresponding to the first language to obtain the predicted prosody information of the first training text; and the loss function value of the prosody prediction model corresponding to the first language is obtained according to the difference between the predicted prosody information and the first label. Based on the loss function value of the prosodic prediction model corresponding to each of the first languages, determine the first target loss function value of the target prosodic prediction model; The parameters of the target prosodic prediction model are iteratively updated based on the first target loss function value to reduce the first target loss function value until the trained target prosodic prediction model is obtained.
4. The speech synthesis method according to claim 1, characterized in that, The text in the first language includes Chinese text and non-Chinese text. After obtaining the initial prosodic information of the text to be synthesized, the method further includes: When the non-Chinese text is followed by the Chinese text, the boundary tone corresponding to the non-Chinese text in the initial prosodic information of the text to be synthesized is determined to be a falling tone.
5. The speech synthesis method according to claim 2, characterized in that, The steps of obtaining audio features are performed based on the trained speech synthesis model, namely, determining the duration of each phoneme in the phoneme sequence according to the speech vector of the first speaker in the first language, and obtaining the target duration sequence. The trained speech synthesis model includes an encoding model and a duration prediction model. The step of determining the duration of each phoneme in the phoneme sequence based on the speech vector of the first speaker in the first language to obtain a target duration sequence includes: The target phoneme sequence is obtained by fusing the phoneme sequence and the target prosodic information; The target phoneme sequence is encoded according to the encoding model to obtain a first vector; The speech vector and the first vector are processed according to the pitch length prediction model to obtain the target duration sequence.
6. The speech synthesis method according to claim 5, characterized in that, The trained speech synthesis model also includes a decoding model, wherein the audio features are obtained based on the target duration sequence and the timbre vector of the second speaker, including: The target duration sequence and the timbre vector are decoded according to the decoding model to obtain the audio features.
7. The speech synthesis method according to claim 5, characterized in that, The trained speech synthesis model is obtained based on the following method: Obtain multiple training samples corresponding to the first language, each training sample including a second training text and sample audio of the second training text; For each training sample, obtain the sample phoneme sequence of the second training text and the sample target prosodic information corresponding to the sample phoneme sequence, as well as the sample audio features of the sample audio. For each training sample, the speech synthesis model performs the following processing to obtain the predicted audio features corresponding to that training sample, the processing including: Based on the sample speech vector of the sample speaker in the first language, the pronunciation duration of each phoneme in the sample phoneme sequence is determined to obtain the first sample duration sequence. The predicted audio features are obtained by processing the first sample duration sequence and the sample speech vector; Based on the difference between the first sample duration sequence and the second sample duration sequence, and the difference between the predicted audio features and the sample audio features, the second objective loss function value of the speech synthesis model is obtained; the second sample duration sequence is the pronunciation duration of each phoneme in the sample phoneme sequence in the sample audio. The parameters of the speech synthesis model are iteratively updated based on the value of the second objective loss function to reduce the value of the second objective loss function until a well-trained speech synthesis model is obtained.
8. A speech synthesis device, characterized in that, include: The acquisition module is configured to acquire the phoneme sequence of the text to be synthesized and the target prosodic information corresponding to the phoneme sequence, wherein the target prosodic information is the prosodic information of the first language to which the text to be synthesized belongs; The synthesis module is configured to synthesize target speech based on the phoneme sequence, the target prosodic information, the speech vector of the first speaker in the first language, and the timbre vector of the second speaker, wherein the target speech represents the speech of the second speaker speaking the text to be synthesized in the first language; The first language includes multiple languages, and the acquisition module is further configured to: extract the phonemes of the text of each first language in the text to be synthesized, and obtain the phoneme sequence; process the text to be synthesized according to the trained target prosodic prediction model to obtain the prosodic information of the text to be synthesized under the first language output by the prosodic prediction model corresponding to each first language; The target prosodic prediction model includes a prosodic prediction model corresponding to each of the first languages; in the prosodic information of the text to be synthesized under each of the first languages, the prosodic information of each of the first languages in the text to be synthesized is extracted to obtain the initial prosodic information of the text to be synthesized. Based on the phoneme sequence, the initial prosodic information is extended to the corresponding phoneme level to obtain the target prosodic information corresponding to the phoneme sequence.
9. A computer-readable medium having a computer program stored thereon, characterized in that, When executed by the processing device, the program implements the steps of the method described in any one of claims 1-7.
10. An electronic device, characterized in that, include: A storage device having at least one computer program stored thereon; At least one processing device is configured to execute the at least one computer program in the storage device to implement the steps of the method according to any one of claims 1-7.
Citation Information
Patent Citations
Multilingual mixed-language text character-pronunciation conversion method and system
CN105989833A
Speech synthesis device supporting styles of multiple speakers, language switching and controllable rhythm
CN112863483A
Speech synthesis method and device and device for speech synthesis
CN113889070A