A speech synthesis method, apparatus and electronic device
By using a shared encoder and intermediate decoder structure, the problem of low speech synthesis accuracy for languages with small data volumes is solved, and high-precision speech information generation is achieved.
Patent Information
- Application Number
- CN202411882752.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2044-12-19
AI Technical Summary
Existing technologies struggle to improve the accuracy of speech synthesis for languages with limited data, such as the Min dialect, especially due to data scarcity and the difficulty in capturing language features.
It employs a shared encoder and language-specific intermediate and decoder structures. The shared encoder learns language-independent abstract feature representations, and the language-specific intermediate and decoder further learn the features and pronunciation characteristics of each language to generate high-precision speech information.
It improves the speech synthesis accuracy of languages with small data volumes, such as the Min dialect, ensuring that high-precision speech information can be generated through this model regardless of the size of the language data.
Smart Images

Figure CN119943025B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the technical field of speech synthesis, and in particular to a speech synthesis method, device and electronic equipment. BACKGROUND
[0002] TTS (text-to-speech) technology is a technology for converting text information into speech, so that computers, smart devices or other application programs can broadcast text in a form that humans can understand. The application of TTS systems is very wide, including intelligent broadcasting, navigation systems, voice assistants, audio books, etc. It can be said that speech synthesis technology is involved in everyone's life to some extent, such as mobile phone assistant Siri, smart speaker XiaoDu, airport and high-speed rail broadcasting, and map navigation, etc.
[0003] Taking Chinese as an example, the speech synthesis effect of Chinese Mandarin has generally reached the expected requirements, and the naturalness and intelligibility are relatively high, but for minority languages, such as dialects, there are still many problems. Taking Min dialect as an example, Min dialect is divided into 7 areas, and each area is further divided into multiple pieces (for example, the Minnan area can be divided into Zhangquan piece, Datian piece and Chaoshan piece), so for a specific local dialect, the data is relatively scarce. Also due to the very complex language phenomenon (such as Min dialects cannot be communicated between north and south, and there are differences between east and west), so it is difficult to capture language characteristics (such as tone, intonation, etc.). Therefore, how to improve the speech synthesis precision for small data languages is a problem to be solved by the present application. SUMMARY
[0004] Embodiments of the present application provide a speech synthesis method, device and electronic equipment, which generates the corresponding spectral features of each language by sharing the encoder and the language-specific intermediate layer and decoder structure, and then obtains the corresponding speech information of each language, thereby improving the speech synthesis precision of small data languages.
[0005] The first aspect of the embodiments of the present application provides a speech synthesis method, which comprises:
[0006] processing a target text to obtain target phoneme information, wherein the target text comprises one or more texts to be processed;
[0007] inputting the target phoneme information and target language information into a speech synthesis model to obtain target spectral features corresponding to the target text, wherein the target language information is one or more language information corresponding to the target text;
[0008] obtaining speech information of a target language based on the target spectral features;
[0009] The speech synthesis model at least comprises: a shared encoder, a plurality of intermediate layers corresponding to a plurality of languages respectively, and a plurality of decoders corresponding to the plurality of languages respectively; the shared encoder is used for processing text conversion tasks of the plurality of languages to generate abstract feature representations irrelevant to the languages; the plurality of intermediate layers corresponding to the plurality of languages respectively are used for enhancing respective characteristics of the plurality of languages; and the plurality of decoders corresponding to the plurality of languages respectively are used for learning pronunciation features corresponding to the plurality of languages respectively.
[0010] The second aspect of the embodiment of the present application provides a speech synthesis device, and the device comprises:
[0011] A target phoneme obtaining module is configured to process target text to obtain target phoneme information, wherein the target text comprises one or more texts to be processed.
[0012] A first model processing module is configured to input the target phoneme information and target language information into a speech synthesis model to obtain target spectral features corresponding to the target text, wherein the target language information is one or more language information corresponding to the target text.
[0013] A target speech synthesis module is configured to obtain speech information of a target language based on the target spectral features.
[0014] The speech synthesis model at least comprises: a shared encoder, a plurality of intermediate layers corresponding to a plurality of languages respectively, and a plurality of decoders corresponding to the plurality of languages respectively; the shared encoder is used for processing text conversion tasks of the plurality of languages to generate abstract feature representations irrelevant to the languages; the plurality of intermediate layers corresponding to the plurality of languages respectively are used for enhancing respective characteristics of the plurality of languages; and the plurality of decoders corresponding to the plurality of languages respectively are used for learning pronunciation features corresponding to the plurality of languages respectively.
[0015] The third aspect of the embodiment of the present application provides an electronic device, which comprises a memory, a processor, and a computer program stored in the memory and running on the processor, and the computer program is executed by the processor to implement the speech synthesis method of the first aspect of the embodiment of the present application.
[0016] In the speech synthesis method provided by the embodiment of the present application, target phoneme information and target language information corresponding to target text are input into a pre-trained speech synthesis model to obtain target spectral information corresponding to the target text output by the speech synthesis model, so as to obtain speech information of a target language. The speech synthesis model at least comprises: a shared encoder used for generating abstract feature representations irrelevant to languages; a plurality of intermediate layers corresponding to a plurality of languages respectively used for enhancing respective characteristics of the plurality of languages; and a plurality of decoders corresponding to the plurality of languages respectively used for learning pronunciation features corresponding to the plurality of languages respectively.
[0017] In the speech synthesis model of the embodiment, the abstract feature representation irrelevant to the language is learned based on the related information of multiple languages by the shared encoder, so that each language can obtain the corresponding abstract feature representation irrelevant to the language based on the shared encoder, and then the corresponding feature and pronunciation characteristic of each language are further learned based on the corresponding related information of each language through the corresponding intermediate layer and the decoder of each language, so that the abstract feature representation irrelevant to the language can be obtained through the shared encoder in the speech synthesis model regardless of the data size (i.e. the size) of the language, and then the accurate pronunciation feature corresponding to the language is obtained through the corresponding intermediate layer and the decoder of the language, and then the high-precision speech information is obtained, thereby improving the speech synthesis precision of the language with small data size. BRIEF DESCRIPTION OF DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the description of the embodiments of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0019] Figure 1 is a structure diagram of a tacotron2 model shown in the related art;
[0020] Figure 2 is a flowchart of a speech synthesis method according to an embodiment of the present application;
[0021] Figure 3 is a structure diagram of a speech synthesis model according to an embodiment of the present application;
[0022] Figure 4 is a structure block diagram of a speech synthesis device according to an embodiment of the present application;
[0023] Figure 5 is a schematic diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0024] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0025] Generally speaking, the working principle of a TTS system can be divided into the following aspects: text processing (receiving input text and pre-processing it, including text standardization such as date, amount, number, symbol conversion, punctuation processing and word segmentation), text analysis (such as part-of-speech tagging), phoneme conversion (text conversion into phoneme sequence) and sound synthesis (conversion of phoneme sequence into waveform).
[0026] Commonly used TTS neural network architectures include tacotron, fast speech and vits, etc. Taking tacotron2 as an example, as shown in Figure 1 Figure 1 is a structural schematic diagram of a tacotron2 model in the related art. In Figure 1 , the main components of the tacotron2 model are an encoder, an attention mechanism-based decoder and a post-processing network, and the tacotron2 model can realize the process from character (text) input to spectrum output (and then converted into waveform). Among them, the main role of the encoder is to convert the character input into a set of vectors, that is, the encoder aims to convert the input text into a set of high-robustness sequence representations. The attention mechanism-based decoder realizes the frame-by-frame prediction of the encoded input sequence to the spectrum by learning the alignment process of the text and the speech signal. The spectrum predicted by the decoder is finally converted into a waveform through a series of linear network layers.
[0027] However, if the above tacotron2 model has a large amount of data of a single language (such as Chinese Mandarin, more than 10 hours for a single person), the tacotron2 model can completely complete the speech synthesis task. However, for small dialects and small languages with limited data (the data required for speech synthesis is annotated and relatively clean in voice quality), the amount of data that can be obtained is very limited. Therefore, it is not very realistic to use the above model structure to complete this task.
[0028] Another solution is the traditional multi-language speech synthesis solution, that is, all language data is given to the model without discrimination, but for small languages and dialects with scarce data, the synthesis quality is usually not good. In addition, such systems often cannot effectively distinguish the subtle differences between different languages, which also leads to the fact that the output of speech synthesis cannot accurately reflect the speech characteristics of a specific language.
[0029] Another common solution is to assume that there is a certain correlation between the two languages, and if one language has a large amount of data, a base model is trained using the language with a large amount of data, and then the second language with a small amount of data is trained using the transfer learning method. For example, a base model is first trained using Chinese, and then migrated to Min dialect. However, this method requires high similarity between the two languages. For example, if the base is Chinese, it is feasible to migrate to Northeast dialect, but it is not realistic to migrate to Shanghai dialect.
[0030] Therefore, in order to at least partially solve one or more of the above problems and other potential problems, embodiments of the present application propose a speech synthesis method, a pre-trained speech synthesis model processes target language and target phoneme information to obtain target spectral features, and then obtains speech information of the target language based on the target spectral features. The speech synthesis model comprises at least: a shared encoder, a plurality of language-specific intermediate layers, and a plurality of language-specific decoders; the shared encoder is used to learn abstract feature representations independent of the language based on relevant information of a plurality of languages, so that each language can obtain its own abstract feature representation independent of the language based on the shared encoder, and then further learn language-specific features and pronunciation characteristics based on the relevant information of each language through the language-specific intermediate layer and the decoder of each language, so that the speech synthesis model can effectively utilize all available data, and each component (such as the plurality of language-specific intermediate layers and the plurality of language-specific decoders) in the model can fully play its role, regardless of the size of the language data (i.e. the size of the scale), the shared encoder in the speech synthesis model can obtain abstract feature representations independent of the language, and then through the language-specific intermediate layer and the decoder, accurate pronunciation features corresponding to the language can be obtained, and high-precision speech information can be obtained, thereby improving the speech synthesis accuracy of the language with small amount of data.
[0031] In the following, specific examples of the present solution will be described in more detail with reference to the accompanying drawings.
[0032] Reference Figure 2 , Figure 2 is a flowchart of a speech synthesis method according to an embodiment of the present application. As shown in Figure 2 , the speech synthesis method according to the present embodiment can include the following steps:
[0033] Step S11: processing the target text to obtain target phoneme information, wherein the target text comprises one or more texts to be processed.
[0034] In this embodiment, the target text can be processed to obtain target phoneme information. For example, the target text can be processed by text processing, text analysis, and phoneme conversion to obtain the target phoneme information. The target phoneme information of this embodiment is the phoneme information corresponding to the target text, and the target text includes one or more texts to be processed, i.e., the text to be processed is the text to be synthesized by voice. That is, the target text in this embodiment can be one target text or multiple target texts.
[0035] This embodiment does not limit the specific way of processing the target text to obtain the target phoneme information, and any method that can convert text into phoneme information is within the protection scope of this embodiment.
[0036] Step S12: inputting the target phoneme information and target language information into a voice synthesis model to obtain target spectral features corresponding to the target text, the target language information being one or more language information corresponding to the target text.
[0037] In this embodiment, the target language information corresponding to the target text also needs to be determined according to the voice synthesis requirement, and the target language information is the related information of the target language, and the target language is the language corresponding to the voice to be synthesized in the voice synthesis requirement of the target text. The target language information of this embodiment is one or more language information corresponding to the target text, that is, one target text in this embodiment can correspond to one or more target languages.
[0038] In this embodiment, the target phoneme information and the target language information can be input into the pre-trained voice synthesis model to obtain the target spectral features corresponding to the target text output by the voice synthesis model, and the target spectral features are the spectral features corresponding to the target text.
[0039] The pre-trained voice synthesis model of this embodiment at least includes a shared encoder, a plurality of language-specific intermediate layers, and a plurality of language-specific decoders. The shared encoder is used to process the text conversion task of multiple languages to generate abstract feature representation independent of the language; the plurality of language-specific intermediate layers are respectively used to enhance the characteristics of the plurality of languages; and the plurality of language-specific decoders are respectively used to learn the pronunciation features corresponding to the plurality of languages.
[0040] In this embodiment, the pre-trained speech synthesis model can be used to generate the spectral features (acoustic features) required for synthesizing speech information in multiple languages, and does not limit the size or data volume of the languages, which can include any one language. Therefore, when the target phoneme information corresponding to one or more target texts and the target language information corresponding to the one or more target texts are input into the speech synthesis model, the corresponding processing can be performed based on the shared encoder, the intermediate layer corresponding to the target language, and the decoder in the speech synthesis model, thereby outputting the target spectral features corresponding to the one or more target texts.
[0041] Step S13: obtaining speech information of the target language based on the target spectral features.
[0042] In this embodiment, after obtaining the target spectral features, the speech information of the target language can be obtained based on the target spectral features. The specific manner of processing the target spectral features to obtain the speech information of the target language in this embodiment is not limited in any way, and any method that can convert spectral features into speech information is within the protection scope of this embodiment.
[0043] In addition, since the target text in this embodiment can be one or more, and the target language information is one or more language information corresponding to the target text, based on this, the possible situations of this embodiment are described below, but are not limited to the following situations:
[0044] In the case where the target text is one and the target language information corresponding to the target text is also one, i.e., the target text is one and the target language is one, for example, the target language corresponding to the target text A is Minnan, at this time, the target phoneme information of the target text A and the language information of Minnan are input into the speech synthesis model, the spectral features A corresponding to the target text A are obtained, and the speech information A of Minnan is obtained based on the spectral features A.
[0045] In the case where the target text is one and the target language information corresponding to the target text is multiple, i.e., the target text is one and the target language is multiple, for example, the target language corresponding to the target text A is Northeastern Mandarin and Mandarin, at this time, the target phoneme information of the target text A and the language information of Northeastern Mandarin can be input into the speech synthesis model at the same time, and the target phoneme information of the target text A and the language information of Mandarin are input into the speech synthesis model, the spectral features A1 and A2 corresponding to the target text A output by the speech synthesis model are obtained, the speech information A1 of Northeastern Mandarin is obtained based on the spectral features A1, and the speech information A2 of Mandarin is obtained based on the spectral features A2.
[0046] In the case that the target text is multiple, and the target language information corresponding to each target text is also multiple, i.e., the target text is multiple, and the target language corresponding to each target text is multiple, for example, the target language corresponding to target text A is Tianjin dialect and Mandarin, and the target language corresponding to target text B is Sichuan dialect and Northeast dialect, at this time, the target phoneme information of target text A and the language information of Tianjin dialect can be input into the speech synthesis model, the target phoneme information of target text A and the language information of Mandarin can be input into the speech synthesis model, the target phoneme information of target text B and the language information of Sichuan dialect can be input into the speech synthesis model, and the target phoneme information of target text B and the language information of Northeast dialect can be input into the speech synthesis model, to obtain the spectral features A1 and A2 corresponding to target text A output by the speech synthesis model, and to obtain the spectral features B1 and B2 corresponding to target text B, so as to obtain the speech information A of Tianjin dialect based on the spectral feature A1, to obtain the speech information A2 of Mandarin based on the spectral feature A2, to obtain the speech information B1 of Sichuan dialect based on the spectral feature B1, and to obtain the speech information B2 of Northeast dialect based on the spectral feature B2.
[0047] In the case that the target text is multiple, and the target language information corresponding to each target text is also multiple, i.e., the target text is multiple, and the target language corresponding to each target text is multiple, for example, the target language corresponding to target text A is Tianjin dialect and Mandarin, and the target language corresponding to target text B is Sichuan dialect and Northeast dialect, at this time, the target phoneme information of target text A and the language information of Tianjin dialect can be input into the speech synthesis model, the target phoneme information of target text A and the language information of Mandarin can be input into the speech synthesis model, the target phoneme information of target text B and the language information of Sichuan dialect can be input into the speech synthesis model, and the target phoneme information of target text B and the language information of Northeast dialect can be input into the speech synthesis model, to obtain the spectral features A1 and A2 corresponding to target text A output by the speech synthesis model, and to obtain the spectral features B1 and B2 corresponding to target text B, so as to obtain the speech information A of Tianjin dialect based on the spectral feature A1, to obtain the speech information A2 of Mandarin based on the spectral feature A2, to obtain the speech information B1 of Sichuan dialect based on the spectral feature B1, and to obtain the speech information B2 of Northeast dialect based on the spectral feature B2.
[0048] In the embodiment, the shared encoder is fully learned based on the related information of multiple languages, so that each language can obtain the corresponding abstract feature representation independent of language based on the shared encoder, and then the corresponding feature and pronunciation characteristic of each language are further learned based on the corresponding related information of each language through the corresponding intermediate layer and decoder of each language, so that the voice synthesis model can effectively utilize all available data, and each component (such as the corresponding intermediate layer of each language and the corresponding decoder of each language) in the model can fully play its role, and the abstract feature representation independent of language can be obtained through the shared encoder in the voice synthesis model regardless of the data size (i.e. the size) of the language, and then the accurate pronunciation feature corresponding to the language can be obtained through the corresponding intermediate layer and decoder of the language, the high-precision voice information is obtained, and the voice synthesis precision of the language with small data size is improved.
[0049] In combination with the above embodiments, in an implementation, the embodiment of the present application further provides a voice synthesis method. Specifically, in the embodiment, the step S12 of “inputting the target phoneme information and target language information into a voice synthesis model to obtain target spectrum features corresponding to the target text” can specifically include steps S21 to S23:
[0050] Step S21: inputting the target phoneme information and the target language information into the shared encoder to obtain target general information.
[0051] In the embodiment, the target phoneme information and the target language information can be input into the shared encoder in advance to obtain the target general information output by the shared encoder. The target general information is general information corresponding to the target text, and the general information of the embodiment represents abstract feature representation independent of language.
[0052] Step S22: inputting the target general information and the target language information into the target language corresponding intermediate layer to obtain intermediate features.
[0053] In the embodiment, the target language corresponding intermediate layer is included in the multiple language corresponding intermediate layers, and after obtaining the target general information output by the shared encoder, the target general information and the target language information can be input into the target language corresponding intermediate layer to obtain the intermediate features of the target text corresponding to the target language corresponding intermediate layer.
[0054] The intermediate layer corresponding to each of the plurality of languages is used to perform specific processing on different languages, and the structure of the intermediate layer can be a fully connected layer, a convolutional layer, or another type of neural network layer. The intermediate layer corresponding to the target language can perform more detailed specific processing on the voice or dialect based on the target universal feature and the target language information, so as to enhance the characteristics of the target language. In an optional example, the specific processing can include, but is not limited to, filtering, compressing, and / or refining the target universal feature and the target language information.
[0055] Step S23: inputting at least the intermediate feature into the decoder corresponding to the target language to obtain a target spectral feature corresponding to the target text.
[0056] In the embodiment, the decoder corresponding to each of the plurality of languages at least includes a decoder corresponding to the target language. After obtaining the intermediate feature output by the intermediate layer corresponding to the target language, at least the intermediate feature is input into the decoder corresponding to the target language to obtain a target spectral feature corresponding to the target text output by the decoder corresponding to the target language.
[0057] The decoder corresponding to each of the plurality of languages is used to learn the unique pronunciation characteristics of each language, and generate a specific voice output suitable for the language based on the features generated from the shared encoder and the intermediate layer. The decoder corresponding to each of the plurality of languages can be based on LSTM+attention mechanism, which is not limited to a specific network architecture. The decoder corresponding to each of the plurality of languages in the speech synthesis model of the embodiment can be optimized for the pronunciation characteristics and tone system of the specific language.
[0058] In combination with the above embodiment, in an implementation, the embodiment of the present application further provides a speech synthesis method. Specifically, in the method, the above step S23 can specifically include steps S31 and S32:
[0059] Step S31: concatenating the intermediate feature, the target phoneme information, and the conditional information corresponding to the target phoneme information to obtain a first concatenated feature.
[0060] In the embodiment, in order to ensure the generation ability and synthesis effect of the speech synthesis model, in addition to using the intermediate feature output by the intermediate layer corresponding to each language as the input of the decoder corresponding to each language, the corresponding phoneme sequence and the conditional information of the phoneme sequence should also be added as additional inputs of the decoder corresponding to each language.
[0061] Specifically, for the decoder corresponding to the target language, the intermediate feature output by the intermediate layer corresponding to the target language, the target phoneme information corresponding to the target text, and the conditioning information corresponding to the target phoneme information are spliced to obtain the first spliced feature. The conditioning information of the embodiment represents the sound production condition information corresponding to the language, for example, the conditioning information can include: tone information, i.e., intonation information, etc.
[0062] Further, in an embodiment, the target phoneme information corresponding to the target text can be encoded to obtain encoded target phoneme information, and the conditioning information corresponding to the target phoneme information can be encoded to obtain encoded conditioning information. Then, the intermediate feature output by the intermediate layer corresponding to the target language, the encoded target phoneme information, and the encoded conditioning information are spliced to obtain the first spliced feature. For example, the target phoneme information is embedded to obtain embedded target phoneme information, and the conditioning information corresponding to the target phoneme information is embedded to obtain embedded conditioning information.
[0063] Step S32: inputting the first spliced feature into the decoder corresponding to the target language to obtain the target spectral feature corresponding to the target text.
[0064] In the embodiment, after obtaining the first spliced feature, the first spliced feature can be input into the decoder corresponding to the target language to obtain the target spectral feature corresponding to the target text output by the decoder corresponding to the target language.
[0065] In combination with the above embodiment, in another implementation, the embodiment of the present application further provides a speech synthesis method. Specifically, in the embodiment, the above step S21 can specifically include steps S41 to S44:
[0066] Step S41: encoding the target phoneme information to obtain encoded target phoneme information.
[0067] In the embodiment, the speech synthesis model further includes an encoding layer, and the target phoneme information corresponding to the target text can be input into the encoding layer for encoding to obtain encoded target phoneme information.
[0068] Step S42: encoding the target language information to obtain encoded target language information.
[0069] In the embodiment, the target language information corresponding to the target text can also be input into the encoding layer for encoding to obtain encoded target language information.
[0070] In an optional embodiment, the encoding layer can be an embed layer, the target phoneme information corresponding to the target text can be input into the embed layer for embed embedding processing to obtain an embedded representation of the embed target phoneme information, and the target language information corresponding to the target text can be input into the embed layer for embed embedding processing to obtain an embedded representation of the embed target language information.
[0071] Step S43: splicing the encoded target phoneme information and the encoded target language information to obtain a second spliced feature.
[0072] In this embodiment, the encoded target phoneme information and the encoded target language information need to be spliced to obtain a second spliced feature.
[0073] Step S44: inputting the second spliced feature into the shared encoder to obtain the target general information.
[0074] In this embodiment, after obtaining the second spliced feature, the second spliced feature can be input into the shared encoder to obtain the target general information corresponding to the target text output by the shared encoder. The target general information at least includes the duration of the pronunciation unit of the target phoneme information and the pause rhythm of the target phoneme information.
[0075] In this embodiment, the target phoneme sequence needs to combine some professional phonetics knowledge, transcribe the target text into target phoneme information, and then input the target phoneme information into the encoding layer (such as the embed layer). In terms of Min dialect, a syllable is composed of an initial, a final, and a tone, and the final is composed of a head, a body, and a tail. There are 15 initials, 82 finals, and 8 tones, which need to be distinguished from Chinese Mandarin, that is, the encoding cannot be mixed (Chinese Mandarin and Min dialect have the same tone class, such as yin ping, which is actually different and cannot use the same encoding; in addition, because the combination of the final body and the final tail of the Min dialect is very complex, the final body and the final tail need to be modeled separately, and the corresponding features are trained, and cannot be mixed with Chinese Mandarin). In addition, the target language information also needs to be encoded, and then spliced with the embedding result of the phoneme sequence (i.e., the encoded target phoneme information), and then the spliced result (such as the second spliced feature) is input into the shared encoder.
[0076] The reason why the shared encoder is designed in this embodiment is to enable the shared encoder to learn general features of multiple languages, such as the duration and rhythm (pause rhythm) of the basic pronunciation unit. It should be noted that the shared encoder can output different general features for different languages, such as different basic pronunciation unit durations and / or rhythms for different languages.
[0077] In combination with the above embodiments, in an embodiment, the plurality of languages in the embodiment include any one or more of the following: a standard Chinese, one or more Chinese dialects (such as Min dialect, Tianjin dialect, Northeastern dialect, Sichuan dialect, etc.), and one or more foreign languages (such as English, Thai, Malay, Spanish, etc.).
[0078] In combination with the above embodiments, in another embodiment, the embodiment also provides a speech synthesis method. In the embodiment, the speech synthesis model is obtained based on an initial speech synthesis model, and the initial speech synthesis model at least includes: an initial shared encoder, an initial intermediate layer corresponding to each of a plurality of languages, and an initial decoder corresponding to each of the plurality of languages. The training steps of the initial speech synthesis model at least include steps S51 to S57:
[0079] Step S51: determining a plurality of sample texts and a plurality of sample audio information of a plurality of sample languages corresponding to the plurality of sample texts.
[0080] The initial speech synthesis model of the embodiment is a multi-task learning model, and the speech synthesis tasks of the plurality of languages can be learned simultaneously during the model training. It can be understood that the initial speech synthesis model of the embodiment is a combined structure of a shared encoder and a language-specific decoder of a multi-task learning and conditional speech synthesis strategy. When training the initial speech synthesis model, a plurality of sample texts need to be determined, and the sample texts are texts of the speech to be synthesized for model training, and a plurality of sample audio information of a plurality of sample languages corresponding to the plurality of sample texts need to be determined.
[0081] In the embodiment, each sample text corresponds to a sample language, and each sample text corresponds to sample audio information of the sample language. The sample language corresponding to each sample text can be the same or different, and the sample language is a sample language corresponding to the sample text determined according to the speech synthesis requirement, such as Min dialect, standard Chinese, etc., and the speech synthesis model is not limited to this. The sample audio information is correct audio information of the sample language corresponding to the sample text, for example, a sample text A corresponds to a Min dialect, and the sample text A corresponds to a sample audio information of the Min dialect.
[0082] Step S52: processing the plurality of sample texts to obtain a plurality of sample phoneme information.
[0083] In this embodiment, the plurality of sample texts can be processed respectively to obtain a plurality of sample phoneme information. For example, the sample phoneme information can be obtained after text processing, text analysis and phoneme conversion of the sample text. The sample phoneme information of this embodiment is the phoneme information corresponding to the sample text. The specific manner of processing the sample text to obtain the sample phoneme information is not limited in this embodiment.
[0084] Step S53: inputting the plurality of sample phoneme information respectively and the sample language information of the sample language corresponding to each of the plurality of sample phoneme information into the initial shared encoder to obtain a plurality of sample general information.
[0085] In this embodiment, the sample language information of the sample language corresponding to each of the plurality of sample texts needs to be determined according to the voice synthesis requirement, that is, the sample language information of the sample language corresponding to each of the plurality of sample phoneme information needs to be determined. The sample language information is the related information of the sample language.
[0086] After obtaining the plurality of sample phoneme information and the sample language information of the sample language corresponding to each of the plurality of sample phoneme information, the plurality of sample phoneme information can be inputted respectively and the sample language information of the sample language corresponding to each of the plurality of sample phoneme information into the initial shared encoder to obtain a plurality of sample general information corresponding to each of the plurality of sample phoneme information. The sample general information is the general information corresponding to the sample text, and the general information of this embodiment represents the abstract feature representation irrelevant to the language.
[0087] For example, for the sample phoneme information A, the sample phoneme information A and the sample language information A corresponding to the sample phoneme information A are inputted into the initial shared encoder to obtain the sample general information A corresponding to the sample phoneme information A. For the sample phoneme information B, the sample phoneme information B and the sample language information B corresponding to the sample phoneme information B are inputted into the initial shared encoder to obtain the sample general information B corresponding to the sample phoneme information B.
[0088] In an optional embodiment, the plurality of sample phoneme information is encoded respectively to obtain a plurality of encoded sample phoneme information, the sample language information corresponding to each of the plurality of sample phoneme information is encoded respectively to obtain a plurality of encoded sample language information, then the plurality of encoded sample phoneme information is spliced respectively with the encoded sample language information corresponding to each of the plurality of encoded sample phoneme information to obtain a plurality of second sample splicing features, and the plurality of first sample splicing features are inputted respectively into the initial shared encoder to obtain a plurality of sample general information corresponding to each of the plurality of sample phoneme information.
[0089] Step S54: inputting the plurality of sample general information respectively and the sample language information corresponding to each of the plurality of sample general information into the initial intermediate layer corresponding to each of the plurality of sample language to obtain a plurality of sample intermediate features.
[0090] In this embodiment, after obtaining the plurality of sample general information, the plurality of sample general information and the sample language information corresponding to each of the plurality of sample general information are input into the initial intermediate layer corresponding to each of the plurality of sample language, to obtain the plurality of sample intermediate features output by the initial intermediate layer corresponding to each of the plurality of sample language. The initial intermediate layer corresponding to each of the plurality of language in this embodiment includes the initial intermediate layer corresponding to each of the plurality of sample language.
[0091] For example, for sample general information A, sample general information A and sample language information A corresponding to sample general information A are input into the initial intermediate layer corresponding to sample language A, to obtain sample intermediate feature A. For sample general information B, sample general information B and sample language information B corresponding to sample general information B are input into the initial intermediate layer corresponding to sample language B, to obtain sample intermediate feature B.
[0092] The initial intermediate layer corresponding to each of the plurality of language is used to perform specific processing on different languages. The structure of the initial intermediate layer can be a fully connected layer, a convolutional layer or other types of neural network layer. The initial intermediate layer can perform more detailed specific processing on the language or dialect based on the sample general feature and the sample language information, so as to enhance the characteristics of the corresponding sample language. In an optional example, the specific processing can include but is not limited to screening, compression and / or refining of the sample general feature and the sample language information.
[0093] Step S55: inputting the plurality of sample intermediate features, the sample phoneme information corresponding to each of the plurality of sample intermediate features and the sample conditioning information corresponding to the sample phoneme information into the initial decoder corresponding to each of the plurality of sample language, to obtain the sample spectrum feature corresponding to each of the plurality of sample text.
[0094] In this embodiment, after obtaining the plurality of sample intermediate features output by the initial intermediate layer corresponding to each of the plurality of sample language, the plurality of sample intermediate features, the sample phoneme information corresponding to each of the plurality of sample intermediate features and the sample conditioning information corresponding to the sample phoneme information are input into the initial decoder corresponding to each of the plurality of sample language, to obtain the plurality of sample spectrum features output by the initial decoder corresponding to each of the plurality of sample language, and the plurality of sample spectrum features are the sample spectrum features corresponding to the plurality of sample text. The sample spectrum feature is the spectrum feature corresponding to the sample text. The initial decoder corresponding to each of the plurality of language in this embodiment includes the initial decoder corresponding to each of the plurality of sample language.
[0095] For example, for the sample intermediate feature A, the sample intermediate feature A, the sample phoneme information A corresponding to the sample intermediate feature A, and the sample conditioning information A corresponding to the sample phoneme information A corresponding to the sample intermediate feature A can be input into the initial decoder corresponding to the sample language A to obtain the sample spectral feature A output by the initial decoder corresponding to the sample language A; for the sample intermediate feature B, the sample intermediate feature B, the sample phoneme information B corresponding to the sample intermediate feature B, and the sample conditioning information B corresponding to the sample phoneme information B corresponding to the sample intermediate feature B can be input into the initial decoder corresponding to the sample language B to obtain the sample spectral feature B output by the initial decoder corresponding to the sample language B.
[0096] The initial decoder corresponding to each of the plurality of languages is used to learn the pronunciation characteristics unique to each language, and generate a voice output specific to the language based on the features generated from the initial shared encoder and the initial intermediate layer. The initial decoder corresponding to each of the plurality of languages can be based on LSTM+attention mechanism, which is not limited to a specific network architecture. The initial decoder corresponding to each language in the initial speech synthesis model of the embodiment can be optimized for the pronunciation characteristics and tone system of the specific language.
[0097] In an optional embodiment, the plurality of sample intermediate features, the sample phoneme information corresponding to each of the plurality of sample intermediate features, and the sample conditioning information corresponding to the sample phoneme information corresponding to each of the plurality of sample intermediate features are spliced to obtain a plurality of first sample spliced features, and then the plurality of first sample spliced features are input into the initial decoder corresponding to each of the plurality of sample languages. The sample conditioning information in the embodiment is the conditioning information corresponding to the sample phoneme information.
[0098] Further, the sample phoneme information corresponding to each of the plurality of sample intermediate features can be encoded to obtain a plurality of encoded sample phoneme information, and the sample conditioning information corresponding to the sample phoneme information corresponding to each of the plurality of sample intermediate features can be encoded to obtain a plurality of encoded sample conditioning information. Then, the plurality of sample intermediate features are spliced with the corresponding encoded sample phoneme information and encoded sample conditioning information to obtain a plurality of first sample spliced features. For example, the sample phoneme information corresponding to each of the plurality of sample intermediate features is embedded to obtain a plurality of embedded sample phoneme information, and the sample conditioning information corresponding to the sample phoneme information corresponding to each of the plurality of sample intermediate features is embedded to obtain a plurality of embedded sample conditioning information.
[0099] Step S56: Based on the plurality of sample spectral features, a plurality of sample voice information corresponding to a plurality of sample languages is obtained.
[0100] In this embodiment, after obtaining the plurality of sample spectral features, the sample speech information of the plurality of sample languages can be obtained based on the plurality of sample spectral features. The specific manner in which the sample spectral features are processed to obtain the sample speech information of the sample language is not limited in this embodiment.
[0101] Step S57: Based on the sample speech information and the sample audio information of the plurality of sample languages respectively, the model parameters of the initial shared encoder, the initial intermediate layer corresponding to each of the plurality of sample languages, and the initial decoder corresponding to each of the plurality of sample languages are updated to obtain the speech synthesis model.
[0102] In this embodiment, the model parameters of the initial shared encoder, the initial intermediate layer corresponding to each of the plurality of sample languages, and the initial decoder corresponding to each of the plurality of sample languages can be updated based on the sample speech information and the sample audio information of the plurality of sample languages respectively to obtain the trained shared encoder, the intermediate layer corresponding to each of the plurality of languages, and the decoder corresponding to each of the plurality of languages, thereby obtaining the trained speech synthesis model. The speech synthesis model of this embodiment can improve the synthesis quality of dialects with small data volume and improve the overall model performance.
[0103] Specifically, the model parameters of the initial shared encoder, the initial intermediate layer corresponding to the sample language, and the initial decoder corresponding to the sample language can be updated based on the sample speech information and the sample audio information corresponding to each of the plurality of sample languages respectively until the model parameters of the initial shared encoder, the initial intermediate layer corresponding to each of the plurality of sample languages, and the initial decoder corresponding to each of the plurality of sample languages are updated to obtain the trained speech synthesis model.
[0104] In combination with the above embodiments, in an implementation, the embodiment of the present application further provides a speech synthesis method. Specifically, in the method, the above step S57 can specifically include step S61:
[0105] Step S61: Based on the sample speech information and the sample audio information of the plurality of sample languages respectively, the model parameters of the initial shared encoder and the initial decoder corresponding to each of the plurality of sample languages are updated to obtain the speech synthesis model.
[0106] In the embodiment, the intermediate layers corresponding to each language are trained separately, so that the characteristics of each language can be better represented and processed in the model. That is, the initial intermediate layers corresponding to each of the multiple languages in the initial speech synthesis model are the pre-trained intermediate layers corresponding to each of the multiple languages, and the initial intermediate layers corresponding to each of the multiple languages do not need to be updated.
[0107] In the embodiment, the model parameters of the initial shared encoder and the initial decoder corresponding to each of the multiple sample languages are updated based on the sample speech information and the sample audio information corresponding to each of the multiple sample languages, to obtain the trained shared encoder and the decoder corresponding to each of the multiple languages, thereby obtaining the trained speech synthesis model.
[0108] In an embodiment, as shown in Figure 3 , Figure 3 is a structural diagram of a speech synthesis model according to an embodiment of the present application. In Figure 3 , the subject part of the speech synthesis model is: a shared encoder, an intermediate layer of each language (such as the intermediate layer-1 and the intermediate layer-2 in Figure 3 ), and a decoder of each language (such as the decoder-1 and the decoder-2 in Figure 3 ), wherein the number of intermediate layers and decoders of the speech synthesis model is not specifically limited in the embodiment taking 2 intermediate layers and 2 decoders as an example, but the intermediate layers and the decoders are one-to-one corresponding.
[0109] wherein langID-1 is language information-1 and langID-2 is language information-2. First, the language information-1 and the phoneme sequence-1, and the language information-2 and the phoneme sequence-2 are input into the speech synthesis model, and the language information-1 and the phoneme sequence-1, and the language information-2 and the phoneme sequence-2 are embedded and processed (i.e., encoded) by the embed layers (i.e., embedding in Figure 3 ) corresponding to each of them, respectively, to obtain the embedded language information-1 and the embedded phoneme sequence-1, and the embedded language information-2 and the embedded phoneme sequence-2.
[0110] Secondly, the embedded language information-1 and the embedded phoneme sequence-1 are spliced and input into the shared encoder to obtain the general feature-1 corresponding to the phoneme sequence-1; the embedded language information-2 and the embedded phoneme sequence-2 are spliced and input into the shared encoder to obtain the general feature-2 corresponding to the phoneme sequence-2.
[0111] Then, the general feature-1 and the language information-1 (which can be the embedded language information-1) are input into the intermediate layer-1 corresponding to the language information-1 to obtain the intermediate feature-1 output by the intermediate layer-1; the general feature-2 and the language information-2 (which can be the embedded language information-2) are input into the intermediate layer-2 corresponding to the language information-2 to obtain the intermediate feature-2 output by the intermediate layer-2.
[0112] Finally, the intermediate feature-1, the phoneme sequence-1 (which can be the embedded phoneme sequence-1) and the tone sequence-1 (i.e. the conditioning information, which can be the embedded tone sequence-1) are spliced to input the decoder-1 corresponding to the language information-1 to obtain the mel_out1 output by the decoder-1, so as to finally obtain the speech information-1 and realize the speech synthesis of the language corresponding to the language information-1. The intermediate feature-2, the phoneme sequence-2 (which can be the embedded phoneme sequence-2) and the tone sequence-2 (i.e. the conditioning information, which can be the embedded tone sequence-2) are spliced to input the decoder-2 corresponding to the language information-2 to obtain the mel_out2 output by the decoder-2, so as to finally obtain the speech information-2 and realize the speech synthesis of the language corresponding to the language information-2.
[0113] In combination with the above embodiments, in an implementation, the embodiments of the present application also provide a speech synthesis method. Specifically, in the method, the initial speech synthesis model further comprises an initial language classifier; in addition to the steps S51 to S57 described above, the training step of the initial speech synthesis model can further comprise a step S71, and the step S57 described above can specifically comprise steps S72 to S74:
[0114] Step S71: inputting the plurality of sample general information into the initial language classifier to obtain a plurality of sample language prediction results.
[0115] In the present embodiment, a language classification task is added in the multi-task learning framework, that is, in addition to connecting the language-specific intermediate layer and the initial decoder for speech synthesis, an initial language classifier can also be added in parallel on the basis of the initial shared encoder. The initial language classifier can be a simple fully connected network, and its task is to predict the language of the input speech data. After obtaining the plurality of sample general information output by the initial shared encoder, the plurality of sample general information can be input into the initial language classifier to obtain a plurality of sample language prediction results output by the initial language classifier.
[0116] Step S72: calculating a first loss based on the sample speech information and the sample audio information corresponding to the plurality of sample languages respectively.
[0117] In this embodiment, the first loss needs to be calculated based on the sample speech information and the sample audio information corresponding to the plurality of sample languages respectively. For example, the audio loss can be calculated based on the sample speech information and the sample audio information corresponding to each sample language respectively, and then the audio losses corresponding to the plurality of sample languages are weighted and summed to obtain the first loss. The first loss can be an MSE loss.
[0118] Step S73: Calculate a second loss based on the plurality of sample language prediction results and the plurality of sample language information.
[0119] In this embodiment, the second loss needs to be calculated based on the sample language prediction results and the sample language information corresponding to the plurality of sample languages respectively. For example, the language loss can be calculated based on the sample language prediction results and the sample language information corresponding to each sample language respectively, and then the language losses corresponding to the plurality of sample languages are weighted and summed to obtain the second loss. The second loss can be a classification loss (such as a cross-entropy loss) to ensure that the model can correctly distinguish different languages, thereby strengthening the accuracy of language recognition and assisting speech synthesis.
[0120] Step S74: Update the model parameters of the initial shared encoder, the initial language classifier, the initial intermediate layers corresponding to the plurality of sample languages respectively, and the initial decoders corresponding to the plurality of sample languages respectively based on the first loss and the second loss, to obtain the speech synthesis model.
[0121] In this embodiment, the model parameters of the initial shared encoder, the initial language classifier, the initial intermediate layers corresponding to the plurality of sample languages respectively, and the initial decoders corresponding to the plurality of sample languages respectively can be updated based on the first loss and the second loss (such as the sum of the two, or the weighted sum of the two) to obtain the trained shared encoder, the language classifier, the intermediate layers corresponding to the plurality of languages respectively, and the decoders corresponding to the plurality of languages respectively, thereby obtaining the trained speech synthesis model.
[0122] In combination with the above embodiments, in an implementation, the embodiments of the present application also provide a speech synthesis method. Specifically, in addition to the above steps, the method can further include steps S81 to S84:
[0123] Step S81: Process a plurality of outbound texts to obtain a plurality of outbound phoneme information.
[0124] The application environment of this embodiment is an intelligent outbound call system in the Min dialect area, and the intelligent outbound call system needs to provide a multi-language outbound call voice library including Mandarin and Min dialect. Based on this, the embodiment can process a plurality of outbound call texts to obtain a plurality of outbound call phoneme information. The outbound call text is a text that needs to be called, including a voice broadcast text and / or a navigation text, etc., such as "Hello, tourists", or "Please pay attention to safety" and the like. The outbound call phoneme information is the phoneme information corresponding to the outbound call text. This step is the same as or similar to the above step S11, and will not be described herein.
[0125] Step S82: input each of the plurality of outbound call phoneme information and the outbound language information of the plurality of outbound languages corresponding to each of the outbound call phoneme information into the speech synthesis model to obtain the outbound spectral features of the plurality of outbound languages corresponding to each of the outbound call texts.
[0126] In this embodiment, it can be determined that each outbound call is not corresponding to the outbound language, and the outbound language is the language corresponding to the outbound text specified in the outbound speech synthesis requirement. Each outbound text can correspond to a plurality of outbound languages, such as Min dialect, Mandarin, English, Korean, etc., which are not limited.
[0127] In this embodiment, the plurality of outbound call phoneme information and the outbound language information of the plurality of outbound languages corresponding to each of the outbound call phoneme information can be input into the pre-trained speech synthesis model to obtain the outbound spectral features of the plurality of outbound languages corresponding to each of the outbound call texts output by the speech synthesis model. The outbound spectral feature is the spectral feature corresponding to the outbound text, and this step is the same as or similar to the above step S12, and will not be described herein.
[0128] Step S83: based on the outbound spectral features of the plurality of outbound languages corresponding to each of the outbound call texts, obtain the speech information of the plurality of outbound languages corresponding to each of the outbound call texts.
[0129] In this embodiment, after obtaining the outbound spectral features of the plurality of outbound languages corresponding to each of the outbound call texts, the speech information of the plurality of outbound languages corresponding to each of the outbound call texts can be obtained based on the outbound spectral features of the plurality of outbound languages corresponding to each of the outbound call texts. This step is the same as or similar to the above step S13, and will not be described herein.
[0130] Step S84: based on the speech information of the plurality of outbound languages corresponding to the plurality of outbound texts respectively, generate a multi-language outbound call voice library.
[0131] In this embodiment, after obtaining the voice information of the plurality of outbound call languages corresponding to the plurality of outbound call texts respectively, a multi-language outbound call audio library can be generated based on the voice information of the plurality of outbound call languages corresponding to the plurality of outbound call texts respectively. The plurality of outbound call languages in this embodiment at least include a plurality of dialects including Mandarin and Min dialect. In addition, the multi-language outbound call audio library generated in this embodiment can be applied to an intelligent outbound call system, which can train a high-quality Min dialect audio library with fewer Min dialect samples (i.e. in the scene of small data amount Min dialect voice synthesis), while maintaining the synthesis quality of Mandarin.
[0132] It should be noted that, for the method embodiments, in order to simply describe, they are all described as a series of action combinations, but those skilled in the art should know that the embodiments of the present application are not limited to the action sequence described, because according to the embodiments of the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions involved are not necessarily necessary for the embodiments of the present application.
[0133] Based on the same inventive concept, an embodiment of the present application provides a speech synthesis device 400. Referring to Figure 4 , Figure 4 is a structural block diagram of the speech synthesis device provided by an embodiment of the present application. As Figure 4 shown, the speech synthesis device 400 includes:
[0134] The target phoneme acquisition module 401 is configured to process a target text to obtain target phoneme information, wherein the target text includes one or more to-be-processed texts.
[0135] The first model processing module 402 is configured to input the target phoneme information and target language information into a speech synthesis model to obtain target spectral features corresponding to the target text, wherein the target language information is one or more language information corresponding to the target text.
[0136] The target speech synthesis module 403 is configured to obtain voice information of a target language based on the target spectral features.
[0137] The speech synthesis model at least includes a shared encoder, a plurality of language-specific intermediate layers, and a plurality of language-specific decoders; the shared encoder is configured to process text conversion tasks of a plurality of languages to generate abstract feature representations independent of the languages; the plurality of language-specific intermediate layers are respectively configured to enhance the characteristics of the plurality of languages; and the plurality of language-specific decoders are respectively configured to learn pronunciation features corresponding to the plurality of languages.
[0138] Optionally, the model processing module 402 includes:
[0139] a shared encoder processing module, configured to input the target phoneme information and the target language information into the shared encoder to obtain target general information;
[0140] an intermediate layer processing module, configured to input the target general information and the target language information into the intermediate layer corresponding to the target language to obtain intermediate features;
[0141] a decoder processing module, configured to input at least the intermediate features into the decoder corresponding to the target language to obtain target spectral features corresponding to the target text.
[0142] Optionally, the decoder processing module comprises:
[0143] a first splicing module, configured to splice the intermediate features, the target phoneme information, and the conditional information corresponding to the target phoneme information to obtain first spliced features;
[0144] a first processing module, configured to input the first spliced features into the decoder corresponding to the target language to obtain the target spectral features corresponding to the target text.
[0145] Optionally, the shared encoder processing module comprises:
[0146] a first encoding module, configured to encode the target phoneme information to obtain encoded target phoneme information;
[0147] a second encoding module, configured to encode the target language information to obtain encoded target language information;
[0148] a second splicing module, configured to splice the encoded target phoneme information and the encoded target language information to obtain second spliced features;
[0149] a second processing module, configured to input the second spliced features into the shared encoder to obtain the target general information, the target general information at least comprising a duration of a pronunciation unit of the target phoneme information and a pause rhythm of the target phoneme information.
[0150] Optionally, the speech synthesis model is obtained by training an initial speech synthesis model, the initial speech synthesis model at least comprising an initial shared encoder, initial intermediate layers corresponding to a plurality of languages respectively, and initial decoders corresponding to the plurality of languages respectively; the speech synthesis apparatus 400 further comprises a model training module, the model training module being configured to train the initial speech synthesis model; the model training module at least comprises:
[0151] determine a plurality of sample audio information of a plurality of sample texts and a plurality of sample languages respectively corresponding to the plurality of sample texts;
[0152] obtain a plurality of sample phoneme information by processing the plurality of sample texts;
[0153] input the plurality of sample phoneme information and sample language information of a sample language respectively corresponding to the plurality of sample phoneme information into the initial shared encoder to obtain a plurality of sample general information;
[0154] input the plurality of sample general information and sample language information respectively corresponding to the plurality of sample general information into an initial intermediate layer respectively corresponding to a plurality of sample languages to obtain a plurality of sample intermediate features;
[0155] input the plurality of sample intermediate features, sample phoneme information respectively corresponding to the plurality of sample intermediate features, and sample conditioning information corresponding to the sample phoneme information into an initial decoder respectively corresponding to a plurality of sample languages to obtain sample spectrum features respectively corresponding to the plurality of sample texts;
[0156] obtain a plurality of sample voice information corresponding to a plurality of sample languages based on a plurality of sample spectrum features;
[0157] update model parameters of the initial shared encoder, the initial intermediate layer respectively corresponding to the plurality of sample languages, and the initial decoder respectively corresponding to the plurality of sample languages based on at least sample voice information and sample audio information respectively corresponding to a plurality of sample languages to obtain the voice synthesis model.
[0158] Optionally, the initial intermediate layer respectively corresponding to the plurality of languages is a plurality of intermediate layers respectively corresponding to the plurality of languages pre-trained.
[0159] The parameter updating module comprises:
[0160] update model parameters of the initial shared encoder and the initial decoder respectively corresponding to the plurality of sample languages based on sample voice information and sample audio information respectively corresponding to a plurality of sample languages to obtain the voice synthesis model.
[0161] Optionally, the initial voice synthesis model further comprises an initial language classifier; and the model training module further comprises:
[0162] The language prediction module is used to input the general information of the multiple samples into the initial language classifier to obtain the language prediction results of the multiple samples.
[0163] The parameter update module includes:
[0164] The first loss calculation module is used to calculate the first loss based on the sample speech information and sample audio information corresponding to multiple sample languages.
[0165] The second loss calculation module is used to calculate the second loss based on the prediction results of multiple sample languages and information of multiple sample languages;
[0166] The second parameter update submodule is used to update the model parameters of the initial shared encoder, the initial language classifier, the initial intermediate layers corresponding to each of the multiple sample languages, and the initial decoders corresponding to each of the multiple sample languages based on the first loss and the second loss, so as to obtain the speech synthesis model.
[0167] Optionally, the speech synthesis device 400 further includes:
[0168] The outbound call phoneme determination module is used to process multiple outbound call texts to obtain multiple outbound call phoneme information;
[0169] The second model processing module is used to input each outgoing phoneme information and the outgoing language information of each outgoing phoneme information corresponding to the multiple outgoing languages into the speech synthesis model to obtain the outgoing spectral features of the multiple outgoing languages corresponding to each outgoing text.
[0170] The outbound call speech synthesis module is used to obtain the speech information of multiple outbound languages corresponding to each outbound call text based on the outbound call spectral features of multiple outbound languages corresponding to each outbound call text;
[0171] The outbound call voice library generation module is used to generate a multilingual outbound call voice library based on the voice information of multiple outbound call texts corresponding to multiple outbound call languages.
[0172] The multiple outbound calling languages include at least several dialects, including Mandarin and Min dialect, and the multilingual outbound calling voice library is applied to the intelligent outbound calling system.
[0173] Based on the same inventive concept, another embodiment of the present invention provides an electronic device 500, such as... Figure 5 As shown. Figure 5 This is a schematic diagram of an electronic device according to an embodiment of the present invention. The electronic device includes a memory 502, a processor 501, and a computer program stored in the memory and executable on the processor. When executed by the processor, the program implements the steps of the speech synthesis method described in any of the above embodiments of the present invention.
[0174] For apparatus embodiments, since they are substantially similar to the method embodiments, the description is relatively simple, and the relevant parts are referred to the part of the method embodiments.
[0175] Each of the embodiments in the specification is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the embodiments can be referred to each other.
[0176] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, an apparatus, or a computer program product. Therefore, the embodiments of the present application can be in the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present application can be in the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0177] The embodiments of the present application are described with reference to flowcharts and / or block diagrams of the method, terminal device (system), and computer program product according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to produce a machine, so that the instructions executed by the computer or other programmable data processing terminal device produce a device implemented in the flowcharts and / or block diagrams. Figure 1 one or more flows and / or blocks Figure 1 an apparatus that performs the functions specified in the flow or flows and / or block or blocks.
[0178] These computer program instructions can also be stored in a computer-readable memory that can direct the computer or other programmable data processing terminal device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including instruction apparatus, which implements the flowcharts and / or block diagrams. Figure 1 one or more flows and / or blocks Figure 1 an apparatus that performs the functions specified in the flow or flows and / or block or blocks.
[0179] These computer program instructions can also be loaded into a computer or other programmable data processing terminal device, so that a series of operation steps are performed on the computer or other programmable terminal device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable terminal device provide a process for implementing the flowcharts and / or block diagrams. Figure 1 one or more flows and / or blocks Figure 1steps of a function specified in one or more blocks.
[0180] While the preferred embodiments of the application have been described above, it should be understood that many modifications and adaptations will occur to those skilled in the art upon the reading and understanding of the foregoing description. For example, the embodiments described above are preferred embodiments of the application, and other embodiments of the application can be made without departing from the scope of the application. Therefore, the scope of the application is intended to cover all reasonable adaptations and modifications to the preferred embodiments. Accordingly, the appended claims are intended to encompass within their scope all adaptations and modifications of the preferred embodiments as
[0181] Finally, it should be noted that the terms "first", "second", and the like, herein do not denote any order, quantity, combination, or importance, but rather are used to distinguish one element from another, and do not imply singular or plural. Moreover, the term "include", "have", or "contain" or any other variant thereof, are intended to encompass non-exclusive inclusions, such that processes, methods, articles, or apparatuses that comprise a set of elements not expressly listed are still within the scope of the claims. The term "comprise", "comprises", or "comprising", or any other variant thereof, are intended to encompass the presence of stated elements, but not to preclude the presence of additional elements, or the presence of additional elements not explicitly stated. The term "consist of", "consists of", or "consisting of", or any other variant thereof, are intended to encompass a process, method, article, or apparatus that consists of the stated elements, but not to preclude the presence of additional elements not explicitly stated, or the presence of additional elements inherent in such process, method, article, or apparatus.
[0182] The above provides a voice synthesis method, device and electronic equipment, and the principles and implementation manners of the present application are described by using specific examples. The above description of the embodiments is only used to help understand the method and core idea of the present application. For those skilled in the art, the specific implementation manners and application scope can be changed according to the idea of the present application. The above description of the present application should not be understood as a limitation.
Claims
1. A speech synthesis method characterized by, The method comprises: processing a target text to obtain target phoneme information, the target text comprising one or more texts to be processed; inputting the target phoneme information and target language information into a speech synthesis model to obtain target spectral features corresponding to the target text, the target language information being one or more language information corresponding to the target text; obtaining speech information of the target language based on the target spectral features; wherein the speech synthesis model comprises at least a shared encoder, a plurality of language-specific intermediate layers, and a plurality of language-specific decoders; the shared encoder is used for processing text conversion tasks of multiple languages to generate abstract feature representations independent of language; the plurality of language-specific intermediate layers are used to enhance the characteristics of the plurality of languages respectively; and the plurality of language-specific decoders are used to learn the pronunciation features corresponding to the plurality of languages respectively; inputting the target phoneme information and target language information into a speech synthesis model to obtain target spectral features corresponding to the target text, comprising: inputting the target phoneme information and the target language information into the shared encoder to obtain target general information; inputting the target general information and the target language information into the intermediate layer corresponding to the target language to obtain intermediate features; inputting at least the intermediate features into the decoder corresponding to the target language to obtain target spectral features corresponding to the target text; inputting at least the intermediate features into the decoder corresponding to the target language to obtain target spectral features corresponding to the target text, comprising: concatenating the intermediate features, the target phoneme information, and the conditional information corresponding to the target phoneme information to obtain first concatenated features; inputting the first concatenated features into the decoder corresponding to the target language to obtain target spectral features corresponding to the target text.
2. The speech synthesis method of claim 1, wherein, inputting the target phoneme information and the target language information into the shared encoder to obtain target general information, comprising: encoding the target phoneme information to obtain encoded target phoneme information; encoding the target language information to obtain encoded target language information; concatenating the encoded target phoneme information and the encoded target language information to obtain second concatenated features; inputting the second concatenated features into the shared encoder to obtain the target general information, the target general information comprising at least the duration of the pronunciation unit of the target phoneme information and the pause rhythm of the target phoneme information.
3. The speech synthesis method according to claim 1 or 2, characterized by, The speech synthesis model is trained based on an initial speech synthesis model, the initial speech synthesis model comprising at least an initial shared encoder, a plurality of language-specific initial intermediate layers, and a plurality of language-specific initial decoders; the training steps of the initial speech synthesis model comprise at least: determining a plurality of sample texts and a plurality of sample audio information corresponding to a plurality of sample languages of the plurality of sample texts; processing the plurality of sample texts to obtain a plurality of sample phoneme information; inputting the plurality of sample phoneme information and sample language information corresponding to the sample language of the plurality of sample phoneme information into the initial shared encoder respectively to obtain a plurality of sample general information; inputting the plurality of sample general information and sample language information corresponding to the sample language of the plurality of sample general information into the initial intermediate layer corresponding to the sample language respectively to obtain a plurality of sample intermediate features; inputting the plurality of sample intermediate features, sample phoneme information corresponding to the plurality of sample intermediate features, and sample conditioning information corresponding to the sample phoneme information into the initial decoder corresponding to the sample language respectively to obtain sample spectrum features corresponding to the plurality of sample texts; obtaining sample voice information corresponding to the plurality of sample languages based on the plurality of sample spectrum features; updating model parameters of the initial shared encoder, the initial intermediate layer corresponding to the plurality of sample languages, and the initial decoder corresponding to the plurality of sample languages based on at least sample voice information and sample audio information corresponding to the plurality of sample languages respectively to obtain the voice synthesis model.
4. The speech synthesis method according to claim 3, characterized by, The initial intermediate layer corresponding to the plurality of languages is a pre-trained intermediate layer corresponding to the plurality of languages; updating model parameters of the initial shared encoder, the initial intermediate layer corresponding to the plurality of sample languages, and the initial decoder corresponding to the plurality of sample languages based on at least sample voice information and sample audio information corresponding to the plurality of sample languages respectively to obtain the voice synthesis model, including: updating model parameters of the initial shared encoder and the initial decoder corresponding to the plurality of sample languages based on sample voice information and sample audio information corresponding to the plurality of sample languages respectively to obtain the voice synthesis model.
5. The speech synthesis method of claim 3, wherein, The initial voice synthesis model further includes an initial language classifier; the training step of the initial voice synthesis model further includes: inputting the plurality of sample general information into the initial language classifier to obtain a plurality of sample language prediction results; updating model parameters of the initial shared encoder, the initial intermediate layer corresponding to the plurality of sample languages, and the initial decoder corresponding to the plurality of sample languages based on at least sample voice information and sample audio information corresponding to the plurality of sample languages respectively to obtain the voice synthesis model, including: calculating a first loss based on sample voice information and sample audio information corresponding to the plurality of sample languages; calculating a second loss based on a plurality of sample language prediction results and a plurality of sample language information; updating model parameters of the initial shared encoder, the initial language classifier, the initial intermediate layer corresponding to the plurality of sample languages, and the initial decoder corresponding to the plurality of sample languages based on the first loss and the second loss to obtain the voice synthesis model.
6. The speech synthesis method according to claim 1 or 2, characterized by, The method further includes: processing a plurality of outbound texts to obtain a plurality of outbound phoneme information; Input each outbound phoneme information in the plurality of outbound phoneme information, and, outbound language information of a plurality of outbound languages corresponding to each outbound phoneme information, into the speech synthesis model to obtain outbound spectral features of a plurality of outbound languages corresponding to each outbound text; Based on the outbound spectral features of a plurality of outbound languages corresponding to each outbound text, obtain the speech information of a plurality of outbound languages corresponding to each outbound text; Based on the speech information of a plurality of outbound languages corresponding to a plurality of outbound texts respectively, generate a multilingual outbound voice library; Wherein, the plurality of outbound languages at least includes: Mandarin and a plurality of dialects including Min dialect, the multilingual outbound voice library is applied to an intelligent outbound system.
7. A speech synthesis apparatus characterized by comprising: The device comprises: A target phoneme acquisition module is configured to process a target text to obtain target phoneme information, wherein the target text comprises one or more texts to be processed; A first model processing module is configured to input the target phoneme information and target language information into a speech synthesis model to obtain target spectral features corresponding to the target text, wherein the target language information is one or more language information corresponding to the target text; A target speech synthesis module is configured to obtain speech information of a target language based on the target spectral features; Wherein, the speech synthesis model at least includes: a shared encoder, a plurality of language corresponding intermediate layers, and a plurality of language corresponding decoders; the shared encoder is configured to process text conversion tasks of a plurality of languages to generate abstract feature representations independent of language; the plurality of language corresponding intermediate layers are respectively configured to enhance the characteristics of the plurality of languages; and the plurality of language corresponding decoders are respectively configured to learn the pronunciation features corresponding to the plurality of languages; The model processing module comprises: A shared encoder processing module is configured to input the target phoneme information and the target language information into the shared encoder to obtain target general information; An intermediate layer processing module is configured to input the target general information and the target language information into the intermediate layer corresponding to the target language to obtain intermediate features; A decoder processing module is configured to input at least the intermediate features into the decoder corresponding to the target language to obtain target spectral features corresponding to the target text; The decoder processing module comprises: A first splicing module is configured to splice the intermediate features, the target phoneme information, and the conditional information corresponding to the target phoneme information to obtain first splicing features; A first processing module is configured to input the first splicing features into the decoder corresponding to the target language to obtain target spectral features corresponding to the target text.
8. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The computer program is executed by the processor to implement the speech synthesis method of any one of claims 1 to 6.
Citation Information
Patent Citations
Voice processing method and related device
CN112397083A
Femule language transfer learning speech synthesis method based on implicit phoneme conversion
CN114822488A