Speech synthesis method and device and electronic equipment

By introducing shared encoder and language-specific intermediate layers and decoders into the speech synthesis model, the spectrum characteristics of each language are generated, and the problem of low speech synthesis accuracy in languages ​​with small data volume is solved, and high-precision speech synthesis is achieved.

CN119943025AActive Publication Date: 2025-05-06BEIJING SINOVOICE TECH CO LTD

Patent Information

Application Number
CN202411882752.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-19
Publication Date
2025-05-06
Estimated Expiration
2044-12-19

AI Technical Summary

Technical Problem

The prior art is difficult to effectively improve the accuracy of speech synthesis in languages ​​with small data volumes, such as Min dialect.

Method used

Spectral features of each language are generated by sharing encoder and language-specific intermediate layers and decoder structures, thereby improving speech synthesis accuracy.

Benefits of technology

High-precision speech synthesis for languages ​​with small data volumes is realized, improving the naturalness and intelligibility of speech synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943025A_ABST
    Figure CN119943025A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a speech synthesis method and device and electronic equipment, and relates to the technical field of speech synthesis. The method comprises the following steps: processing a target text to obtain target phoneme information; inputting the target phoneme information and the target language information into a speech synthesis model to obtain a target spectrum feature; obtaining voice information of a target language based on the target spectrum feature; the speech synthesis model at least comprises a text conversion task for processing a plurality of languages, a shared encoder for generating abstract feature representations irrelevant to the languages, intermediate layers corresponding to the plurality of languages for enhancing respective characteristics of the plurality of languages, and a text conversion task for processing the plurality of languages. And decoders respectively corresponding to the plurality of languages, which are used for respectively learning the pronunciation characteristics respectively corresponding to the plurality of languages. According to the speech synthesis method provided by the embodiment of the invention, the speech synthesis precision corresponding to a language with a small data volume can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of speech synthesis technology, and in particular to a speech synthesis method, device and electronic device. Background Art

[0002] TTS (text-to-speech) technology is a technology that converts text information into speech, enabling computers, smart devices or other applications to broadcast text in a form that humans can understand. TTS systems are widely used, including smart broadcasts, navigation systems, voice assistants, audio books, etc. It can be said that everyone's life is involved in a certain degree of speech synthesis technology, such as mobile assistant Siri, smart speaker Xiaodu, airport high-speed rail broadcasts and map navigation.

[0003] Taking Chinese as an example, the speech synthesis effect of Mandarin Chinese has generally reached the expected requirements, and its naturalness and comprehensibility are both high, but for small languages, such as dialects, there are still many problems. Taking the Min dialect as an example, the Min dialect is divided into 7 areas, and each area is further divided into multiple slices (for example, the southern Fujian area can be divided into Zhangquan slice, Datian slice and Chaoshan slice), so, for a specific dialect, its data is relatively scarce. It is also because its language phenomenon is very complicated (such as the Min dialect cannot communicate with the north and south, and there are also differences between the east and the west), so language characteristics (such as tone, intonation, etc.) are difficult to capture. Therefore, how to improve the speech synthesis accuracy for languages ​​with small data volume is a problem to be solved in the present invention. Summary of the invention

[0004] The embodiments of the present invention provide a speech synthesis method, device and electronic device, which generate spectral features corresponding to each language by sharing an encoder and a language-specific intermediate layer and decoder structure, and then obtain speech information corresponding to each language, thereby improving the speech synthesis accuracy corresponding to languages ​​with small data volumes.

[0005] A first aspect of an embodiment of the present invention provides a speech synthesis method, the method comprising:

[0006] Processing a target text to obtain target phoneme information, wherein the target text includes: one or more texts to be processed;

[0007] Inputting the target phoneme information and the target language information into a speech synthesis model to obtain target frequency spectrum features corresponding to the target text, wherein the target language information is one or more language information corresponding to the target text;

[0008] Obtaining speech information of a target language based on the target frequency spectrum features;

[0009] Among them, the speech synthesis model at least includes: a shared encoder, an intermediate layer corresponding to each of the multiple languages, and a decoder corresponding to each of the multiple languages; the shared encoder is used to process text conversion tasks in multiple languages ​​and generate language-independent abstract feature representations; the intermediate layers corresponding to each of the multiple languages ​​are used to enhance the characteristics of each of the multiple languages; the decoders corresponding to each of the multiple languages ​​are used to learn the pronunciation features corresponding to each of the multiple languages.

[0010] A second aspect of an embodiment of the present invention provides a speech synthesis device, the device comprising:

[0011] A target phoneme acquisition module is used to process a target text to obtain target phoneme information, wherein the target text includes: one or more texts to be processed;

[0012] A first model processing module is used to input the target phoneme information and the target language information into a speech synthesis model to obtain a target spectrum feature corresponding to the target text, wherein the target language information is one or more language information corresponding to the target text;

[0013] A target speech synthesis module, used to obtain speech information of a target language based on the target spectral features;

[0014] Among them, the speech synthesis model at least includes: a shared encoder, an intermediate layer corresponding to each of the multiple languages, and a decoder corresponding to each of the multiple languages; the shared encoder is used to process text conversion tasks in multiple languages ​​and generate language-independent abstract feature representations; the intermediate layers corresponding to each of the multiple languages ​​are used to enhance the characteristics of each of the multiple languages; the decoders corresponding to each of the multiple languages ​​are used to learn the pronunciation features corresponding to each of the multiple languages.

[0015] A third aspect of an embodiment of the present invention provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and running on the processor, wherein when the computer program is executed by the processor, the speech synthesis method of the first aspect of the embodiment of the present invention is implemented.

[0016] In the speech synthesis method provided by the embodiment of the present invention, the target phoneme information and target language information corresponding to the target text are input into a pre-trained speech synthesis model, and the target spectrum information corresponding to the target text output by the speech synthesis model is obtained, thereby obtaining the speech information of the target language. The speech synthesis model at least includes: a shared encoder for generating language-independent abstract feature representations; a plurality of corresponding intermediate layers for enhancing the characteristics of the plurality of languages; and a plurality of corresponding decoders for learning the pronunciation features of the plurality of languages.

[0017] In the speech synthesis model of the present embodiment, a shared encoder is used to fully learn language-independent abstract feature representations based on relevant information of multiple languages, so that each language can obtain its own corresponding language-independent abstract feature representation based on the shared encoder, and then further learn the corresponding features and pronunciation characteristics of each language through the corresponding intermediate layer and decoder based on the relevant information corresponding to each language, so that regardless of the amount of data of the language (i.e., the scale), the shared encoder in the speech synthesis model can obtain the language-independent abstract feature representation, and then the corresponding intermediate layer and decoder of the language can obtain the accurate pronunciation characteristics of the language, and then obtain high-precision speech information, thereby improving the speech synthesis accuracy corresponding to languages ​​with small data volumes. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.

[0019] Figure 1 It is a structural schematic diagram of a tacotron2 model shown in the related art;

[0020] Figure 2 is a flow chart of a speech synthesis method shown in one embodiment of the present invention;

[0021] Figure 3 is a structural schematic diagram of a speech synthesis model shown in one embodiment of the present invention;

[0022] Figure 4 is a structural block diagram of a speech synthesis device provided by an embodiment of the present invention;

[0023] Figure 5 It is a schematic diagram of an electronic device shown in an embodiment of the present invention. DETAILED DESCRIPTION

[0024] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0025] Generally speaking, the working principle of the TTS system is divided into the following steps: text processing (receiving input text and preprocessing it, including text standardization, such as conversion of dates, amounts, numbers, symbols, punctuation processing and word segmentation), text analysis (such as part-of-speech tagging, etc.), phoneme conversion (converting text into phoneme sequences) and sound synthesis (converting phoneme sequences into waveforms).

[0026] Commonly used TTS neural network architectures include tacotron, fast speech, and vits. Taking tacotron2 as an example, Figure 1 As shown, Figure 1 FIG. 1 is a schematic diagram of the structure of a tacotron2 model shown in the related art. Figure 1 In the tacotron2 model, the main components are the encoder, the decoder based on the attention mechanism, and the post-processing network. The tacotron2 model can realize the process from character (text) input to spectrum output (and then converted into waveform). Among them, the main function of the encoder is to convert the character input into a set of vectors, that is, the encoder's goal is to convert the input text into a set of highly robust sequence representations. The decoder based on the attention mechanism realizes the frame-by-frame prediction of the encoded input sequence to the spectrum by learning the alignment process of text and speech signals. The spectrum predicted by the decoder passes through a series of linear network layers and is finally converted into a waveform.

[0027] However, if the above tacotron2 model can have a large amount of data in a single language (such as Mandarin Chinese, more than 10 hours per person), then this tacotron2 model can fully complete the speech synthesis task. However, for dialects and small languages ​​with small data volumes, such as the synthesis of the Min dialect, the amount of data that can be obtained is very limited (the data required for speech synthesis must be annotated and have relatively clean sound quality). Therefore, it is not realistic to use the above model structure to complete this task.

[0028] Another solution is the traditional multilingual speech synthesis solution, which is to feed all language data into the model and train it indiscriminately. However, for small languages ​​and dialects with scarce data, its synthesis quality is usually poor. In addition, such systems are often unable to effectively distinguish the subtle differences between different languages, which means that the output of speech synthesis cannot accurately reflect the speech characteristics of a specific language.

[0029] Another common solution is to assume that there is a certain correlation between the two languages. If the amount of data for one language is large, a base model is trained using the language with large data, and then the second language with smaller data is trained using transfer learning. For example, a base model is first trained in Chinese, and then transferred to the Min dialect. However, this method has high requirements for the similarity between the two languages. For example, if the base is Mandarin Chinese, it is feasible to transfer to the Northeast dialect, but it is not realistic to transfer to Shanghainese.

[0030] Therefore, in order to at least partially solve one or more of the above-mentioned problems and other potential problems, an embodiment of the present invention proposes a speech synthesis method, in which a pre-trained speech synthesis model processes the target language and target phoneme information to obtain target spectral features, and then obtains the speech information of the target language based on the target spectral features. The speech synthesis model at least includes: a shared encoder, intermediate layers corresponding to multiple languages, and decoders corresponding to multiple languages; the shared encoder fully learns language-independent abstract feature representations based on relevant information of multiple languages, so that each language can obtain its own language-independent abstract feature representation based on the shared encoder, and then further learn the corresponding features and pronunciation characteristics of each language through the intermediate layers and decoders corresponding to each language based on the relevant information corresponding to each language, so that the speech synthesis model can not only effectively utilize all available data, but also enable each component inside the model (such as the intermediate layers corresponding to multiple languages ​​and the decoders corresponding to multiple languages) to fully play their role, regardless of the data volume (i.e., scale) of the language, the shared encoder in the speech synthesis model can obtain language-independent abstract feature representations, and then the corresponding intermediate layers and decoders of the languages ​​can obtain accurate pronunciation features of the languages, and obtain high-precision speech information, thereby improving the speech synthesis accuracy of languages ​​with small data volumes.

[0031] Hereinafter, specific examples of the present solution will be described in more detail with reference to the accompanying drawings.

[0032] refer to Figure 2 , Figure 2 FIG. 1 is a flow chart of a speech synthesis method according to an embodiment of the present invention. Figure 2 As shown, the speech synthesis method of this embodiment may include the following steps:

[0033] Step S11: Process the target text to obtain target phoneme information, wherein the target text includes: one or more texts to be processed.

[0034] In this embodiment, the target text may be processed to obtain target phoneme information. For example, the target text may be processed, analyzed, and converted to phonemes to obtain the target phoneme information. The target phoneme information in this embodiment is the phoneme information corresponding to the target text, and the target text includes: one or more texts to be processed, and the text to be processed is the text to be subjected to speech synthesis. That is, the target text in this embodiment may be one target text or multiple target texts.

[0035] This embodiment does not impose any restrictions on the specific method of processing the target text to obtain the target phoneme information, and all methods that can convert text into phoneme information are within the protection scope of this embodiment.

[0036] Step S12: inputting the target phoneme information and the target language information into a speech synthesis model to obtain target frequency spectrum features corresponding to the target text, wherein the target language information is one or more language information corresponding to the target text.

[0037] In this embodiment, the target language information corresponding to the target text needs to be determined according to the speech synthesis requirement. The target language information is the relevant information of the target language. The target language is the language corresponding to the speech to be synthesized in the speech synthesis requirement of the target text. The target language information of this embodiment is one or more language information corresponding to the target text. That is, one target text in this embodiment can correspond to one or more target languages.

[0038] In this embodiment, the target phoneme information and the target language information may be input into a pre-trained speech synthesis model to obtain a target spectrum feature corresponding to the target text output by the speech synthesis model, where the target spectrum feature is a spectrum feature corresponding to the target text.

[0039] The pre-trained speech synthesis model of this embodiment includes at least: a shared encoder, an intermediate layer corresponding to each of the multiple languages, and a decoder corresponding to each of the multiple languages. The shared encoder is used to process text conversion tasks in multiple languages ​​and generate language-independent abstract feature representations; the intermediate layers corresponding to each of the multiple languages ​​are used to enhance the characteristics of each of the multiple languages; and the decoders corresponding to each of the multiple languages ​​are used to learn the pronunciation features corresponding to each of the multiple languages.

[0040] In this embodiment, the pre-trained speech synthesis model can be used to generate the spectral features (acoustic features) required for synthesizing speech information of multiple languages, and the scale of languages ​​or the amount of language data is not limited, and the multiple languages ​​can include any language. Therefore, when the target phoneme information corresponding to one or more target texts and the target language information corresponding to the one or more target texts are input into the speech synthesis model, corresponding processing can be performed based on the shared encoder, the intermediate layer corresponding to the target language, and the decoder in the speech synthesis model, so as to output the target spectral features corresponding to the one or more target texts.

[0041] Step S13: obtaining speech information of the target language based on the target spectrum features.

[0042] In this embodiment, after obtaining the target spectrum features, the speech information of the target language can be obtained based on the target spectrum features. Among them, this embodiment does not impose any restrictions on the specific method of processing the target spectrum features to obtain the speech information of the target language, and all methods that can convert spectrum features into speech information are within the protection scope of this embodiment.

[0043] In addition, since the target text in this embodiment can be one or more, and the target language information is one or more language information corresponding to the target text, based on this, the following describes the possible situations in this embodiment, but is not limited to the following situations:

[0044] When there is one target text and the target language information corresponding to the target text is also one, that is, there is one target text and one target language, for example, the target language corresponding to the target text A is Minnan dialect, at this time, the target phoneme information of the target text A and the language information of Minnan dialect are input into the speech synthesis model to obtain the spectral feature A corresponding to the target text A, and then the speech information A of Minnan dialect is obtained based on the spectral feature A.

[0045] When there is one target text and the target language information corresponding to the target text is multiple, that is, when there is one target text and the target languages ​​are multiple, for example, the target languages ​​corresponding to the target text A are Northeastern dialect and Mandarin, at this time, the target phoneme information of the target text A and the language information of Northeastern dialect can be simultaneously input into the speech synthesis model, and the target phoneme information of the target text A and the language information of Mandarin can be input into the speech synthesis model to obtain the spectral feature A1 and spectral feature A2 corresponding to the target text A output by the speech synthesis model, so as to obtain the speech information A1 of Northeastern dialect based on the spectral feature A1, and obtain the speech information A2 of Mandarin based on the spectral feature A2.

[0046] When there are multiple target texts, and each target text corresponds to one target language information, that is, there are multiple target texts, and each target text corresponds to one target language (including the case where multiple target texts correspond to the same target language), for example, the target language corresponding to target text A is Minnan dialect, and the target language corresponding to target text B is Sichuan dialect, at this time, the target phoneme information of target text A and the language information of Minnan dialect can be input into the speech synthesis model at the same time, and the target phoneme information of target text B and the language information of Sichuan dialect can be input into the speech synthesis model to obtain the spectral feature A corresponding to the target text A and the spectral feature B corresponding to the target text B output by the speech synthesis model, so as to obtain the speech information A of Minnan dialect based on the spectral feature A, and the speech information B of Sichuan dialect based on the spectral feature B.

[0047] When there are multiple target texts and each target text corresponds to multiple target language information, that is, when there are multiple target texts and each target text corresponds to multiple target languages, for example, the target languages ​​corresponding to target text A are Tianjin dialect and Mandarin, and the target languages ​​corresponding to target text B are Sichuan dialect and Northeastern dialect, then the target phoneme information of target text A and the language information of Tianjin dialect can be input into the speech synthesis model at the same time, the target phoneme information of target text A and the language information of Mandarin can be input into the speech synthesis model, the target phoneme information of target text B and the language information of Sichuan dialect can be input into the speech synthesis model, and the target phoneme information of target text B and the language information of Northeastern dialect can be input into the speech synthesis model to obtain the spectral feature A1 and spectral feature A2 corresponding to the target text A output by the speech synthesis model, and the spectral feature B1 and spectral feature B2 corresponding to the target text B, so as to obtain the speech information A of Tianjin dialect based on the spectral feature A1, the speech information A2 of Mandarin can be obtained based on the spectral feature A2, the speech information B1 of Sichuan dialect can be obtained based on the spectral feature B1, and the speech information B2 of Northeastern dialect can be obtained based on the spectral feature B2.

[0048] In this embodiment, a shared encoder is used to fully learn language-independent abstract feature representations based on relevant information of multiple languages, so that each language can obtain its own corresponding language-independent abstract feature representation based on the shared encoder, and then further learn the corresponding features and pronunciation characteristics of each language through the corresponding intermediate layer and decoder based on the relevant information corresponding to each language, so that this speech synthesis model can not only effectively utilize all available data, but also enable each component within the model (such as the intermediate layers corresponding to multiple languages ​​and the decoders corresponding to multiple languages) to fully play their roles. Regardless of the amount of data (i.e., scale) of the language, the shared encoder in the speech synthesis model can obtain language-independent abstract feature representations, and then the corresponding intermediate layer and decoder of the language can obtain the accurate pronunciation features of the language, and obtain high-precision speech information, thereby improving the speech synthesis accuracy of languages ​​with small data volumes.

[0049] In combination with the above embodiments, in one implementation, an embodiment of the present invention further provides a speech synthesis method. Specifically, in this embodiment, the above step S12 of "inputting the target phoneme information and the target language information into the speech synthesis model to obtain the target spectral features corresponding to the target text" may specifically include steps S21 to S23:

[0050] Step S21: inputting the target phoneme information and the target language information into the shared encoder to obtain target general information.

[0051] In this embodiment, the target phoneme information and the target language information can be input into the shared encoder in advance to obtain the target general information output by the shared encoder. The target general information is the general information corresponding to the target text, and the general information in this embodiment represents an abstract feature representation that is independent of the language.

[0052] Step S22: inputting the target general information and the target language information into the intermediate layer corresponding to the target language to obtain intermediate features.

[0053] In this embodiment, the intermediate layers corresponding to the multiple languages ​​include at least: an intermediate layer corresponding to the target language. After obtaining the target general information output by the shared encoder, the target general information and the target language information can be input into the intermediate layer corresponding to the target language to obtain the intermediate features corresponding to the target text output by the intermediate layer corresponding to the target language.

[0054] Among them, the intermediate layers corresponding to the multiple languages ​​are for specific processing of different languages, and their structures can be fully connected layers, convolutional layers or other types of neural network layers. The intermediate layer corresponding to the target language can perform more detailed speech or dialect specific processing based on the target general features and target language information by receiving the target general features transmitted by the shared encoder to enhance the characteristics of the target language. In an optional example, the specific processing may include but is not limited to: screening, compressing and / or refining the target general features and target language information.

[0055] Step S23: at least input the intermediate features into a decoder corresponding to the target language to obtain target spectrum features corresponding to the target text.

[0056] In this embodiment, the decoders corresponding to the multiple languages ​​include at least: a decoder corresponding to the target language. After obtaining the intermediate features output by the intermediate layer corresponding to the target language, at least the intermediate features are input into the decoder corresponding to the target language to obtain the target spectrum features corresponding to the target text output by the decoder corresponding to the target language.

[0057] Among them, the decoders corresponding to each of the multiple languages ​​are to learn the unique pronunciation features of each language, and generate unique speech output adapted to the language based on the features generated from the shared encoder and the intermediate layer. The decoders corresponding to each of the multiple languages ​​can be based on the LSTM+attention mechanism, and are not limited to a specific network architecture. The decoders corresponding to each language in the speech synthesis model of this embodiment can optimize it according to the pronunciation features and tone system of its specific language.

[0058] In combination with the above embodiments, in one implementation, the present invention further provides a speech synthesis method. Specifically, in the method, the above step S23 may include step S31 and step S32:

[0059] Step S31: concatenate the intermediate feature, the target phoneme information, and the conditional information corresponding to the target phoneme information to obtain a first concatenated feature.

[0060] In this embodiment, in order to ensure the generation ability and synthesis effect of the speech synthesis model, in addition to using the intermediate features output by the intermediate layer corresponding to each language as the input of the decoder corresponding to each language, the corresponding phoneme sequence and the conditioning information of the phoneme sequence should also be added as additional input of the decoder corresponding to each language.

[0061] Specifically, for the decoder corresponding to the target language, the intermediate features output by the intermediate layer corresponding to the target language, the target phoneme information corresponding to the target text, and the conditional information corresponding to the target phoneme information are concatenated to obtain a first concatenated feature. The conditional information of this embodiment represents the pronunciation condition information corresponding to the language, for example, the conditional information may include: tone information, i.e., tone information, etc.

[0062] Furthermore, in one embodiment, the target phoneme information corresponding to the target text may be encoded to obtain the encoded target phoneme information, and the conditional information corresponding to the target phoneme information may be encoded to obtain the encoded conditional information; then, the intermediate features output by the intermediate layer corresponding to the target language, the encoded target phoneme information, and the encoded conditional information may be concatenated to obtain the first concatenated features. For example, the target phoneme information may be embedded to obtain the embedded target phoneme information, and the conditional information corresponding to the target phoneme information may be embedded to obtain the embedded conditional information.

[0063] Step S32: inputting the first concatenation feature into a decoder corresponding to the target language to obtain a target spectrum feature corresponding to the target text.

[0064] In this embodiment, after obtaining the first concatenation feature, the first concatenation feature may be input into a decoder corresponding to the target language to obtain a target spectrum feature corresponding to the target text output by the decoder corresponding to the target language.

[0065] In combination with the above embodiments, in another implementation, the present invention also provides a speech synthesis method. Specifically, in this embodiment, the above step S21 may include steps S41 to S44:

[0066] Step S41: Encode the target phoneme information to obtain encoded target phoneme information.

[0067] In this embodiment, the speech synthesis model further includes a coding layer, and the target phoneme information corresponding to the target text can be input into the coding layer for coding to obtain the encoded target phoneme information.

[0068] Step S42: Encode the target language information to obtain encoded target language information.

[0069] In this embodiment, the target language information corresponding to the target text may also be input into the encoding layer for encoding to obtain the encoded target language information.

[0070] In an optional embodiment, the encoding layer may be an embed layer, and the target phoneme information corresponding to the target text may be input into the embed layer for embedding processing to obtain an embedded representation of the embedded target phoneme information, and the target language information corresponding to the target text may be input into the embed layer for embedding processing to obtain an embedded representation of the embedded target language information.

[0071] Step S43: concatenating the encoded target phoneme information and the encoded target language information to obtain a second concatenated feature.

[0072] In this embodiment, the encoded target phoneme information and the encoded target language information need to be concatenated to obtain a second concatenated feature.

[0073] Step S44: input the second splicing feature into the shared encoder to obtain the target general information.

[0074] In this embodiment, after obtaining the second splicing feature, the second splicing feature can be input into the shared encoder to obtain the target general information corresponding to the target text output by the shared encoder, wherein the target general information at least includes: the duration of the pronunciation unit of the target phoneme information and the pause rhythm of the target phoneme information.

[0075] In this embodiment, the target phoneme sequence needs to be combined with some professional phonetics knowledge, and the target text is transcribed to obtain the target phoneme information, and then the target phoneme information is accessed to the encoding layer (such as the embed layer). As for the Min dialect, its syllables are composed of three parts: initial consonants, finals and tones, and the finals are composed of three parts: rhyme head, rhyme nucleus and rhyme coda. 15 initial consonants, 82 finals and 8 tones, all of which need to be distinguished from Mandarin Chinese, that is, their encoding cannot be mixed (the same tone categories of Mandarin Chinese and Min dialect, such as Yinping, are actually different, and the same encoding cannot be used; in addition, since the rhyme nucleus and rhyme coda combination of the Min dialect is very complex, it is necessary to model the rhyme nucleus and rhyme coda separately, train the corresponding features, and cannot be mixed with Mandarin Chinese). And, the target language information must also be encoded, and then spliced ​​with the embedding result of the phoneme sequence (i.e., the encoded target phoneme information), and then the splicing result (such as the second splicing feature) is input into the shared encoder.

[0076] The reason why the shared encoder is designed in this embodiment is to hope that the shared encoder can learn the common features of multiple languages, such as the duration and rhythm (pause rhythm) of basic pronunciation units. It should be noted that for different languages, the shared encoder may output different common features, such as the duration and / or rhythm of different basic pronunciation units corresponding to different languages.

[0077] In combination with the above embodiments, in one implementation, the multiple languages ​​in this embodiment include any one or more of the following: Mandarin, one or more Chinese dialects (such as Min dialect, Tianjin dialect, Northeastern dialect, Sichuan dialect, etc.), one or more foreign languages ​​(such as English, Thai, Malay, Spanish, etc.).

[0078] In combination with the above embodiments, in another implementation, an embodiment of the present invention further provides a speech synthesis method. In this embodiment, the speech synthesis model is obtained by training based on an initial speech synthesis model, and the initial speech synthesis model at least includes: an initial shared encoder, initial intermediate layers corresponding to multiple languages, and initial decoders corresponding to multiple languages. The training steps of the initial speech synthesis model may at least include steps S51 to S57:

[0079] Step S51: determining a plurality of sample texts and a plurality of sample audio information of a plurality of sample languages ​​corresponding to the plurality of sample texts respectively.

[0080] The initial speech synthesis model of this embodiment is a multi-task learning model. During model training, speech synthesis tasks of multiple languages ​​can be learned simultaneously. It can be understood that the initial speech synthesis model of this embodiment is a synthetic structure of a shared encoder and a language-specific decoder of a multi-task learning and conditional speech synthesis strategy. When training the initial speech synthesis model, it is necessary to determine a plurality of sample texts, which are texts of speech to be synthesized for model training, and to determine a plurality of sample audio information of a plurality of sample languages ​​corresponding to each of the plurality of sample texts.

[0081] In this embodiment, each sample text corresponds to a sample language, and each sample text corresponds to a sample audio information of the sample language. The sample language corresponding to each sample text may be the same or different. The sample language is the language corresponding to the sample text determined according to the speech synthesis requirements, such as the Min dialect, Mandarin, etc., for training the speech synthesis model, without limitation. The sample audio information is the correct audio information of the sample language corresponding to the sample text. For example, the sample language corresponding to a sample text A is the Min dialect, and the sample text A corresponds to a sample audio information of the Min dialect.

[0082] Step S52: Process the multiple sample texts to obtain multiple sample phoneme information.

[0083] In this embodiment, multiple sample texts may be processed separately to obtain multiple sample phoneme information. For example, the sample texts may be processed, analyzed, and converted to phonemes to obtain the sample phoneme information. The sample phoneme information in this embodiment is the phoneme information corresponding to the sample text. This embodiment does not impose any restrictions on the specific method of processing the sample text to obtain the sample phoneme information.

[0084] Step S53: Inputting the plurality of sample phoneme information and the sample language information of the sample languages ​​corresponding to the plurality of sample phoneme information into the initial shared encoder to obtain a plurality of sample common information.

[0085] In this embodiment, it is necessary to determine the sample language information of the sample languages ​​corresponding to the multiple sample texts according to the speech synthesis requirements, that is, to determine the sample language information of the sample languages ​​corresponding to the multiple sample phoneme information, and the sample language information is the relevant information of the sample languages.

[0086] After obtaining a plurality of sample phoneme information and sample language information of the sample languages ​​corresponding to the plurality of sample phoneme information, the plurality of sample phoneme information and the sample language information of the sample languages ​​corresponding to the plurality of sample phoneme information can be respectively input into the initial shared encoder to obtain a plurality of sample general information corresponding to the plurality of sample phoneme information. The sample general information is general information corresponding to the sample text, and the general information of this embodiment represents an abstract feature representation that is independent of the language.

[0087] For example, for sample phoneme information A, the sample phoneme information A and the sample language information A corresponding to the sample phoneme information A are input into the initial shared encoder to obtain sample general information A corresponding to the sample phoneme information A. For sample phoneme information B, the sample phoneme information B and the sample language information B corresponding to the sample phoneme information B are input into the initial shared encoder to obtain sample general information B corresponding to the sample phoneme information B.

[0088] In an optional embodiment, a plurality of sample phoneme information are respectively encoded to obtain a plurality of encoded sample phoneme information, and the sample language information corresponding to each of the plurality of sample phoneme information are respectively encoded to obtain a plurality of encoded sample language information, and then the plurality of encoded sample phoneme information are respectively concatenated with the encoded sample language information corresponding to each of them to obtain a plurality of second sample concatenation features, and then the plurality of first sample concatenation features are respectively input into the initial shared encoder to obtain a plurality of sample common information corresponding to each of the plurality of sample phoneme information.

[0089] Step S54: inputting the plurality of sample general information and the sample language information corresponding to each of the plurality of sample general information into the initial intermediate layers corresponding to each of the plurality of sample languages ​​to obtain a plurality of sample intermediate features.

[0090] In this embodiment, after obtaining a plurality of sample general information, the plurality of sample general information and the sample language information corresponding to each of the plurality of sample general information can be respectively input into the initial intermediate layer corresponding to each of the plurality of sample languages ​​to obtain a plurality of sample intermediate features output by the initial intermediate layer corresponding to each of the plurality of sample languages. The initial intermediate layer corresponding to each of the plurality of languages ​​in this embodiment includes: the initial intermediate layer corresponding to each of the plurality of sample languages.

[0091] For example, for sample general information A, sample general information A and sample language information A corresponding to sample general information A are input into the initial intermediate layer corresponding to sample language A to obtain sample intermediate feature A. For sample general information B, sample general information B and sample language information B corresponding to sample general information B are input into the initial intermediate layer corresponding to sample language B to obtain sample intermediate feature B.

[0092] Among them, the initial intermediate layer corresponding to each of the multiple languages ​​is for specific processing of different languages, and its structure can be a fully connected layer, a convolutional layer or other types of neural network layers. The initial intermediate layer receives the sample common features of the corresponding sample language transmitted by the initial shared encoder, and can perform more detailed specific processing of the speech or dialect based on the sample common features and sample language information to enhance the characteristics of the corresponding sample language. In an optional example, the specific processing may include but is not limited to: screening, compressing and / or refining the sample common features and sample language information.

[0093] Step S55: Input the plurality of sample intermediate features, the sample phoneme information corresponding to each of the plurality of sample intermediate features, and the sample conditioning information corresponding to the sample phoneme information into initial decoders corresponding to each of the plurality of sample languages ​​to obtain sample spectral features corresponding to each of the plurality of sample texts.

[0094] In this embodiment, after obtaining multiple sample intermediate features outputted by the initial intermediate layer corresponding to each of the multiple sample languages, the multiple sample intermediate features, the sample phoneme information corresponding to each of the multiple sample intermediate features, and the sample conditioning information corresponding to the sample phoneme information corresponding to each of the multiple sample intermediate features can be inputted into the initial decoders corresponding to each of the multiple sample languages ​​to obtain multiple sample spectrum features outputted by the initial decoders corresponding to each of the multiple sample languages, and the multiple sample spectrum features are the sample spectrum features corresponding to each of the multiple sample texts. The sample spectrum features are the spectrum features corresponding to the sample texts. Among them, the initial decoders corresponding to each of the multiple languages ​​in this embodiment include: initial decoders corresponding to each of the multiple sample languages.

[0095] For example, for the sample intermediate feature A, the sample intermediate feature A, the sample phoneme information A corresponding to the sample intermediate feature A, and the sample conditioning information A corresponding to the sample phoneme information A corresponding to the sample intermediate feature A can be input into the initial decoder corresponding to the sample language A to obtain the sample spectral feature A output by the initial decoder corresponding to the sample language A; for the sample intermediate feature B, the sample intermediate feature B, the sample phoneme information B corresponding to the sample intermediate feature B, and the sample conditioning information B corresponding to the sample phoneme information B corresponding to the sample intermediate feature B can be input into the initial decoder corresponding to the sample language B to obtain the sample spectral feature B output by the initial decoder corresponding to the sample language B.

[0096] Among them, the initial decoders corresponding to each of the multiple languages ​​are to learn the unique pronunciation features of each language, and generate unique speech output adapted to the language based on the features generated from the initial shared encoder and the initial intermediate layer. The initial decoders corresponding to each of the multiple languages ​​can be based on the LSTM+attention mechanism, and are not limited to a specific network architecture. The initial decoder corresponding to each language in the initial speech synthesis model of this embodiment can optimize it according to the pronunciation features and tone system of its specific language.

[0097] In an optional embodiment, multiple sample intermediate features, sample phoneme information corresponding to each of the multiple sample intermediate features, and sample conditioning information corresponding to the sample phoneme information corresponding to each of the multiple sample intermediate features are concatenated to obtain multiple first sample concatenated features, and then the multiple first sample concatenated features are respectively input into initial decoders corresponding to each of the multiple sample languages. The sample conditioning information in this embodiment is conditioning information corresponding to the sample phoneme information.

[0098] Furthermore, the sample phoneme information corresponding to each of the multiple sample intermediate features may be encoded respectively to obtain multiple encoded sample phoneme information; and the sample conditioning information corresponding to the sample phoneme information corresponding to each of the multiple sample intermediate features may be encoded to obtain multiple encoded sample conditioning information; then, the multiple sample intermediate features may be spliced ​​with the encoded sample phoneme information and the encoded sample conditioning information corresponding to them to obtain multiple first sample splicing features. For example, the sample phoneme information corresponding to each of the multiple sample intermediate features may be embedded respectively to obtain multiple embedded sample phoneme information; and the sample conditioning information corresponding to the sample phoneme information corresponding to each of the multiple sample intermediate features may be embedded respectively to obtain multiple embedded sample conditioning information.

[0099] Step S56: Based on the plurality of sample frequency spectrum features, a plurality of sample speech information corresponding to a plurality of sample languages ​​is obtained.

[0100] In this embodiment, after obtaining multiple sample spectrum features, sample speech information of multiple sample languages ​​can be obtained based on the multiple sample spectrum features. In this embodiment, the specific method of processing the sample spectrum features to obtain the sample speech information of the sample language is not limited.

[0101] Step S57: Based at least on the sample speech information and sample audio information corresponding to the multiple sample languages ​​respectively, the model parameters of the initial shared encoder, the initial intermediate layer corresponding to the multiple sample languages ​​respectively, and the initial decoder corresponding to the multiple sample languages ​​respectively are updated to obtain the speech synthesis model.

[0102] In this embodiment, based on the sample speech information and sample audio information corresponding to the multiple sample languages, the model parameters of the initial shared encoder, the initial intermediate layer corresponding to the multiple sample languages, and the initial decoder corresponding to the multiple sample languages ​​can be updated to obtain a trained shared encoder, an intermediate layer corresponding to the multiple languages, and a decoder corresponding to the multiple languages, thereby obtaining a trained speech synthesis model. The speech synthesis model of this embodiment can improve the synthesis quality of dialects with small data volume and enhance the overall model performance.

[0103] Specifically, the model parameters of the initial shared encoder, the initial intermediate layer corresponding to the sample language, and the initial decoder corresponding to the sample language may be updated based on the sample speech information and the sample audio information corresponding to each of the multiple sample languages, until the model parameters of the initial shared encoder, the initial intermediate layer corresponding to the multiple sample languages, and the initial decoder corresponding to the multiple sample languages ​​are updated, thereby obtaining a trained speech synthesis model.

[0104] In combination with the above embodiments, in one implementation, the present invention further provides a speech synthesis method. Specifically, in the method, the above step S57 may include step S61:

[0105] Step S61: based on sample speech information and sample audio information respectively corresponding to a plurality of sample languages, model parameters of the initial shared encoder and initial decoders respectively corresponding to the plurality of sample languages ​​are updated to obtain the speech synthesis model.

[0106] In this embodiment, the intermediate layer corresponding to each language is trained separately, so that the characteristics of each language can be better represented and processed in the model. In other words, the initial intermediate layers corresponding to the multiple languages ​​in the initial speech synthesis model are pre-trained intermediate layers corresponding to the multiple languages, and the initial intermediate layers corresponding to the multiple languages ​​do not need to update the model parameters.

[0107] In this embodiment, based on the sample speech information and sample audio information corresponding to the multiple sample languages, the model parameters of the initial shared encoder and the initial decoders corresponding to the multiple sample languages ​​are updated to obtain a trained shared encoder and decoders corresponding to the multiple languages, thereby obtaining a trained speech synthesis model.

[0108] In one embodiment, if Figure 3 As shown, Figure 3 FIG. 1 is a schematic diagram of a structure of a speech synthesis model according to an embodiment of the present invention. Figure 3 The main parts of the speech synthesis model are: a shared encoder, separate intermediate layers for each language (such as Figure 3 The middle layer-1, middle layer-2 in the code and the separate decoders for each language (such as Figure 3 Decoder-1, decoder-2), wherein this embodiment takes 2 intermediate layers and 2 decoders as an example, and there is no specific restriction on the number of intermediate layers and decoders of the speech synthesis model, but the intermediate layers correspond one to one to the decoders.

[0109] Among them, langID-1 is language information-1, and langID-2 is language information-2. First, the language information-1 and phoneme sequence-1, as well as the language information-2 and phoneme sequence-2 are input into the speech synthesis model. The language information-1 and phoneme sequence-1, as well as the language information-2 and phoneme sequence-2 are respectively passed through their corresponding embedding layers (i.e. Figure 3 The embedding in the text file is embedded (i.e., encoded) to obtain the embedded language information-1 and the embedded phoneme sequence-1, as well as the embedded language information-2 and the embedded phoneme sequence-2.

[0110] Secondly, the embedded language information-1 and the embedded phoneme sequence-1 are concatenated and input into the shared encoder to obtain the universal feature-1 corresponding to the phoneme sequence-1; the embedded language information-2 and the embedded phoneme sequence-2 are concatenated and input into the shared encoder to obtain the universal feature-2 corresponding to the phoneme sequence-2.

[0111] Then, the general feature-1 and the language information-1 (such as the language information-1 after embedding) are input into the intermediate layer-1 corresponding to the language information-1 to obtain the intermediate feature-1 output by the intermediate layer-1; the general feature-2 and the language information-2 (such as the language information-2 after embedding) are input into the intermediate layer-2 corresponding to the language information-2 to obtain the intermediate feature-2 output by the intermediate layer-2.

[0112] Finally, the intermediate feature-1 is concatenated with the phoneme sequence-1 (such as the embedded phoneme sequence-1) and the tone sequence-1 (i.e., the conditional information, such as the embedded tone sequence-1) and input into the decoder-1 corresponding to the language information-1, and mel_out1 output by the decoder-1 is obtained, thereby finally obtaining the speech information-1, and realizing the speech synthesis of the language corresponding to the language information-1. The intermediate feature-2 is concatenated with the phoneme sequence-2 (such as the embedded phoneme sequence-2) and the tone sequence-2 (i.e., the conditional information, such as the embedded tone sequence-2) and input into the decoder-2 corresponding to the language information-2, and mel_out2 output by the decoder-2 is obtained, thereby finally obtaining the speech information-2, and realizing the speech synthesis of the language corresponding to the language information-2.

[0113] In combination with the above embodiments, in one implementation, the embodiment of the present invention further provides a speech synthesis method. Specifically, in the method, the initial speech synthesis model further includes: an initial language classifier; the training steps of the initial speech synthesis model may include step S71 in addition to the above steps S51 to S57, and the above step S57 may specifically include steps S72 to S74:

[0114] Step S71: inputting the plurality of sample common information into the initial language classifier to obtain a plurality of sample language prediction results.

[0115] In this embodiment, a language classification task is added to the multi-task learning framework, that is, on the basis of the initial shared encoder, in addition to connecting the language-specific intermediate layer and the initial decoder for speech synthesis, an initial language classifier can also be added in parallel. Among them, the initial language classifier can be a simple fully connected network, and its task is to predict the language of the input speech data. After obtaining the general information of multiple samples output by the initial shared encoder, this embodiment can input the general information of multiple samples into the initial language classifier to obtain the language prediction results of multiple samples output by the initial language classifier.

[0116] Step S72: Calculate a first loss based on sample speech information and sample audio information corresponding to a plurality of sample languages.

[0117] In this embodiment, it is necessary to calculate the first loss for the sample speech information and sample audio information corresponding to the multiple sample languages. For example, the audio loss may be calculated for the sample speech information and sample audio information corresponding to each sample language, and then the audio losses corresponding to the multiple sample languages ​​are weighted and summed to obtain the first loss. The first loss may be the MSE loss.

[0118] Step S73: Calculate a second loss based on the multiple sample language prediction results and the multiple sample language information.

[0119] In this embodiment, it is necessary to calculate the second loss using the sample language prediction results and sample language information corresponding to the multiple sample languages. For example, the language loss may be calculated using the sample language prediction results and sample language information corresponding to each sample language, and then the language losses corresponding to the multiple sample languages ​​are weighted and summed to obtain the second loss. The second loss may be a classification loss (such as a cross entropy loss) to ensure that the model can correctly distinguish different languages, to enhance the accuracy of language recognition, and to assist speech synthesis.

[0120] Step S74: Based on the first loss and the second loss, the model parameters of the initial shared encoder, the initial language classifier, the initial intermediate layers corresponding to the multiple sample languages, and the initial decoders corresponding to the multiple sample languages ​​are updated to obtain the speech synthesis model.

[0121] In this embodiment, based on the first loss and the second loss (such as the sum of the two, or the weighted sum of the two), the model parameters of the initial shared encoder, the initial language classifier, the initial intermediate layers corresponding to the multiple sample languages, and the initial decoders corresponding to the multiple sample languages ​​can be updated to obtain the trained shared encoder, language classifier, the intermediate layers corresponding to the multiple languages, and the decoders corresponding to the multiple languages, thereby obtaining a trained speech synthesis model.

[0122] In combination with the above embodiments, in one implementation, the present invention further provides a speech synthesis method. Specifically, in addition to the above steps, the method may further include steps S81 to S84:

[0123] Step S81: Process multiple outbound call texts to obtain multiple outbound call phoneme information.

[0124] The application environment of this embodiment is an intelligent outbound call system in the Min dialect area, and the intelligent outbound call system needs to provide a multi-language outbound call voice library including Mandarin and Min dialect. Based on this, this embodiment can process multiple outbound call texts to obtain multiple outbound call phoneme information. Among them, the outbound call text is the text that needs to make an outbound call, including: voice broadcast text and / or navigation text, such as "Hello, tourist friends", or "Please pay attention to safety", etc. The outbound call phoneme information is the phoneme information corresponding to the outbound call text. This step is the same or similar to the above-mentioned step S11, and will not be repeated here.

[0125] Step S82: input each outbound call phoneme information of the multiple outbound call phoneme information and the outbound call language information of the multiple outbound call languages ​​corresponding to each outbound call phoneme information into the speech synthesis model to obtain the outbound call frequency spectrum features of the multiple outbound call languages ​​corresponding to each outbound call text.

[0126] In this embodiment, the unmatched outbound language of each outbound call can be determined. The outbound language is the language corresponding to the outbound text specified in the outbound speech synthesis requirement. Each outbound text can correspond to multiple outbound languages, such as the Fujian dialect, Mandarin, English, Korean, etc., without limitation.

[0127] In this embodiment, each outbound call phoneme information of the multiple outbound call phoneme information and the outbound call language information of the multiple outbound call languages ​​corresponding to each outbound call phoneme information can be input into a pre-trained speech synthesis model to obtain the outbound call spectrum features of the multiple outbound call languages ​​corresponding to each outbound call text output by the speech synthesis model. Among them, the outbound call spectrum features are the spectrum features corresponding to the outbound call text. This step is the same or similar to the above step S12, and will not be described in detail.

[0128] Step S83: based on the outbound call spectrum features of the multiple outbound call languages ​​corresponding to each outbound call text, obtain the voice information of the multiple outbound call languages ​​corresponding to each outbound call text.

[0129] In this embodiment, after obtaining the outbound call frequency spectrum features of the multiple outbound call languages ​​corresponding to each outbound call text, the voice information of the multiple outbound call languages ​​corresponding to each outbound call text can be obtained based on the outbound call frequency spectrum features of the multiple outbound call languages ​​corresponding to each outbound call text. This step is the same or similar to the above step S13, and will not be described in detail.

[0130] Step S84: Generate a multi-language outbound call voice library based on the voice information of multiple outbound call languages ​​corresponding to the multiple outbound call texts.

[0131] In this embodiment, after obtaining the voice information of multiple outbound languages ​​corresponding to the multiple outbound texts, a multilingual outbound voice library can be generated based on the voice information of the multiple outbound languages ​​corresponding to the multiple outbound texts. The multiple outbound languages ​​of this embodiment include at least multiple dialects including Mandarin and Min dialect. In addition, the multilingual outbound voice library generated by this embodiment can be applied to an intelligent outbound call system, which can train a high-quality Min dialect voice library with fewer Min dialect samples (i.e., in the scenario of Min dialect speech synthesis with a small amount of data) while maintaining the synthesis quality of Mandarin.

[0132] It should be noted that, for the sake of simplicity, the method embodiments are described as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.

[0133] Based on the same inventive concept, an embodiment of the present invention provides a speech synthesis device 400. Figure 4 , Figure 4 FIG. 1 is a structural block diagram of a speech synthesis device provided by an embodiment of the present invention. Figure 4 As shown, the speech synthesis device 400 includes:

[0134] The target phoneme acquisition module 401 is used to process the target text to obtain target phoneme information, wherein the target text includes: one or more texts to be processed;

[0135] The first model processing module 402 is used to input the target phoneme information and the target language information into a speech synthesis model to obtain a target spectrum feature corresponding to the target text, wherein the target language information is one or more language information corresponding to the target text;

[0136] A target speech synthesis module 403, used to obtain speech information of a target language based on the target spectral features;

[0137] Among them, the speech synthesis model at least includes: a shared encoder, an intermediate layer corresponding to each of the multiple languages, and a decoder corresponding to each of the multiple languages; the shared encoder is used to process text conversion tasks in multiple languages ​​and generate language-independent abstract feature representations; the intermediate layers corresponding to each of the multiple languages ​​are used to enhance the characteristics of each of the multiple languages; the decoders corresponding to each of the multiple languages ​​are used to learn the pronunciation features corresponding to each of the multiple languages.

[0138] Optionally, the model processing module 402 includes:

[0139] A shared encoder processing module, used for inputting the target phoneme information and the target language information into the shared encoder to obtain target general information;

[0140] An intermediate layer processing module, used for inputting the target general information and the target language information into the intermediate layer corresponding to the target language to obtain intermediate features;

[0141] The decoder processing module is used to at least input the intermediate features into a decoder corresponding to the target language to obtain target spectrum features corresponding to the target text.

[0142] Optionally, the decoder processing module includes:

[0143] A first concatenation module, configured to concatenate the intermediate feature, the target phoneme information, and the conditional information corresponding to the target phoneme information to obtain a first concatenation feature;

[0144] The first processing module is used to input the first concatenation feature into a decoder corresponding to the target language to obtain a target spectrum feature corresponding to the target text.

[0145] Optionally, a shared encoder processing module comprises:

[0146] A first encoding module, used for encoding the target phoneme information to obtain encoded target phoneme information;

[0147] A second encoding module is used to encode the target language information to obtain the encoded target language information;

[0148] A second concatenation module, configured to concatenate the encoded target phoneme information and the encoded target language information to obtain a second concatenation feature;

[0149] The second processing module is used to input the second splicing feature into the shared encoder to obtain the target general information, where the target general information at least includes: the duration of the pronunciation unit of the target phoneme information and the pause rhythm of the target phoneme information.

[0150] Optionally, the speech synthesis model is obtained by training an initial speech synthesis model, wherein the initial speech synthesis model at least includes: an initial shared encoder, initial intermediate layers corresponding to each of the multiple languages, and initial decoders corresponding to each of the multiple languages; the speech synthesis device 400 further includes a model training module, wherein the model training module is used to train the initial speech synthesis model; the model training module at least includes:

[0151] A sample information determination module, used to determine a plurality of sample texts and a plurality of sample audio information of a plurality of sample languages ​​corresponding to the plurality of sample texts;

[0152] A sample phoneme acquisition module, used for processing the plurality of sample texts to obtain a plurality of sample phoneme information;

[0153] A first sample processing module, configured to input the plurality of sample phoneme information and sample language information of the sample languages ​​corresponding to the plurality of sample phoneme information into the initial shared encoder to obtain a plurality of sample general information;

[0154] A second sample processing module, configured to input the plurality of sample general information and the sample language information corresponding to each of the plurality of sample general information into initial intermediate layers corresponding to each of the plurality of sample languages, to obtain a plurality of sample intermediate features;

[0155] A third sample processing module is used to input the plurality of sample intermediate features, the sample phoneme information corresponding to each of the plurality of sample intermediate features, and the sample conditioning information corresponding to the sample phoneme information into initial decoders corresponding to each of the plurality of sample languages, to obtain sample spectrum features corresponding to each of the plurality of sample texts;

[0156] A sample speech synthesis module, used for obtaining a plurality of sample speech information corresponding to a plurality of sample languages ​​based on a plurality of sample spectrum features;

[0157] A parameter updating module is used to update the model parameters of the initial shared encoder, the initial intermediate layer corresponding to each of the multiple sample languages, and the initial decoder corresponding to each of the multiple sample languages ​​based on at least the sample speech information and the sample audio information corresponding to each of the multiple sample languages, so as to obtain the speech synthesis model.

[0158] Optionally, the initial intermediate layers corresponding to the multiple languages ​​are pre-trained intermediate layers corresponding to the multiple languages;

[0159] The parameter updating module comprises:

[0160] The first parameter updating submodule is used to update the model parameters of the initial shared encoder and the initial decoders corresponding to the multiple sample languages ​​based on the sample speech information and the sample audio information corresponding to the multiple sample languages ​​respectively, so as to obtain the speech synthesis model.

[0161] Optionally, the initial speech synthesis model further includes: an initial language classifier; and the model training module further includes:

[0162] A language prediction module, used for inputting the plurality of sample common information into the initial language classifier to obtain a plurality of sample language prediction results;

[0163] The parameter updating module comprises:

[0164] A first loss calculation module, used to calculate a first loss based on sample speech information and sample audio information corresponding to a plurality of sample languages ​​respectively;

[0165] A second loss calculation module, used to calculate a second loss based on multiple sample language prediction results and multiple sample language information;

[0166] The second parameter updating submodule is used to update the model parameters of the initial shared encoder, the initial language classifier, the initial intermediate layers corresponding to the multiple sample languages, and the initial decoders corresponding to the multiple sample languages ​​based on the first loss and the second loss to obtain the speech synthesis model.

[0167] Optionally, the speech synthesis device 400 further includes:

[0168] An outbound call phoneme determination module, used for processing a plurality of outbound call texts to obtain a plurality of outbound call phoneme information;

[0169] The second model processing module is used to input each outbound call phoneme information of the multiple outbound call phoneme information and the outbound call language information of the multiple outbound call languages ​​corresponding to each outbound call phoneme information into the speech synthesis model to obtain the outbound call frequency spectrum features of the multiple outbound call languages ​​corresponding to each outbound call text;

[0170] An outbound call speech synthesis module, used to obtain speech information of multiple outbound call languages ​​corresponding to each outbound call text based on the outbound call spectrum features of multiple outbound call languages ​​corresponding to each outbound call text;

[0171] An outbound call voice library generation module, used to generate a multi-language outbound call voice library based on voice information of multiple outbound call languages ​​corresponding to multiple outbound call texts;

[0172] The multiple outbound call languages ​​include at least multiple dialects including Mandarin and the Fujian dialect, and the multi-language outbound call voice library is applied to the intelligent outbound call system.

[0173] Based on the same inventive concept, another embodiment of the present invention provides an electronic device 500, such as Figure 5 shown. Figure 5 1 is a schematic diagram of an electronic device according to an embodiment of the present invention. The electronic device comprises a memory 502, a processor 501 and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the speech synthesis method according to any of the above embodiments of the present invention when executing the computer program.

[0174] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0175] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0176] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, devices, or computer program products. Therefore, embodiments of the present invention may take the form of complete hardware embodiments, complete software embodiments, or embodiments combining software and hardware. Furthermore, embodiments of the present invention may take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0177] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0178] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0179] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable terminal device. Figure 1 A process or multiple processes and / or boxes Figure 1The steps for the functions specified in one or more boxes.

[0180] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.

[0181] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements.

[0182] The above is a detailed introduction to a speech synthesis method, device and electronic device provided by the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for a person skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.

Claims

1. A speech synthesis method, characterized in that: The method comprises: Processing a target text to obtain target phoneme information, wherein the target text includes: one or more texts to be processed; Inputting the target phoneme information and the target language information into a speech synthesis model to obtain target spectral features corresponding to the target text, wherein the target language information is one or more language information corresponding to the target text; Obtaining speech information of a target language based on the target frequency spectrum features; Among them, the speech synthesis model at least includes: a shared encoder, an intermediate layer corresponding to each of the multiple languages, and a decoder corresponding to each of the multiple languages; the shared encoder is used to process text conversion tasks in multiple languages ​​and generate language-independent abstract feature representations; the intermediate layers corresponding to each of the multiple languages ​​are used to enhance the characteristics of each of the multiple languages; the decoders corresponding to each of the multiple languages ​​are used to learn the pronunciation features corresponding to each of the multiple languages.

2. The speech synthesis method according to claim 1, characterized in that: Inputting the target phoneme information and the target language information into a speech synthesis model to obtain target spectrum features corresponding to the target text includes: Inputting the target phoneme information and the target language information into the shared encoder to obtain target general information; Inputting the target general information and the target language information into the intermediate layer corresponding to the target language to obtain intermediate features; At least the intermediate features are input into a decoder corresponding to the target language to obtain target spectrum features corresponding to the target text.

3. The speech synthesis method according to claim 2, characterized in that: At least inputting the intermediate features into a decoder corresponding to the target language to obtain target spectrum features corresponding to the target text includes: Concatenating the intermediate feature, the target phoneme information, and the conditional information corresponding to the target phoneme information to obtain a first concatenated feature; The first concatenated feature is input into a decoder corresponding to the target language to obtain a target spectrum feature corresponding to the target text.

4. The speech synthesis method according to claim 2, characterized in that: Inputting the target phoneme information and the target language information into the shared encoder to obtain target general information includes: Encoding the target phoneme information to obtain encoded target phoneme information; Encoding the target language information to obtain encoded target language information; Concatenating the encoded target phoneme information and the encoded target language information to obtain a second concatenated feature; The second concatenation feature is input into the shared encoder to obtain the target general information, which at least includes: the duration of the pronunciation unit of the target phoneme information and the pause rhythm of the target phoneme information.

5. The speech synthesis method according to any one of claims 1 to 4, characterized in that: The speech synthesis model is obtained by training based on an initial speech synthesis model, wherein the initial speech synthesis model at least includes: an initial shared encoder, initial intermediate layers corresponding to each of the multiple languages, and initial decoders corresponding to each of the multiple languages; the training steps of the initial speech synthesis model at least include: Determine a plurality of sample texts and a plurality of sample audio information of a plurality of sample languages ​​corresponding to the plurality of sample texts respectively; Processing the plurality of sample texts to obtain a plurality of sample phoneme information; Inputting the plurality of sample phoneme information and sample language information of the sample languages ​​corresponding to the plurality of sample phoneme information into the initial shared encoder to obtain a plurality of sample general information; Inputting the plurality of sample general information and the sample language information corresponding to each of the plurality of sample general information into the initial intermediate layers corresponding to each of the plurality of sample languages ​​to obtain a plurality of sample intermediate features; Inputting the plurality of sample intermediate features, the sample phoneme information corresponding to each of the plurality of sample intermediate features, and the sample conditioning information corresponding to the sample phoneme information into initial decoders corresponding to each of the plurality of sample languages, to obtain sample spectral features corresponding to each of the plurality of sample texts; Based on the plurality of sample spectral features, obtaining a plurality of sample speech information corresponding to a plurality of sample languages; At least based on the sample speech information and sample audio information corresponding to the multiple sample languages ​​respectively, the model parameters of the initial shared encoder, the initial intermediate layer corresponding to the multiple sample languages ​​respectively, and the initial decoder corresponding to the multiple sample languages ​​respectively are updated to obtain the speech synthesis model.

6. The speech synthesis method according to claim 5, characterized in that: The initial intermediate layers corresponding to the multiple languages ​​are pre-trained intermediate layers corresponding to the multiple languages; The method includes updating the model parameters of the initial shared encoder, the initial intermediate layer corresponding to each of the multiple sample languages, and the initial decoder corresponding to each of the multiple sample languages ​​based on at least the sample speech information and the sample audio information respectively corresponding to the multiple sample languages ​​to obtain the speech synthesis model, including: Based on sample speech information and sample audio information respectively corresponding to a plurality of sample languages, model parameters of the initial shared encoder and the initial decoders respectively corresponding to the plurality of sample languages ​​are updated to obtain the speech synthesis model.

7. The speech synthesis method according to claim 5, characterized in that: The initial speech synthesis model further includes: an initial language classifier; the training step of the initial speech synthesis model further includes: Inputting the plurality of sample common information into the initial language classifier to obtain a plurality of sample language prediction results; The method includes updating the model parameters of the initial shared encoder, the initial intermediate layer corresponding to each of the multiple sample languages, and the initial decoder corresponding to each of the multiple sample languages ​​based on at least the sample speech information and the sample audio information respectively corresponding to the multiple sample languages ​​to obtain the speech synthesis model, including: Calculating a first loss based on sample speech information and sample audio information respectively corresponding to a plurality of sample languages; Calculating a second loss based on the multiple sample language prediction results and the multiple sample language information; Based on the first loss and the second loss, the model parameters of the initial shared encoder, the initial language classifier, the initial intermediate layers corresponding to the multiple sample languages, and the initial decoders corresponding to the multiple sample languages ​​are updated to obtain the speech synthesis model.

8. The speech synthesis method according to any one of claims 1 to 4, characterized in that: The method further comprises: Processing multiple outbound call texts to obtain multiple outbound call phoneme information; Input each outbound call phoneme information of the plurality of outbound call phoneme information and the outbound call language information of the plurality of outbound call languages ​​corresponding to each outbound call phoneme information into the speech synthesis model to obtain outbound call frequency spectrum features of the plurality of outbound call languages ​​corresponding to each outbound call text; Based on the outbound call spectrum features of the multiple outbound call languages ​​corresponding to each outbound call text, voice information of the multiple outbound call languages ​​corresponding to each outbound call text is obtained; Generate a multilingual outbound call voice library based on voice information of multiple outbound call languages ​​corresponding to multiple outbound call texts; The multiple outbound call languages ​​include at least multiple dialects including Mandarin and the Fujian dialect, and the multi-language outbound call voice library is applied to the intelligent outbound call system.

9. A speech synthesis device, characterized in that: The device comprises: A target phoneme acquisition module is used to process a target text to obtain target phoneme information, wherein the target text includes: one or more texts to be processed; A first model processing module is used to input the target phoneme information and the target language information into a speech synthesis model to obtain a target spectrum feature corresponding to the target text, wherein the target language information is one or more language information corresponding to the target text; A target speech synthesis module, used to obtain speech information of a target language based on the target spectral features; Among them, the speech synthesis model at least includes: a shared encoder, an intermediate layer corresponding to each of the multiple languages, and a decoder corresponding to each of the multiple languages; the shared encoder is used to process text conversion tasks in multiple languages ​​and generate language-independent abstract feature representations; the intermediate layers corresponding to each of the multiple languages ​​are used to enhance the characteristics of each of the multiple languages; the decoders corresponding to each of the multiple languages ​​are used to learn the pronunciation features corresponding to each of the multiple languages.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the computer program is executed by the processor, the speech synthesis method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Voice processing method and related device

    CN112397083A

  • Femule language transfer learning speech synthesis method based on implicit phoneme conversion

    CN114822488A

  • Speech synthesis method and system for minority language, electronic equipment and storage medium

    CN116453500A

  • Cross-language speech synthesis method and device, electronic equipment and storage medium

    CN116825084A

  • Multilingual speech synthesis and cross-language voice cloning

    US20200380952A1

Cited By

  • Voice generation method and device, computer readable storage medium and electronic equipment

    CN120600001A