Cross-lingual corpus synthesis method, speech synthesis model training method, and related devices

By acquiring cross-lingual text and language embedding vectors, synthesizing cross-lingual corpora, and constructing a more natural speech synthesis model, the problem of corpus scarcity in cross-lingual speech synthesis is solved, and natural speech synthesis of cross-lingual speech is realized.

CN115985283BActive Publication Date: 2026-05-29UNIV OF SCI & TECH OF CHINA

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
UNIV OF SCI & TECH OF CHINA
Filing Date
2022-08-31
Publication Date
2026-05-29

Smart Images

  • Figure CN115985283B_ABST
    Figure CN115985283B_ABST
Patent Text Reader

Abstract

The application provides a cross-language corpus synthesis method, a speech synthesis model training method and related equipment. The cross-language corpus synthesis method comprises the following steps: obtaining a cross-language text, a target speaker embedding vector, and a language embedding vector corresponding to each language included in the cross-language text; determining a character embedding vector corresponding to each character included in the cross-language text; determining a speech spectrum corresponding to the cross-language text according to the language embedding vector corresponding to each language included in the cross-language text, the character embedding vector corresponding to each character included in the cross-language text, and the target speaker embedding vector; and composing a cross-language corpus from the speech spectrum and the cross-language text. The application can synthesize speech spectrums of the same speaker switching between languages, thereby obtaining a cross-language corpus composed of the speech spectrums and the cross-language text, so that a speech synthesis model with higher naturalness of synthesized speech can be constructed based on the obtained cross-language corpus subsequently.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech signal processing technology, and in particular to a cross-language corpus synthesis method, a speech synthesis model training method, and related equipment. Background Technology

[0002] Speech synthesis models are systems that convert text into corresponding speech. In related technologies, multilingual monolingual corpora are typically used to train speech synthesis models, resulting in models that demonstrate good speech synthesis performance for monolingual text.

[0003] However, in cross-language speech synthesis scenarios, it is often necessary to synthesize cross-language text into speech from the same speaker. However, it is difficult to obtain corpora of the same speaker in different languages, and cross-language corpora of the same speaker switching languages ​​are even scarcer, thus posing a challenge to existing cross-language speech synthesis methods. Currently, in cross-language speech synthesis scenarios, only speech synthesis models trained with multilingual monolingual corpora can be used for cross-language speech synthesis. However, this often results in synthesis failures or situations where the same sentence contains multiple different timbres, leading to low naturalness of the synthesized speech.

[0004] Therefore, there is an urgent need for a method to obtain cross-lingual corpora in order to construct a more natural speech synthesis model based on cross-lingual corpora. Summary of the Invention

[0005] In view of this, this application provides a cross-lingual corpus synthesis method, a speech synthesis model training method, and related equipment for synthesizing cross-lingual corpora, so as to train a speech synthesis model with higher naturalness based on the cross-lingual corpora. The technical solution is as follows:

[0006] A method for cross-linguistic corpus synthesis, comprising:

[0007] Obtain cross-language text, target speaker embedding vector, and language embedding vectors corresponding to each language contained in the cross-language text. The cross-language text is composed of word segments in multiple languages, each word segment contains at least one character, the language embedding vector represents the language information of the corresponding character and / or word segment, and the target speaker embedding vector represents the timbre information of the target speaker.

[0008] Determine the character embedding vectors corresponding to each character in the cross-language text;

[0009] The speech spectrum corresponding to the cross-language text is determined based on the language embedding vectors corresponding to each language contained in the cross-language text, the character embedding vectors corresponding to each character contained in the cross-language text, and the target speaker embedding vector.

[0010] The cross-linguistic corpus consists of speech spectrum and cross-linguistic text.

[0011] Optionally, based on the language embedding vectors corresponding to each language in the cross-language text, the character embedding vectors corresponding to each character in the cross-language text, and the target speaker embedding vector, the speech spectrum corresponding to the cross-language text is determined, including:

[0012] Based on the language embedding vectors corresponding to each language in the cross-language text and the character embedding vectors corresponding to each character in the cross-language text, the character encoding vectors corresponding to each character in the cross-language text are determined, where the character encoding vectors represent the semantics and language information of the corresponding characters.

[0013] Based on the character encoding vectors corresponding to each character in the cross-language text, the text encoding vectors corresponding to each character in the cross-language text are determined. The text encoding vectors represent the semantic information of the corresponding character and its preceding and following adjacent characters, as well as the language information of the corresponding character.

[0014] Determine the cross-language context-related vectors corresponding to each word in the cross-language text, where the cross-language context-related vectors represent the semantic information of the corresponding word and its preceding and following adjacent words;

[0015] Based on the text encoding vectors corresponding to each character in the cross-language text and the cross-language context-related vectors corresponding to each word segment in the cross-language text, the multi-scale text encoding vectors corresponding to each character in the cross-language text are determined. The multi-scale text encoding vectors represent the semantic information of the corresponding character and its adjacent characters, the semantic information of the word segment to which the corresponding character is located and its adjacent words, and the language information of the corresponding character.

[0016] The speech spectrum corresponding to the cross-language text is determined based on the multi-scale text encoding vectors and target speaker embedding vectors corresponding to each character in the cross-language text.

[0017] Optionally, based on the language embedding vectors corresponding to each language in the cross-language text and the character embedding vectors corresponding to each character in the cross-language text, the character encoding vectors corresponding to each character in the cross-language text are determined, including:

[0018] Based on the language embedding vectors of each language contained in the cross-language text, determine the convolution parameters corresponding to each language in the cross-language text.

[0019] By using the convolution parameters corresponding to each language in the cross-language text, the character embedding vectors corresponding to each language in the cross-language text are convolved to obtain the character encoding vectors corresponding to each character in the cross-language text.

[0020] Optionally, based on the text encoding vectors corresponding to each character in the cross-language text and the cross-language context-related vectors corresponding to each word segment in the cross-language text, a multi-scale text encoding vector corresponding to each character in the cross-language text is determined, including:

[0021] For each word contained in the cross-language text:

[0022] Copy the cross-language context-related vector corresponding to the word segment to the number of cross-language context-related vectors contained in the word segment;

[0023] The cross-language context-related vectors containing the number of characters in the word segment are concatenated with the text encoding vectors corresponding to each character in the word segment to obtain the multi-scale text encoding vectors corresponding to each character in the word segment.

[0024] This allows us to obtain the multi-scale text encoding vectors corresponding to each character in the cross-language text.

[0025] Optionally, based on the multi-scale text encoding vectors and target speaker embedding vectors corresponding to each character in the cross-lingual text, the speech spectrum corresponding to the cross-lingual text is determined, including:

[0026] Traverse the words in the cross-language text in the order they appear. For the currently traversed word:

[0027] Based on the multi-scale text encoding vectors corresponding to all characters in the segment and the spectral encoding vector at the decoding time corresponding to the previous segment, the alignment weights corresponding to all characters in the segment are determined. If the segment is the first segment in a cross-language text, the spectral encoding vector at the decoding time corresponding to the previous segment is a preset value. The alignment weights represent the position information of the corresponding characters in the speech spectrum.

[0028] The multi-scale text encoding vectors corresponding to all characters in the segmented word are concatenated with the target speaker embedding vector to obtain the concatenated vectors corresponding to all characters in the segmented word.

[0029] The concatenation vectors corresponding to all characters in the segmented word and the alignment weights corresponding to all characters in the segmented word are weighted and summed to obtain the spectral encoding vector at the decoding time corresponding to the segmented word.

[0030] Based on the spectral encoding vector at the decoding time corresponding to the word segment, determine the speech spectrum frame at the decoding time corresponding to the word segment;

[0031] At the end of the traversal, the speech spectrum frames at the decoding time corresponding to each word in the cross-language text are combined to form the speech spectrum corresponding to the cross-language text.

[0032] Optionally, obtain the cross-language text, the target speaker embedding vector, and the language embedding vectors corresponding to each language contained in the cross-language text; determine the character embedding vectors corresponding to each character contained in the cross-language text; and determine the speech spectrum corresponding to the cross-language text based on the language embedding vectors corresponding to each language contained in the cross-language text, the character embedding vectors corresponding to each character contained in the cross-language text, and the target speaker embedding vector, including:

[0033] The pre-trained cross-linguistic spectrum synthesis model is used to process cross-linguistic text, target speaker embedding vector, and language embedding vectors corresponding to each language contained in the cross-linguistic text to obtain the speech spectrum corresponding to the cross-linguistic text. The cross-linguistic spectrum synthesis model is obtained by using multilingual monolingual text, the labeled speech spectrum corresponding to the multilingual monolingual text, and the speaker embedding vector.

[0034] A method for training a speech synthesis model includes:

[0035] Multiple cross-linguistic corpora are obtained by using any of the cross-linguistic corpus synthesis methods mentioned above, and used as the first training corpus;

[0036] Obtain the second training corpus, which includes multilingual monolingual texts and the labeled speech spectra corresponding to the multilingual monolingual texts;

[0037] The pre-constructed acoustic model is trained using the first and second training corpora to obtain the trained target acoustic model.

[0038] The initial vocoder is trained using the speech spectrum corresponding to multilingual monolingual text and the labeled speech corresponding to multilingual monolingual text, and the target vocoder is obtained.

[0039] The speech synthesis model consists of a target acoustic model and a target vocoder.

[0040] Optionally, a pre-built acoustic model is trained using a first training corpus and a second training corpus, including:

[0041] Abnormal speech spectra are determined from the speech spectra contained in the first training corpus. The cross-language corpus corresponding to the abnormal speech spectra is then removed from the first training corpus to obtain the first training corpus after removal.

[0042] The texts in the first and second training corpora after filtering are converted into phoneme sequences, respectively, to obtain the first and second training corpora after phoneme sequence conversion;

[0043] The acoustic model is trained using the first and second training corpora, which are converted into phoneme sequences.

[0044] A cross-language corpus synthesis device, comprising:

[0045] The information acquisition module is used to acquire cross-language text, target speaker embedding vector, and language embedding vectors corresponding to each language contained in the cross-language text. The cross-language text is composed of word segments in multiple languages, each word segment contains at least one character, the language embedding vector represents the language information of the corresponding character and / or word segment, and the target speaker embedding vector represents the timbre information of the target speaker.

[0046] The character embedding vector determination module is used to determine the character embedding vector corresponding to each character contained in cross-language text;

[0047] The speech spectrum determination module is used to determine the speech spectrum of the cross-language text based on the language embedding vectors corresponding to each language contained in the cross-language text, the character embedding vectors corresponding to each character contained in the cross-language text, and the target speaker embedding vector.

[0048] The cross-linguistic corpus determination module is used to determine cross-linguistic corpora composed of speech spectra and cross-linguistic text.

[0049] A speech synthesis model training device, comprising:

[0050] The first training corpus acquisition module is used to obtain multiple cross-language corpora using any of the cross-language corpus synthesis methods mentioned above, and use them as the first training corpus.

[0051] The second training corpus acquisition module is used to acquire the second training corpus, which includes multilingual monolingual text and the labeled speech spectrum corresponding to the multilingual monolingual text;

[0052] The acoustic model training module is used to train a pre-constructed acoustic model using the first training corpus and the second training corpus to obtain a trained target acoustic model.

[0053] The vocoder training module is used to train an initial vocoder using the speech spectrum corresponding to multilingual monolingual text and the labeled speech corresponding to multilingual monolingual text, and to obtain a target vocoder.

[0054] The speech synthesis model determination module is used to compose a speech synthesis model consisting of a target acoustic model and a target vocoder.

[0055] As described above, the cross-language corpus synthesis method provided in this application first obtains the cross-language text, the target speaker embedding vector, and the language embedding vectors corresponding to each language contained in the cross-language text. Then, it determines the character embedding vectors corresponding to each character contained in the cross-language text. Next, based on the language embedding vectors corresponding to each language contained in the cross-language text, the character embedding vectors corresponding to each character contained in the cross-language text, and the target speaker embedding vector, it determines the speech spectrum corresponding to the cross-language text. Finally, the speech spectrum and the cross-language text constitute the cross-language corpus. This application can synthesize the speech spectrum of the same speaker switching languages ​​based on the character embedding vectors corresponding to each character contained in the cross-language text, the character embedding vectors corresponding to each character contained in the cross-language text, and the target speaker embedding vector, thereby obtaining the cross-language corpus composed of the speech spectrum and the cross-language text. Subsequently, a speech synthesis model with higher naturalness of synthesized speech can be constructed based on the obtained cross-language corpus. Attached Figure Description

[0056] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0057] Figure 1 A flowchart illustrating the cross-language corpus synthesis method provided in this application embodiment;

[0058] Figure 2 A flowchart illustrating the speech synthesis model training method provided in this application embodiment;

[0059] Figure 3 This is a schematic diagram of the cross-language corpus synthesis device provided in the embodiments of this application;

[0060] Figure 4 This is a schematic diagram of the structure of the speech synthesis model training device provided in the embodiments of this application;

[0061] Figure 5 Hardware structure block diagram of the cross-language corpus synthesis device provided in the embodiments of this application

[0062] Figure 6 This is a hardware structure block diagram of the speech synthesis model training device provided in an embodiment of this application. Detailed Implementation

[0063] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0064] The existing training corpus is a multilingual monolingual corpus, that is, there are multiple languages in the corpus, but each sentence included in the corpus has only one language. The speech synthesis model trained with this training corpus has a relatively good speech synthesis effect on monolingual texts. However, when performing speech synthesis on cross-lingual texts, synthesis failures often occur, or there are multiple different timbres for the same sentence, resulting in a relatively low naturalness of the synthesized speech.

[0065] In view of this, the inventors of this case conducted in-depth research and finally proposed a cross-lingual corpus synthesis method capable of synthesizing the speech spectrum corresponding to cross-lingual texts, and proposed a method for training a speech synthesis model with a cross-lingual corpus and a multilingual monolingual corpus obtained based on the cross-lingual corpus synthesis method. Next, the cross-lingual corpus synthesis method provided by the present application will be introduced in detail through the following embodiments.

[0066] Please refer to Figure 1 , which shows a schematic flowchart of the cross-lingual corpus synthesis method provided by the embodiments of the present application. The cross-lingual corpus synthesis method may include:

[0067] Step S101, obtain a cross-lingual text, a target speaker embedding vector, and language embedding vectors corresponding to each language included in the cross-lingual text. [[ID= sixteen]]

[0068] Among them, the cross-lingual text is composed of word segments in multiple languages, each word segment contains at least one character, the language embedding vector represents the language information of the corresponding character and / or word segment belonging to the language, and the target speaker embedding vector represents the timbre information of the target speaker.

[0069] For example, the cross-lingual text is "hello world", this cross-lingual text includes 3 word segments: "hello", "世", and "界". Taking the word segment "hello" as an example, this word segment contains 5 characters: "h", "e", "l", "l", and "o". Taking the word segment "世" as an example, this word segment contains 4 characters: "s", "h", "i", and "4 (representing the fourth tone of "shi").

[0070] In "hello world", there are two languages, Chinese and English. Among them, the segmented words "世界" and the characters contained therein correspond to the Chinese language embedding vector, and the segmented word "hello" and the characters contained therein correspond to the English language embedding vector.

[0071] Step S102: Determine the character embedding vectors corresponding to each character included in the cross-lingual text.

[0072] Optionally, the process of "determining the character embedding vectors corresponding to each character included in the cross-lingual text" in this step includes: Latinizing the cross-lingual text. For example, mapping the Chinese text to pinyin and tones, and keeping the English unchanged, to obtain the Latinized cross-lingual text, and converting the Latinized cross-lingual text into vector form to obtain the character embedding vectors corresponding to each character included in the cross-lingual text.

[0073] For example, if the cross-lingual text is "hello world", the text obtained after Latinizing it is: "hello shi4jie4". There are 13 characters in this Latinized cross-lingual text. This step can convert these 13 characters into vector form respectively to obtain 13 character embedding vectors.

[0074] Step S103: Determine the speech spectrum corresponding to the cross-lingual text according to the language embedding vectors corresponding to each language included in the cross-lingual text, the character embedding vectors corresponding to each character included in the cross-lingual text, and the target speaker embedding vector.

[0075] This step can combine the language embedding vector, the character embedding vector and the target speaker embedding vector to obtain the speech spectrum of language switching of the same speaker (i.e., the target speaker). Here, language switching means that there are multiple languages in a sentence and different languages can be switched arbitrarily.

[0076] Step S104: Compose a cross-lingual corpus from the speech spectrum and the cross-lingual text.

[0077] The cross-language corpus synthesis method provided in this application first obtains the cross-language text, the target speaker's embedding vector, and the language embedding vectors corresponding to each language contained in the cross-language text. Then, it determines the character embedding vectors corresponding to each character in the cross-language text. Next, based on the language embedding vectors corresponding to each language in the cross-language text, the character embedding vectors corresponding to each character in the cross-language text, and the target speaker's embedding vector, it determines the speech spectrum corresponding to the cross-language text. Finally, the speech spectrum and the cross-language text are combined to form the cross-language corpus. This application can synthesize the speech spectrum of the same speaker switching languages ​​based on the character embedding vectors corresponding to each character in the cross-language text, the character embedding vectors corresponding to each character in the cross-language text, and the target speaker's embedding vector. This yields the cross-language corpus composed of the speech spectrum and the cross-language text, which can then be used to construct a more natural-sounding speech synthesis model.

[0078] In one embodiment of this application, the process of "step S103, determining the speech spectrum corresponding to the cross-language text based on the language embedding vectors corresponding to each language contained in the cross-language text, the character embedding vectors corresponding to each character contained in the cross-language text, and the target speaker embedding vector" is described in detail.

[0079] Specifically, the process of "step S103, determining the speech spectrum corresponding to the cross-language text based on the language embedding vectors corresponding to each language in the cross-language text, the character embedding vectors corresponding to each character in the cross-language text, and the target speaker embedding vector" includes:

[0080] Step A1: Based on the language embedding vectors corresponding to each language in the cross-language text and the character embedding vectors corresponding to each character in the cross-language text, determine the character encoding vectors corresponding to each character in the cross-language text.

[0081] Among them, the character encoding vector represents the semantics and language information of the corresponding character.

[0082] In this step, the character embedding vectors corresponding to each character in the cross-language text can be combined with their corresponding language embedding vectors to obtain character encoding vectors that contain character semantics and language information.

[0083] In an optional embodiment, this step may include: determining the convolution parameters corresponding to each language in the cross-language text based on the language embedding vectors corresponding to each language in the cross-language text; and performing convolution processing on the character embedding vectors corresponding to each character in the cross-language text using the convolution parameters corresponding to each language in the cross-language text to obtain the character encoding vectors corresponding to each character in the cross-language text.

[0084] Specifically, this step can be implemented based on a first preprocessing module consisting of several layers of one-dimensional convolution and a high-speed network. This first preprocessing module includes a parameter generation network for generating convolution parameters.

[0085] Here, the process by which the parameter generation network "determines the convolution parameters corresponding to each language in the cross-language text based on the language embedding vectors corresponding to each language in the cross-language text" includes: In order to achieve controllable cross-language parameter sharing and prevent the generation of highly language-specific parameters, the parameter generation network first reduces the dimensionality of the language embedding vectors corresponding to each language using a first fully connected layer, and then uses two second fully connected layers to map them to the convolution kernels and bias values ​​corresponding to each language, respectively. The convolution kernels and bias values ​​corresponding to each language are the convolution parameters corresponding to each language.

[0086] The first preprocessing module in this step can perform convolution processing on the character embedding vectors corresponding to the characters in each language of the cross-language text, based on the convolution parameters corresponding to each language, to obtain the character encoding vectors corresponding to each character in the cross-language text. For example, based on the convolution parameters corresponding to Chinese, the character embedding vectors corresponding to the characters in the word segments “shi4” and “jie4” are convolved to obtain the character encoding vectors corresponding to the characters in the word segments “shi4” and “jie4”. Similarly, based on the convolution parameters corresponding to English, the character embedding vectors corresponding to the characters in “hello” are convolved to obtain the character encoding vectors corresponding to the characters in “hello”. Thus, a total of 13 character encoding vectors are obtained.

[0087] Step A2: Determine the text encoding vector corresponding to each character in the cross-language text based on the character encoding vector corresponding to each character.

[0088] The text encoding vector represents the semantic information of the corresponding character and its preceding and following adjacent characters, as well as the language information of the corresponding character.

[0089] This step can be implemented based on a recurrent neural network shared by various languages. Specifically, the character encoding vectors corresponding to each character included in the cross-language text are input into the recurrent neural network. In the recurrent neural network, the output at each moment is used as the final text encoding vector for the corresponding character, thereby obtaining the text encoding vectors corresponding to each character included in the cross-language text.

[0090] Step A3: Determine the cross-language context-related vectors corresponding to each token included in the cross-language text.

[0091] Among them, the cross-language context-related vector represents the semantic information of the corresponding token and its adjacent tokens before and after.

[0092] Optionally, this step can be implemented based on the pre-trained cross-language language model BERT. This model is a pre-trained sequence-to-sequence model that can be trained using a large number of multi-language texts including language switches, as well as the cross-language context-related vectors corresponding to each character included in the annotated multi-language texts. During the training process, some texts are randomly masked to force the model to generate context-related vectors with certain predictive capabilities.

[0093] After the model training is completed, this step performs token segmentation on the cross-language text to obtain the latinized text segmented into tokens. For example, "hello世界" is segmented into three tokens: "hello", "世", and "界". Then, the tokens are input into the BERT model to encode each token included in the cross-language text into a cross-language context-related vector. For example, the cross-language context-related vectors corresponding to the three tokens "hello", "世", and "界" are obtained.

[0094] Optionally, this step can perform token segmentation on the cross-language text through a vocabulary and Byte Pair Encoding (BPE) algorithm. Of course, this tokenization method is only an example. In addition, other tokenization methods can also be used in this embodiment for token segmentation.

[0095] Step A4: Determine the multi-scale text encoding vectors corresponding to each character included in the cross-language text according to the text encoding vectors corresponding to each character included in the cross-language text and the cross-language context-related vectors corresponding to each token included in the cross-language text.

[0096] Among them, the multi-scale text encoding vector represents the semantic information of the corresponding character and its adjacent characters before and after, the semantic information of the token where the corresponding character is located and its adjacent tokens before and after, and the language information of the corresponding character.

[0097] It should be noted that, in this embodiment, "adjacent characters" can be one character that is adjacent to the front and one character that is adjacent to the back, or preferably, multiple characters that are adjacent to the front and multiple characters that are adjacent to the back; similarly, "adjacent word segments" can be one word segment that is adjacent to the front and one word segment that is adjacent to the back, or preferably, multiple words segment that are adjacent to the front and multiple words that are adjacent to the back.

[0098] In this embodiment, the cross-language context-related vectors are at the word segmentation level; that is, the number of cross-language context-related vectors is equal to the number of words contained in the cross-language text. The text encoding vectors, on the other hand, are at the character level; that is, the number of text encoding vectors is equal to the number of characters contained in the cross-language text. Since the number of cross-language context-related vectors and text encoding vectors is unequal, this step can be performed on each word in the cross-language text. First, the cross-language context-related vector corresponding to that word is copied to the number of cross-language context-related vectors corresponding to the number of characters contained in that word. Then, these number of cross-language context-related vectors corresponding to the number of characters contained in that word are concatenated with the text encoding vectors corresponding to each character in that word, resulting in multi-scale text encoding vectors corresponding to each character in that word. This process is repeated sequentially to obtain the multi-scale text encoding vectors corresponding to each character in the cross-language text.

[0099] For example, for the word segment "hello", the cross-language context-related vector corresponding to the whole "hello" can be repeated 5 times to obtain 5 cross-language context-related vectors, which serve as the cross-language context-related vectors corresponding to "h", "e", "l", "l", and "o" respectively. Then, the cross-language context-related vector corresponding to "h" is concatenated to the text encoding vector corresponding to "h" to obtain the multi-scale text encoding vector corresponding to "h". Similarly, the cross-language context-related vector corresponding to "e" is concatenated to the text encoding vector corresponding to "e" to obtain the multi-scale text encoding vector corresponding to "e". The cross-language context-related vector corresponding to the first "l" is concatenated to the text encoding vector corresponding to the first "l" to obtain the multi-scale text encoding vector corresponding to the first "l". The cross-language context-related vector corresponding to the second "l" is concatenated to the text encoding vector corresponding to the second "l" to obtain the multi-scale text encoding vector corresponding to the second "l" to obtain the multi-scale text encoding vector corresponding to the second "l" to obtain the multi-scale text encoding vector corresponding to the second "l" to obtain the multi-scale text encoding vector corresponding to the second "l" to obtain the multi-scale text encoding vector corresponding to the second "l" to obtain the multi-scale text encoding vector corresponding to the second "l" to obtain the multi-scale text encoding vector corresponding to the second "o" to obtain the multi-scale text encoding vector corresponding to the second "o" to obtain the multi-scale text encoding vector corresponding to the second "o".

[0100] Step A5: Determine the speech spectrum corresponding to the cross-language text based on the multi-scale text encoding vector and the target speaker embedding vector corresponding to each character in the cross-language text.

[0101] In this step, the multi-scale text encoding vector is a text-level vector, while the target speaker embedding vector is a speech-level vector. Combining these two vectors yields a speech spectrum that contains only the timbre of the target speaker.

[0102] In an optional embodiment, this step can traverse each word segment in the cross-language text in the order of their occurrence. For the currently traversed word segment, the multi-scale text encoding vectors corresponding to all characters contained in the word segment are used as the multi-scale text encoding vectors at the encoding time corresponding to the word segment. Based on the multi-scale text encoding vectors at the encoding time corresponding to the word segment and the spectral encoding vector at the decoding time corresponding to the previous word segment, the alignment weights corresponding to all characters contained in the word segment are determined.

[0103] If the segmented word is the first segmented word contained in the cross-language text, such as "hello" in "hello world", then the spectral encoding vector at the previous decoding time is a preset value, and the alignment weight represents the position information of the corresponding character in the speech spectrum.

[0104] After obtaining the alignment weights corresponding to all characters in the segmented word, this step concatenates the multi-scale text encoding vectors corresponding to all characters in the segmented word with the target speaker embedding vector to obtain the concatenated vectors corresponding to all characters in the segmented word.

[0105] Then, the concatenation vectors corresponding to all characters in the segment and the alignment weights corresponding to all characters in the segment can be weighted and summed to obtain the spectral encoding vector at the decoding time corresponding to the segment.

[0106] Finally, this step determines the speech spectrum frame at the decoding time corresponding to the segmented word based on the spectrum encoding vector at the decoding time corresponding to the segmented word.

[0107] Optionally, the spectral encoding vector at the decoding time corresponding to the word segment can be input into two stacked recurrent neural networks for processing to obtain the speech spectrum frame at the decoding time corresponding to the word segment.

[0108] For each word segment contained in the cross-language text, the processing procedure described in this embodiment is applied sequentially to obtain the speech spectrum frames at the decoding time corresponding to each word segment contained in the cross-language text. Thus, at the end of the traversal, a speech spectrum composed of the speech spectrum frames at the decoding time corresponding to each word segment contained in the cross-language text is obtained, which is the speech spectrum corresponding to the cross-language text.

[0109] In one possible implementation of this application, this embodiment proposes a multilingual monolingual corpus training method that can generate a cross-lingual spectrum synthesis model for cross-lingual speech spectra, and uses the cross-lingual spectrum synthesis model to generate corpus with language switching.

[0110] Specifically, a cross-lingual spectral synthesis model can be pre-trained based on existing multilingual monolingual texts and their corresponding annotated speech spectra, supplemented by speaker embedding vectors. Here, the monolingual text can be, for example, texts containing only European Spanish, or only Chinese, or only English, etc.

[0111] Therefore, the above steps S101 to S103 may include: using a pre-trained cross-language spectrum synthesis model to process the cross-language text, the target speaker embedding vector, and the language embedding vectors corresponding to each language contained in the cross-language text, to obtain the speech spectrum corresponding to the cross-language text.

[0112] In an optional embodiment, the above-described cross-lingual spectral synthesis model includes a language embedding encoder, a cross-lingual pre-trained language model, and a decoder.

[0113] The aforementioned language embedding encoder includes a character embedding layer, a first preprocessing module in step A1, and a recurrent neural network in step A2. The character embedding layer is used to determine the character embedding vector corresponding to each character in the cross-language text. The first preprocessing module is used to determine the character encoding vector corresponding to each character in the cross-language text based on the language embedding vectors corresponding to each language and the character embedding vectors corresponding to each character in the cross-language text. The recurrent neural network is used to determine the text encoding vector corresponding to each character in the cross-language text based on the character encoding vectors corresponding to each character in the cross-language text.

[0114] Specifically, the cross-language text is first Latinized and then directly input into the language embedding encoder as characters. In the language embedding encoder, each character in the cross-language text is first mapped to a character embedding vector through a character embedding layer, and then enters the first preprocessing module. The first preprocessing module encodes the character embedding vectors corresponding to each character into character encoding vectors, and then enters the recurrent neural network shared by all languages ​​to obtain the text encoding vectors corresponding to each character in the cross-language text.

[0115] The aforementioned cross-language pre-trained language model includes the BERT model in step A3. In this embodiment, the ability of the BERT model to generate context-related vectors with certain predictive capabilities can be utilized to improve the language switching and transfer capabilities of the cross-language spectrum synthesis model.

[0116] The cross-language pre-trained language model first uses the BERT model to determine the cross-language context-related vectors corresponding to each word in the cross-language text. Then, based on the text encoding vectors corresponding to each character in the cross-language text and the cross-language context-related vectors corresponding to each word in the cross-language text, it determines the multi-scale text encoding vectors corresponding to each character in the cross-language text.

[0117] It is worth noting that when training a cross-lingual spectrum synthesis model, all weights in the cross-lingual pre-trained language model need to be frozen. That is, in the cross-lingual pre-trained language model, only the forward propagation process is performed, and no back propagation or weight updates are performed.

[0118] The decoder described above consists of a position-sensitive attention module and two stacked recurrent network layers.

[0119] The sensitive attention module includes a second preprocessing module and an attention recurrent network. For the traversal process mentioned in step A5, the sensitive attention module first uses the second preprocessing module to determine the alignment weights corresponding to all characters in the segmented word based on the multi-scale text encoding vectors corresponding to all characters in the segmented word and the spectral encoding vector at the decoding time corresponding to the previous segmented word. Then, it concatenates the multi-scale text encoding vectors corresponding to all characters in the segmented word with the target speaker embedding vector to obtain the concatenated vector corresponding to all characters in the segmented word. Finally, the attention recurrent network performs a weighted summation of the concatenated vectors corresponding to all characters in the segmented word and the alignment weights corresponding to all characters in the segmented word to obtain the spectral encoding vector at the decoding time corresponding to the segmented word.

[0120] The recurrent network layer is used to determine the speech spectrum frame at the decoding time corresponding to the word segment based on the spectrum encoding vector at the decoding time corresponding to the word segment.

[0121] Specifically, in the decoder, for the currently traversed word segment, the multi-scale text encoding vectors corresponding to all characters contained in the word segment, and the spectral encoding vector at the decoding time corresponding to the previous word segment, are first processed in the second preprocessing module to obtain the alignment weights corresponding to all characters contained in the word segment. After concatenating the multi-scale text encoding vectors corresponding to all characters contained in the word segment with the target speaker embedding vector, the vectors are then fed into the attention recurrent network to generate the spectral encoding vector.

[0122] Similar to the existing Tacotron-2 decoder, the position-sensitive attention module is responsible for determining the information content of each multi-scale text encoding vector at each decoding time step. It takes the spectral encoding vector from the previous time step (i.e., the spectral encoding vector at the decoding time step corresponding to the word segment preceding this one) as input and outputs the alignment weight of each multi-scale text encoding vector, that is:

[0123]

[0124] in, W, V, and U are the parameters that need to be trained in the cross-language spectrum synthesis model (these four parameters are fixed values ​​during application after the model is trained), b is the bias value, i and j are the decoding and encoding times respectively, and s i-1 h represents the spectral encoding vector of the previous time step i-1. j The multi-scale text encoding vector at the current encoding time j is represented by score(s). i-1 ,h j ) represents the alignment weight of the multi-scale text encoding vector j at decoding time i (in this embodiment, score(s) i-1 ,h j ) can also be represented as α i,j That is, score(s) i-1 ,h j )=α i,j ), f i =F*α i-1 The total alignment weight α represents the total alignment weights at the previous time step. i-1 The positional features obtained after convolution F, α i-1 This refers to the alignment weights of all characters contained in the preceding segment of the word.

[0125] After obtaining the alignment weights corresponding to all characters in the segmented word, the multi-scale text encoding vectors corresponding to all characters in the segmented word are weighted and summed based on the alignment weights, and then concatenated with the target speaker embedding vector as the input to the subsequent recurrent network layers. Two stacked recurrent network layers in the decoder are used for the aforementioned position-sensitive attention and to generate the decoder output at the current moment. The decoder output is the speech spectrum frame corresponding to the segmented word. The decoder generates the complete speech spectrum time-by-time in an autoregressive manner.

[0126] Through the above process, cross-language corpus consisting of cross-language text and corresponding speech spectra can be obtained. Experiments have shown that the cross-language spectrum synthesis model trained in this application embodiment can synthesize speech spectra with language switching relatively stably when cross-language text is input.

[0127] This application enables the distillation of a speech synthesis model (a lightweight end-to-end model) by collecting multiple cross-linguistic corpora.

[0128] For details, see Figure 2 The diagram shown is a flowchart illustrating the speech synthesis model training method provided in this application embodiment. The speech synthesis model training method may include:

[0129] Step S201: Use a cross-language corpus synthesis method to obtain multiple cross-language corpora, which are used as the first training corpus.

[0130] This step can use any of the above-mentioned cross-lingual corpus synthesis methods to obtain multiple cross-lingual corpora. It should be understood that multiple cross-lingual corpora here refer to different cross-lingual corpora. For example, multiple cross-lingual corpora may include cross-lingual corpora 1 to 3, wherein the speakers in the speech spectra corresponding to cross-lingual corpora 1 and 2 are the same, but the speakers in the speech spectra corresponding to cross-lingual corpora 3 are different; the cross-lingual texts corresponding to cross-lingual corpora 1 and 3 are the same, but the cross-lingual texts corresponding to cross-lingual corpora 2 are different; or, the speakers and cross-lingual texts corresponding to cross-lingual corpora 1 to 3 are different respectively.

[0131] Step S202: Obtain the second training corpus.

[0132] The second training corpus includes multilingual monolingual texts and the annotated speech spectra corresponding to the multilingual monolingual texts.

[0133] The multilingual monolingual text used here can be the same as or different from the multilingual monolingual text used when training the cross-lingual spectrum synthesis model.

[0134] Step S203: Train the pre-constructed acoustic model using the first training corpus and the second training corpus to obtain the trained target acoustic model.

[0135] In this step, the pre-built acoustic model is a Tacotron-2-like acoustic model, which includes an encoder, a decoder, and an attention mechanism. It takes the phonemes corresponding to the text as input and outputs the corresponding speech spectrum.

[0136] In one possible implementation, this step can convert the text in the first training corpus and the second training corpus into phoneme sequences, respectively, to obtain the first training corpus and the second training corpus after being converted into phoneme sequences. Then, the acoustic model is trained using the first training corpus and the second training corpus after being converted into phoneme sequences.

[0137] Optionally, an existing text-to-phoneme system can be used to convert the text in the first and second training corpora into phoneme sequences, respectively.

[0138] Optionally, to prevent the acoustic model from overfitting and to prevent the same combination of samples from appearing repeatedly, causing the acoustic model to remember the order of these samples and thus affecting its generalization ability, the overall training process of the acoustic model can be divided into multiple rounds. Each round trains the acoustic model using the first training corpus and the second training corpus. However, before training the acoustic model based on the first training corpus and the second training corpus in each round, the order is shuffled, and then the acoustic model is trained based on the shuffled first training corpus and the second training corpus.

[0139] In a preferred implementation, considering that the cross-lingual spectral synthesis model uses characters directly as input, which may produce pronunciation errors, and that the synthesized cross-lingual corpus may also contain synthesis defects such as word skipping, repetition, and failure to stop properly, in order to avoid cross-lingual corpus with synthesis defects or synthesis errors affecting the quality of the distilled speech synthesis model, the cross-lingual corpus can be automatically or manually screened by certain means before training the acoustic model based on the first training corpus and the second training corpus.

[0140] Specifically, the process of "training the pre-built acoustic model using the first training corpus and the second training corpus" can include:

[0141] Step B1: Determine abnormal speech spectra from the speech spectra contained in the first training corpus, and remove the cross-language corpus corresponding to the abnormal speech spectra from the first training corpus to obtain the first training corpus after removal.

[0142] Optionally, two screening metrics, namely the proportion of strip alignment weights and the monotonic alignment distance, can be used to determine abnormal speech spectra from the speech spectra contained in the first training corpus.

[0143] Specifically, for each speech spectrum contained in the first training corpus, two screening indicators are set with corresponding acceptable ranges (i.e., thresholds). If the speech spectrum exceeds the acceptable range, it is determined to be an abnormal speech spectrum.

[0144] Here, the proportion of alignment weight (PAW) is defined as:

[0145]

[0146] In the formula, PAW represents the calculated strip alignment weight ratio, 0≤i≤I-1, 0≤j≤J-1, c=J / I, d is a hyperparameter representing the width of the strip, which is optional, b=5, and I and J represent the total number of decoding and encoding times, respectively.

[0147] Optionally, the threshold corresponding to the strip alignment weight ratio can be set to 0.5. That is, for each speech spectrum contained in the first training corpus, if the strip alignment weight ratio (PAW) corresponding to the speech spectrum is less than the threshold of 0.5, then the cross-lingual corpus corresponding to that speech spectrum is deleted. Of course, the above threshold is only an example and is not intended to limit this application.

[0148] Referring to the soft monotonic alignment loss function in the prior art, the alignment distance set in this embodiment is defined as:

[0149] MAD=|||Δπ|+Δπ||+|||Δπ-1|+(Δπ-1)|| Formula (3)

[0150] Where, Δπ=π i -π i-1 , π i The calculation formula is as follows:

[0151]

[0152] Where, p i = {0, 1, 2, ..., J-1} represents the indices of the phonemes contained in the cross-language text, and Δπ refers to the difference between the decoding time i and the corresponding soft (soft) time of the encoding at decoding time i-1. i It refers to the soft (soft,) moment of the encoding corresponding to the decoding moment i (i.e., the phoneme index, obtained through attention).

[0153] Formula (3) can be used to calculate the continuity and monotonicity of the alignment of the cross-lingual spectrum synthesis model. Optionally, the threshold corresponding to the alignment distance can be set to 0.1. That is, for each speech spectrum contained in the first training corpus, if the alignment distance MAD corresponding to the speech spectrum is greater than the threshold of 0.1, the cross-lingual corpus corresponding to that speech spectrum is deleted. Of course, the above threshold is only an example and is not intended to limit this application.

[0154] This step, through the calculation of the two indicators mentioned above, can filter out higher-quality synthesized data for distillation of subsequent speech synthesis models.

[0155] Step B2: Convert the text in the first and second training corpora after filtering into phoneme sequences, respectively, to obtain the first and second training corpora after phoneme sequence conversion.

[0156] Step B3: Train the acoustic model using the first and second training corpora after converting the data into phoneme sequences.

[0157] The procedures for steps B2 and B3 can be referred to the description in the previous steps, and will not be repeated here.

[0158] Step S204: Train the initial vocoder using the speech spectrum corresponding to the multilingual monolingual text and the labeled speech corresponding to the multilingual monolingual text to obtain the target vocoder.

[0159] Optionally, the vocoder can be a HiFi-GAN, which takes the speech spectrum as input and outputs the final speech waveform.

[0160] The multilingual monolingual text in this step is the same as the multilingual monolingual text in step S202.

[0161] Step S205: A speech synthesis model is composed of a target acoustic model and a target vocoder.

[0162] Because the training corpus of acoustic models and vocoders contains cross-language data with language switching capabilities, ordinary acoustic models and vocoders can achieve cross-language speech synthesis, ultimately completing the distillation of synthesized speech with language switching capabilities.

[0163] Experiments have shown that the speech synthesis model provided in this application has the ability to synthesize speech across languages. When inputting cross-language text, it can synthesize speech that switches languages ​​within a single sentence, and the synthesized speech has a high degree of naturalness.

[0164] This application also provides a cross-language corpus synthesis device; please refer to [link to relevant documentation]. Figure 3 The diagram shows a schematic representation of the cross-language corpus synthesis device provided in an embodiment of this application. Figure 3 As shown, the cross-language corpus synthesis device may include: an information acquisition module 301, a character embedding vector determination module 302, a speech spectrum determination module 303, and a cross-language corpus determination module 304.

[0165] The information acquisition module 301 is used to acquire cross-language text, target speaker embedding vector, and language embedding vectors corresponding to each language contained in the cross-language text. The cross-language text is composed of word segments in multiple languages, each word segment contains at least one character, the language embedding vector represents the language information of the corresponding character and / or word segment, and the target speaker embedding vector represents the timbre information of the target speaker.

[0166] The character embedding vector determination module 302 is used to determine the character embedding vector corresponding to each character contained in the cross-language text.

[0167] The speech spectrum determination module 303 is used to determine the speech spectrum corresponding to the cross-language text based on the language embedding vectors corresponding to each language contained in the cross-language text, the character embedding vectors corresponding to each character contained in the cross-language text, and the target speaker embedding vector.

[0168] The cross-language corpus determination module 304 is used to determine cross-language corpora composed of speech spectrum and cross-language text.

[0169] In summary, the working principle of the cross-language corpus synthesis device disclosed in this embodiment is the same as that of the cross-language corpus synthesis method described in the above embodiments. For details, please refer to the description in the foregoing embodiments, which will not be repeated here.

[0170] This application also provides a speech synthesis model training device. Please refer to [link to relevant documentation]. Figure 4 The diagram shows a schematic representation of the speech synthesis model training device provided in an embodiment of this application. Figure 4 As shown, the speech synthesis model training device may include: a first training corpus acquisition module 401, a second training corpus acquisition module 402, an acoustic model training module 403, a vocoder training module 404, and a speech synthesis model determination module 405.

[0171] The first training corpus acquisition module 401 is used to obtain multiple cross-language corpora using the aforementioned cross-language corpus synthesis method, which are then used as the first training corpus.

[0172] The second training corpus acquisition module 402 is used to acquire the second training corpus, which includes multilingual monolingual text and the labeled speech spectrum corresponding to the multilingual monolingual text.

[0173] The acoustic model training module 403 is used to train a pre-constructed acoustic model using a first training corpus and a second training corpus to obtain a trained target acoustic model.

[0174] The vocoder training module 404 is used to train an initial vocoder using the speech spectrum corresponding to the multilingual monolingual text and the labeled speech corresponding to the multilingual monolingual text, and to obtain a target vocoder.

[0175] The speech synthesis model determination module 405 is used to compose a speech synthesis model consisting of a target acoustic model and a target vocoder.

[0176] In summary, the working principle of the speech synthesis model training device disclosed in this embodiment is the same as that of the speech synthesis model training method described in the previous embodiment. For details, please refer to the description in the foregoing embodiment, which will not be repeated here.

[0177] This application also provides a cross-language corpus synthesis device. Optionally, Figure 5 The hardware structure block diagram of the cross-language corpus synthesis device is shown, with reference to Figure 5 The hardware structure of the cross-language corpus synthesis device may include: at least one processor 501, at least one communication interface 502, at least one memory 503 and at least one communication bus 504.

[0178] In this embodiment of the application, the number of processor 501, communication interface 502, memory 503 and communication bus 504 is at least one, and processor 501, communication interface 502 and memory 503 communicate with each other through communication bus 504.

[0179] The processor 501 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention.

[0180] The memory 503 may include high-speed RAM, and may also include non-volatile memory, such as at least one disk storage device;

[0181] The memory 503 stores a program, and the processor 501 can call the program stored in the memory 503. The program is used for:

[0182] Obtain cross-language text, target speaker embedding vector, and language embedding vectors corresponding to each language contained in the cross-language text. The cross-language text is composed of word segments in multiple languages, each word segment contains at least one character, the language embedding vector represents the language information of the corresponding character and / or word segment, and the target speaker embedding vector represents the timbre information of the target speaker.

[0183] Determine the character embedding vectors corresponding to each character in the cross-language text;

[0184] The speech spectrum corresponding to the cross-language text is determined based on the language embedding vectors corresponding to each language contained in the cross-language text, the character embedding vectors corresponding to each character contained in the cross-language text, and the target speaker embedding vector.

[0185] The cross-linguistic corpus consists of speech spectrum and cross-linguistic text.

[0186] Optionally, the refined and extended functions of the program can be found in the description above.

[0187] This application also provides a speech synthesis model training device. Optionally, Figure 6 The hardware structure block diagram of the speech synthesis model training device is shown below. Figure 6 The hardware structure of the speech synthesis model training device may include: at least one processor 601, at least one communication interface 602, at least one memory 603 and at least one communication bus 604.

[0188] In this embodiment of the application, the number of processor 601, communication interface 602, memory 603 and communication bus 604 is at least one, and processor 601, communication interface 602 and memory 603 communicate with each other through communication bus 604.

[0189] The processor 601 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention.

[0190] The memory 603 may include high-speed RAM, and may also include non-volatile memory, such as at least one disk storage device;

[0191] The memory 603 stores a program, and the processor 601 can call the program stored in the memory 603. The program is used for:

[0192] Multiple cross-linguistic corpora were obtained using a cross-linguistic corpus synthesis method and used as the first training corpus;

[0193] Obtain the second training corpus, which includes multilingual monolingual texts and the labeled speech spectra corresponding to the multilingual monolingual texts;

[0194] The pre-constructed acoustic model is trained using the first and second training corpora to obtain the trained target acoustic model.

[0195] The initial vocoder is trained using the speech spectrum corresponding to multilingual monolingual text and the labeled speech corresponding to multilingual monolingual text, and the target vocoder is obtained.

[0196] The speech synthesis model consists of a target acoustic model and a target vocoder.

[0197] Optionally, the refined and extended functions of the program can be found in the description above.

[0198] This application also provides a readable storage medium storing a computer program thereon, which, when executed by a processor, implements the cross-language corpus synthesis method described above.

[0199] Optionally, the refined and extended functions of the program can be found in the description above.

[0200] This application also provides a readable storage medium storing a computer program thereon, which, when executed by a processor, implements the speech synthesis model training method described above.

[0201] Optionally, the refined and extended functions of the program can be found in the description above.

[0202] Finally, it should be noted that in this document, relational terms such as "second" and "etc." are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0203] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0204] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for synthesizing cross-linguistic corpora, characterized in that, include: Obtain cross-language text, target speaker embedding vector, and language embedding vectors corresponding to each language contained in the cross-language text. The cross-language text is composed of word segments under multiple languages, each word segment contains at least one character, the language embedding vector represents the language information of the corresponding character and / or word segment, and the target speaker embedding vector represents the timbre information of the target speaker. Determine the character embedding vector corresponding to each character contained in the cross-language text; The speech spectrum corresponding to the cross-language text is determined based on the language embedding vectors corresponding to each language contained in the cross-language text, the character embedding vectors corresponding to each character contained in the cross-language text, and the target speaker embedding vector. The cross-language corpus is composed of the speech spectrum and the cross-language text; Specifically, determining the speech spectrum corresponding to the cross-language text based on the language embedding vectors corresponding to each language contained in the cross-language text, the character embedding vectors corresponding to each character contained in the cross-language text, and the target speaker embedding vector includes: Based on the language embedding vectors corresponding to each language contained in the cross-language text and the character embedding vectors corresponding to each character contained in the cross-language text, the character encoding vectors corresponding to each character contained in the cross-language text are determined, wherein the character encoding vectors represent the semantics and language information of the corresponding characters. Based on the character encoding vectors corresponding to each character in the cross-language text, the text encoding vectors corresponding to each character in the cross-language text are determined, wherein the text encoding vectors represent the semantic information of the corresponding character and its preceding and following adjacent characters, as well as the language information of the corresponding character. Determine the cross-language context-related vectors corresponding to each word segment contained in the cross-language text, wherein the cross-language context-related vectors represent the semantic information of the corresponding word segment and the adjacent words before and after it; Based on the text encoding vectors corresponding to each character in the cross-language text and the cross-language context-related vectors corresponding to each word segment in the cross-language text, the multi-scale text encoding vectors corresponding to each character in the cross-language text are determined. The multi-scale text encoding vectors represent the semantic information of the corresponding character and its adjacent characters, the semantic information of the word segment in which the corresponding character is located and its adjacent words, and the language information of the corresponding character. The speech spectrum corresponding to the cross-language text is determined based on the multi-scale text encoding vectors corresponding to each character in the cross-language text and the target speaker embedding vector.

2. The cross-linguistic corpus synthesis method according to claim 1, characterized in that, The step of determining the character encoding vector corresponding to each character in the cross-language text based on the language embedding vector corresponding to each language in the cross-language text and the character embedding vector corresponding to each character in the cross-language text includes: Based on the language embedding vectors corresponding to each language contained in the cross-language text, determine the convolution parameters corresponding to each language contained in the cross-language text. By using the convolution parameters corresponding to each language in the cross-language text, the character embedding vectors corresponding to the characters in each language in the cross-language text are convolved to obtain the character encoding vectors corresponding to each character in the cross-language text.

3. The cross-linguistic corpus synthesis method according to claim 1, characterized in that, The step of determining the multi-scale text encoding vector corresponding to each character in the cross-language text based on the text encoding vector corresponding to each character in the cross-language text and the cross-language context-related vector corresponding to each word segment in the cross-language text includes: For each word segment contained in the cross-language text: Copy the cross-language context-related vector corresponding to the word segment to the number of cross-language context-related vectors contained in the word segment; The cross-language context-related vectors containing the number of characters in the word segment are concatenated with the text encoding vectors corresponding to each character in the word segment to obtain the multi-scale text encoding vectors corresponding to each character in the word segment. This allows us to obtain the multi-scale text encoding vectors corresponding to each character in the cross-language text.

4. The cross-linguistic corpus synthesis method according to claim 1, characterized in that, The step of determining the speech spectrum corresponding to the cross-language text based on the multi-scale text encoding vectors corresponding to each character in the cross-language text and the target speaker embedding vector includes: The text is traversed according to the order of the word segments contained in the cross-language text. For the currently traversed word segment: Based on the multi-scale text encoding vectors corresponding to all characters in the segmented word and the spectral encoding vector at the decoding time corresponding to the previous segmented word, the alignment weights corresponding to all characters in the segmented word are determined. If the segmented word is the first segmented word in the cross-language text, the spectral encoding vector at the decoding time corresponding to the previous segmented word is a preset value. The alignment weights represent the position information of the corresponding characters in the speech spectrum. The multi-scale text encoding vectors corresponding to all characters in the segmented word are concatenated with the target speaker embedding vector to obtain the concatenated vectors corresponding to all characters in the segmented word. The concatenation vectors corresponding to all characters in the segmented word and the alignment weights corresponding to all characters in the segmented word are weighted and summed to obtain the spectral encoding vector at the decoding time corresponding to the segmented word. Based on the spectral encoding vector at the decoding time corresponding to the word segment, determine the speech spectrum frame at the decoding time corresponding to the word segment; At the end of the traversal, the speech spectrum frames at the decoding time corresponding to each word in the cross-language text are combined to form the speech spectrum corresponding to the cross-language text.

5. The cross-linguistic corpus synthesis method according to claim 1, characterized in that, The process of obtaining cross-language text, target speaker embedding vector, and language embedding vectors corresponding to each language contained in the cross-language text, determining character embedding vectors corresponding to each character contained in the cross-language text, and determining the speech spectrum corresponding to the cross-language text based on the language embedding vectors corresponding to each language contained in the cross-language text, the character embedding vectors corresponding to each character contained in the cross-language text, and the target speaker embedding vector, includes: The cross-language spectrum synthesis model is pre-trained to process the cross-language text, the target speaker embedding vector, and the language embedding vectors corresponding to each language contained in the cross-language text to obtain the speech spectrum corresponding to the cross-language text. The cross-language spectrum synthesis model is obtained by training the multilingual monolingual text, the labeled speech spectrum corresponding to the multilingual monolingual text, and the speaker embedding vector.

6. A method for training a speech synthesis model, characterized in that, include: Multiple cross-linguistic corpora are obtained using the cross-linguistic corpus synthesis method as described in any one of claims 1 to 5, and used as the first training corpus; Obtain a second training corpus, wherein the second training corpus includes multilingual monolingual text and the labeled speech spectrum corresponding to the multilingual monolingual text; The pre-constructed acoustic model is trained using the first training corpus and the second training corpus to obtain the trained target acoustic model. The initial vocoder is trained using the speech spectrum corresponding to the multilingual monolingual text and the labeled speech corresponding to the multilingual monolingual text, and the target vocoder is obtained. The speech synthesis model consists of the target acoustic model and the target vocoder.

7. The speech synthesis model training method according to claim 6, characterized in that, Training a pre-built acoustic model using the first training corpus and the second training corpus includes: Abnormal speech spectra are determined from the speech spectra contained in the first training corpus, and the cross-language corpus corresponding to the abnormal speech spectra is screened out from the first training corpus to obtain the screened first training corpus. The text in the first and second training corpora after filtering is converted into phoneme sequences, respectively, to obtain the first and second training corpora after being converted into phoneme sequences; The acoustic model is trained using the first and second training corpora, which have been converted into phoneme sequences.

8. A cross-language corpus synthesis device, characterized in that, include: The information acquisition module is used to acquire cross-language text, target speaker embedding vector, and language embedding vectors corresponding to each language contained in the cross-language text. The cross-language text is composed of word segments under multiple languages, each word segment contains at least one character, the language embedding vector represents the language information of the corresponding character and / or word segment, and the target speaker embedding vector represents the timbre information of the target speaker. The character embedding vector determination module is used to determine the character embedding vector corresponding to each character contained in the cross-language text. The speech spectrum determination module is used to determine the character encoding vector corresponding to each character in the cross-language text based on the language embedding vectors corresponding to each language and the character embedding vectors corresponding to each character in the cross-language text. The character encoding vector represents the semantic and language information of the corresponding character. Based on the character encoding vectors corresponding to each character in the cross-language text, the module determines the text encoding vector corresponding to each character in the cross-language text. The text encoding vector represents the semantic information of the corresponding character and its preceding and following characters, as well as the language information of the corresponding character. Finally, the module determines the cross-language context-related vector corresponding to each word segment in the cross-language text. The cross-language context-related vector represents the semantic information of the corresponding word segment and its preceding and following words. Based on the text encoding vectors corresponding to each character in the cross-language text and the cross-language context-related vectors corresponding to each word in the cross-language text, the multi-scale text encoding vectors corresponding to each character in the cross-language text are determined. The multi-scale text encoding vectors represent the semantic information of the corresponding character and its preceding and following characters, the semantic information of the word segment in which the corresponding character is located and its preceding and following words, and the language information of the corresponding character. Based on the multi-scale text encoding vectors corresponding to each character in the cross-language text and the target speaker embedding vector, the speech spectrum corresponding to the cross-language text is determined. A cross-language corpus determination module is used to compose the cross-language corpus from the speech spectrum and the cross-language text.

9. A speech synthesis model training device, characterized in that, include: The first training corpus acquisition module is used to obtain multiple cross-language corpora using the cross-language corpus synthesis method as described in any one of claims 1 to 5, and use them as the first training corpus; The second training corpus acquisition module is used to acquire the second training corpus, wherein the second training corpus includes multilingual monolingual text and the labeled speech spectrum corresponding to the multilingual monolingual text; The acoustic model training module is used to train a pre-constructed acoustic model using the first training corpus and the second training corpus to obtain a trained target acoustic model. The vocoder training module is used to train an initial vocoder using the speech spectrum corresponding to the multilingual monolingual text and the labeled speech corresponding to the multilingual monolingual text, and to obtain a target vocoder. A speech synthesis model determination module is used to compose the speech synthesis model from the target acoustic model and the target vocoder.