A cross-lingual speech cloning method, device and network-attached storage device

CN122531356APending Publication Date: 2026-08-07SHENZHEN GREEN CONNECTION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN GREEN CONNECTION TECH CO LTD
Filing Date
2026-05-08
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0003]目前,跨语种声音克隆主要是通过多语种语音大模型(如cosyvoice)实现,然而,实践发现,当前语音大模型的推理过程主要通过speaker-embedding和语言种类标记来实现跨语言声音克隆,音色相似度欠佳,或者,采用zero-short形式通过LLM自回归续写的方式得到更强的克隆音色相似度,但其对语种的扩展支持性不佳,难以兼顾目标音色的音色相似度和语种多样性,使得目标音色的克隆准确性和克隆可靠性较低

Benefits of technology

本发明能够将语种覆盖语料及音色覆盖语料输入至后端模型训练器中训练出多语种及多音色的后端基础模型,并根据目标音色语料通过后端模型训练器对后端基础模型微调训练出所述目标音色的多语种模型,在多语种语音合成引擎中将待克隆语种和待合成文本输入至多语种前端文本分析器中分析出发音符号序列从而输入至多语种模型中克隆出目标音色的语音波形,能够提高多语种模型的语种覆盖语料及音色覆盖语料的多样性、针对性及全面性,有利于实现兼顾目标音色的克隆音色相似度和语种多样性,从而有利于提高目标音色的克隆准确性和克隆可靠性,进而有利于提升目标音色的多语种声音克隆效果,此外,通过先训练多语种多音色后端基础模型再微调训练针对目标音色的语音克隆模型的方式,还能够减少模型训练数据量以及降低算力需求,从而有利于提高语音克隆模型的训练效率和推理速度,进而有利于提高目标音色的克隆效率和克隆速度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122531356A_ABST
    Figure CN122531356A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of speech processing, and discloses a cross-language speech cloning method and device and a network-attached storage device, comprising: inputting language coverage corpus and timbre coverage corpus into a backend model trainer to train a multilingual and multi-timbre backend base model, and fine-tuning the backend base model through the backend model trainer according to target timbre corpus to train a multilingual model of the target timbre; inputting a to-be-cloned language and to-be-synthesized text into a multilingual front-end text analyzer in a multilingual speech synthesis engine to analyze phonetic symbol sequences and then input the phonetic symbol sequences into the multilingual model to clone speech waveforms of the target timbre, thereby improving the diversity, pertinence and comprehensiveness of the language coverage corpus and the timbre coverage corpus of the multilingual model, achieving a balance between the cloning timbre similarity and the language diversity of the target timbre, improving the cloning accuracy and cloning reliability of the target timbre, and further improving the multilingual sound cloning effect of the target timbre.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech processing technology, and in particular to a cross-language speech cloning method, apparatus, and network-attached storage device. Background Technology

[0002] Voice cloning is a technique that generates other speech content by providing a small amount of corpus, combining it with model fine-tuning, or using it as a prompt for a large model; further, cross-language voice cloning refers to the technique of providing a target timbre corpus in one language to generate speech in other languages ​​with that timbre.

[0003] Currently, cross-lingual voice cloning is mainly achieved through multilingual speech models (such as cosyvoice). However, practice has shown that the inference process of current speech models primarily relies on speaker-embedding and language tagging to achieve cross-lingual voice cloning, resulting in poor timbre similarity. Alternatively, a zero-short approach using LLM autoregressive continuation can achieve stronger timbre similarity, but its support for language expansion is poor, making it difficult to balance timbre similarity and language diversity, leading to low accuracy and reliability in target timbre cloning. Therefore, providing a new cross-lingual voice cloning method to improve the accuracy and reliability of target timbre cloning is particularly important. Summary of the Invention This invention provides a cross-language speech cloning method and apparatus, which can improve the accuracy and reliability of target timbre cloning, thereby enhancing the multilingual sound cloning effect of the target timbre.

[0004] To address the aforementioned technical problems, the first aspect of this invention discloses a cross-language voice cloning method, applied to a network-attached storage device, the method comprising: The determined language coverage corpus and timbre coverage corpus are input into the pre-trained model of the preset back-end model trainer for training, so as to obtain a multilingual and multi-timbre back-end basic model. When it is detected that speech cloning of the target speech corpus of the target timbre is required, the back-end basic model is adjusted and trained through the back-end model trainer according to the target speech corpus of the target timbre to obtain the multilingual model of the target timbre. In a pre-built multilingual speech synthesis engine, the multilingual text to be synthesized is input into a preset multilingual front-end text analyzer for analysis, and the phonetic symbol sequence of the multilingual text to be synthesized is obtained. The phonetic symbol sequence of the multilingual text to be synthesized is input into the multilingual model for cloning to obtain the speech cloning result of the target timbre, and the speech cloning result includes a speech waveform.

[0005] As an optional implementation, in the first aspect of the present invention, the language coverage corpus includes multilingual speech corpus for each of a plurality of similar timbre groups, and multilingual speech corpus for each of a plurality of identical timbre groups. The timbre-covered corpus includes the same language speech corpus of each of the multiple speakers; Furthermore, the process of inputting the determined language coverage corpus and timbre coverage corpus into the pre-trained model of the preset backend model trainer for training, to obtain a multilingual and multi-timbre backend basic model, includes: The determined language coverage corpus and timbre coverage corpus are input into the pre-trained model in the preset back-end model trainer. The pre-trained model learns the correlation between different language pronunciation symbols through the language coverage corpus and the pronunciation ability of different speakers through the timbre coverage corpus. Based on the correlation and the pronunciation ability, the common features of pronunciation symbols that are independent of timbre characteristics are abstracted to obtain the multilingual and multi-timbre back-end basic model.

[0006] As an optional implementation, in the first aspect of the present invention, the language coverage corpus is determined by the following method: Multiple similar timbre groups and multiple identical timbre groups are obtained; wherein, all timbres in each of the similar timbre groups are different and the similarity between all timbres is greater than or equal to a preset similarity, and all timbres in each of the identical timbre groups are the same; Obtain a first voice file for each group of similar timbres and a second voice file for each group of the same timbres, wherein the first voice file includes voice files with a number of languages ​​equal to a first preset number of languages, and the second voice file includes voice files with a number of languages ​​equal to a second preset number of languages, wherein the first preset number of languages ​​is greater than the second preset number of languages. For each group of similar timbres, based on the first speech file of the group of similar timbres, the multilingual speech corpus of the group of similar timbres is determined through the phonetic system of the language corresponding to the first speech file and the preset corpus format. For each group of the same timbre, based on the second speech file of the same timbre group, and through the phonetic system of the language corresponding to the second speech file, and the form of the speech corpus, the multilingual speech corpus of the same timbre group is determined; The multilingual speech corpora of all the aforementioned similar timbre groups, as well as the multilingual speech corpora of all the aforementioned identical timbre groups, are identified as language-covered corpora.

[0007] As an optional implementation, in the first aspect of the invention, the timbre coverage corpus is determined by the following method: Obtain a third speech file for each of the multiple speakers, and each of the third speech files corresponds to a language; For each speaker, based on the third speech file of that speaker, the same language speech corpus of that speaker is determined through the phonetic transcription system of the language corresponding to the third speech file and the corpus format; The same language speech data of all the speakers were identified as timbre-covered data.

[0008] As an optional implementation, in the first aspect of the invention, determining the same-language speech corpus for each speaker, based on a third speech file of that speaker and through the phonetic transcription system of the language corresponding to the third speech file, and the corpus format, includes: For each speaker, a corresponding speaker number is assigned, and the speaker's pronunciation symbol sequence is determined based on the third speech file of that speaker and the phonetic symbol system of the language corresponding to the third speech file. By integrating the speaker's speaker number, the speaker's pronunciation symbol sequence, and the speaker's third speech file using the corpus format, a speech corpus of the same language for that speaker is obtained.

[0009] As an optional implementation, in the first aspect of the present invention, the step of performing adjustment training operations on the back-end basic model through the back-end model trainer based on the target speech corpus of the target timbre to obtain a multilingual model of the target timbre includes: Based on the target speech corpus with the target timbre, the backend basic model is adjusted and trained using the backend model trainer; During the process of adjusting and training the backend basic model, it is determined whether the current conditions of the backend basic model meet the preset convergence conditions. When it is determined that the current conditions of the backend basic model do not meet the convergence condition, the target speech corpus of the target timbre is updated, and the operation of adjusting and training the backend basic model through the backend model trainer is triggered until the current conditions of the backend basic model meet the convergence condition. When it is determined that the current conditions of the backend basic model meet the convergence condition, the backend basic model is determined as the multilingual model of the target timbre.

[0010] As an optional implementation, in the first aspect of the present invention, the step of adjusting and training the back-end basic model using the back-end model trainer based on the target speech corpus with the target timbre includes: Based on the accessibility of each of the languages ​​covered by the determined language coverage corpus, at least one language whose accessibility is greater than or equal to a preset accessibility is selected from all the languages ​​as the target language to be adjusted for training. Based on the target language, a pre-built data annotation system is used to perform speech data processing operations on the target speech corpus of the target timbre to obtain speech corpus data pairs of the target timbre with respect to the target language. Based on the target speech corpus data pair of the target timbre, the backend basic model is adjusted and trained through the backend model trainer.

[0011] A second aspect of the present invention discloses a cross-language voice cloning device, the device being applied to a network-attached storage device, the device comprising: The training module is used to input the determined language coverage corpus and timbre coverage corpus into the pre-trained model of the preset back-end model trainer for training, so as to obtain a multilingual and multi-timbre back-end basic model. The adjustment module is used to perform adjustment training operations on the back-end basic model according to the target speech corpus of the target timbre when it is detected that speech cloning of the target speech corpus of the target timbre is required, so as to obtain the multilingual model of the target timbre. The analysis module is used to input the multilingual text to be synthesized into a preset multilingual front-end text analyzer in a pre-built multilingual speech synthesis engine to obtain the pronunciation symbol sequence of the multilingual text to be synthesized. The cloning module is used to input the pronunciation symbol sequence of the multilingual text to be synthesized into the multilingual model for cloning, and to obtain the speech cloning result of the target timbre, wherein the speech cloning result includes a speech waveform.

[0012] As an optional implementation, in a second aspect of the present invention, the language coverage corpus includes multilingual speech corpus for each of a plurality of similar timbre groups, and multilingual speech corpus for each of a plurality of identical timbre groups. The timbre-covered corpus includes the same language speech corpus of each of the multiple speakers; Furthermore, the specific methods by which the training module inputs the determined language coverage corpus and timbre coverage corpus into the pre-trained model of the preset backend model trainer for training, to obtain the multilingual and multi-timbre backend basic model, include: The determined language coverage corpus and timbre coverage corpus are input into the pre-trained model in the preset back-end model trainer. The pre-trained model learns the correlation between different language pronunciation symbols through the language coverage corpus and the pronunciation ability of different speakers through the timbre coverage corpus. Based on the correlation and the pronunciation ability, the common features of pronunciation symbols that are independent of timbre characteristics are abstracted to obtain the multilingual and multi-timbre back-end basic model.

[0013] As an optional implementation, in the second aspect of the invention, the language coverage corpus is determined in the following manner: Multiple similar timbre groups and multiple identical timbre groups are obtained; wherein, all timbres in each of the similar timbre groups are different and the similarity between all timbres is greater than or equal to a preset similarity, and all timbres in each of the identical timbre groups are the same; Obtain a first voice file for each group of similar timbres and a second voice file for each group of the same timbres, wherein the first voice file includes voice files with a number of languages ​​equal to a first preset number of languages, and the second voice file includes voice files with a number of languages ​​equal to a second preset number of languages, wherein the first preset number of languages ​​is greater than the second preset number of languages. For each group of similar timbres, based on the first speech file of the group of similar timbres, the multilingual speech corpus of the group of similar timbres is determined through the phonetic system of the language corresponding to the first speech file and the preset corpus format. For each group of the same timbre, based on the second speech file of the same timbre group, and through the phonetic system of the language corresponding to the second speech file, and the form of the speech corpus, the multilingual speech corpus of the same timbre group is determined; The multilingual speech corpora of all the aforementioned similar timbre groups, as well as the multilingual speech corpora of all the aforementioned identical timbre groups, are identified as language-covered corpora.

[0014] As an optional implementation, in a second aspect of the invention, the timbre coverage corpus is determined by the following method: Obtain a third speech file for each of the multiple speakers, and each of the third speech files corresponds to a language; For each speaker, based on the third speech file of that speaker, the same language speech corpus of that speaker is determined through the phonetic transcription system of the language corresponding to the third speech file and the corpus format; The same language speech data of all the speakers were identified as timbre-covered data.

[0015] As an optional implementation, in the second aspect of the present invention, the specific method for determining the same language speech corpus of each speaker, based on the third speech file of that speaker, through the phonetic transcription system of the language corresponding to the third speech file, and the corpus format, includes: For each speaker, a corresponding speaker number is assigned, and the speaker's pronunciation symbol sequence is determined based on the third speech file of that speaker and the phonetic symbol system of the language corresponding to the third speech file. By integrating the speaker's speaker number, the speaker's pronunciation symbol sequence, and the speaker's third speech file using the corpus format, a speech corpus of the same language for that speaker is obtained.

[0016] As an optional implementation, in the second aspect of the present invention, the adjustment module, based on the target speech corpus of the target timbre, performs adjustment training operations on the backend basic model through the backend model trainer to obtain a multilingual model of the target timbre. The specific methods include: Based on the target speech corpus with the target timbre, the backend basic model is adjusted and trained using the backend model trainer; During the process of adjusting and training the backend basic model, it is determined whether the current conditions of the backend basic model meet the preset convergence conditions. When it is determined that the current conditions of the backend basic model do not meet the convergence condition, the target speech corpus of the target timbre is updated, and the operation of adjusting and training the backend basic model through the backend model trainer is triggered until the current conditions of the backend basic model meet the convergence condition. When it is determined that the current conditions of the backend basic model meet the convergence condition, the backend basic model is determined as the multilingual model of the target timbre.

[0017] As an optional implementation, in a second aspect of the present invention, the adjustment module adjusts and trains the backend basic model based on the target speech corpus of the target timbre through the backend model trainer in the following specific ways: Based on the accessibility of each of the languages ​​covered by the determined language coverage corpus, at least one language whose accessibility is greater than or equal to a preset accessibility is selected from all the languages ​​as the target language to be adjusted for training. Based on the target language, a pre-built data annotation system is used to perform speech data processing operations on the target speech corpus of the target timbre to obtain speech corpus data pairs of the target timbre with respect to the target language. Based on the target speech corpus data pair of the target timbre, the backend basic model is adjusted and trained through the backend model trainer.

[0018] A third aspect of the present invention discloses a network-attached storage device, the network-attached storage device comprising: Memory containing executable program code; A processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the cross-language speech cloning method disclosed in the first aspect of the present invention.

[0019] The fourth aspect of the present invention discloses a computer storage medium storing computer instructions, which, when invoked, are used to execute the cross-language speech cloning method disclosed in the first aspect of the present invention.

[0020] Compared with the prior art, the embodiments of the present invention have the following beneficial effects: This invention enables the input of language-covered and timbre-covered corpora into a backend model trainer to train a multilingual and multi-timbre backend basic model. Based on the target timbre corpus, the backend model trainer fine-tunes the backend basic model to train a multilingual model for the target timbre. In the multilingual speech synthesis engine, the language to be cloned and the text to be synthesized are input into a multilingual frontend text analyzer to analyze the pronunciation symbol sequence, which is then input into the multilingual model to clone the speech waveform of the target timbre. This improves the diversity, relevance, and comprehensiveness of the language-covered and timbre-covered corpora of the multilingual model, facilitating the achievement of balancing the similarity of the cloned timbre with language diversity. This improves the accuracy and reliability of the target timbre cloning, thereby enhancing the multilingual sound cloning effect of the target timbre. Furthermore, by first training the multilingual and multi-timbre backend basic model and then fine-tuning the training of the speech cloning model for the target timbre, the amount of training data and computational power requirements are reduced, thus improving the training efficiency and inference speed of the speech cloning model, and consequently, the cloning efficiency and speed of the target timbre. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is a flowchart illustrating a cross-language voice cloning method disclosed in an embodiment of the present invention; Figure 2 This is a flowchart illustrating another cross-language speech cloning method disclosed in an embodiment of the present invention; Figure 3 This is a flowchart illustrating another cross-language speech cloning method disclosed in an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of a cross-language voice cloning device disclosed in an embodiment of the present invention; Figure 5 This is a schematic diagram of the structure of a network-attached storage device disclosed in an embodiment of the present invention. Detailed Implementation

[0023] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0024] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, apparatus, product, or end that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or ends.

[0025] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0026] This invention discloses a cross-lingual speech cloning method and apparatus. It can input language-covered and timbre-covered corpora into a backend model trainer to train a multilingual and multi-timbre backend basic model. Based on the target timbre corpus, the backend model trainer fine-tunes the backend basic model to train a multilingual model for the target timbre. In a multilingual speech synthesis engine, the language to be cloned and the text to be synthesized are input into a multilingual frontend text analyzer to analyze the phonetic symbol sequence, which is then input into the multilingual model to clone the speech waveform of the target timbre. This improves the language coverage and timbre accuracy of the multilingual model. The diversity, specificity, and comprehensiveness of the color-covered corpus are beneficial for achieving both similarity to the target timbre in cloning and linguistic diversity. This improves the accuracy and reliability of target timbre cloning, and consequently enhances the multilingual sound cloning effect. Furthermore, by first training a multilingual, multi-timbre backend base model and then fine-tuning the target timbre-specific speech cloning model, the amount of training data and computational requirements can be reduced. This improves the training efficiency and inference speed of the speech cloning model, further enhancing the cloning efficiency and speed of the target timbre. These points will be explained in detail below.

[0027] Example 1 Please see Figure 1 , Figure 1 This is a flowchart illustrating a cross-language speech cloning method disclosed in an embodiment of the present invention. Figure 1 The described cross-language speech cloning method can be applied to a cross-language speech cloning device, which can be applied to a network-attached storage device. This device may include a server or platform for cloning speech data, and the server may be a local server or a cloud server; this embodiment of the invention is not limited to this. Figure 1 As shown, this cross-language speech cloning method may include the following operations: 101. Input the determined language coverage corpus and timbre coverage corpus into the pre-trained model of the preset back-end model trainer for training to obtain a multilingual and multi-timbre back-end basic model.

[0028] In this embodiment of the invention, training a multilingual and multi-timbre backend base model (wherein the backend base model can be referred to as the base model) requires multilingual training corpus and multi-timbre training corpus, which respectively correspond to the aforementioned language-covered corpus and timbre-covered corpus.

[0029] The language-covered corpus includes multilingual speech data for each of several similar timbre groups, as well as multilingual speech data for each of several identical timbre groups. For example, the language-covered corpus includes similar timbre data and identical timbre Chinese-English data. The similar timbre data can include data from four languages ​​(e.g., Chinese, English, Japanese, German), with 5,000 sentences for each language, totaling 20,000 sentences and approximately 20 hours of data. The identical timbre Chinese-English data can include bilingual Chinese and English data, with 20,000 Chinese sentences, 5,000 English sentences, and 5,000 mixed Chinese-English sentences, totaling approximately 30 hours of data. The language-covered corpus requires long-term, precisely labeled data for each language (preferably identical or similar timbres) for the backend acoustic model to learn the acoustic features corresponding to the phonetic symbols of each language and the relationships between various phonetic symbols. The relationships between phonetic symbols are the foundation for cross-language pronunciation.

[0030] In this embodiment of the invention, the timbre coverage corpus includes the same language speech data of each of multiple speakers. The timbre coverage corpus requires that the data cover speakers with various timbres as much as possible, such as: different ages, genders, speaking speeds, and characteristics (deep, sharp), etc. Based on one minute of data, the model can learn and summarize the pronunciation characteristics of the corresponding speaker; therefore, the data for each speaker does not need to be too large, approximately one minute is sufficient, totaling about 20 hours (or more). For example, the timbre coverage corpus can include Chinese data from 2000 speakers and English data from 100 speakers. Each speaker's Chinese data consists of 10 sentences, totaling approximately 20,000 sentences of multi-timbre data, approximately 20 hours of Chinese data; each speaker's English data consists of 400 sentences, totaling approximately 40,000 sentences, approximately 40 hours of English data. The timbre coverage corpus is used to enable the backend basic model to learn and simulate the ability of various vocal organs. This is very important in the later cloning and fine-tuning of the target timbre, so that after the base model has seen multiple timbres, it can quickly converge to the target speaker's timbre by fine-tuning with a small amount of target speaker data.

[0031] In this embodiment of the invention, both the language-covered corpus and the timbre-covered corpus have corresponding speech waveforms, which can be determined manually or obtained from the output of other trained large-scale speech models. The pre-trained model can be a neural network model. Specifically, the language-covered corpus is used as input and its speech waveform is used as output, and the timbre-covered corpus is used as input and its speech waveform is used as output to train the pre-trained model of the back-end model trainer, resulting in a converged pre-trained model. This converged pre-trained model is then used as the multilingual and multi-timbre back-end base model. The converged pre-trained model can refer to a situation where the loss reduction during the current training period is less than a preset threshold. The loss reduction is calculated by comparing the loss value between the predicted and actual speech waveforms.

[0032] 102. When it is detected that speech cloning of the target speech corpus of the target timbre is required, the backend basic model is adjusted and trained through the backend model trainer based on the target speech corpus of the target timbre to obtain the multilingual model of the target timbre.

[0033] In this embodiment of the invention, a backend model is trained on the full base corpus using a backend model trainer to obtain a backend base model, and then fine-tuned on the corpus containing the timbre of the target speaker to be cloned to obtain a multilingual model of the target timbre. Specifically, for the full base corpus, such as hundreds of hours of data, the backend model trainer iterates multiple times until the model converges to obtain the backend base model. This backend base model has learned from language-covered corpora and possesses multilingual pronunciation capabilities, and has also learned from timbre-covered corpora, possessing the ability or potential to pronounce different timbres. For the corpus containing the timbre of the target speaker to be cloned, such as minutes of data, the backend model trainer fine-tunes and trains on the base model using the timbre data to be cloned to obtain a multilingual model of the target timbre.

[0034] Specifically, the training process of the base model is as follows: The entire prepared base training data (in the form of <speaker number, pronunciation symbol sequence, speech file>) is fed into the backend model for training. The fine-tuning process for the target timbre is as follows: First, the pre-trained base model is loaded. A new speaker number (e.g., 0) is set for the target timbre. The target timbre corpus <0, pronunciation symbol sequence, speech file> is constructed. The base model is then fine-tuned on a small amount of the target timbre corpus until convergence.

[0035] 103. In the pre-built multilingual speech synthesis engine, the multilingual text to be synthesized is input into the preset multilingual front-end text analyzer for analysis, and the pronunciation symbol sequence of the multilingual text to be synthesized is obtained.

[0036] In this embodiment of the invention, optionally, the multilingual speech synthesis engine integrates functional modules for at least one operation, such as preprocessing the input text, language recognition, and text segmentation. The preprocessing operation may include removing irrelevant symbols (such as HTML tags and extra spaces), standardizing the encoding format (such as UTF-8), and standardizing numerical symbols / special symbols, for example, converting "123" to "one hundred and twenty-three" and "%" to "hundred percent". The language recognition operation may involve determining the language corresponding to each paragraph and / or sentence and / or phrase in the input text. Text segmentation may involve dividing the input text according to sentences, paragraphs, or phrases to ensure that each segmented text corresponds to a single language.

[0037] In this embodiment of the invention, the multilingual front-end text analyzer can integrate multiple language phonetic symbol systems to generate a sequence of phonetic symbols for each language.

[0038] Specifically, the multilingual speech synthesis engine receives multilingual text input by the user and divides it into blocks according to language type and text order, resulting in multiple text blocks. Each text block corresponds to a language and has a corresponding text order, which can be sentence-based or paragraph-based. The text blocks corresponding to each language are then input into the multilingual front-end text analyzer. The multilingual front-end text analyzer generates a block pronunciation symbol sequence for each text block and its corresponding language, and integrates the block pronunciation symbol sequences for all text blocks according to their respective language orders to obtain the pronunciation symbol sequence for the multilingual text to be synthesized.

[0039] 104. Input the pronunciation symbol sequence of the multilingual text to be synthesized into the multilingual model for cloning to obtain the speech cloning result of the target timbre.

[0040] In this embodiment of the invention, optionally, the speech cloning result includes a speech waveform. Further optionally, the speech cloning result of the target timbre, and / or the speech cloning result of the target timbre, is provided to a speech playback device, causing the speech playback device to play the cloned speech of the target timbre related to the aforementioned multilingual text to be synthesized. The speech playback device is communicatively connected to a network attached storage device. Specifically, the network attached storage device clones the target speech corpus of the target timbre in the network attached storage device using the aforementioned cross-language speech cloning method to obtain the speech cloning result of the target timbre, and then sends the speech cloning result to the speech playback device, causing the speech playback device to play the cloned speech of the target timbre.

[0041] It is evident that implementation Figure 1The described cross-lingual speech cloning method can input language-covered and timbre-covered corpora into a backend model trainer to train a multilingual and multi-timbre backend base model. Based on the target timbre corpus, the backend model trainer fine-tunes the backend base model to train a multilingual model for the target timbre. In the multilingual speech synthesis engine, the language to be cloned and the text to be synthesized are input into a multilingual frontend text analyzer to analyze the phonetic symbol sequences, which are then input into the multilingual model to clone the speech waveform of the target timbre. This method can improve the language-covered and timbre-covered corpora of the multilingual model. The diversity, specificity, and comprehensiveness of the technology facilitate the achievement of cloning timbre similarity and language diversity while taking into account the target timbre. This improves the accuracy and reliability of target timbre cloning, and consequently enhances the multilingual sound cloning effect of the target timbre. Furthermore, by first training a multilingual, multi-timbre backend basic model and then fine-tuning the training of the speech cloning model for the target timbre, the amount of training data and computational power requirements can be reduced. This improves the training efficiency and inference speed of the speech cloning model, and consequently, the cloning efficiency and speed of the target timbre.

[0042] In an optional embodiment, step 101 above, which involves inputting the determined language coverage corpus and timbre coverage corpus into a pre-trained model of a preset backend model trainer for training, yields a multilingual and multi-timbre backend basic model, including: The determined language coverage corpus and timbre coverage corpus are input into the pre-trained model in the preset back-end model trainer. The pre-trained model learns the correlation between different language pronunciation symbols through the language coverage corpus and the pronunciation ability of different speakers through the timbre coverage corpus. Based on the correlation and pronunciation ability, the common features of pronunciation symbols that are independent of timbre characteristics are abstracted to obtain the multilingual and multi-timbre back-end basic model.

[0043] In this embodiment of the invention, the backend basic model can be a seq2seq neural network acoustic backend model. The input of this model is a sequence of phonetic symbols from various languages ​​and the ID of the target timbre, and the output is speech acoustic features (such as Mel spectrograms) or speech waveforms. Furthermore, the phonetic symbols supported by this backend basic model are a collection of phonetic symbols from various languages. During multilingual corpus training, the model can learn the relationships between various phonetic symbols; and during training with timbre-covered corpora, the model can learn the pronunciation abilities of multiple speakers, enabling the model to simulate various forms of vocal organs, thereby abstracting the common features of phonetic symbols beyond timbre characteristics. For example, the model selection can be a multi-speaker VITS (Variational Inference with Textual Supervision) model framework, or a multi-speaker Tacotron network; this embodiment of the invention is not limited to any particular model.

[0044] As can be seen, this optional embodiment can train the backend basic model using language-covered corpora and timbre-covered corpora, enabling the model to learn the correlation between pronunciation symbols of different languages, and at the same time learn the pronunciation ability of different speakers' timbres. This allows the model to abstract the common features of pronunciation symbols that are independent of timbre characteristics, thereby improving the training accuracy and reliability of the backend basic model and enhancing the comprehensiveness of the trained backend basic model for multiple languages ​​and timbres. This, in turn, helps to improve the efficiency and speed of subsequent model fine-tuning.

[0045] In another optional embodiment, the language coverage corpus described above is determined in the following manner: Get multiple similar timbre groups and multiple identical timbre groups; wherein, all timbres in each similar timbre group are different and the similarity between all timbres is greater than or equal to the preset similarity, and all timbres in each identical timbre group are the same. Obtain the first speech file for each group of similar timbres, and the second speech file for each group of identical timbres; For each similar timbre group, based on the first speech file of the similar timbre group, the multilingual speech corpus of the similar timbre group is determined through the phonetic system of the language corresponding to the first speech file and the preset corpus format; For each group of the same timbre, based on the second speech file of the same timbre group, the multilingual speech corpus of the same timbre group is determined through the phonetic system of the language corresponding to the second speech file and the corpus format. The multilingual speech corpora of all similar timbre groups and the multilingual speech corpora of all the same timbre groups were identified as language coverage corpora.

[0046] Optionally, the preset similarity can be 85%, 70%, or any other pre-defined value; this embodiment of the invention does not limit this. The first audio file includes audio files in a number of languages ​​equal to the first preset number of languages, and the second audio file includes audio files in a number of languages ​​equal to the second preset number of languages, where the first preset number of languages ​​is greater than the second preset number of languages. For example, the first audio file in a similar timbre group includes audio files in four languages, and the second audio file in the same timbre group includes audio files in two languages.

[0047] In this embodiment of the invention, the method for determining multilingual speech corpora of similar timbre groups and the method for determining multilingual speech corpora of the same timbre group can refer to the method for determining the same language speech corpora of the speakers described below, and will not be repeated in this embodiment of the invention.

[0048] As can be seen, this optional embodiment can obtain a large number of speech files with similar timbres in various languages ​​and a relatively concise speech file with the same timbre in various languages. By using the phonetic symbol system of the corresponding language and the preset corpus format, similar timbre corpus and the same timbre corpus are respectively determined as language coverage corpus. This can improve the accuracy and reliability of determining language coverage corpus, which is beneficial for providing diverse training corpus for the backend model trainer and for providing an accurate and reliable data foundation for the subsequent model to learn the relationship between pronunciation symbols.

[0049] In yet another optional embodiment, the aforementioned timbre coverage corpus is determined in the following manner: Obtain the third audio file for each of the multiple speakers, with each third audio file corresponding to a language; For each speaker, based on the third speech file of that speaker, the phonetic transcription system of the language corresponding to the third speech file and the form of the corpus are used to determine the speech corpus of that speaker in the same language; Identify the same language speech data of all speakers as timbre coverage data.

[0050] In this embodiment of the invention, different languages ​​have different phonetic systems. For example, Chinese has the Pinyin system, English has the CMU phonetic system, Japanese has the Romanization phonetic system, and German uses the internationally common IPA phonetic system, etc. For instance, for each speaker speaking Chinese, the Chinese speech data is determined based on the speaker's audio file using the Pinyin system and the aforementioned data format; similarly, for each speaker speaking English, the English speech data is determined based on the speaker's audio file using the CMU phonetic system and the aforementioned data format.

[0051] It can be seen that this optional embodiment can obtain the speech files of each speaker among multiple speakers, and thus determine the same-language speech corpus of each speaker as the timbre coverage corpus through the phonetic symbol system and the preset corpus form of the corresponding language. This can improve the accuracy and reliability of determining the timbre coverage corpus, is conducive to providing a training corpus with diverse timbres for the backend model trainer, and is also conducive to providing an accurate and reliable data basis for the subsequent model to learn the pronunciation ability of each timbre.

[0052] In this optional embodiment, as an optional implementation manner, for each speaker, according to the third speech file of this speaker, through the phonetic symbol system of the language corresponding to the third speech file and the corpus form, the same-language speech corpus of this speaker is determined, including: For each speaker, assign a corresponding speaker serial number to this speaker, and according to the third speech file of this speaker, determine the pronunciation symbol sequence of this speaker through the phonetic symbol system of the language corresponding to the third speech file; Through the corpus form, integrate the speaker serial number of this speaker, the pronunciation symbol sequence of this speaker, and the third speech file of this speaker to obtain the same-language speech corpus of this speaker.

[0053] In the embodiments of the present invention, specifically, each speaker is assigned a digital serial number, and the speech of each speaker needs to be marked with a pronunciation symbol sequence. For example, the speech file "你好" needs to be marked as Chinese pinyin "ni2 hao3" or IPA phonetic symbols "n'i35_| X'Au214_|"; and the corpus form of each speaker is <speaker serial number, pronunciation symbol sequence, speech file>, such as <1, ni2 hao3, 你好.wav>.

[0054] It can be seen that this optional implementation manner can assign a corresponding speaker serial number to each speaker, determine the pronunciation symbol sequence of this speaker according to the speech file of this speaker through the phonetic symbol system of its corresponding language, and then integrate the speaker serial number, the pronunciation symbol sequence, and the speech file through the corpus form to obtain the same-language speech corpus of the corresponding speaker, which can improve the integration accuracy and reliability of the same-language speech corpus of each speaker and is conducive to improving the timbre diversity of the same-language speech corpus.

[0055] Embodiment 2 Please refer to Figure 2 , Figure 2 which is a schematic flowchart of a cross-language voice cloning method disclosed in the embodiments of the present invention. Among them, Figure 2The described cross-language speech cloning method can be applied to a cross-language speech cloning device, which can be applied to a network-attached storage device. This device may include a server or platform for cloning speech data, and the server may be a local server or a cloud server; this embodiment of the invention is not limited to this. Figure 2 As shown, this cross-language speech cloning method may include the following operations: 201. Input the determined language coverage corpus and timbre coverage corpus into the pre-trained model of the preset back-end model trainer for training to obtain a multilingual and multi-timbre back-end basic model.

[0056] 202. When it is detected that speech cloning of the target speech corpus with the target timbre is required, the backend basic model is adjusted and trained according to the target speech corpus with the target timbre through the backend model trainer.

[0057] In this embodiment of the invention, the target speech corpus of the target timbre may include multiple speech corpora, and each speech corpus corresponds to a speech waveform. Specifically, each speech corpus of the target timbre is taken as input, and the speech waveform corresponding to the speech corpus is taken as output. The backend basic model is adjusted and trained through the backend model trainer.

[0058] 203. During the process of adjusting and training the backend basic model, determine whether the current conditions of the backend basic model meet the preset convergence conditions.

[0059] In this embodiment of the invention, specifically, after the end of each training cycle, the loss reduction magnitude between the predicted target timbre speech waveform and the actual speech waveform in the current training cycle is calculated, and it is determined whether the loss reduction magnitude is less than a preset threshold. If it is determined to be less than the preset threshold, it is determined that the current conditions of the backend basic model meet the preset convergence condition; if it is determined to be greater than or equal to the preset threshold, it is determined that the current conditions of the backend basic model do not meet the preset convergence condition.

[0060] In this embodiment of the invention, when the judgment result of step 203 is negative, that is, when it is determined that the current condition of the backend basic model does not meet the convergence condition, step 204 is triggered; when the judgment result of step 203 is positive, that is, when it is determined that the current condition of the backend basic model meets the convergence condition, step 205 is triggered.

[0061] 204. Update the target speech corpus for the target timbre.

[0062] Specifically, text generation technology is used to generate a synthetic speech corpus of the target timbre, and the synthetic speech corpus of the target timbre is integrated into the target speech corpus of the target timbre to update the target speech corpus of the target timbre.

[0063] In this embodiment of the invention, after executing step 204, the operation of adjusting and training the backend basic model through the backend model trainer in step 202 is triggered until the current conditions of the backend basic model meet the convergence condition.

[0064] 205. Determine the backend basic model as the multilingual model for the target timbre.

[0065] 206. In the pre-built multilingual speech synthesis engine, the multilingual text to be synthesized is input into the preset multilingual front-end text analyzer for analysis, and the pronunciation symbol sequence of the multilingual text to be synthesized is obtained.

[0066] 207. Input the pronunciation symbol sequence of the multilingual text to be synthesized into the multilingual model for cloning to obtain the speech cloning result of the target timbre.

[0067] In this embodiment of the invention, for other descriptions of steps 201, 206 and 207, please refer to the detailed description of steps 101, 103 and 104 in Embodiment 1. These descriptions will not be repeated in this embodiment of the invention.

[0068] As can be seen, the embodiments of the present invention can input language-covered corpora and timbre-covered corpora into a backend model trainer to train a multilingual and multi-timbre backend basic model. Based on the target timbre corpus, the backend model trainer fine-tunes the backend basic model to train a multilingual model for the target timbre. In the multilingual speech synthesis engine, the language to be cloned and the text to be synthesized are input into a multilingual frontend text analyzer to analyze the phonetic symbol sequence, which is then input into the multilingual model to clone the speech waveform of the target timbre. This can improve the diversity of language-covered and timbre-covered corpora in the multilingual model. The targeted, specific, and comprehensive nature of this approach facilitates the achievement of cloning timbres that balance similarity with target timbres and linguistic diversity. This improves the accuracy and reliability of target timbre cloning, thereby enhancing the multilingual sound cloning effect. Furthermore, by first training a multilingual, multi-timbre backend base model and then fine-tuning the training of the target timbre speech cloning model, the amount of training data and computational power required can be reduced. This improves the training efficiency and inference speed of the speech cloning model, further enhancing the cloning efficiency and speed. Additionally, the backend base model can be trained using the target speech corpus for the desired timbre speech, adjusting its training through the backend model trainer. During this adjustment process, convergence is assessed. If convergence is not observed, the target speech corpus is updated and retrained until convergence occurs. If convergence is achieved, the backend base model is designated as the multilingual model for the target timbre. This improves the efficiency, accuracy, and reliability of the multilingual model obtained through fine-tuning, further enhancing the cloning efficiency, accuracy, and reliability of the target timbre.

[0069] In an optional embodiment, step 202 above, which involves adjusting and training the backend base model using a backend model trainer based on the target speech corpus with the target timbre, includes: Based on the accessibility of each language among all languages ​​covered by the determined language coverage corpus, at least one language with an accessibility greater than or equal to the preset accessibility is selected as the target language to be adjusted for training. Based on the target language, a pre-built data annotation system is used to perform speech data processing operations on the target speech corpus of the target timbre to obtain speech corpus data pairs of the target timbre with respect to the target language; Based on the target speech corpus data pairs of the target timbre, the backend basic model is adjusted and trained through the backend model trainer.

[0070] In this embodiment of the invention, the target language can be designed for easily accessible monolingual corpora, such as Chinese. The data annotation system can integrate functional modules for at least one speech data processing operation, such as noise reduction, segmentation (sentence segmentation), recognition, and phonetic transcription; the data annotation system is used to input speech data to perform noise reduction, segmentation (sentence segmentation), recognition, and phonetic transcription to obtain a data pair of <pronunciation symbol sequence, speech file> for a specific language.

[0071] Optionally, a flowchart of this solution can also be found in the appendix to the instruction manual. Figure 3 The flowchart shown is a cross-language voice cloning method, and the embodiments of the present invention are not limited thereto.

[0072] As can be seen, this optional embodiment can select a target language with an accessibility greater than or equal to a preset accessibility level from all languages ​​covered by the determined language coverage corpus, based on the accessibility of each language. Then, through a data annotation system, speech data processing operations are performed on the target speech corpus of the target timbre to obtain speech corpus data pairs of the target timbre with respect to the target language. Subsequently, fine-tuning training is performed through a back-end model trainer, which can improve the accessibility of the speech corpus of the target timbre with respect to the target language, and can also improve the efficiency, rationality, and relevance of the target timbre corpus. This can improve the efficiency, timeliness, and relevance of the multilingual model of the target timbre obtained through fine-tuning, and further improve the cloning efficiency, timeliness, and relevance of the target timbre.

[0073] Example 3 Please see Figure 4 , Figure 4 This is a schematic diagram of the structure of a cross-language voice cloning device disclosed in an embodiment of the present invention. Figure 4 The described cross-language voice cloning device can be applied to a network-attached storage device. This device can include a server or platform for cloning audio corpora, where the server may be a local server or a cloud server; this embodiment of the invention is not limited to this. Figure 4 As shown, the cross-language voice cloning device may include: The training module 301 is used to input the determined language coverage corpus and timbre coverage corpus into the pre-trained model of the preset back-end model trainer for training, so as to obtain a multilingual and multi-timbre back-end basic model.

[0074] The adjustment module 302 is used to perform adjustment training operations on the backend basic model based on the target speech corpus of the target timbre when it is detected that speech cloning of the target timbre is required, so as to obtain a multilingual model of the target timbre.

[0075] The analysis module 303 is used to input the multilingual text to be synthesized into a preset multilingual front-end text analyzer in a pre-built multilingual speech synthesis engine to obtain the pronunciation symbol sequence of the multilingual text to be synthesized.

[0076] The cloning module 304 is used to input the pronunciation symbol sequence of the multilingual text to be synthesized into the multilingual model for cloning, and to obtain the speech cloning result of the target timbre, which includes the speech waveform.

[0077] It is evident that implementation Figure 4 The described cross-lingual speech cloning device can input language-covered and timbre-covered corpora into a backend model trainer to train a multilingual and multi-timbre backend base model. Based on the target timbre corpus, the backend model trainer fine-tunes the backend base model to train a multilingual model for the target timbre. In the multilingual speech synthesis engine, the language to be cloned and the text to be synthesized are input into a multilingual frontend text analyzer to analyze the phonetic symbol sequence, which is then input into the multilingual model to clone the speech waveform of the target timbre. This improves the language-covered and timbre-covered corpora of the multilingual model. The diversity, specificity, and comprehensiveness of the technology facilitate the achievement of cloning timbre similarity and language diversity while taking into account the target timbre. This improves the accuracy and reliability of target timbre cloning, and consequently enhances the multilingual sound cloning effect of the target timbre. Furthermore, by first training a multilingual, multi-timbre backend basic model and then fine-tuning the training of the speech cloning model for the target timbre, the amount of training data and computational power requirements can be reduced. This improves the training efficiency and inference speed of the speech cloning model, and consequently, the cloning efficiency and speed of the target timbre.

[0078] In an optional embodiment, the language-covered corpus includes multilingual speech data for each of multiple similar timbre groups, and multilingual speech data for each of multiple identical timbre groups; the timbre-covered corpus includes same-language speech data for each of multiple speakers. Furthermore, the training module 301 inputs the determined language-covered corpus and timbre-covered corpus into a pre-trained model of a preset backend model trainer for training, obtaining a multilingual and multi-timbre backend basic model in the following specific ways: The determined language coverage corpus and timbre coverage corpus are input into the pre-trained model in the preset back-end model trainer. The pre-trained model learns the correlation between different language pronunciation symbols through the language coverage corpus and the pronunciation ability of different speakers through the timbre coverage corpus. Based on the correlation and pronunciation ability, the common features of pronunciation symbols that are independent of timbre characteristics are abstracted to obtain the multilingual and multi-timbre back-end basic model.

[0079] As can be seen, this optional embodiment can train the backend basic model using language-covered corpora and timbre-covered corpora, enabling the model to learn the correlation between pronunciation symbols of different languages, and at the same time learn the pronunciation ability of different speakers' timbres. This allows the model to abstract the common features of pronunciation symbols that are independent of timbre characteristics, thereby improving the training accuracy and reliability of the backend basic model and enhancing the comprehensiveness of the trained backend basic model for multiple languages ​​and timbres. This, in turn, helps to improve the efficiency and speed of subsequent model fine-tuning.

[0080] In this optional embodiment, as an optional implementation method, the language coverage corpus is determined in the following way: Get multiple similar timbre groups and multiple identical timbre groups; wherein, all timbres in each similar timbre group are different and the similarity between all timbres is greater than or equal to the preset similarity, and all timbres in each identical timbre group are the same. Obtain a first audio file for each similar timbre group and a second audio file for each identical timbre group, wherein the first audio file includes audio files with a number of languages ​​equal to a first preset number of languages, and the second audio file includes audio files with a number of languages ​​equal to a second preset number of languages, wherein the first preset number of languages ​​is greater than the second preset number of languages. For each similar timbre group, based on the first speech file of the similar timbre group, the multilingual speech corpus of the similar timbre group is determined through the phonetic system of the language corresponding to the first speech file and the preset corpus format; For each group of the same timbre, based on the second speech file of the same timbre group, the multilingual speech corpus of the same timbre group is determined through the phonetic system of the language corresponding to the second speech file and the corpus format. The multilingual speech corpora of all similar timbre groups and the multilingual speech corpora of all the same timbre groups were identified as language coverage corpora.

[0081] As can be seen, this optional implementation can obtain a large number of speech files with similar timbres in various languages ​​and a relatively concise speech file with the same timbre in various languages. By using the phonetic symbol system of the corresponding language and the preset corpus format, similar timbre corpus and the same timbre corpus are respectively determined as language coverage corpus. This can improve the accuracy and reliability of determining language coverage corpus, which is beneficial for providing diverse training corpus for the backend model trainer and for providing an accurate and reliable data foundation for the subsequent model to learn the association relationship between pronunciation symbols.

[0082] In this optional embodiment, as another optional implementation, the timbre coverage corpus is determined in the following manner: Obtain the third audio file for each of the multiple speakers, with each third audio file corresponding to a language; For each speaker, based on the third speech file of that speaker, the phonetic transcription system of the language corresponding to the third speech file and the form of the corpus are used to determine the speech corpus of that speaker in the same language; Identify the same language speech data of all speakers as timbre coverage data.

[0083] As can be seen, this optional implementation can acquire the speech files of each of the multiple speakers, and then determine the same language speech data of each speaker as timbre coverage data by using the phonetic symbol system of the corresponding language and the preset data format. This can improve the accuracy and reliability of the determination of timbre coverage data, which is beneficial for providing the backend model trainer with training data with diverse timbres, and also beneficial for providing an accurate and reliable data foundation for the subsequent model to learn the pronunciation ability of each timbre.

[0084] In this optional implementation, optionally, for each speaker, the specific method for determining the speaker's same-language speech corpus based on the speaker's third speech file, through the phonetic transcription system of the language corresponding to the third speech file, and the corpus format, includes: For each speaker, a corresponding speaker number is assigned, and the speaker's pronunciation symbol sequence is determined based on the third speech file of that speaker and the phonetic symbol system of the language corresponding to the third speech file. By integrating the speaker's speaker number, the speaker's pronunciation symbol sequence, and the speaker's third speech file in the form of a corpus, a speech corpus of the same language for that speaker is obtained.

[0085] As can be seen, this optional implementation can also assign a corresponding speaker number to each speaker, and determine the speaker's pronunciation symbol sequence through the phonetic symbol system of the corresponding language based on the speaker's voice file. Then, the speaker number, pronunciation symbol sequence, and voice file are integrated in the form of a corpus to obtain the corresponding speaker's same language voice corpus. This can improve the accuracy and reliability of integrating the same language voice corpus of each speaker, and is conducive to improving the timbre diversity of the same language voice corpus.

[0086] In another optional embodiment, the adjustment module 302, based on the target speech corpus of the target timbre, performs adjustment training operations on the backend basic model through the backend model trainer to obtain a multilingual model of the target timbre. The specific methods include: Based on the target speech corpus with the target timbre, the backend basic model is adjusted and trained through the backend model trainer; During the process of adjusting and training the backend basic model, it is determined whether the current conditions of the backend basic model meet the preset convergence conditions. When it is determined that the current conditions of the backend basic model do not meet the convergence condition, the target speech corpus of the target timbre is updated, and the operation of adjusting and training the backend basic model through the backend model trainer is triggered until the current conditions of the backend basic model meet the convergence condition. When it is determined that the current conditions of the backend basic model meet the convergence condition, the backend basic model is determined as the multilingual model of the target timbre.

[0087] As can be seen, this optional embodiment can, as needed, use the target speech corpus for speech cloning, adjust and train the backend basic model through the backend model trainer, and during the adjustment and training process, determine whether the backend basic model has converged. If it is determined that it has not converged, the target speech corpus is updated and retrained until convergence is achieved. If it is determined that it has converged, the backend basic model is determined as the multilingual model of the target speech corpus. This can improve the efficiency, accuracy and reliability of fine-tuning the multilingual model of the target speech corpus, thereby helping to further improve the cloning efficiency, accuracy and reliability of the target speech corpus.

[0088] In this optional embodiment, as an optional implementation method, the adjustment module 302 adjusts and trains the backend basic model according to the target speech corpus of the target timbre through the backend model trainer in the following specific ways: Based on the accessibility of each language among all languages ​​covered by the determined language coverage corpus, at least one language with an accessibility greater than or equal to the preset accessibility is selected as the target language to be adjusted for training. Based on the target language, a pre-built data annotation system is used to perform speech data processing operations on the target speech corpus of the target timbre to obtain speech corpus data pairs of the target timbre with respect to the target language; Based on the target speech corpus data pairs of the target timbre, the backend basic model is adjusted and trained through the backend model trainer.

[0089] As can be seen, this optional implementation method can select a target language with an accessibility greater than or equal to a preset accessibility level from all languages ​​covered by the determined language coverage corpus, and then perform speech data processing operations on the target speech corpus of the target timbre through a data annotation system to obtain speech corpus data pairs of the target timbre with respect to the target language. Then, fine-tuning training is performed through a back-end model trainer, which can improve the accessibility of the speech corpus of the target timbre with respect to the target language, and can improve the efficiency, rationality and relevance of the target timbre corpus. This can improve the efficiency, timeliness and relevance of the multilingual model of the target timbre obtained by fine-tuning, and further improve the cloning efficiency, timeliness and relevance of the target timbre.

[0090] Example 4 Please see Figure 5 , Figure 5 This is a schematic diagram of the structure of a network-attached storage device disclosed in an embodiment of the present invention. Figure 5 As shown, the network-attached storage device may include: Memory 401 storing executable program code; Processor 402 coupled to memory 401; The processor 402 calls the executable program code stored in the memory 401 to execute the steps in the cross-language speech cloning method described in Embodiment 1 or Embodiment 2 of the present invention.

[0091] Example 5 This invention discloses a computer storage medium storing computer instructions, which, when invoked, are used to execute the steps of the cross-language speech cloning method described in Embodiment 1 or Embodiment 2 of this invention.

[0092] Example 6 This invention discloses a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to perform the steps in the cross-language speech cloning method described in Embodiment 1 or Embodiment 2.

[0093] The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0094] Through the detailed description of the above embodiments, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, including read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-Erasable Programmable Read-Only Memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, disk storage, magnetic tape storage, or any other computer-readable medium that can be used to carry or store data.

[0095] Finally, it should be noted that the cross-language voice cloning method and apparatus disclosed in the embodiments of the present invention are merely preferred embodiments of the present invention and are only used to illustrate the technical solutions of the present invention, not to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for cross-language speech cloning, characterized in that, The method is applied to a network-attached storage device, and the method includes: The determined language coverage corpus and timbre coverage corpus are input into the pre-trained model of the preset back-end model trainer for training, so as to obtain a multilingual and multi-timbre back-end basic model. When it is detected that speech cloning of the target speech corpus of the target timbre is required, the back-end basic model is adjusted and trained through the back-end model trainer according to the target speech corpus of the target timbre to obtain a multilingual model of the target timbre. In a pre-built multilingual speech synthesis engine, the multilingual text to be synthesized is input into a preset multilingual front-end text analyzer for analysis, and the phonetic symbol sequence of the multilingual text to be synthesized is obtained. The phonetic symbol sequence of the multilingual text to be synthesized is input into the multilingual model for cloning to obtain the speech cloning result of the target timbre, and the speech cloning result includes a speech waveform.

2. The cross-language speech cloning method according to claim 1, characterized in that, The language-covered corpus includes multilingual speech corpus for each of the multiple similar timbre groups, and multilingual speech corpus for each of the multiple identical timbre groups. The timbre-covered corpus includes the same language speech corpus of each of the multiple speakers; Furthermore, the process of inputting the determined language coverage corpus and timbre coverage corpus into the pre-trained model of the preset backend model trainer for training, to obtain a multilingual and multi-timbre backend basic model, includes: The determined language coverage corpus and timbre coverage corpus are input into the pre-trained model in the preset back-end model trainer. The pre-trained model learns the correlation between different language pronunciation symbols through the language coverage corpus and the pronunciation ability of different speakers through the timbre coverage corpus. Based on the correlation and the pronunciation ability, the common features of pronunciation symbols that are independent of timbre characteristics are abstracted to obtain the multilingual and multi-timbre back-end basic model.

3. The cross-language speech cloning method according to claim 2, characterized in that, The language coverage corpus was determined in the following way: Multiple similar timbre groups and multiple identical timbre groups are obtained; wherein, all timbres in each of the similar timbre groups are different and the similarity between all timbres is greater than or equal to a preset similarity, and all timbres in each of the identical timbre groups are the same; Obtain a first voice file for each group of similar timbres and a second voice file for each group of the same timbres, wherein the first voice file includes voice files with a number of languages ​​equal to a first preset number of languages, and the second voice file includes voice files with a number of languages ​​equal to a second preset number of languages, wherein the first preset number of languages ​​is greater than the second preset number of languages. For each group of similar timbres, based on the first speech file of the group of similar timbres, the multilingual speech corpus of the group of similar timbres is determined through the phonetic system of the language corresponding to the first speech file and the preset corpus format. For each group of the same timbre, based on the second speech file of the same timbre group, and through the phonetic system of the language corresponding to the second speech file, and the form of the speech corpus, the multilingual speech corpus of the same timbre group is determined; The multilingual speech corpora of all the aforementioned similar timbre groups, as well as the multilingual speech corpora of all the aforementioned identical timbre groups, are identified as language-covered corpora.

4. The cross-language speech cloning method according to claim 3, characterized in that, The timbre coverage corpus was determined in the following way: Obtain a third speech file for each of the multiple speakers, and each of the third speech files corresponds to a language; For each speaker, based on the third speech file of that speaker, the same language speech corpus of that speaker is determined through the phonetic transcription system of the language corresponding to the third speech file and the corpus format; The same language speech data of all the speakers were identified as timbre-covered data.

5. The cross-language speech cloning method according to claim 4, characterized in that, For each speaker, determining the same language speech corpus for that speaker based on the speaker's third speech file, using the phonetic transcription system of the language corresponding to the third speech file, and the corpus format, includes: For each speaker, a corresponding speaker number is assigned, and the speaker's pronunciation symbol sequence is determined based on the third speech file of that speaker and the phonetic symbol system of the language corresponding to the third speech file. By integrating the speaker's speaker number, the speaker's pronunciation symbol sequence, and the speaker's third speech file using the corpus format, a speech corpus of the same language for that speaker is obtained.

6. The cross-lingual speech cloning method according to any one of claims 1-5, characterized in that, The process of obtaining a multilingual model of the target timbre by performing adjustment training operations on the backend basic model through the backend model trainer based on the target speech corpus of the target timbre includes: Based on the target speech corpus with the target timbre, the backend basic model is adjusted and trained using the backend model trainer; During the process of adjusting and training the backend basic model, it is determined whether the current conditions of the backend basic model meet the preset convergence conditions. When it is determined that the current conditions of the backend basic model do not meet the convergence condition, the target speech corpus of the target timbre is updated, and the operation of adjusting and training the backend basic model through the backend model trainer is triggered until the current conditions of the backend basic model meet the convergence condition. When it is determined that the current conditions of the backend basic model meet the convergence condition, the backend basic model is determined as the multilingual model of the target timbre.

7. The cross-language speech cloning method according to claim 6, characterized in that, The step of adjusting and training the backend base model using the backend model trainer based on the target speech corpus with the target timbre includes: Based on the accessibility of each of the languages ​​covered by the determined language coverage corpus, at least one language whose accessibility is greater than or equal to a preset accessibility is selected from all the languages ​​as the target language to be adjusted for training. Based on the target language, a pre-built data annotation system is used to perform speech data processing operations on the target speech corpus of the target timbre to obtain speech corpus data pairs of the target timbre with respect to the target language. Based on the target speech corpus data pair of the target timbre, the backend basic model is adjusted and trained through the backend model trainer.

8. A cross-language voice cloning device, characterized in that, The apparatus is used in a network-attached storage device, and the apparatus includes: The training module is used to input the determined language coverage corpus and timbre coverage corpus into the pre-trained model of the preset back-end model trainer for training, so as to obtain a multilingual and multi-timbre back-end basic model. The adjustment module is used to perform adjustment training operations on the back-end basic model according to the target speech corpus of the target timbre when it is detected that speech cloning of the target speech corpus of the target timbre is required, so as to obtain the multilingual model of the target timbre. The analysis module is used to input the multilingual text to be synthesized into a preset multilingual front-end text analyzer in a pre-built multilingual speech synthesis engine to obtain the pronunciation symbol sequence of the multilingual text to be synthesized. The cloning module is used to input the pronunciation symbol sequence of the multilingual text to be synthesized into the multilingual model for cloning, and to obtain the speech cloning result of the target timbre, wherein the speech cloning result includes a speech waveform.

9. A network-attached storage device, characterized in that, The network-attached storage device includes: Memory containing executable program code; A processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the cross-language speech cloning method as described in any one of claims 1-7.

10. A computer storage medium, characterized in that, The computer storage medium stores computer instructions, which, when invoked, are used to execute the cross-language speech cloning method as described in any one of claims 1-7.