Method for voice conversion, method and apparatus for training a voice synthesis model

By training the speech synthesis model for users with different language proficiency levels and utilizing samples and feature information from multiple languages, the problem of inaccurate timbre and pronunciation in speech synthesis was solved, achieving a higher speech conversion accuracy.

CN115359776BActive Publication Date: 2026-03-27阳光保险集团股份有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-24
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing speech synthesis technology results in inaccurate voice timbre or pronunciation, affecting the accuracy of speech conversion.

Method used

By training a speech synthesis model, the model is trained using speech samples, word speech samples, and pronunciation factor samples from multiple languages ​​at a preset ratio corresponding to the target user's language proficiency level. Gender information, timbre features, and boundary information are embedded to improve the model's ability to imitate the user's pronunciation.

Benefits of technology

It improves the accuracy of speech conversion, making the synthesized speech more accurately reflect the user's pronunciation timbre.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115359776B_ABST
    Figure CN115359776B_ABST
Patent Text Reader

Abstract

The application provides a speech conversion method and a speech synthesis model training method and device. The method comprises: obtaining a text to be converted of a target user; and converting the text to be converted by using a speech synthesis model to obtain a speech corresponding to the text to be converted. The speech synthesis model is obtained by training a basic model according to a language ability level of the target user, speech samples of multiple languages in a preset proportion corresponding to the language ability level, word speech samples of the multiple languages, and pronunciation factor samples of the multiple languages. The preset proportion corresponding to different language ability levels is different. The basic model is obtained by training a general model by using speech samples of multiple languages in a basic preset proportion and mixed speech samples of the multiple languages. The method can improve the accuracy of speech conversion.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of voice conversion, in particular, to a voice conversion method, a method and device for training a voice synthesis model. BACKGROUND

[0002] At present, with the vigorous development of artificial intelligence technology, automatic voice synthesis has been widely applied. Voice synthesis technology based on deep neural network technology has achieved good results, and automatic voice synthesis technology has been applied in various fields, greatly promoting the development of intelligence.

[0003] The existing voice synthesis mainly relies on a certain degree of post-processing, and uses similar sounds to replace the user's voice to let the model complete the voice synthesis task.

[0004] However, when using the above method for voice synthesis, the voice timbre output by the final model is inaccurate or the synthesized voice pronunciation is inaccurate.

[0005] Therefore, how to improve the accuracy of voice conversion is a technical problem to be solved. SUMMARY

[0006] The purpose of the embodiments of the present application is to provide a voice conversion method, which can improve the accuracy of voice conversion through the technical solutions of the embodiments of the present application.

[0007] In a first aspect, the embodiments of the present application provide a voice conversion method, comprising: obtaining a target user's to-be-converted text; converting the to-be-converted text through a voice synthesis model to obtain a voice corresponding to the to-be-converted text, wherein the voice synthesis model is obtained by training a base model according to a language ability level of the target user, through a plurality of language samples of a plurality of languages in a preset proportion corresponding to the language ability level, a plurality of word voice samples of a plurality of languages, and a plurality of pronunciation factor samples of a plurality of languages, wherein the preset proportion corresponding to different language ability levels is different, and the base model is obtained by training a general model through a plurality of language samples of a plurality of languages in a base preset proportion and a plurality of mixed voice samples of a plurality of languages.

[0008] In the above embodiments, the model is trained through the plurality of language samples of a plurality of languages, the plurality of word voice samples of a plurality of languages, and the plurality of pronunciation factor samples of a plurality of languages, so that the voice synthesis model can learn the user's pronunciation, and then the voice synthesis model can convert the to-be-converted text into corresponding voice with the user's pronunciation timbre. Through this method, the accuracy of voice conversion can be improved.

[0009] In some embodiments, the language ability level comprises:

[0010] The first language proficiency level, the second language proficiency level, and the third language proficiency level, wherein the first language proficiency level indicates that the target user masters a plurality of languages, the second language proficiency level indicates that the target user masters part of the plurality of languages, and the third language proficiency level indicates that the target user does not master any of the plurality of languages.

[0011] In the above embodiment, the user's different language proficiency levels can be set to correspond to the preset proportion of samples, the corresponding speech synthesis model can be trained for different users, and the resources for model training can be saved.

[0012] In some embodiments, the text to be converted of the target user is obtained, comprising:

[0013] The initial text to be converted is preprocessed to obtain the text to be converted, wherein the preprocessing includes at least one of the following processing methods: cleaning, deleting, intercepting, and completing.

[0014] In the above embodiment, the text obtained after preprocessing is relatively standard, and the speech conversion is also more accurate.

[0015] In some embodiments, before obtaining the text to be converted of the target user, further comprising:

[0016] Obtaining the language proficiency level of the target user and the speech samples of a plurality of languages corresponding to the preset proportion of the language proficiency level, the word speech samples of a plurality of languages, and the pronunciation factor samples of a plurality of languages;

[0017] Training the general model using the speech samples of a plurality of languages and the mixed speech samples of a plurality of languages to obtain a basic model;

[0018] Training the basic model using the speech samples of a plurality of languages, the word speech samples of a plurality of languages, and the pronunciation factor samples of a plurality of languages to obtain a speech synthesis model.

[0019] In the above embodiment, the speech synthesis model is trained using the speech samples of a plurality of languages, the word speech samples of a plurality of languages, and the pronunciation factor samples of a plurality of languages, so that the speech synthesis model can learn the pronunciation of the user.

[0020] In a second aspect, the embodiments of the present application provide a method for training a speech synthesis model, comprising: training a general model by speech samples of multiple languages in a preset proportion and mixed speech samples of multiple languages to obtain a basic model; training the basic model by speech samples of multiple languages in a preset proportion corresponding to a language proficiency level of a target user, word speech samples of multiple languages and pronunciation factor samples of multiple languages to obtain the speech synthesis model, wherein the preset proportion corresponding to different language proficiency levels is different.

[0021] In the above embodiments, the model is trained by speech samples of multiple languages, word speech samples of multiple languages and pronunciation factor samples of multiple languages of the user, so that the speech synthesis model can learn the pronunciation of the user.

[0022] In some embodiments, after the basic model is trained by speech samples of multiple languages in a preset proportion corresponding to a language proficiency level of a target user, word speech samples of multiple languages and pronunciation factor samples of multiple languages to obtain the speech synthesis model, the method further comprises:

[0023] Embedding gender information, timbre feature information and boundary information of the target user into the speech synthesis model, wherein the boundary information is used to identify whether the speech languages in the speech samples are the same.

[0024] In the above embodiments, the model can synthesize more accurate speech by embedding the features of the gender information, timbre feature information and boundary information.

[0025] In a third aspect, the embodiments of the present application provide a speech conversion device, comprising:

[0026] The acquisition module is configured to acquire a text to be converted of a target user.

[0027] The conversion module is configured to convert the text to be converted by the speech synthesis model to obtain speech corresponding to the text to be converted, wherein the speech synthesis model is obtained by training a basic model by speech samples of multiple languages in a preset proportion corresponding to a language proficiency level of a target user, word speech samples of multiple languages and pronunciation factor samples of multiple languages, different language proficiency levels correspond to different preset proportions, and the basic model is obtained by training a general model by speech samples of multiple languages in a preset proportion and mixed speech samples of multiple languages.

[0028] Optionally, the language proficiency level comprises:

[0029] The first language proficiency level, the second language proficiency level, and the third language proficiency level, wherein the first language proficiency level indicates that the target user masters a plurality of languages, the second language proficiency level indicates that the target user masters part of the plurality of languages, and the third language proficiency level indicates that the target user does not master any of the plurality of languages.

[0030] Optionally, the obtaining module is specifically configured to:

[0031] The initial to-be-converted text is preprocessed to obtain the to-be-converted text, wherein the preprocessing includes at least one of the following processing methods: cleaning, deleting, intercepting, and completing.

[0032] Optionally, the device further comprises:

[0033] The training module is configured to, before the obtaining module obtains the to-be-converted text of the target user, obtain a plurality of language proficiency levels of the target user and a plurality of speech samples of a plurality of languages corresponding to a preset proportion of the language proficiency levels, a plurality of word speech samples of the plurality of languages, and a plurality of pronunciation factor samples of the plurality of languages.

[0034] The universal model is trained by using the plurality of speech samples of the plurality of languages corresponding to the preset proportion and the plurality of mixed speech samples of the plurality of languages to obtain a basic model.

[0035] The basic model is trained by using the plurality of speech samples of the plurality of languages corresponding to the preset proportion, the plurality of word speech samples of the plurality of languages, and the plurality of pronunciation factor samples of the plurality of languages to obtain a speech synthesis model.

[0036] In a fourth aspect, an embodiment of the present application provides a device for training a speech synthesis model, comprising:

[0037] The first training module is configured to train the universal model by using the plurality of speech samples of the plurality of languages corresponding to the preset proportion and the plurality of mixed speech samples of the plurality of languages to obtain a basic model.

[0038] The second training module is configured to train the basic model by using the plurality of speech samples of the plurality of languages corresponding to the preset proportion of the language proficiency levels of the target user, the plurality of word speech samples of the plurality of languages, and the plurality of pronunciation factor samples of the plurality of languages to obtain a speech synthesis model, wherein the preset proportions corresponding to different language proficiency levels are different.

[0039] Optionally, the device further comprises:

[0040] The embedding module is configured to embed gender information, timbre feature information and boundary information of the target user into the speech synthesis model after the second training module trains the base model by using speech samples of a plurality of languages in a preset proportion corresponding to the language proficiency level of the target user, word speech samples of the plurality of languages and pronunciation factor samples of the plurality of languages, to obtain a speech synthesis model, wherein the boundary information is used to identify whether the speech languages in the speech samples are the same.

[0041] In a fifth aspect, an electronic device is provided, which includes a processor and a memory. The memory stores computer readable instructions. When the computer readable instructions are executed by the processor, the steps in the method provided in the first aspect or the second aspect are performed.

[0042] In a sixth aspect, a readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps in the method provided in the first aspect or the second aspect are performed.

[0043] Other features and advantages of the present application will be described in the following description, and in part will become apparent from the description, or be learned from the practice of the application. The purposes and other advantages of the present application can be realized and obtained by the structure particularly pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF DRAWINGS

[0044] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments of the present application. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be considered as limiting the scope. For those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.

[0045] Figure 1 A flowchart of a speech conversion method provided by the embodiments of the present application;

[0046] Figure 2 A flowchart of a method for training a speech synthesis model provided by the embodiments of the present application;

[0047] Figure 3 A flowchart of an implementation method for training a speech synthesis model provided by the embodiments of the present application;

[0048] Figure 4 A schematic block diagram of a speech conversion device provided by the embodiments of the present application;

[0049] Figure 5 A schematic block diagram of a device for training a speech synthesis model provided by the embodiments of the present application;

[0050] Figure 6 A structural schematic diagram of a device for voice conversion provided by an embodiment of the present application is shown in FIG. 1.

[0051] Figure 7 A structural schematic diagram of a device for training a speech synthesis model provided by an embodiment of the present application is shown in FIG. 2. DETAILED DESCRIPTION

[0052] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings of the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments of the present application. The components of the embodiments of the present application described and shown in the accompanying drawings can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of the present application.

[0053] First, some terms involved in the embodiments of the present application will be described to facilitate understanding by those skilled in the art.

[0054] Basic model: a model trained with general data, which can be understood as a speech synthesis model in the traditional sense.

[0055] Multiple languages: here refers to the native language of the user, for example, we customize a voice of a user who does not speak English, and make the model speak English. At this time, the native language is Chinese, and the target language is English. If we customize a voice of an English speaker to make him speak Chinese, then the native language is English, and the target language is Chinese.

[0056] It should be noted that similar reference numerals and letters refer to similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. Meanwhile, in the description of the present application, the terms “first”, “second”, etc. are only used for differentiation, and cannot be understood as indicating or implying relative importance.

[0057] The present application is applied to the scenario of voice conversion, and the specific scenario is to train a model by obtaining different kinds of speech of a user, to make the model learn the pronunciation of the user, and to convert the text input into the model to obtain the same pronunciation as the user.

[0058] However, the existing voice synthesis mainly relies on a certain degree of post-processing and uses similar sounds to replace the user's voice to let the model complete the task of voice synthesis. When using the above method for voice synthesis, the voice timbre output by the final model is inaccurate or the pronunciation of the synthesized voice is inaccurate.

[0059] To this end, the application takes the target user's text to be converted, converts the text to be converted by a voice synthesis model to obtain the voice corresponding to the text to be converted, wherein the voice synthesis model is obtained by training a base model according to the language ability level of the target user, through a plurality of voice samples of a plurality of languages in a preset proportion corresponding to the language ability level, a plurality of word voice samples of a plurality of languages, and a plurality of pronunciation factor samples of a plurality of languages. Different language ability levels correspond to different preset proportions, and the base model is obtained by training a general model through a plurality of voice samples of a plurality of languages in a base preset proportion and a plurality of mixed voice samples of a plurality of languages. Training the model through the user's plurality of voice samples of a plurality of languages, a plurality of word voice samples of a plurality of languages, and a plurality of pronunciation factor samples of a plurality of languages can make the voice synthesis model learn the user's pronunciation, and then the voice synthesis model can convert the text to be converted into corresponding voice with the user's pronunciation timbre. Through this method, the accuracy of voice conversion can be improved.

[0060] In the embodiments of the application, the execution subject can be a voice conversion device in a voice conversion system. In actual applications, the voice conversion device can be an electronic device such as a terminal device and a server, which is not limited here.

[0061] The voice conversion method of the embodiments of the application will be described in detail below. Figure 1 The voice conversion method of the embodiments of the application will be described in detail below.

[0062] Please refer to Figure 1 , Figure 1 The flowchart of the voice conversion method provided in the embodiments of the application is shown in Figure 1 The voice conversion method includes the following steps.

[0063] Step 110: Obtain the text to be converted of the target user.

[0064] The text to be converted can be a plurality of language texts.

[0065] In some embodiments of the application, the language ability level includes a first language ability level, a second language ability level, and a third language ability level, wherein the first language ability level indicates that the target user masters a plurality of languages, the second language ability level indicates that the target user masters part of the plurality of languages, and the third language ability level indicates that the target user does not master any of the plurality of languages.

[0066] The application can set the sample corresponding to the preset proportion according to the different language proficiency levels of the user, train the corresponding speech synthesis model for different users, and save the resources for model training.

[0067] The first language proficiency level can represent that the user masters multiple languages, and can be that the target user masters one non-native language. The second language proficiency level can represent that the target user masters part of the multiple languages, or that the target user masters part of the common language in one language. The third language proficiency level can represent that the target user does not master any of the multiple languages, or that the target user only masters simple language in addition to the native language.

[0068] In some embodiments of the application, before obtaining the text to be converted of the target user, Figure 1 The method also includes obtaining the language proficiency level of the target user and the speech samples of multiple languages corresponding to the preset proportion of the language proficiency level, the word speech samples of multiple languages, and the pronunciation factor samples of multiple languages; training the general model by using the speech samples of multiple languages corresponding to the basic preset proportion and the mixed speech samples of multiple languages to obtain a basic model; and training the basic model by using the speech samples of multiple languages corresponding to the preset proportion, the word speech samples of multiple languages, and the pronunciation factor samples of multiple languages to obtain a speech synthesis model.

[0069] In the above process, the application trains the model by using the speech samples of multiple languages, the word speech samples of multiple languages, and the pronunciation factor samples of multiple languages, so that the speech synthesis model can learn the pronunciation of the user.

[0070] The language proficiency level and the preset proportion corresponding to the language proficiency level are, for example, the speech sample of a certain language in the first language proficiency level can be 30%-50%, the speech sample of multiple languages can be 20%-40%, the word speech sample of multiple languages can be 5%-15%, and the pronunciation factor sample of multiple languages can be 5%-10%. The speech sample of a certain language in the second language proficiency level can be 5%-20%, the speech sample of multiple languages can be 20%-30%, the word speech sample of multiple languages can be 25%-45%, and the pronunciation factor sample of multiple languages can be 15%-25%. The speech sample of a certain language in the third language proficiency level can be 0%, the speech sample of multiple languages can be 5%-10%, the word speech sample of multiple languages can be 10%-20%, and the pronunciation factor sample of multiple languages can be 70%-90%. Each sample is composed of speech and text, for example, the speech sample of the English language is, for example, "what is your name?" and the corresponding standard speech, the speech sample of multiple languages is, for example, "we really long time no see." and the corresponding standard speech, the word speech sample of multiple languages is, for example, "happy, name, how are you, nice" and the corresponding standard speech, and the pronunciation factor sample of multiple languages is, for example, "A, B, X, ei" and the corresponding standard speech. Phonemes are the basic units of pronunciation, for example, the pronunciation of our letters, A is a single phoneme: ei can emit A. For the letter X, multiple phoneme combinations are needed, such as "ai k si". It can be understood that the phoneme is the most fine-grained element of a speech.

[0071] In addition, when training the base model, the function is replaced from relu (linear rectifier function) to gelu (neural network activation function). In the training process of the base model, in addition to the embedding of speech information, gender information embedding and user voice ID (identification) embedding are also added, and when training the base model, a redundant number of ID information is set to ensure the common use of multiple users. Data preparation, for example, multiple language speech samples, such as Chinese speech data: 30%-45%, English speech data: 30%-45%, and mixed language speech samples, which can also be Chinese-English mixed speech: 10%-40%.

[0072] In some embodiments of the present application, the target user's text to be converted is obtained, including: preprocessing the initial text to be converted to obtain the text to be converted, wherein the preprocessing includes at least one of the following processing methods: cleaning, deleting, intercepting and completing.

[0073] In the above process, the text obtained after preprocessing is relatively standard, and the speech conversion is also more accurate.

[0074] Step 120: converting the text to be converted by the speech synthesis model to obtain the speech corresponding to the text to be converted.

[0075] The speech synthesis model is obtained by training a basic model according to a language proficiency level of the target user, through speech samples of multiple languages in a preset proportion corresponding to the language proficiency level, word speech samples of multiple languages, and pronunciation factor samples of multiple languages. Different language proficiency levels correspond to different preset proportions, and the basic model is obtained by training a general model through speech samples of multiple languages in a basic preset proportion and mixed speech samples of multiple languages.

[0076] In the above Figure 1 The method comprises the following steps: taking the text to be converted of the target user; converting the text to be converted by the speech synthesis model to obtain the speech corresponding to the text to be converted. The speech synthesis model is obtained by training a basic model according to a language proficiency level of the target user, through speech samples of multiple languages in a preset proportion corresponding to the language proficiency level, word speech samples of multiple languages, and pronunciation factor samples of multiple languages. Different language proficiency levels correspond to different preset proportions, and the basic model is obtained by training a general model through speech samples of multiple languages in a basic preset proportion and mixed speech samples of multiple languages. Training the model through speech samples of multiple languages, word speech samples of multiple languages, and pronunciation factor samples of multiple languages of the user can enable the speech synthesis model to learn the pronunciation of the user, and then the text to be converted can be converted into corresponding speech with the pronunciation and tone of the user by the speech synthesis model. This method can improve the accuracy of speech conversion.

[0077] The following will be described in detail Figure 2 The method for training the speech synthesis model of the embodiment of the present application is described in detail.

[0078] Please refer to Figure 2 , Figure 2 A flowchart of the method for training the speech synthesis model provided by the embodiment of the present application is shown in Figure 2 The method for training the speech synthesis model comprises the following steps:

[0079] Step 210: training a general model through speech samples of multiple languages in a basic preset proportion and mixed speech samples of multiple languages to obtain a basic model.

[0080] Step 220: training the basic model through speech samples of multiple languages in a preset proportion corresponding to the language proficiency level of the target user, word speech samples of multiple languages, and pronunciation factor samples of multiple languages to obtain the speech synthesis model.

[0081] Different language proficiency levels correspond to different preset proportions.

[0082] In the above process, the model is trained by the user's voice samples in multiple languages, word voice samples in multiple languages, and pronunciation factor samples in multiple languages, so that the voice synthesis model can learn the user's pronunciation.

[0083] In some embodiments of the present application, after training the base model by voice samples in multiple languages, word voice samples in multiple languages, and pronunciation factor samples in multiple languages corresponding to the preset proportion of the target user's language proficiency level, the voice synthesis model is obtained, Figure 2 The method also includes embedding the target user's gender information, timbre feature information, and boundary information into the voice synthesis model, wherein the boundary information is used to identify whether the voice languages in the voice samples are the same.

[0084] In the above process, the model can synthesize more accurate voice through feature embedding of gender information, timbre feature information, and boundary information.

[0085] The boundary information is used to indicate whether the next unit and the current unit are the same language in the smallest unit. If they are, the boundary information is 0 (indicating that there is no language boundary conversion between the two smallest units), and if they are not, the boundary information is 1 (indicating that there is a language boundary conversion between the two smallest units). The smallest unit is each word in a sentence, and the English word is a word, such as "we really long time no see ah", which is converted into the smallest unit set: ["I", "we", "really", "is", "long", "time", "no", "see", "ah"]. At this time, the corresponding boundary information is [0, 0, 0, 1, 0, 0, 0, 1, 0], which corresponds to the above smallest unit set one by one, as follows: I: 0, we: 0, really: 0, is: 1, long: 0, time: 0, no: 0, see: 1, and ah: 0. As can be seen, the boundary information at "is" and "see" is 1, indicating that the language changes from them.

[0086] The following will be described in conjunction with Figure 3 The implementation method for training the voice synthesis model of the embodiments of the present application is described in detail.

[0087] Please refer to Figure 3 , Figure 3 A flowchart of an implementation method for training a voice synthesis model provided by the embodiments of the present application is shown in Figure 3 The implementation method for training a voice synthesis model includes:

[0088] Step 310: Select the base data.

[0089] Specifically, the voice data of multiple languages specified by the user is selected.

[0090] Step 320: embedding feature information.

[0091] Specifically, the gender information, voice timbre information and basic data of the user.

[0092] Step 330: training a basic model.

[0093] Specifically, the basic model is obtained by training the general model through the voice samples of multiple languages in the preset proportion and the mixed voice samples of multiple languages.

[0094] Step 340: determining a language proficiency level.

[0095] Specifically, the language proficiency level of the user can be selected by the user or determined by testing.

[0096] Step 350: selecting model migration data.

[0097] Specifically, the voice samples of multiple languages, the word voice samples of multiple languages and the pronunciation factor samples of multiple languages are selected in the preset proportion corresponding to the language proficiency level of the user.

[0098] Step 360: embedding feature information.

[0099] Specifically, the model migration data, the gender and voice timbre information of the user are embedded.

[0100] Step 370: whether it is a pure language.

[0101] Specifically, if it is not a pure language, go to step 380, and if it is a pure language, go to step 390.

[0102] Step 380: embedding boundary information.

[0103] Specifically, the voice samples of multiple languages, the word voice samples of multiple languages and the pronunciation factor samples of multiple languages are marked with boundaries in the preset proportion corresponding to the language proficiency level of the user.

[0104] Step 390: training a voice synthesis model.

[0105] Specifically, the basic model is trained through the voice samples of multiple languages, the word voice samples of multiple languages and the pronunciation factor samples of multiple languages in the preset proportion corresponding to the language proficiency level.

[0106] In addition, Figure 3 The methods and steps shown can refer to Figure 1 and Figure 2The methods and steps shown will not be elaborated further here.

[0107] The previous text passed Figures 1-3 The methods for speech conversion and training speech synthesis models are described below. Figures 4-7 A device for speech conversion.

[0108] Please refer to Figure 4 This is a schematic block diagram of a speech conversion device 400 provided in an embodiment of this application. The device 400 may be a module, program segment, or code on an electronic device. This device 400 is related to the above... Figure 1 The method implementation corresponds to this and can be executed. Figure 1 The various steps involved in the method embodiment, and the specific functions of the device 400, can be found in the following description. To avoid repetition, detailed descriptions are omitted here.

[0109] Optionally, the device 400 includes:

[0110] Module 410 is used to acquire the text to be converted from the target user;

[0111] The conversion module 420 is used to convert the text to be converted using a speech synthesis model to obtain the speech corresponding to the text to be converted. The speech synthesis model is trained on a basic model based on the target user's language ability level by using speech samples from multiple languages, word speech samples from multiple languages, and pronunciation factor samples from multiple languages ​​at a preset ratio corresponding to the language ability level. Different preset ratios correspond to different language ability levels. The basic model is trained on a general model using speech samples from multiple languages ​​at a preset ratio and mixed speech samples from multiple languages.

[0112] Optional language proficiency levels include:

[0113] The levels are divided into three categories: first language proficiency level, second language proficiency level, and third language proficiency level. The first language proficiency level indicates that the target user has mastered multiple languages, the second language proficiency level indicates that the target user has mastered some of the multiple languages, and the third language proficiency level indicates that the target user has not mastered any of the multiple languages.

[0114] Optionally, the acquisition module is specifically used for:

[0115] The initial text to be converted is preprocessed to obtain the text to be converted. The preprocessing includes at least one of the following methods: cleaning, deletion, truncation and completion.

[0116] Optionally, the device further includes:

[0117] The training module is configured to, before the acquisition module acquires the to-be-converted text of the target user, acquire the language proficiency level of the target user and the preset proportion of the plurality of languages corresponding to the language proficiency level, and acquire the plurality of language phonetic samples, the plurality of language word phonetic samples and the plurality of language pronunciation factor samples; train a general model by using the plurality of language phonetic samples of the preset proportion and the mixed language phonetic samples to obtain a basic model; and train the basic model by using the plurality of language phonetic samples, the plurality of language word phonetic samples and the plurality of language pronunciation factor samples to obtain a speech synthesis model.

[0118] Please refer to Figure 5 The device 500 provided in the embodiment of the present application is a schematic block diagram of a device for training a speech synthesis model, which can be a module, a program segment or code on an electronic device. The device 500 corresponds to the method embodiment described above and can perform each step involved in the method embodiment. The specific functions of the device 500 can be found in the description below, and the detailed description is appropriately omitted here to avoid repetition. Figure 2 The device 500 corresponds to the method embodiment described above and can perform each step involved in the method embodiment. The specific functions of the device 500 can be found in the description below, and the detailed description is appropriately omitted here to avoid repetition. Figure 2 The device 500 corresponds to the method embodiment described above and can perform each step involved in the method embodiment. The specific functions of the device 500 can be found in the description below, and the detailed description is appropriately omitted here to avoid repetition.

[0119] Optionally, the device 500 comprises:

[0120] The first training module 510 is configured to train a general model by using the plurality of language phonetic samples of the preset proportion and the mixed language phonetic samples to obtain a basic model.

[0121] The second training module 520 is configured to train the basic model by using the plurality of language phonetic samples, the plurality of language word phonetic samples and the plurality of language pronunciation factor samples of the preset proportion corresponding to the language proficiency level of the target user to obtain a speech synthesis model, wherein the preset proportions corresponding to different language proficiency levels are different.

[0122] Optionally, the device further comprises:

[0123] The embedding module is configured to, after the second training module trains the basic model by using the plurality of language phonetic samples, the plurality of language word phonetic samples and the plurality of language pronunciation factor samples of the preset proportion corresponding to the language proficiency level of the target user to obtain a speech synthesis model, embed the gender information, the timbre feature information and the boundary information of the target user into the speech synthesis model, wherein the boundary information is used to identify whether the language of the speech samples is the same.

[0124] Please refer to Figure 6 The device provided in the embodiment of the present application is a structural schematic block diagram of a device for speech conversion, which can comprise a memory 610 and a processor 620. Optionally, the device can further comprise a communication interface 630 and a communication bus 640. The device corresponds to the method embodiment described aboveFigure 1 The method embodiments correspond to, and can be executed by Figure 1 The various steps involved in the method embodiments, and the specific functions of the device, can be seen from the description below.

[0125] Specifically, the memory 610 is configured to store computer readable instructions.

[0126] The processor 620 is configured to process the computer readable instructions stored in the memory, and can execute Figure 1 The various steps in the method.

[0127] The communication interface 630 is configured to communicate signaling or data with other node devices. For example, the communication interface 630 can be configured to communicate with a server or a terminal, or communicate with other device nodes, and the embodiments of the present application are not limited thereto.

[0128] The communication bus 640 is configured to realize direct connection communication among the above components.

[0129] In the embodiments of the present application, the communication interface 630 of the device is configured to communicate signaling or data with other node devices. The memory 610 can be a high-speed RAM memory, or a non-volatile memory such as at least one disk memory. The memory 610 can optionally be at least one storage device located away from the aforementioned processor. The memory 610 stores computer readable instructions, and when the computer readable instructions are executed by the processor 620, the electronic device executes the method process shown above. The processor 620 can be used in the device 400, and is configured to execute the functions in the present application. For example, the processor 620 described above can be a general processor, a digital signal processor (Digital Signal Processor, DSP), an application specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field programmable gate array (Field Programmable Gate Array, FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, and the embodiments of the present application are not limited thereto. Figure 1 The processor 620 can be used in the device 400, and is configured to execute the functions in the present application. For example, the processor 620 described above can be a general processor, a digital signal processor (Digital Signal Processor, DSP), an application specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field programmable gate array (Field Programmable Gate Array, FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, and the embodiments of the present application are not limited thereto.

[0130] Please refer to Figure 7 A structure schematic block diagram of a device 700 for training a speech synthesis model is provided in the embodiments of the present application. The device can include a memory 710 and a processor 720. Optionally, the device can further include a communication interface 730 and a communication bus 740. The device corresponds to the above Figure 2 The method embodiments correspond to, and can be executed by Figure 2The method embodiment relates to each step, and the specific function of the device can be referred to the description below.

[0131] Specifically, the memory 710 is used for storing computer readable instructions.

[0132] The processor 720 is used for processing the readable instructions stored in the memory, and can execute Figure 2 Each step in the method.

[0133] The communication interface 730 is used for signaling or data communication with other node devices. For example, the communication interface 730 is used for communication with a server or a terminal, or communication with other device nodes, and the embodiment of the present application is not limited to this.

[0134] The communication bus 740 is used for realizing direct connection communication of the above-mentioned components.

[0135] In the embodiment of the present application, the communication interface 730 of the device is used for signaling or data communication with other node devices. The memory 710 can be a high-speed RAM memory, or a non-volatile memory such as at least one disk memory. The memory 710 can also be at least one storage device located away from the above-mentioned processor. The memory 710 stores computer readable instructions, and when the computer readable instructions are executed by the processor 720, the electronic device executes the above-mentioned Figure 2 The processor 720 can be used on the device 500, and is used for executing the functions in the present application. The processor 720 described above can be a general processor, a digital signal processor (Digital Signal Processor, DSP), an application specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field programmable gate array (Field Programmable Gate Array, FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and the embodiment of the present application is not limited to this.

[0136] The embodiment of the present application also provides a readable storage medium, and when the computer program is executed by the processor, the method process executed by the electronic device in the method embodiment as Figure 1 Or Figure 2 The method process executed by the electronic device in the method embodiment as

[0137] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the device described above can refer to the corresponding process in the foregoing method, and will not be described in detail here.

[0138] To sum up, the embodiment of the present application provides a speech conversion method, a speech synthesis model training method and device. The method comprises the following steps: obtaining a text to be converted of a target user; converting the text to be converted by using a speech synthesis model to obtain a speech corresponding to the text to be converted, wherein the speech synthesis model is obtained by training a basic model by using speech samples of multiple languages in a preset proportion corresponding to a language ability level of the target user, word speech samples of the multiple languages and pronunciation factor samples of the multiple languages, wherein the preset proportion corresponding to different language ability levels is different, and the basic model is obtained by training a general model by using speech samples of multiple languages in a basic preset proportion and mixed speech samples of the multiple languages. The method can improve the accuracy of speech conversion.

[0139] In several embodiments provided in the present application, it should be understood that the disclosed apparatus and method can also be implemented by other manners. The apparatus embodiments described above are only schematic, for example, the flow charts and block diagrams in the drawings show the possible implementation architectures, functions and operations of the apparatus, method and computer program product according to the embodiments of the present application. In this regard, each block in the flow charts or block diagrams can represent a module, a program segment or a part of code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that, in some alternative implementation manners, the functions noted in the blocks can occur in different orders from that noted in the drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and sometimes they can be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flow charts, and the combination of blocks in the block diagrams and / or flow charts, can be implemented by a dedicated hardware-based system for implementing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0140] In addition, each functional module in the embodiments of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0141] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the parts of the prior art that make contributions or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0142] The above merely provides an example of the present application and is not intended to limit the protection scope of the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application. It should be noted that similar reference numbers and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0143] The above merely provides an example of the present application and is not intended to limit the protection scope of the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application. It should be noted that similar reference numbers and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0144] It should be noted that, in this document, the terms such as first and second are used merely to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the processes, methods, articles or devices including a series of elements not only include those elements, but also include other elements not explicitly listed or inherent to such processes, methods, articles or devices. Without more limitations, the element defined by the statement "including a" does not exclude the presence of another identical element in the process, method, article or device including the element.

Claims

1. A method for speech conversion, characterized in that, include: Obtain the text to be converted from the target user; The text to be converted is converted using a speech synthesis model to obtain the corresponding speech. The speech synthesis model is trained on a base model based on the target user's language proficiency level using speech samples from multiple languages ​​at a preset ratio corresponding to the language proficiency level, word speech samples from the multiple languages, and pronunciation factor samples from the multiple languages. The preset ratios are different for different language proficiency levels. The base model is trained on a general model using speech samples from the multiple languages ​​at a preset ratio and mixed speech samples from the multiple languages.

2. The method according to claim 1, characterized in that, The language proficiency level includes: The system comprises a first language proficiency level, a second language proficiency level, and a third language proficiency level, wherein the first language proficiency level indicates that the target user has mastered the multiple languages, the second language proficiency level indicates that the target user has mastered some of the multiple languages, and the third language proficiency level indicates that the target user has not mastered any of the multiple languages.

3. The method according to claim 1 or 2, characterized in that, The process of obtaining the target user's text to be converted includes: The initial text to be converted is preprocessed to obtain the text to be converted, wherein the preprocessing includes at least one of the following processing methods: cleaning, deletion, truncation and completion.

4. The method according to claim 1 or 2, characterized in that, Before obtaining the text to be converted from the target user, the method further includes: Obtain the target user's language proficiency level and the corresponding preset proportion of speech samples, word speech samples, and pronunciation factor samples of the multiple languages ​​in the multiple languages; The general model is trained using speech samples of the multiple languages ​​and mixed speech samples of the multiple languages ​​at a predetermined ratio to obtain the basic model. The speech synthesis model is obtained by training the base model using speech samples of the multiple languages, word speech samples of the multiple languages, and pronunciation factor samples of the multiple languages ​​at the preset ratio.

5. A method for training a speech synthesis model, characterized in that, include: The general model is trained by using speech samples from multiple languages ​​at a predetermined ratio and mixed speech samples from the multiple languages ​​to obtain the basic model. The basic model is trained by using speech samples, word speech samples, and pronunciation factor samples of the multiple languages ​​at a preset ratio corresponding to the target user's language proficiency level to obtain a speech synthesis model. The preset ratios are different for different language proficiency levels.

6. The method according to claim 5, characterized in that, After training the base model with speech samples from multiple languages, word speech samples from multiple languages, and pronunciation factor samples from multiple languages ​​at a preset ratio corresponding to the target user's language proficiency level to obtain a speech synthesis model, the method further includes: The gender information, timbre feature information, and boundary information of the target user are embedded into the speech synthesis model, wherein the boundary information is used to identify whether the speech samples are of the same language.

7. A speech conversion device, characterized in that, include: The acquisition module is used to acquire the text to be converted from the target user; The conversion module is used to convert the text to be converted using a speech synthesis model to obtain the speech corresponding to the text to be converted. The speech synthesis model is trained on a base model based on the target user's language proficiency level using speech samples from multiple languages ​​at a preset ratio corresponding to the language proficiency level, word speech samples from the multiple languages, and pronunciation factor samples from the multiple languages. The preset ratios are different for different language proficiency levels. The base model is trained on a general model using speech samples from the multiple languages ​​at a preset ratio and mixed speech samples from the multiple languages.

8. An apparatus for training a speech synthesis model, characterized in that, include: The first training module is used to train the general model using speech samples from multiple languages ​​and mixed speech samples from the multiple languages ​​at a preset ratio, so as to obtain the basic model. The second training module is used to train the base model using speech samples, word speech samples, and pronunciation factor samples of the multiple languages ​​at a preset ratio corresponding to the target user's language proficiency level, to obtain a speech synthesis model, wherein the preset ratios corresponding to different language proficiency levels are different.

9. An electronic device, characterized in that, include: A memory and a processor, the memory storing computer-readable instructions that, when executed by the processor, perform the steps of the method as described in any one of claims 1-4 or 5-6.

10. A computer-readable storage medium, characterized in that, include: A computer program that, when run on a computer, causes the computer to perform the method as described in any one of claims 1-4 or 5-6.

Citation Information

Patent Citations

  • Speech synthesis method and speech synthesis device

    CN105845125A

  • Voice synthesis related system, method and device and equipment

    CN113870833A