Speech synthesis related systems, methods, apparatuses, and devices

By generating multilingual speech datasets through cross-language speech conversion algorithms, a multilingual speech synthesizer was constructed, solving the problems of inconsistent timbre and inaccurate pronunciation in multilingual speech synthesis, improving user experience and reducing costs.

CN113870833BActive Publication Date: 2025-12-09ALIBABA GROUP HOLDING LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202010617107.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-06-30
Publication Date
2025-12-09
Estimated Expiration
2041-06-08

AI Technical Summary

Technical Problem

Existing multilingual speech synthesis technologies suffer from inconsistent timbre and inaccurate pronunciation, resulting in a poor user experience, especially when switching between multiple languages. Furthermore, hiring professional multilingual speakers is costly.

Method used

A cross-language speech conversion algorithm is used to generate a multilingual speech dataset with the target user's voice timbre. A multilingual speech synthesizer for the target user is then constructed. The speech synthesizer generates high-quality multilingual speech data, avoiding the problem of inconsistent voice timbre when switching languages.

Benefits of technology

It improves the quality and user experience of multilingual speech synthesis, saves the cost of hiring professional multilingual speakers, and the synthesized speech retains the timbre of the target speaker, approaching the pronunciation level of a native speaker of the foreign language.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113870833B_ABST
    Figure CN113870833B_ABST
Patent Text Reader

Abstract

The application discloses a speech interaction related system, method, device and equipment. Among them, the speech synthesis method generates a second speech data set of a first language of a first user with a first user voice by a cross-language speech conversion algorithm of the first user according to a first speech data set of the first language of a second user; generates a speech synthesizer with multi-language capability of the first user according to the second speech data set and a third speech data set of a second language of the first user; and generates speech synthesis data of the first user corresponding to a first multi-language mixed text through the speech synthesizer. This processing method can effectively improve the speech synthesis quality of multi-language text and improve user experience.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of speech synthesis, in particular to a speech synthesis method and device, an online text-to-speech synthesis system, method and device, a speech interaction system, method and device, a news broadcast system, method and device, and an electronic device. BACKGROUND

[0002] With the rapid development of speech synthesis technology and the increasing popularity of applications, speech synthesis services are rapidly expanding and are increasingly accepted and used by users. With the improvement of users' education, more and more application scenarios involve multilingual content, especially mixed reading of Chinese and English. Therefore, there is a demand for multilingual speech synthesis services, which has driven the development of related technologies.

[0003] A typical multilingual speech synthesis system adopts the following processing method: first, based on different languages, a plurality of sets of speech synthesizers are established using different speaker data modeling methods; then, according to different languages in the text, the corresponding synthesizers are switched for use to complete the synthesis task. Another typical multilingual speech synthesis system adopts a scheme of directly mapping non-native phonetic symbols into a native phonetic symbol system, and then using a native speech synthesizer to synthesize speech. The currently popular solution is to collect target speaker multilingual data for modeling.

[0004] However, in the process of implementing the present application, the inventors have found that the above technical solutions at least have the following problems: 1) the first method above usually causes problems of inconsistent tone and prosody, which affects the naturalness of speech synthesis and user experience; 2) in the second method above, since the mapping relationship of phonetic symbols is only based on simple pronunciation similarity, the synthesized non-native speech will have obvious pronunciation inaccuracies or even errors, resulting in unnatural overall effect; 3) in the third method above, most speakers are not fluent in languages other than their native language, and have a heavy accent. The model trained using such data will not be standard in synthesizing speech in non-native languages, which reduces user experience. In addition, hiring professional multilingual speakers is also costly. Moreover, some popular speakers may not be proficient in bilingual or multilingual, which makes it very difficult to obtain a high-quality multilingual speech synthesizer for a specific target speaker.

[0005] In summary, how to improve the quality of multilingual speech synthesis to synthesize natural, accurate and tone-unified multilingual speech from multilingual text is still a problem to be solved. SUMMARY

[0006] The application provides a speech synthesis method to solve the problem of low speech synthesis quality of multi-lingual text in the prior art. The application further provides a speech synthesis device, an online text speech synthesis system, method and device, a speech interaction system, method and device, a news broadcast system, method and device, and an electronic device.

[0007] The application provides a speech interaction system, comprising:

[0008] The smart speaker is configured to collect user voice data, send the user voice data to a server, and play response voice data returned by the server.

[0009] The server is configured to generate a second voice data set of a first language of a target user with a target user voice by a cross-language speech conversion algorithm according to a first voice data set of the first language of the target user, generate a speech synthesizer with multi-language capability of the target user according to the second voice data set and a third voice data set of a second language of the target user, determine a response text in the same language as the user voice data, and generate response voice data corresponding to the response text by the speech synthesizer.

[0010] The application further provides an online text speech synthesis system, comprising:

[0011] The terminal device is configured to send a first user speech synthesis request for a target multi-lingual mixed text to a server.

[0012] The server is configured to generate a second voice data set of a first language of a first user with a first user voice by a cross-language speech conversion algorithm according to a first voice data set of the first language of the second user, generate a speech synthesizer with multi-language capability of the first user according to the second voice data set and a third voice data set of a second language of the first user, and generate speech synthesis data corresponding to the mixed text by the speech synthesizer.

[0013] The application further provides a news broadcast system, comprising:

[0014] The terminal device is configured to send a request for a multi-lingual broadcast text to a server, and play multi-lingual voice data corresponding to the text to be broadcast and broadcasted by a target user returned by the server.

[0015] The server is configured to generate, by a cross-lingual speech conversion algorithm, a second speech data set of a first language with a target user's voice timbre according to a first speech data set of the first language; generate a multi-lingual speech synthesizer of the target user according to the second speech data set and a third speech data set of a second language of the target user; and generate, by the speech synthesizer, multi-lingual speech data corresponding to a text to be broadcast by the target user.

[0016] The application also provides a speech synthesis method, comprising:

[0017] A second speech data set of a first language of a first user with a first user's voice timbre is generated by a cross-lingual speech conversion algorithm according to a first speech data set of the first language of the second user;

[0018] A multi-lingual speech synthesizer of the first user is generated according to the second speech data set and a third speech data set of a second language of the first user;

[0019] Speech synthesis data of the first user corresponding to a first multi-lingual mixed text is generated by the speech synthesizer.

[0020] Optionally, the speech synthesis data of the first user corresponding to the first multi-lingual mixed text is generated by the speech synthesizer, comprising:

[0021] A pronunciation unit sequence of the first multi-lingual mixed text is determined by a text input module included in the speech synthesizer, wherein the pronunciation units of the text segments of different languages are pronunciation units of the corresponding languages;

[0022] An acoustic feature sequence with the first user's voice timbre is determined by an acoustic feature synthesis network included in the speech synthesizer according to the pronunciation unit sequence;

[0023] The speech synthesis data is generated by a vocoder included in the speech synthesizer according to the acoustic feature sequence.

[0024] Optionally, the Chinese pronunciation unit comprises an initial and a final of Chinese Pinyin, and a tone;

[0025] The English pronunciation unit comprises an English phoneme and a stress;

[0026] Optionally, the English pronunciation unit sequence is determined in the following manner:

[0027] Spaces are inserted between the pronunciation units, and punctuation marks are inserted according to the inter-word speech pause length.

[0028] Optionally, the generating the speech synthesizer with multi-language capability of the first user according to the second speech data set and the third speech data set of the second language of the first user comprises:

[0029] generating a fourth speech data set of the first user in a mixed language according to the second speech data set and the third speech data set;

[0030] generating the speech synthesizer according to the second speech data set, the third speech data set and the fourth speech data set.

[0031] Optionally, the generating the fourth speech data set of the first user in a mixed language according to the second speech data set and the third speech data set of the second language of the first user comprises:

[0032] generating the speech synthesizer of the first user according to the second speech data set and the third speech data set;

[0033] determining a second multi-language mixed text set;

[0034] determining, for each second multi-language mixed text, a pronunciation unit sequence of the second multi-language mixed text by a text input module included in the speech synthesizer of the first user, wherein the pronunciation units of the text segments in different languages are pronunciation units in corresponding languages;

[0035] determining, by an acoustic feature synthesis network included in the speech synthesizer of the first user, an acoustic feature sequence with a voice color of the first user according to the pronunciation unit sequence;

[0036] generating, by a vocoder included in the speech synthesizer of the first user, speech synthesis data of the first user corresponding to the second multi-language mixed text according to the acoustic feature sequence;

[0037] determining the fourth speech data set according to the speech synthesis data of the first user corresponding to the second multi-language mixed text.

[0038] Optionally, the speech synthesizer of the first user comprises a speech synthesizer based on a Transformer model;

[0039] the generating the speech synthesizer according to the second speech data set, the third speech data set and the fourth speech data set comprises:

[0040] generating the acoustic feature synthesis network;

[0041] the generating the acoustic feature synthesis network comprises:

[0042] According to the fourth voice data set, an acoustic feature synthesis network based on a Transformer model is optimized.

[0043] Optionally, the generating the speech synthesizer according to the second voice data set, the third voice data set and the fourth voice data set comprises:

[0044] The acoustic feature synthesis network is generated.

[0045] The generating the acoustic feature synthesis network comprises:

[0046] According to the second voice data set, the third voice data set and the fourth voice data set, an acoustic feature synthesis network based on a Tacotron2 model or a FastSpeech model is generated.

[0047] Optionally, the generating the speech synthesizer with multi-language capability of the first user according to the second voice data set and the third voice data set of the second language of the first user comprises:

[0048] The vocoder is generated according to the third voice data set.

[0049] Optionally, the cross-language speech conversion algorithm comprises a cross-language speech conversion algorithm based on a speech posterior probability graph (PPG).

[0050] The application further provides a speech interaction method, comprising:

[0051] According to at least one first language first voice data set, a target user's second voice data set of the first language with a target user's voice color is generated through a cross-language speech conversion algorithm.

[0052] According to the second voice data set and a third voice data set of a second language of the target user, a speech synthesizer with multi-language capability of the target user is generated.

[0053] For the user voice data sent by the client, a response text in the same language corresponding to the user voice data is determined.

[0054] Through the speech synthesizer, response voice data corresponding to the response text is generated.

[0055] The application further provides a speech interaction method, comprising:

[0056] User voice data is collected, and the user voice data is sent to a server.

[0057] playing the response voice data sent back by the service end; the response voice data is determined in the following manner: the service end generates a second voice data set of a first language of a target user with a target user voice by using a cross-language voice conversion algorithm according to a first voice data set of the first language; generates a multi-language-capable voice synthesizer of the target user according to the second voice data set and a third voice data set of a second language of the target user; and determines a response text in the same language as the user voice data; and generates response voice data corresponding to the response text by using the voice synthesizer.

[0058] The application further provides an online text-to-speech synthesis method, comprising:

[0059] sending a first user voice synthesis request for a target multi-language mixed text to a service end, so that the service end generates a second voice data set of a first language of a first user with a first user voice by using a cross-language voice conversion algorithm according to a first voice data set of the first language of a second user; generates a multi-language-capable voice synthesizer of the first user according to the second voice data set and a third voice data set of a second language of the first user; and generates voice synthesis data corresponding to the mixed text by using the voice synthesizer.

[0060] The application further provides a news broadcasting method, comprising:

[0061] generating a second voice data set of at least one first language with a target user voice by using a cross-language voice conversion algorithm according to a first voice data set of the at least one first language;

[0062] generating a multi-language-capable voice synthesizer of the target user according to the second voice data set and a third voice data set of a second language of the target user;

[0063] generating multi-language voice data corresponding to the text to be broadcast by the target user by using the voice synthesizer according to a request for broadcasting the text in multiple languages sent by a client.

[0064] The application further provides a news broadcasting method, comprising:

[0065] sending a request for broadcasting a text in multiple languages to a service end;

[0066] playing back the multi-lingual voice data corresponding to the text to be broadcasted and broadcasted by the target user, which is returned by the server; the multi-lingual voice data is generated in the following manner: the server generates a second voice data set of at least one first language with the voice color of the target user according to a first voice data set of the at least one first language through a cross-language voice conversion algorithm; generates a voice synthesizer with multi-lingual capability of the target user according to the second voice data set and a third voice data set of a second language of the target user; and generates the multi-lingual voice data corresponding to the text to be broadcasted and broadcasted by the target user through the voice synthesizer.

[0067] The application further provides a voice synthesizer construction method, comprising:

[0068] generating a second voice data set of at least one first language with the voice color of the first user according to a first voice data set of at least one first language of at least one second user through a cross-language voice conversion algorithm;

[0069] generating a voice synthesizer with multi-lingual capability of the first user according to the second voice data set and a third voice data set of a second language of the first user.

[0070] The application further provides a cross-language voice generation method, comprising:

[0071] determining a text to be processed, and sending a voice generation request of the text read by a first user to a server;

[0072] playing back voice data of the text read by the first user returned by the server; the text to be processed comprises a text in a first language or a text mixed with the first language and a second language, and the first language is the native language of the first user.

[0073] The application further provides a cross-dialect voice generation system, comprising:

[0074] a terminal device configured to determine a text to be processed, send a voice generation request of the text read by a first user in a first dialect to a server, and play back first voice data of the text read by the first user in the first dialect returned by the server;

[0075] a server configured to generate a third voice data set of the first dialect with the voice color of the first user according to a second voice data set of the first dialect of a second user through a cross-dialect voice conversion algorithm, generate a voice synthesizer with multi-dialect capability of the first user according to the third voice data set and a fourth voice data set of a second dialect of the first user, and generate the first voice data through the voice synthesizer for the request.

[0076] The application further provides a cross-dialect voice generation method, comprising:

[0077] determining a text to be processed, sending a speech generation request for the text read by a first user in a first dialect to a server;

[0078] playing first speech data of the text read by the first user in the first dialect returned by the server.

[0079] The application also provides a cross-dialect speech generation method, comprising:

[0080] generating, by a cross-dialect speech conversion algorithm, a third speech data set of the first dialect with the voice of the first user according to a second speech data set of the first dialect of a second user;

[0081] generating a multi-dialect capable speech synthesizer of the first user according to the third speech data set and a fourth speech data set of the second dialect of the first user;

[0082] generating first speech data by the speech synthesizer for a speech generation request sent by the client for the text read by the first user in the first dialect.

[0083] The application also provides a speech synthesizer construction method, comprising:

[0084] generating, by a cross-dialect speech conversion algorithm, a third speech data set of the first dialect with the voice of the first user according to a second speech data set of the first dialect of a second user;

[0085] generating a multi-dialect capable speech synthesizer of the first user according to the third speech data set and a fourth speech data set of the second dialect of the first user.

[0086] The application also provides a speech synthesis device, comprising:

[0087] a training data generation unit configured to generate, by a cross-language speech conversion algorithm, a second speech data set of a first language of a first user with the voice of the first user according to a first speech data set of the first language of a second user;

[0088] a speech synthesizer training unit configured to generate a multi-language capable speech synthesizer of the first user according to the second speech data set and a third speech data set of a second language of the first user;

[0089] a speech synthesis unit configured to generate speech synthesis data of the first user corresponding to a first multi-language mixed text by the speech synthesizer.

[0090] The application also provides an electronic device, comprising:

[0091] a processor; and

[0092] a memory for storing a program implementing the speech synthesis method, the device, after being powered on and running the program of the method by the processor, performs the following steps: generating, by a cross-language speech conversion algorithm, a second speech data set of a first language of a first user with a first user voice according to a first speech data set of the first language of the second user; generating a multi-lingual speech synthesizer of the first user according to the second speech data set and a third speech data set of a second language of the first user; and generating speech synthesis data of the first user corresponding to a first multi-lingual mixed text by the speech synthesizer.

[0093] The application also provides a speech synthesizer construction device, comprising:

[0094] a training data generation unit for generating, by a cross-language speech conversion algorithm, a second speech data set of a first language of a first user with a first user voice according to a first speech data set of at least one first language of at least one second user;

[0095] a speech synthesizer training unit for generating a multi-lingual speech synthesizer of the first user according to the second speech data set and a third speech data set of a second language of the first user.

[0096] The application also provides an electronic device, comprising:

[0097] a processor; and

[0098] a memory for storing a program implementing the speech synthesizer construction method, the device, after being powered on and running the program of the method by the processor, performs the following steps: generating, by a cross-language speech conversion algorithm, a second speech data set of a first language of a first user with a first user voice according to a first speech data set of at least one first language of at least one second user; and generating a multi-lingual speech synthesizer of the first user according to the second speech data set and a third speech data set of a second language of the first user.

[0099] The application also provides a speech synthesizer construction device, comprising:

[0100] a training data generation unit for generating, by a cross-dialect speech conversion algorithm, a third speech data set of a first dialect of a first user with a first user voice according to a second speech data set of a second dialect of a second user;

[0101] a speech synthesizer training unit for generating a multi-dialect speech synthesizer of the first user according to the third speech data set and a fourth speech data set of the second dialect of the first user.

[0102] The application also provides an electronic device, comprising:

[0103] a processor; and

[0104] a memory for storing a program implementing a method for constructing a voice synthesizer, the device, after being powered on and running the program of the method by the processor, performs the following steps: generating a third speech data set of a first dialect with a first user's voice timbre according to a second speech data set of a second dialect of a second user by a cross-dialect speech conversion algorithm; and generating a multi-dialect capable voice synthesizer of the first user according to the third speech data set and a fourth speech data set of the second dialect of the first user.

[0105] The application also provides a computer readable storage medium, which stores instructions, and when the instructions are run on a computer, the computer executes the various methods described above.

[0106] The application also provides a computer program product comprising instructions, and when the instructions are run on a computer, the computer executes the various methods described above.

[0107] Compared with the prior art, the application has the following advantages:

[0108] The speech synthesis method provided by the embodiment of the present application comprises the following steps: a first user's cross-language speech conversion algorithm is used to generate a second speech data set of a first language of the first user with a first user's voice according to a first speech data set of the first language of a second user; a speech synthesizer with multi-language capability of the first user is generated according to the second speech data set and a third speech data set of a second language of the first user; and speech synthesis data of the first user corresponding to a first multi-language mixed text is generated through the speech synthesizer. This processing manner uses the cross-language speech conversion technology to generate high-quality non-native and mixed language data with a target speaker's voice, which is combined with the original native speech data to serve as training data, thereby obtaining a speech synthesizer with bilingual / multi-lingual / hybrid language capability and target speaker's voice, avoiding the problem of inconsistent timbre and unnatural effect caused by switching between different language synthesizers in cross-language and mixed language speech synthesis. Therefore, the speech synthesis quality of multi-lingual text can be effectively improved, thereby improving the user experience. In addition, the system is no longer limited to the native language of the speaker, but only focuses on the timbre of the speaker. As long as the timbre of the speaker is selected and the native speech of the speaker is recorded, the timbre can be extended to other languages, and speech synthesis processing can be performed on any text in other languages. At the same time, this processing manner does not need to use a phonetic symbol mapping method to cross languages, avoiding the problem of inaccurate or even incorrect pronunciation caused by phonetic symbol mapping. In addition, this processing manner makes it possible to build a system only by using single-language databases of different speakers, thereby saving the high cost of hiring professional multi-lingual speakers. Furthermore, this processing manner makes it possible to approach the pronunciation level of a native speaker of a foreign language without affecting the native language performance, and the synthesized speech of different languages well maintains the timbre of the target speaker, so that any (single-language) timbre can be given excellent multi-lingual capability.

[0109] The cross-dialect speech generation system provided in this application uses a terminal device to determine the text to be processed and sends a speech generation request to a server, in which a first user reads the text aloud in a first dialect; and plays the first speech data of the first user reading the text aloud in the first dialect, returned by the server; the server is configured to generate a third speech dataset with the first user's timbre in the first dialect based on a second speech dataset of the first dialect of a second user, using a cross-dialect speech conversion algorithm; generate a speech synthesizer with multi-dialect capability for the first user based on the third speech dataset and a fourth speech dataset of the first user's second dialect; and respond to the request. The system generates first speech data through the speech synthesizer. This processing method utilizes cross-dialect speech conversion technology to generate high-quality language data in one or more dialects with the target speaker's timbre. This data is then merged with the original native dialect recording data and used as training data to obtain a speech synthesizer capable of speaking two, multiple, or mixed dialects. This avoids the inconsistencies and unnatural effects caused by switching between synthesizers for different dialects in cross-dialect and mixed dialect speech synthesis. Therefore, it effectively improves the speech synthesis quality of multi-dialect texts, thereby enhancing the user experience. Furthermore, this system is no longer limited to the speaker's native dialect but focuses solely on the speaker's timbre. By selecting the speaker's timbre and recording their native dialect, the timbre can be extended to other dialects, enabling speech synthesis processing of any text in other dialects. Moreover, this processing method allows for system construction using only single-dialect databases of different speakers, saving the high costs associated with hiring professional multi-dialect speakers. Furthermore, this processing method allows the synthesis of other dialect parts to approach the pronunciation level of the native speaker without affecting the performance of the original dialect. At the same time, the synthesized speech of different dialects maintains the timbre of the target speaker very well. Therefore, it can endow any (single dialect) with excellent timbre and multi-dialect ability. Attached Figure Description

[0110] Figure 1 This application provides a schematic diagram of the structure of an embodiment of a voice interaction system;

[0111] Figure 2 This application provides a scenario diagram of an embodiment of a voice interaction system;

[0112] Figure 3 This application provides a device interaction diagram of an embodiment of a voice interaction system;

[0113] Figure 4 This application provides a flowchart illustrating an embodiment of a speech synthesis method. Detailed Implementation

[0114] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present application. However, the present application can be practiced without the specific details. In other instances, well-known methods have not been described in detail in order not to unnecessarily obscure aspects of the present application.

[0115] In the present application, a speech synthesis method and device, an online text-to-speech synthesis system, method and device, a voice interaction system, method and device, a news broadcast system, method and device, a smart speaker, and an electronic device are provided. Various schemes are described in detail one by one in the following embodiments.

[0116] First embodiment

[0117] Please refer to Figure 1 which is a schematic diagram of an embodiment of the voice interaction system of the present application. The voice interaction system provided in the embodiment includes a server 1 and a smart speaker 2.

[0118] The server 1 can be a server deployed on a cloud server, or a server dedicated to implementing a voice interaction system, which can be deployed in a data center.

[0119] The smart speaker 2 can be a tool for a home consumer to use voice to surf the Internet, such as to play songs on demand, to shop online, or to understand weather forecasts. It can also control smart home devices, such as opening curtains, setting refrigerator temperature, and preheating a water heater, etc.

[0120] Please refer to Figure 2 which is a schematic diagram of a scene of the voice interaction system of the present application. The server 1 and the smart speaker 2 can be connected through a network, such as the smart speaker 2 can be connected through WI FI, etc. The user interacts with the smart speaker through voice. The smart speaker has a dialogue system and can face users from different places or countries and support the user's language during the dialogue. If the user speaks Chinese, the smart speaker can dialogue in Chinese, and if the user speaks English, the smart speaker can dialogue in English. The user issues voice instruction data to the smart speaker 2, the server can determine the response text in the same language corresponding to the user's voice data, and generate response voice data corresponding to the response text through the voice synthesizer of the smart speaker with multi-language capability. The response voice has the same language as the user of the speaker.

[0121] Please refer to Figure 3, which is a device schematic diagram of the voice interaction system of the present application. In the embodiment, the smart speaker is configured to collect user voice data, send the user voice data to the server, and play the response voice data sent back by the server; the server is configured to generate a second voice data set of a first language of a target user with a target user voice color according to a first voice data set of the first language of the target user by using a cross-language voice conversion algorithm, generate a voice synthesizer of the target user with multi-language capability according to the second voice data set and a third voice data set of a second language of the target user, and determine a response text in the same language as the user voice data; and generate response voice data corresponding to the response text by using the voice synthesizer.

[0122] For example, the server of the smart speaker system can first generate a voice synthesizer with Chinese voice synthesis capability according to the Chinese voice data of the target user, so that the server has Chinese response capability and can interact with user A of the speaker end in Chinese; thereafter, user B of another speaker uses English to interact with the speaker, in order to enable the server to have English response capability, the cross-language voice conversion algorithm can be used to generate English voice data (second voice data set) with the target user voice color according to the English voice data (first voice data set) of other users; a voice synthesizer of the target user with Chinese and English voice synthesis capability is generated according to the second voice data set and the Chinese voice data (third voice data set) of the target user; and after determining the English response text corresponding to the English voice data of user B, the English response voice data corresponding to the English response text is generated by using the voice synthesizer; thereafter, user C of another speaker uses French to interact with the speaker, in order to enable the server to have French response capability, the cross-language voice conversion algorithm can be used to generate French voice data (second voice data set) with the target user voice color according to the French voice data (first voice data set) of other users; a voice synthesizer of the target user with Chinese, English and French voice synthesis capability is generated according to the French voice data set, the English voice data set and the Chinese voice data of the target user; and after determining the French response text corresponding to the French voice data of user C, the French response voice data corresponding to the French response text is generated by using the voice synthesizer.

[0123] The specific processing process of the server is described in the second embodiment, which will not be repeated here.

[0124] From the above embodiments, the voice interaction system provided by the embodiments of the present application can collect user voice data through an intelligent sound box, send the user voice data to a server, and play response voice data returned by the server. The server generates a second voice data set of a first language of a target user having a target user voice tone according to a first voice data set of the first language through a cross-language voice conversion algorithm, generates a voice synthesizer of the target user having multi-language capability according to the second voice data set and a third voice data set of a second language of the target user, and determines response text of the same language corresponding to the user voice data. The voice synthesizer generates response voice data corresponding to the response text. This processing manner enables the use of cross-language voice conversion technology to generate high-quality non-native and mixed language data having a target speaker tone, which is combined with original native language recording data as training data, thereby obtaining a voice synthesizer having bilingual / multi-lingual / hybrid language capability, avoiding the problem of inconsistent tone and unnatural effect caused by mutual switching when using different language synthesizers in cross-language and hybrid language voice synthesis. Therefore, the voice synthesis quality of multi-lingual text can be effectively improved, thereby improving user experience. In addition, the system is no longer limited to the native language of the speaker, but only focuses on the tone of the speaker. As long as the tone of the speaker and the native language recording of the speaker are selected, the tone can be expanded to other languages, and any text in other languages can be processed for voice synthesis. At the same time, this processing manner also does not need to use a phonetic symbol mapping method to cross languages, avoiding pronunciation problems caused by phonetic symbol mapping. In addition, this processing manner enables system construction using only single-language databases of different speakers, thereby saving high costs of hiring professional multi-lingual speakers. Furthermore, this processing manner enables the synthesis effect of foreign language parts to approach the pronunciation level of native speakers of foreign languages without affecting the native language performance, and the synthesized voice of different languages well maintains the tone of the target speaker, so that any (single-language) tone can be given excellent multi-lingual capability.

[0125] Second embodiment

[0126] In the above embodiments, a voice interaction system is provided. Correspondingly, the present application also provides a voice synthesis method. The execution subject of the method can be a server or the like. The method corresponds to the above-mentioned system embodiments. The same parts of the present embodiment as the first embodiment are not described again, please refer to the corresponding part in the first embodiment.

[0127] Please refer to Figure 4 which is a flowchart of an embodiment of the voice synthesis method of the present application. In the present embodiment, the method comprises the following steps:

[0128] Step S101: generating, by a cross-language speech conversion algorithm of the first user, a second speech data set of the first language of the first user with the voice color of the first user according to a first speech data set of the first language of the second user.

[0129] The second user can be a plurality of second users other than the first user, that is, the first speech data set can include first speech data of a plurality of second users.

[0130] The method uses a cross-language speech conversion technology to generate non-native and mixed language data with the voice color of a target speaker, and combines the data with the original native language recording data of the user to obtain a speech synthesizer with the voice color of a target speaker with bilingual or multilingual and mixed language capabilities.

[0131] In the embodiment, monolingual recordings of speakers with two different native languages (for example, a Chinese native speaker and an English native speaker) are used to build a Chinese-English bilingual and mixed language speech synthesis system for each speaker, so that Chinese-English bilingual and mixed language speech synthesis tasks can be performed for any of the two speakers, that is, Chinese and English text can be input to synthesize the corresponding speech of the same speaker.

[0132] The cross-language speech conversion algorithm includes but is not limited to a cross-language speech conversion algorithm based on a phonetic posterior probability graph (PPG). In specific implementation, other conventional cross-language speech conversion algorithms can also be used, such as converting the speech signal of the source speaker into corresponding text information, then combining the text information with the speech feature information of the target speaker to synthesize a speech signal with the voice color of the target speaker. Since the cross-language speech conversion algorithm is a relatively mature existing technology, it will not be described here.

[0133] In the embodiment, the cross-language speech conversion algorithm based on PPG can generate high-quality non-native and mixed language data with the voice color of the target speaker. The algorithm can include the following steps: 1) constructing a phonetic posterior probability graph (PPG) feature extractor and a speech synthesis model of the first user; 2) determining PPG feature data of the first speech data according to first acoustic feature data of the first speech data by the PPG feature extractor; the first acoustic feature data includes voiceprint information and speech content information of the second user; 3) generating second speech data of the first language of the first user corresponding to the first speech data by the speech synthesis model of the first user according to the PPG feature data and second acoustic feature data of the first speech data; the second acoustic feature data includes prosody information.

[0134] In particular implementation, the Chinese speaker and the English speaker can be trained respectively using the Chinese and English audio data to obtain a cross-language speech conversion system for each of the Chinese speaker and the English speaker; then the English audio data is converted into the English speech of the Chinese speaker using the cross-language speech conversion system for the Chinese speaker; and the Chinese audio data is converted into the Chinese speech of the English speaker using the cross-language speech conversion system for the English speaker.

[0135] The cross-language speech conversion technology is used in step S101 to generate TTS training corpus of the target speaker (the first user) in a language other than the native language, i.e., the second speech data set.

[0136] In step S103, a multi-language speech synthesizer of the first user is generated according to the second speech data set and a third speech data set of the first user in the second language.

[0137] The method provided in the embodiments of the present application includes two stages of training and speech synthesis, step S103 is the training stage, and step S101 is the training data preparation stage. In the training stage, the model training of the cross-language speech conversion system (cross-language speech conversion algorithm) and the training of the acoustic feature synthesis module and the vocoder in the bilingual and mixed language speech synthesis system (speech synthesizer) of Chinese and English can be performed.

[0138] In the embodiments, the Chinese audio data (the third speech data set) of the Chinese speaker (the first user) and the English speech (the second speech data set) of the Chinese speaker converted by step S101 are used to construct the bilingual and mixed language speech synthesizer of the Chinese speaker.

[0139] In particular implementation, the speech synthesizer can include three main modules: an input representation module, a synthesis network, and a vocoder. The three modules are described below.

[0140] The input representation module is configured to convert normal text into a sequence of pronunciation units, such as the initial and final consonants and the tone of Chinese pronunciation units (pinyin), and the English pronunciation units (phonemes and stress).

[0141] The vocoder is configured to synthesize speech in the form of a waveform from the LPCNet acoustic feature sequence synthesized by the acoustic feature synthesis module. The vocoder can be an LPCNet vocoder or a vocoder of another network. In implementation, the vocoder can be generated based on the third speech data set. That is, the original recording set (native recording set) of the target speaker can be used to train the vocoder.

[0142] The acoustic feature synthesis network is configured to synthesize an LPCNet acoustic feature sequence from the pronunciation sequence processed by the text input module. The module can be based on an existing speech synthesis model structure or designed using other model structures such as Tacotron2, Transformer, and FastSpeech.

[0143] In one example, the speech synthesizer with multilingual capability of the first user is generated based on training data that does not contain Chinese-English mixed text, i.e., based on the second speech data set and the third speech data set. However, experiments show that the three speech synthesis model structures cannot achieve very ideal results using only training data that does not contain Chinese-English mixed text.

[0144] In another example, step S103 can include the following sub-steps:

[0145] Step S1031: generating a fourth speech data set of the first user in mixed languages based on the second speech data set and the third speech data set.

[0146] The fourth speech data set can include speech in multiple languages, and the second speech data set and the third speech data set include speech in a single language.

[0147] In implementation, step S1031 can include the following sub-steps:

[0148] Step S10311: generating a speech synthesizer of the first user based on the second speech data set and the third speech data set.

[0149] The speech synthesizer can be a speech synthesizer based on a Transformer model, i.e., synthesizing Chinese-English mixed speech through a Transformer system. In implementation, the speech synthesizer can also be based on a Tacotron2 model or a FastSpeech model.

[0150] Step S10313: determining a second set of multilingual mixed text;

[0151] Step S10315: For each second multi-language mixed text, determining a sequence of pronunciation units of the second multi-language mixed text by a text input module included in the first user's speech synthesizer, wherein the pronunciation units of the text segments in different languages are pronunciation units in the corresponding languages;

[0152] Step S10317: Determining a sequence of acoustic features with the first user's voice color according to the sequence of pronunciation units by an acoustic feature synthesis network included in the first user's speech synthesizer;

[0153] Step S10319: Generating first user's speech synthesis data corresponding to the second multi-language mixed text according to the sequence of acoustic features by a vocoder included in the first user's speech synthesizer;

[0154] Step S10310: Determining the fourth speech data set according to the first user's speech synthesis data corresponding to the second multi-language mixed text.

[0155] In this embodiment, the training data is expanded based on the Chinese recording of the Chinese speaker and the English speech obtained by conversion, and the speech synthesis system (speech synthesizer) based on the Transformer model. First, the speech synthesis system based on the Transformer model is trained using the training set of the Chinese recording of the Chinese speaker and the English speech obtained by conversion. Then, more than 10,000 Chinese-English mixed texts (second multi-language mixed text set) are prepared, and the Chinese-English mixed texts are synthesized into Chinese-English mixed speech using the above process and the speech synthesis system based on the Transformer model. Then, accurate speech results (fourth speech data set) are selected as new training sets by artificial screening from the synthesized Chinese-English mixed speech, and the screened new Chinese-English mixed training set is added to the training set of the Chinese recording of the Chinese speaker and the English speech obtained by conversion.

[0156] After the fourth speech data set is generated by step S1031, step S1033 can be entered, in which the speech synthesizer is trained using the expanded training set.

[0157] Step S1033: Generating the speech synthesizer according to the second speech data set, the third speech data set and the fourth speech data set.

[0158] With this processing mode, the training of the acoustic feature synthesis network requires three parts of the data set: 1) the original recording of the target speaker (the third speech data set); 2) the speech obtained by cross-language speech conversion (the second speech data set); and 3) the Chinese-English mixed speech (the fourth speech data set).

[0159] In practice, step S1033 can include the following sub-step: generating the acoustic feature synthesis network; the generation of the acoustic feature synthesis network can be achieved in the following manner: generating an acoustic feature synthesis network based on a Transformer model, a Tacotron2 model, or a FastSpeech model, etc., according to the second speech data set, the third speech data set, and the fourth speech data set.

[0160] In this embodiment, the generation of the acoustic feature synthesis network in step S1033 includes: optimizing the acoustic feature synthesis network based on the Transformer model trained in step S1031 according to the fourth speech data set. With this Transformer system, the first training based on the first two parts of the training set (the second speech data set and the third speech data set) can be performed first, and then the third part of the training set (the fourth speech data set) is added for optimization training, so that the construction efficiency of the speech synthesizer can be effectively improved.

[0161] In practice, systems such as Tacotron2 and FastSpeech can be trained at one time using the first three parts of the training set. In this embodiment, three speech synthesis model structures based on Tacotron2, Transformer, and FastSpeech are used respectively, and ideal results are achieved.

[0162] In this embodiment, the English recordings of English speakers and the Chinese speech obtained by conversion can also be used to build a speech synthesis system for English speakers to speak Chinese and mixed languages, and the process is the same as the process of building a speech synthesis system for Chinese speakers to speak Chinese and mixed languages described above, so it will not be repeated here.

[0163] After the training phase is completed through step S103, the system can be used to perform the second phase of the synthesis task.

[0164] Step S105: generating speech synthesis data of the first user corresponding to the first multi-language mixed text by the speech synthesizer.

[0165] In one example, step S105 can include the following sub-steps: determining, by a text input module included in the speech synthesizer, a sequence of pronunciation units of the first multi-language mixed text, wherein the pronunciation units of a text segment in different languages (such as “I am happy” including a Chinese segment “I am” and an English segment “happy”) are pronunciation units in the corresponding language; determining, by an acoustic feature synthesis network included in the speech synthesizer, a sequence of acoustic features with a first user voice color according to the sequence of pronunciation units; and generating, by a vocoder included in the speech synthesizer, the speech synthesis data according to the sequence of acoustic features. The processing manner of each module in the speech synthesizer is described in detail in step S103, and will not be described here again.

[0166] For example, given a pure Chinese, pure English or Chinese-English mixed text, input to the text input module to generate an input sequence of the acoustic feature synthesis network, and then the acoustic feature synthesis network generates an LPCNet acoustic feature sequence, and the LPCNet vocoder synthesizes the Waveform speech (the speech synthesis data) from the LPCNet acoustic feature sequence for playing.

[0167] The Chinese pronunciation unit includes but is not limited to initial and final of Chinese pinyin, and tone; and the English pronunciation unit includes but is not limited to English phoneme and stress. The sequence of English pronunciation units can be determined in the following manner: inserting spaces between pronunciation units, and inserting punctuation symbols according to the length of inter-word speech pauses.

[0168] It should be noted that the three modules in the method: the text input module, the acoustic feature synthesis network, and the LPCNet vocoder can all use other solutions. The text input module can be selected to output other pronunciation sequences, such as using Byte sequence or IPA sequence, the acoustic feature synthesis network can use other model structures, and the alternatives of LPCNet include WaveNet, WaveRNN, etc.

[0169] As can be seen from the above embodiments, the speech synthesis method provided by the embodiments of the present application generates a second speech data set of a first language of a first user with a first user timbre according to a first speech data set of a first language of a second user through a cross-language speech conversion algorithm of the first user; generates a speech synthesizer of the first user with multi-language capability according to the second speech data set and a third speech data set of a second language of the first user; and generates speech synthesis data of the first user corresponding to a first multi-language mixed text through the speech synthesizer. This processing manner enables the use of cross-language speech conversion technology to generate high-quality non-native and mixed language data with a target speaker timbre, and the data is combined with the original native speech data as training data, thereby obtaining a speech synthesizer with bilingual / multi-lingual / hybrid language capability and target speaker timbre, avoiding the problem of inconsistent timbre and unnatural effect caused by switching between different language synthesizers in cross-language and mixed language speech synthesis. Therefore, the speech synthesis quality of multi-lingual text can be effectively improved, thereby improving the user experience. In addition, the system is no longer limited to the native language of the speaker, but only focuses on the timbre of the speaker. As long as the timbre of the speaker is selected and the native language of the speaker is recorded, the timbre can be extended to other languages, and speech synthesis processing can be performed on any text in other languages. At the same time, this processing manner also does not require the use of phonetic symbol mapping and the like to cross languages, avoiding pronunciation problems caused by phonetic symbol mapping. In addition, this processing manner enables system construction using only single-language databases of different speakers, thereby saving the high cost of hiring professional multi-lingual speakers. Furthermore, this processing manner enables the synthesis effect of the foreign language part to approach the pronunciation level of a native speaker of the foreign language without affecting the native language performance, and the synthesized speech of different languages well maintains the timbre of the target speaker, so that any (single-language) timbre can be given excellent multi-lingual capability.

[0170] Third embodiment

[0171] In the above embodiments, a speech synthesis method is provided, and the present application also provides a speech synthesis device corresponding thereto. The device corresponds to the embodiments of the above method. The same parts of the present embodiment as the second embodiment are not described again, please refer to the corresponding part in embodiment two.

[0172] The speech synthesis device provided by the present application comprises:

[0173] The training data generation unit is configured to generate a second speech data set of a first language of a first user with a first user timbre according to a first speech data set of a first language of a second user through a cross-language speech conversion algorithm of the first user;

[0174] a voice synthesizer training unit configured to generate a multi-lingual voice synthesizer of the first user according to the second voice data set and a third voice data set of a second language of the first user;

[0175] a voice synthesis unit configured to generate voice synthesis data of the first user corresponding to the first multi-lingual mixed text by the voice synthesizer.

[0176] a fourth embodiment

[0177] The present application also provides an electronic device. Since the device embodiments are basically similar to the method embodiments, they are described more simply, and the relevant parts are referred to the part of the description of the method embodiments. The device embodiments described below are only illustrative.

[0178] An electronic device of the present embodiment includes a microphone, a processor and a memory, and the memory is configured to store a program for implementing a voice synthesis method. After the device is powered on and the program of the method is run by the processor, the following steps are performed: generating a second voice data set of a first language of a first user with a voice color of the first user according to a first voice data set of the first language of a second user by a cross-lingual voice conversion algorithm; generating a multi-lingual voice synthesizer of the first user according to the second voice data set and a third voice data set of a second language of the first user; and generating voice synthesis data of the first user corresponding to a first multi-lingual mixed text by the voice synthesizer.

[0179] The electronic device can be a smart speaker, a food ordering machine, a vending machine, a ticketing machine, a chat robot, etc.

[0180] a fifth embodiment

[0181] In the above embodiments, a voice interaction system is provided, and the present application also provides a voice interaction method, and the execution subject of the method can be a server, etc. The method corresponds to the above-mentioned system embodiments. The parts of the present embodiment that are the same as those of the first embodiment are not described again, and please refer to the corresponding parts in the first embodiment.

[0182] The voice interaction method provided by the present application can include the following steps:

[0183] Step 1: generating a second voice data set of a first language of a target user with a voice color of the target user according to a first voice data set of at least one first language by a cross-lingual voice conversion algorithm;

[0184] Step 2: generating a multi-lingual voice synthesizer of the target user according to the second voice data set and a third voice data set of a second language of the target user;

[0185] Step 3: determining, for the user voice data sent by the client, a same-language response text corresponding to the user voice data;

[0186] Step 4: generating, by the speech synthesizer, response voice data corresponding to the response text.

[0187] Sixth embodiment

[0188] In the above embodiments, a voice interaction method is provided, and a voice interaction device corresponding to the method is also provided. The device corresponds to the above method embodiments. The same parts of this embodiment as the first embodiment are not described again, please refer to the corresponding parts in the first embodiment.

[0189] The voice interaction device provided in the present application comprises:

[0190] A training data generation unit is configured to generate, by a cross-language speech conversion algorithm, a second speech data set of a first language of a target user with a target user voice tone according to a first speech data set of the first language;

[0191] A speech synthesizer training unit is configured to generate, according to the second speech data set and a third speech data set of a second language of the target user, a speech synthesizer of the target user with multi-language capability;

[0192] A response text determination unit is configured to determine, for user voice data sent by a client, a same-language response text corresponding to the user voice data;

[0193] A speech synthesis unit is configured to generate, by the speech synthesizer, response voice data corresponding to the response text.

[0194] Seventh embodiment

[0195] The present application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, it is described simply, and the related parts can be referred to the part of the method embodiment. The device embodiment described below is only illustrative.

[0196] The electronic device of the present embodiment comprises a processor and a memory. The memory is configured to store a program for implementing a voice interaction method. After the device is powered on and the program of the method is run by the processor, the following steps are performed: generating, by a cross-language speech conversion algorithm, a second speech data set of a first language of a target user with a target user voice tone according to a first speech data set of the first language;

[0197] generate a speech synthesizer with multi-language capability of the target user according to the second speech data set and a third speech data set of a second language of the target user;

[0198] determine a reply text of the same language corresponding to the user speech data according to the user speech data sent by the client;

[0199] generate reply speech data corresponding to the reply text through the speech synthesizer.

[0200] Eighth embodiment

[0201] In the above embodiments, a voice interaction system is provided, and the present application also provides a voice interaction method. The execution subject of the method can be a terminal device or the like. The method corresponds to the above-mentioned embodiments of the system. The same parts of the present embodiment as those of the first embodiment will not be described again, and please refer to the corresponding parts in the first embodiment.

[0202] The voice interaction method provided by the present application can include the following steps:

[0203] Step 1: collect user speech data and send the user speech data to the server;

[0204] Step 2: play the reply speech data sent back by the server. The reply speech data is determined in the following manner: the server generates a second speech data set of a first language of a target user according to a first speech data set of at least one first language through a cross-language speech conversion algorithm; generates a speech synthesizer with multi-language capability of the target user according to the second speech data set and a third speech data set of a second language of the target user; determines a reply text of the same language corresponding to the user speech data; and generates reply speech data corresponding to the reply text through the speech synthesizer.

[0205] Ninth embodiment

[0206] In the above embodiments, a voice interaction method is provided, and the present application also provides a voice interaction device. The device corresponds to the above-mentioned embodiments of the method. The same parts of the present embodiment as those of the first embodiment will not be described again, and please refer to the corresponding parts in the first embodiment.

[0207] The voice interaction device provided by the present application includes:

[0208] a voice collection unit configured to collect user speech data and send the user speech data to the server;

[0209] The voice playing unit is configured to play the response voice data returned by the server. The response voice data is determined in the following manner: the server generates a second voice data set of a first language of a target user having a voice tone of the target user according to a first voice data set of the first language by using a cross-language voice conversion algorithm; generates a voice synthesizer of the target user having a multi-language capability according to the second voice data set and a third voice data set of a second language of the target user; and determines a response text of the same language as the user voice data corresponding to the user voice data. The voice synthesizer is configured to generate the response voice data corresponding to the response text.

[0210] The tenth embodiment

[0211] The present application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, it is described more simply, and the relevant part can be referred to the part of the method embodiment. The device embodiment described below is only illustrative.

[0212] The electronic device of the present embodiment comprises a microphone, a processor and a memory. The memory is configured to store a program for implementing the voice interaction method. After the device is powered on and the program of the method is run by the processor, the following steps are performed: collecting user voice data and sending the user voice data to the server; playing the response voice data returned by the server. The response voice data is determined in the following manner: the server generates a second voice data set of a first language of a target user having a voice tone of the target user according to a first voice data set of the first language by using a cross-language voice conversion algorithm; generates a voice synthesizer of the target user having a multi-language capability according to the second voice data set and a third voice data set of a second language of the target user; and determines a response text of the same language as the user voice data corresponding to the user voice data. The voice synthesizer is configured to generate the response voice data corresponding to the response text.

[0213] The electronic device includes, but is not limited to, a smart speaker, a smart phone, a vending machine, an automatic ordering machine, etc.

[0214] The eleventh embodiment

[0215] In the above-mentioned embodiments, a voice interaction system is provided, and the present application also provides an online text-to-speech interaction system corresponding thereto. The interaction system corresponds to the above-mentioned system embodiments. The parts of the present embodiment that are the same as those of the first embodiment will not be described again, and please refer to the corresponding part in the first embodiment.

[0216] The online text-to-speech interaction system provided by the present application comprises a terminal device and a server.

[0217] The terminal device is configured to send a first user voice synthesis request for a target multi-language mixed text to a server; the server is configured to generate a second voice data set of a first language of a first user with a first user voice by using a cross-language voice conversion algorithm according to a first voice data set of the first language of the first user; generate a multi-language capable voice synthesizer of the first user according to the second voice data set and a third voice data set of a second language of the first user; and generate voice synthesis data corresponding to the mixed text by using the voice synthesizer.

[0218] For example, the first user's mother tongue is Chinese, and the user cannot speak English, so the user's English voice data set cannot be directly obtained, and the user's voice synthesizer with Chinese-English bilingual voice synthesis capability cannot be directly generated. However, by using the system provided in the embodiments of the present application, the English voice data set of the first user can be automatically generated according to the English voice data set of the second user by using the cross-language voice conversion algorithm, and the voice synthesizer with Chinese-English mixed voice synthesis capability of the first user can be generated according to the Chinese and English voice data sets. Then, for the voice synthesis demand of the terminal device for the multi-language mixed text, the corresponding Chinese-English mixed voice data can be generated by using the voice synthesizer.

[0219] As can be seen from the above embodiments, the online text-to-speech synthesis interactive system provided by the application sends a first user speech synthesis request for a target multi-language mixed text to a server through a terminal device; the server is configured to generate a second speech data set of a first language of the first user with a first user voice by using a cross-language speech conversion algorithm according to a first speech data set of the first language of the second user; generate a speech synthesizer with multi-language capability of the first user according to the second speech data set and a third speech data set of a second language of the first user; and generate speech synthesis data corresponding to the mixed text through the speech synthesizer; this processing manner uses cross-language speech conversion technology to generate high-quality non-native and mixed language data with a target speaker voice of a target speaker, and combines the data with native speech data to obtain a speech synthesizer with bilingual / multi-lingual / hybrid language capability, thereby avoiding the problem of inconsistent voice caused by switching between different language synthesizers in cross-language and mixed language speech synthesis, and unnatural effect; therefore, the speech synthesis quality of multi-lingual text can be effectively improved, thereby improving the user experience. In addition, the system is no longer limited to the native language of the speaker, but only focuses on the voice of the speaker. As long as the voice of the speaker and the native language recording of the speaker are selected, the voice can be extended to other languages, and speech synthesis processing can be performed on any text in other languages. At the same time, this processing manner also does not need to use a phonetic symbol mapping method to cross languages, thereby avoiding the problem of inaccurate or even incorrect pronunciation caused by phonetic symbol mapping. In addition, this processing manner also enables the system to be built only by using single-language databases of different speakers, thereby saving the high cost of hiring professional multi-lingual speakers. Furthermore, this processing manner enables the synthesis effect of foreign language parts to approach the pronunciation level of native speakers of foreign languages without affecting the native language performance, and the synthesized speech of different languages well maintains the voice of the target speaker, so that any (single-language) voice can be given excellent multi-lingual capability.

[0220] Twelfth embodiment

[0221] In the above embodiments, an online text-to-speech synthesis interactive system is provided, and the application also provides an online text-to-speech synthesis interactive method. The execution subject of the method can be a server and the like. The method corresponds to the above-mentioned system embodiments. The same parts of this embodiment as the first embodiment are not described again, please refer to the corresponding parts in the first embodiment.

[0222] The online text-to-speech interaction method provided by the application can include the following steps: sending a first user text-to-speech request for a target multi-language mixed text to a server, so that the server generates a second speech data set of a first language of a first user with a first user voice tone according to a first speech data set of the first language of the second user by using a cross-language speech conversion algorithm; generating a multi-language capable speech synthesizer of the first user according to the second speech data set and a third speech data set of a second language of the first user; and generating speech synthesis data corresponding to the mixed text by using the speech synthesizer.

[0223] Thirteenth embodiment

[0224] The application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, it is described more simply, and the relevant parts can be referred to the part of the description of the method embodiment. The device embodiment described below is only illustrative.

[0225] The electronic device of the embodiment includes a microphone, a processor and a memory. The memory is used to store a program for implementing the online text-to-speech interaction method. After the device is powered on and the program of the method is run by the processor, the following steps are performed: sending a first user text-to-speech request for a target multi-language mixed text to a server, so that the server generates a second speech data set of a first language of a first user with a first user voice tone according to a first speech data set of the first language of the second user by using a cross-language speech conversion algorithm; generating a multi-language capable speech synthesizer of the first user according to the second speech data set and a third speech data set of a second language of the first user; and generating speech synthesis data corresponding to the mixed text by using the speech synthesizer.

[0226] Fourteenth embodiment

[0227] In the above-mentioned embodiments, a speech interaction system is provided, and a news broadcast system is also provided by the application. The interaction system corresponds to the above-mentioned system embodiments. The embodiment has the same content as the first embodiment, and the same parts are not described again. Please refer to the corresponding part in the first embodiment.

[0228] The news broadcast system provided by the application includes a terminal device and a server.

[0229] The terminal device is configured to send a request for multi-lingual text broadcast to a service end; the multi-lingual voice data corresponding to the text to be broadcast and broadcast by a target user is returned by the service end; the service end is configured to generate a second voice data set of at least one first language with a target user's voice tone according to a first voice data set of the at least one first language through a cross-language voice conversion algorithm; generate a voice synthesizer with multi-language capability of the target user according to the second voice data set and a third voice data set of a second language of the target user; and generate multi-lingual voice data corresponding to the text to be broadcast and broadcast by the target user through the voice synthesizer.

[0230] The text to be broadcast can be a text mixed with multiple languages, i.e., a text including multiple languages of characters, such as "I am very happy"; accordingly, the multi-lingual voice data is mixed-lingual voice data. The text to be broadcast can also be multiple language versions of a text, such as an English version, a French version, etc., such as "I am very happy" and "I'm very happy"; accordingly, the multi-lingual voice data is multi-lingual voice data.

[0231] For example, a host wants to broadcast multiple language versions of a news, including Chinese, English, German, Vietnamese, etc., but the host can only speak Chinese and English. The prior art only generates a Chinese voice synthesizer and an English voice synthesizer according to the host's own Chinese and English voice data, but cannot generate a multi-lingual voice synthesizer of Chinese, English, German, and Vietnamese. The system provided in the embodiments of the present application can collect voice data of other languages of other users, generate voice data sets of other languages of the host with the voice tone of the host according to voice data sets of other languages of the other users through a cross-language voice conversion algorithm, generate a voice synthesizer with multi-language capability of the host according to the voice data sets of other languages of the host with the voice tone of the host and the Chinese voice data set and the English voice data set of the host, which is a voice synthesizer that can be used for different languages, and generate multi-lingual voice data corresponding to the news broadcast by the host through the voice synthesizer, such as Chinese voice, English voice, German voice, and Vietnamese voice of the news.

[0232] For another example, a host is to broadcast a news including Chinese, English and German, but the host can only speak Chinese and English, and cannot read or read poorly the German, the prior art only trains and generates a Chinese speech synthesizer and an English speech synthesizer according to the host's own Chinese and English speech data, and cannot generate a speech synthesizer with multi-language mixing of Chinese, English and German; and the system provided by the embodiment of the present application can collect other users' German speech data, generate a German speech data set with the host's voice according to the other users' German speech data set through the cross-language speech conversion algorithm, and generate a speech synthesizer with multi-language mixing of Chinese, English and German of the host according to the German speech data set with the host's voice, the Chinese speech data set and the English speech data set of the host, the speech synthesizer can synthesize Chinese speech data, English speech data, German speech data and mixed speech data of the three languages, and through the speech synthesizer, mixed language speech data corresponding to the news including Chinese, English and German broadcasted by the host can be generated.

[0233] As can be seen from the above embodiments, the news broadcast system provided by the application sends a request for multi-lingual broadcast text to the server through the terminal device; plays back the multi-lingual voice data corresponding to the to-be-broadcast text broadcast by the target user returned by the playing server; the server generates a second voice data set of at least one first language with the voice color of the target user according to a first voice data set of at least one first language through a cross-language voice conversion algorithm; generates a voice synthesizer with multi-lingual capability of the target user according to the second voice data set and a third voice data set of a second language of the target user; and generates multi-lingual voice data corresponding to the to-be-broadcast text broadcast by the target user through the voice synthesizer; this processing mode uses cross-language voice conversion technology to generate high-quality non-native and mixed language data with the voice color of the target speaker, which is combined with the original native language recording data as training data, and a voice synthesizer with bilingual / multi-lingual / mixed language capability of the target speaker is obtained, avoiding the problem of inconsistent timbre and unnatural effect caused by mutual switching when using different language synthesizers in cross-language and mixed language voice synthesis; therefore, the voice synthesis quality of multi-lingual text can be effectively improved, thereby improving the user experience. In addition, the system is no longer limited to the native language of the speaker, but only focuses on the timbre of the speaker. As long as the timbre of the speaker and the native language recording of the speaker are selected, the timbre can be extended to other languages, and voice synthesis processing can be performed on any text in other languages. At the same time, this processing mode also does not need to use methods such as phonetic symbol mapping to cross languages, avoiding the problem of inaccurate or even incorrect pronunciation caused by phonetic symbol mapping. In addition, this processing mode also makes it possible to build a system using only single-language databases of different speakers, saving the high cost of hiring professional multi-lingual speakers. Furthermore, this processing mode makes it possible to approach the pronunciation level of a native speaker of a foreign language in the synthesized foreign language part without affecting the native language performance, and the synthesized voice of different languages well maintains the timbre of the target speaker, so it can be endowed with excellent multi-lingual capability of any (single) timbre.

[0234] Fifteenth embodiment

[0235] In the above embodiments, a news broadcast system is provided, and the application also provides a news broadcast method corresponding thereto. The execution subject of the method can be a server and the like. The method corresponds to the above-mentioned system embodiments. The same parts of this embodiment as the first embodiment are not described again, please refer to the corresponding part in embodiment one.

[0236] The news broadcast method provided by the application can include the following steps:

[0237] Step 1: generating, by a cross-language speech conversion algorithm, a second speech data set of at least one first language with a target user's voice tone according to a first speech data set of the at least one first language;

[0238] Step 2: generating, by a speech synthesizer, a multi-lingual speech synthesizer of the target user according to the second speech data set and a third speech data set of a second language of the target user;

[0239] Step 3: generating, by the speech synthesizer, multi-lingual speech data corresponding to the text to be broadcast by the target user according to a request of the client to broadcast the text in multiple languages.

[0240] Sixteenth embodiment

[0241] In the above embodiments, a news broadcast method is provided, and a news broadcast device corresponding to the method is also provided. The device corresponds to the above method embodiments. The same parts of this embodiment as the first embodiment are not described again, please refer to the corresponding parts in the first embodiment.

[0242] The news broadcast device provided by the application comprises:

[0243] A training data generation unit is configured to generate, by a cross-language speech conversion algorithm, a second speech data set of at least one first language with a target user's voice tone according to a first speech data set of the at least one first language;

[0244] A speech synthesizer training unit is configured to generate, by a speech synthesizer, a multi-lingual speech synthesizer of the target user according to the second speech data set and a third speech data set of a second language of the target user;

[0245] A speech synthesis unit is configured to generate, by the speech synthesizer, multi-lingual speech data corresponding to the text to be broadcast by the target user according to a request of the client to broadcast the text in multiple languages.

[0246] Seventeenth embodiment

[0247] The application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, it is described simply, and the related parts are described in the method embodiment. The device embodiment described below is only illustrative.

[0248] The electronic device of the embodiment comprises a microphone, a processor and a memory; the memory is used for storing a program for implementing a news broadcasting method; after the device is powered on and the program of the method is run by the processor, the following steps are executed: a second voice data set of at least one first language with a target user's voice color is generated from a first voice data set of the at least one first language by a cross-language voice conversion algorithm; a multi-language capable voice synthesizer of the target user is generated from the second voice data set and a third voice data set of a second language of the target user; and multi-lingual voice data corresponding to the text to be broadcast and broadcast by the target user is generated by the voice synthesizer for a request of broadcasting text in multiple languages sent by a client.

[0249] The eighteenth embodiment

[0250] In the above-mentioned embodiments, a news broadcasting system is provided, and the present application further provides a news broadcasting method, and the execution subject of the method can be a terminal device or the like. The method corresponds to the embodiments of the above-mentioned system. The same parts of the embodiment as those of the first embodiment will not be described herein again, and please refer to the corresponding parts in the first embodiment.

[0251] The news broadcasting method provided by the present application can comprise the following steps:

[0252] Step 1: sending a request of broadcasting text in multiple languages to a server;

[0253] Step 2: playing multi-lingual voice data corresponding to the text to be broadcast and broadcast by the target user returned by the server; the multi-lingual voice data is generated in the following manner: a second voice data set of at least one first language with a target user's voice color is generated from a first voice data set of the at least one first language by a cross-language voice conversion algorithm; a multi-language capable voice synthesizer of the target user is generated from the second voice data set and a third voice data set of a second language of the target user; and multi-lingual voice data corresponding to the text to be broadcast and broadcast by the target user is generated by the voice synthesizer.

[0254] The nineteenth embodiment

[0255] In the above-mentioned embodiments, a news broadcasting method is provided, and the present application further provides a news broadcasting device. The device corresponds to the embodiments of the above-mentioned method. The same parts of the embodiment as those of the first embodiment will not be described herein again, and please refer to the corresponding parts in the first embodiment.

[0256] The news broadcasting device provided by the present application comprises:

[0257] The request sending unit is configured to send a request for multi-lingual broadcast text to the server.

[0258] The voice playing unit is configured to play multi-lingual voice data corresponding to the text to be broadcast, which is broadcast by the target user and returned by the server. The multi-lingual voice data is generated in the following manner: the server generates a second voice data set of at least one first language with the voice color of the target user according to a first voice data set of the at least one first language through a cross-language voice conversion algorithm; generates a voice synthesizer of the target user with multi-lingual capability according to the second voice data set and a third voice data set of a second language of the target user; and generates the multi-lingual voice data corresponding to the text to be broadcast, which is broadcast by the target user, through the voice synthesizer.

[0259] The twentieth embodiment

[0260] The present application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, it is described more simply, and the related parts can be referred to the part of the description of the method embodiment. The device embodiment described below is only illustrative.

[0261] The electronic device of the present embodiment comprises a microphone, a processor and a memory. The memory is configured to store a program for implementing the news broadcast method. After the device is powered on and the program of the method is run by the processor, the following steps are performed: sending a request for multi-lingual broadcast text to the server; playing multi-lingual voice data corresponding to the text to be broadcast, which is broadcast by the target user and returned by the server. The multi-lingual voice data is generated in the following manner: the server generates a second voice data set of at least one first language with the voice color of the target user according to a first voice data set of the at least one first language through a cross-language voice conversion algorithm; generates a voice synthesizer of the target user with multi-lingual capability according to the second voice data set and a third voice data set of a second language of the target user; and generates the multi-lingual voice data corresponding to the text to be broadcast, which is broadcast by the target user, through the voice synthesizer.

[0262] The twenty-first embodiment

[0263] In the above-mentioned embodiments, a voice synthesis method is provided. Correspondingly, the present application also provides a voice synthesizer construction method. The execution subject of the method can be a terminal device or the like. The method corresponds to the above-mentioned system embodiments. The parts of the present embodiment that are the same as those of the second embodiment will not be described again, and please refer to the corresponding parts in the first embodiment.

[0264] The voice synthesizer construction method provided by the present application can comprise the following steps:

[0265] Step 1: generating a second speech data set of a first language with a first user's voice tone according to a first speech data set of at least one first language of at least one second user by a cross-lingual speech conversion algorithm;

[0266] Step 2: generating a multi-lingual speech synthesizer of the first user according to the second speech data set and a third speech data set of a second language of the first user.

[0267] Twenty-second embodiment

[0268] In the above-mentioned embodiments, a speech synthesizer construction method is provided, and a speech synthesizer construction device is also provided. The device corresponds to the above-mentioned embodiments of the method. The same parts of this embodiment as the first embodiment are not described again, and please refer to the corresponding parts in the first embodiment.

[0269] The speech synthesizer construction device provided by the present application comprises:

[0270] a training data generation unit configured to generate a second speech data set of a first language with a first user's voice tone according to a first speech data set of at least one first language of at least one second user by a cross-lingual speech conversion algorithm;

[0271] a speech synthesizer training unit configured to generate a multi-lingual speech synthesizer of the first user according to the second speech data set and a third speech data set of a second language of the first user.

[0272] Twenty-third embodiment

[0273] The present application also provides an electronic device. Since the device embodiments are basically similar to the method embodiments, they are described more simply, and please refer to the corresponding parts of the method embodiments. The device embodiments described below are only illustrative.

[0274] The electronic device of the present embodiment comprises a microphone, a processor and a memory. The memory is configured to store a program for implementing the speech synthesizer construction method. After the device is powered on and the program of the method is run by the processor, the following steps are performed: generating a second speech data set of a first language with a first user's voice tone according to a first speech data set of at least one first language of at least one second user by a cross-lingual speech conversion algorithm; generating a multi-lingual speech synthesizer of the first user according to the second speech data set and a third speech data set of a second language of the first user.

[0275] Twenty-fourth embodiment

[0276] In the above embodiment, a speech synthesis method is provided, and the application further provides a cross-language speech generation method. The execution subject of the method can be a terminal device or the like. The method corresponds to the above system embodiment. The same parts of the embodiment and the second embodiment are not described again, and please refer to the corresponding part in the first embodiment.

[0277] The cross-language speech generation method provided by the application can include the following steps:

[0278] Step 1: determining a text to be processed, and sending a speech generation request for the first user to read the text to a server.

[0279] Step 2: playing first speech data of the first user reading the text returned by the server; the text to be processed includes a text in a first language or a text mixed with the first language and a second language, and the first user's mother tongue is the second language.

[0280] For example, a first user wants to read an English text, but the user is a Chinese person who cannot speak English or speaks English poorly. In order to achieve the effect of the first user speaking fluent English, the first user can determine an English text to be read through a terminal device, and send a speech generation request for the first user to read the English text to a server. The server can generate speech data of the first user reading the English text through the speech synthesis method provided in the above second embodiment, as if the first user has good English level.

[0281] As can be seen from the above embodiment, the cross-language speech generation method provided by the embodiment of the application determines a text to be processed through a terminal device, and sends a speech generation request for a first user to read the text to a server. Speech data of the first user reading the text returned by the server is played. The text to be processed includes a text in a first language or a text mixed with the first language and a second language, and the first user's mother tongue is the second language. This processing method enables the first user to generate speech data of the first user reading a text in a certain language even if the first user does not have the ability to read the text in the certain language, and realizes cross-language text reading.

[0282] Twenty-fifth embodiment

[0283] In the above embodiment, a cross-language speech generation method is provided, and the application further provides a cross-language speech generation device. The device corresponds to the above method embodiment. The same parts of the embodiment and the first embodiment are not described again, and please refer to the corresponding part in the first embodiment.

[0284] The cross-language speech generation device provided by the application can include:

[0285] The text determining unit is configured to determine a text to be processed, and send a speech generation request for the first user to read the text to a server;

[0286] The speech playing unit is configured to play speech data of the first user reading the text returned by the server; the text to be processed comprises a text in a first language or a text mixed with the first language and a second language, and a native language of the first user is the second language.

[0287] The twenty-sixth embodiment

[0288] The present application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple, and the related parts can be referred to the part of the description of the method embodiment. The device embodiment described below is only illustrative.

[0289] The electronic device of the embodiment comprises a microphone, a processor and a memory; the memory is configured to store a program for implementing the cross-language speech generation method, and after the device is powered on and the program of the method is run by the processor, the following steps are performed: determining a text to be processed, and sending a speech generation request for the first user to read the text to a server; playing speech data of the first user reading the text returned by the server; the text to be processed comprises a text in a first language or a text mixed with the first language and a second language, and a native language of the first user is the second language.

[0290] The twenty-seventh embodiment

[0291] In the above-mentioned embodiments, a speech interaction system is provided, and a cross-dialect speech generation system is also provided in the present application. The interaction system corresponds to the above-mentioned system embodiments. The parts of the present embodiment that are the same as the first embodiment are not described again, and please refer to the corresponding parts in the first embodiment.

[0292] The present application provides a cross-dialect speech generation system, comprising a terminal device and a server.

[0293] The terminal device is configured to determine a text to be processed, and send a speech generation request for the first user to read the text in a first dialect to a server; and play first speech data of the first user reading the text in the first dialect returned by the server; the server is configured to generate a third speech data set of the first dialect with the voice of the first user according to a second speech data set of the first dialect of a second user by a cross-dialect speech conversion algorithm; generate a speech synthesizer of the first user with multi-dialect capability according to the third speech data set and a fourth speech data set of the second dialect of the first user; and generate the first speech data by the speech synthesizer for the request.

[0294] For example, a first user speaking Mandarin wants to speak a paragraph in Cantonese (a first dialect), but the user does not speak Cantonese or speaks Cantonese poorly. In order to achieve the effect of the first user speaking Cantonese fluently, the server can first generate Cantonese speech data (third speech data set) with the first user's voice tone according to the Cantonese speech data (second speech data set) of a second user speaking Cantonese (a second dialect) through a cross-dialect speech conversion algorithm; generate a speech synthesizer with Mandarin and Cantonese capabilities for the first user according to the third speech data set and a fourth speech data set of the first user's Mandarin (a second dialect); and generate speech data (first speech data) of the text read by the first user in Cantonese through the speech synthesizer for the text to be read by the first user determined through a terminal device, as if the first user has good English level.

[0295] The cross-dialect speech conversion algorithm can generate high-quality language data with a certain dialect or a mixture of multiple dialects with the target speaker's voice tone. The algorithm can include the following steps: 1) constructing a speech posterior probability graph (PPG) feature extractor and a speech synthesis model of the first user; 2) determining PPG feature data of the second speech data according to first acoustic feature data of the second speech data of the second user's first dialect through the PPG feature extractor; the first acoustic feature data includes voiceprint information and speech content information of the second user; 3) generating third speech data of the first dialect of the first user corresponding to the second speech data through the speech synthesis model of the first user according to the PPG feature data and second acoustic feature data of the second speech data; the second acoustic feature data includes prosodic information. Since the cross-dialect speech conversion algorithm has a similar processing process to the above-mentioned cross-language speech conversion algorithm, it will not be described here.

[0296] As can be seen from the above embodiments, the cross-dialect speech generation system provided by the application determines the text to be processed by the terminal device, sends a speech generation request for the text read by a first user in a first dialect to a server, and plays the first speech data of the text read by the first user in the first dialect returned by the server; the server generates a third speech data set of the first dialect with the voice color of the first user according to a second speech data set of the first dialect of a second user by using a cross-dialect speech conversion algorithm, generates a speech synthesizer with multi-dialect capability of the first user according to the third speech data set and a fourth speech data set of the second dialect of the first user, and generates the first speech data by using the speech synthesizer for the request. This processing manner uses the cross-dialect speech conversion technology to generate high-quality language data of a certain dialect or a mixed dialect with the voice color of a target speaker, which is combined with the original native dialect recording data as training data, so that a speech synthesizer with the voice color of a target speaker with double-dialect / multi-dialect / mixed-dialect language capability is obtained, and the problem of inconsistent voice color and unnatural effect caused by mutual switching of different dialect synthesizers is avoided. Therefore, the speech synthesis quality of multi-dialect text can be effectively improved, and the user experience is improved. In addition, the system is no longer limited to the native dialect of the speaker, but only focuses on the voice color of the speaker. As long as the voice color of the speaker and the native dialect recording of the speaker are selected, the voice color can be expanded to other dialects, and any text in other dialects can be processed by speech synthesis. In addition, this processing manner also enables the system to be built only by using single-dialect databases of different speakers, thereby saving the high cost of hiring professional multi-dialect speakers. Furthermore, this processing manner enables the synthesis effect of other dialects to approach the pronunciation level of native speakers of the dialects without affecting the performance of the native dialect, and the synthesized speech of different dialects well maintains the voice color of the target speaker, so that any (single-dialect) voice color can be given excellent multi-dialect capability.

[0297] Twenty-eighth embodiment

[0298] In the above embodiments, a cross-dialect speech generation system is provided, and the application also provides a cross-dialect speech generation method. The execution subject of the method can be a server or the like. The method corresponds to the above-mentioned system embodiments. The same parts of this embodiment as those of the first embodiment are not described again, and please refer to the corresponding parts in the first embodiment.

[0299] The cross-dialect speech generation method provided by the application can include the following steps:

[0300] Step 1: generating, by a cross-dialect speech conversion algorithm, a third speech data set of a first dialect with a first user's voice timbre according to a second speech data set of the first dialect of a second user;

[0301] Step 2: generating, by a speech synthesizer, a multi-dialect capable speech synthesizer of the first user according to the third speech data set and a fourth speech data set of a second dialect of the first user;

[0302] Step 3: generating, by the speech synthesizer, first speech data for a speech generation request sent by a client for the first user to read the text in the first dialect.

[0303] The twenty-ninth embodiment

[0304] In the above-mentioned embodiments, a cross-dialect speech generation method is provided, and a cross-dialect speech generation device is also provided. The device corresponds to the above-mentioned method embodiments. The same parts of this embodiment as the first embodiment are not described again, please refer to the corresponding parts in the first embodiment.

[0305] The cross-dialect speech generation device provided by the present application comprises:

[0306] A training data generation unit is configured to generate, by a cross-dialect speech conversion algorithm, a third speech data set of a first dialect with a first user's voice timbre according to a second speech data set of the first dialect of a second user;

[0307] A speech synthesizer training unit is configured to generate, by a speech synthesizer, a multi-dialect capable speech synthesizer of the first user according to the third speech data set and a fourth speech data set of a second dialect of the first user;

[0308] A speech synthesis unit is configured to generate, by the speech synthesizer, first speech data for a speech generation request sent by a client for the first user to read the text in the first dialect.

[0309] The thirtieth embodiment

[0310] The present application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, it is described simply, and the related parts can be referred to the part of the method embodiment. The device embodiment described below is only illustrative.

[0311] The electronic device of the embodiment comprises a microphone, a processor and a memory; the memory is used for storing a program for implementing a cross-dialect speech generation method; after the device is powered on and the program of the method is run by the processor, the following steps are executed: a third speech data set of a first dialect with a first user's voice is generated according to a second speech data set of a second dialect of a second user by a cross-dialect speech conversion algorithm; a multi-dialect capable speech synthesizer of the first user is generated according to the third speech data set and a fourth speech data set of a second dialect of the first user; and first speech data is generated by the speech synthesizer for a speech generation request sent by a client, in which the first user reads the text in the first dialect.

[0312] Thirty-first embodiment

[0313] In the above-mentioned embodiments, a cross-dialect speech generation system is provided, and the present application further provides a cross-dialect speech generation method, the execution subject of which can be a terminal device or the like. The method corresponds to the above-mentioned embodiments of the system. The same parts of the present embodiment as those of the first embodiment will not be described herein again, and please refer to the corresponding parts in the first embodiment.

[0314] The present application provides a cross-dialect speech generation method, which can comprise the following steps:

[0315] Step 1: determining a text to be processed, and sending a speech generation request to a server, in which a first user reads the text in a first dialect;

[0316] Step 2: playing first speech data of the first user reading the text in the first dialect returned by the server.

[0317] Thirty-second embodiment

[0318] In the above-mentioned embodiments, a cross-dialect speech generation method is provided, and the present application further provides a cross-dialect speech generation device. The device corresponds to the above-mentioned embodiments of the method. The same parts of the present embodiment as those of the first embodiment will not be described herein again, and please refer to the corresponding parts in the first embodiment.

[0319] The cross-dialect speech generation device provided by the present application comprises:

[0320] A request sending unit is configured to determine a text to be processed, and send a speech generation request to a server, in which a first user reads the text in a first dialect;

[0321] A speech playing unit is configured to play first speech data of the first user reading the text in the first dialect returned by the server.

[0322] Thirty-third embodiment

[0323] The present application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple, and the relevant part can be referred to the part of the method embodiment. The device embodiment described below is only illustrative.

[0324] The electronic device of the embodiment comprises a microphone, a processor and a memory; the memory is used to store the program for implementing the cross-dialect speech generation method. After the device is powered on and the program of the method is run by the processor, the following steps are performed: determining the text to be processed, sending a speech generation request of the text read by a first user in a first dialect to a server; playing the first speech data of the text read by the first user in the first dialect returned by the server.

[0325] The thirty-fourth embodiment

[0326] In the above-mentioned embodiments, a cross-dialect speech generation method is provided. Correspondingly, the present application also provides a speech synthesizer construction method, and the execution subject of the method can be a server or the like. The method corresponds to the above-mentioned system embodiment. The same parts of the embodiment and the second embodiment are not described again, and please refer to the corresponding part in the first embodiment.

[0327] The speech synthesizer construction method provided by the present application can comprise the following steps:

[0328] Step 1: generating a third speech data set of the first dialect with the first user's voice tone according to a second speech data set of the first dialect of a second user by a cross-dialect speech conversion algorithm;

[0329] Step 2: generating a speech synthesizer with multi-dialect capability of the first user according to the third speech data set and a fourth speech data set of the second dialect of the first user.

[0330] The thirty-fifth embodiment

[0331] In the above-mentioned embodiments, a speech synthesizer construction method is provided. Correspondingly, the present application also provides a speech synthesizer construction device. The device corresponds to the above-mentioned method embodiment. The same parts of the embodiment and the first embodiment are not described again, and please refer to the corresponding part in the first embodiment.

[0332] The speech synthesizer construction device provided by the present application comprises:

[0333] The training data generation unit is used to generate a third speech data set of the first dialect with the first user's voice tone according to a second speech data set of the first dialect of a second user by a cross-dialect speech conversion algorithm;

[0334] a voice synthesizer training unit configured to generate a multi-dialect capable voice synthesizer of the first user according to the third voice data set and a fourth voice data set of a second dialect of the first user.

[0335] thirty-sixth embodiment

[0336] The present application also provides an electronic device. Since the device embodiments are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the description of the method embodiments. The device embodiments described below are only illustrative.

[0337] An electronic device of the embodiment, the electronic device comprising: a microphone, a processor and a memory; the memory is configured to store a program for implementing the voice synthesizer construction method, after the device is powered on and the program of the method is run by the processor, the following steps are performed: generating a third voice data set of a first dialect with a first user's voice timbre according to a second voice data set of a second dialect of a second user by a cross-dialect speech conversion algorithm; generating a multi-dialect capable voice synthesizer of the first user according to the third voice data set and a fourth voice data set of a second dialect of the first user.

[0338] Although the present application is disclosed with the preferred embodiments, it is not intended to limit the present application, any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of the present application, therefore the protection scope of the present application should be subject to the scope defined by the claims of the present application.

[0339] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0340] The memory can include non-persistent memory in computer readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer readable media.

[0341] 1. Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology for storage of information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory computer-readable medium that can be used to store information accessible to computing devices. According to the definition herein, computer-readable media does not include transitory media, such as modulated data signals and carrier waves.

[0342] 2. Those skilled in the art will appreciate that embodiments of the present application can be provided as methods, systems or computer program products. Accordingly, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage etc.) containing computer usable program code.

Claims

1. A voice interaction system, characterized by The method comprises: an intelligent sound box, configured to collect user voice data, and send the user voice data to a server; and play response voice data returned by the server; the server, configured to generate, by using a cross-language voice conversion algorithm, a second voice data set of a first language of a target user from a first voice data set of the first language of the target user, the second voice data set having a target user voice; generate, by using the second voice data set and a third voice data set of a second language of the target user, a voice synthesizer of the target user with multi-language capability; determine response text in the same language as the user voice data; and generate, by using the voice synthesizer, response voice data corresponding to the response text.

2. An online text-to-speech synthesis system, characterized by The method comprises: a terminal device, configured to send, to a server, a first user voice synthesis request for a target multi-language mixed text; the server, configured to generate, by using a cross-language voice conversion algorithm, a second voice data set of a first language of a first user from a first voice data set of the first language of a second user, the second voice data set having a first user voice; generate, by using the second voice data set and a third voice data set of a second language of the first user, a voice synthesizer of the first user with multi-language capability; and generate, by using the voice synthesizer, voice synthesis data corresponding to the mixed text.

3. A news broadcast system characterized by comprising: The method comprises: a terminal device, configured to send, to a server, a request for multi-language broadcast text; and play multi-language voice data corresponding to the text to be broadcast, which is broadcast by a target user and returned by the server; the server, configured to generate, by using a cross-language voice conversion algorithm, a second voice data set of at least one first language from a first voice data set of the at least one first language, the second voice data set having a target user voice; generate, by using the second voice data set and a third voice data set of a second language of the target user, a voice synthesizer of the target user with multi-language capability; and generate, by using the voice synthesizer, multi-language voice data corresponding to the text to be broadcast, which is broadcast by the target user.

4. A speech synthesis method characterized by, The method comprises: generate, by using a cross-language voice conversion algorithm, a second voice data set of a first language of a first user from a first voice data set of the first language of a second user, the second voice data set having a first user voice; generate, by using the second voice data set and a third voice data set of a second language of the first user, a voice synthesizer of the first user with multi-language capability; and generate, by using the voice synthesizer, voice synthesis data of the first user corresponding to a first multi-language mixed text.

5. The method of claim 4, wherein, The method of generating, by using the voice synthesizer, voice synthesis data of the first user corresponding to the first multi-language mixed text comprises: determine, by using a text input module included in the voice synthesizer, a pronunciation unit sequence of the first multi-language mixed text, wherein the pronunciation units of the text segments in different languages are pronunciation units in corresponding languages; determine, by using an acoustic feature synthesis network included in the voice synthesizer, an acoustic feature sequence having the first user voice according to the pronunciation unit sequence; and generate, by using a vocoder included in the voice synthesizer, the voice synthesis data according to the acoustic feature sequence.

6. The method of claim 5, wherein The Chinese pronunciation units include initials and finals of Chinese Pinyin, and tones; The English pronunciation units include English phonemes and stresses.

7. The method of claim 6, wherein, The sequence of English pronunciation units is determined in the following manner: Spaces are inserted between pronunciation units, and punctuation marks are inserted according to the length of inter-word pauses.

8. The method of claim 4, wherein, The generating of the speech synthesizer with multi-language capability of the first user according to the second speech data set and the third speech data set in the second language of the first user comprises: The generating of the fourth speech data set in a mixed language of the first user according to the second speech data set and the third speech data set comprises: The generating of the speech synthesizer according to the second speech data set, the third speech data set and the fourth speech data set comprises:

9. The method of claim 8, wherein, The generating of the fourth speech data set in a mixed language of the first user according to the second speech data set and the third speech data set in the second language of the first user comprises: The generating of the speech synthesizer of the first user according to the second speech data set and the third speech data set comprises: The determining of the second multi-language mixed text set comprises: The determining of the pronunciation unit sequence of the second multi-language mixed text by the text input module included in the speech synthesizer of the first user comprises: The determining of the acoustic feature sequence with the voice color of the first user by the acoustic feature synthesis network included in the speech synthesizer of the first user according to the pronunciation unit sequence comprises: The generating of the speech synthesis data of the first user corresponding to the second multi-language mixed text by the vocoder included in the speech synthesizer of the first user according to the acoustic feature sequence comprises: The determining of the fourth speech data set according to the speech synthesis data of the first user corresponding to the second multi-language mixed text comprises:

10. The method of claim 9, wherein, The speech synthesizer of the first user comprises a speech synthesizer based on a Transformer model; The generating of the speech synthesizer according to the second speech data set, the third speech data set and the fourth speech data set comprises: The generating of the acoustic feature synthesis network comprises: The generating of the acoustic feature synthesis network based on the fourth speech data set comprises: The generating of the speech synthesizer according to the second speech data set, the third speech data set and the fourth speech data set comprises:

11. The method of claim 8, wherein, The generating of the acoustic feature synthesis network comprises: The generating of the acoustic feature synthesis network based on the second speech data set, the third speech data set and the fourth speech data set comprises: The generating of the speech synthesizer with multi-language capability of the first user according to the second speech data set and the third speech data set in the second language of the first user comprises: The generating of the vocoder according to the third speech data set comprises:

12. The method of claim 5, wherein, ​ ​ 13. The method of claim 4, wherein the cross-lingual speech conversion algorithm comprises a speech posterior probability graph (PPG) based cross-lingual speech conversion algorithm. comprises:

14. A voice interaction method, characterized by generating, by a cross-lingual speech conversion algorithm, a second speech data set of a first language of a target user having a target user voice color according to a first speech data set of at least one first language; generating a multi-lingual speech synthesizer of the target user according to the second speech data set and a third speech data set of a second language of the target user; determining, for user speech data sent by the client, a response text in the same language as the user speech data; generating, by the speech synthesizer, response speech data corresponding to the response text. comprises:

15. A voice interaction method, characterized in that, collecting user speech data and sending the user speech data to the server; playing response speech data sent back by the server; the response speech data is determined in the following manner: the server generates, by a cross-lingual speech conversion algorithm, a second speech data set of a first language of a target user having a target user voice color according to a first speech data set of at least one first language; generating a multi-lingual speech synthesizer of the target user according to the second speech data set and a third speech data set of a second language of the target user; and determining a response text in the same language as the user speech data; generating, by the speech synthesizer, response speech data corresponding to the response text. comprises:

16. An online text-to-speech synthesis method, characterized by, sending, to the server, a first user speech synthesis request for a target multi-lingual mixed text, so that the server generates, by a cross-lingual speech conversion algorithm, a second speech data set of a first language of a first user having a first user voice color according to a first speech data set of a second language of the second user; generating a multi-lingual speech synthesizer of the first user according to the second speech data set and a third speech data set of a second language of the first user; and generating, by the speech synthesizer, speech synthesis data corresponding to the mixed text. comprises:

17. A news broadcasting method characterized by comprising: generating, by a cross-lingual speech conversion algorithm, a second speech data set of at least one first language having a target user voice color according to a first speech data set of the at least one first language; generating a multi-lingual speech synthesizer of the target user according to the second speech data set and a third speech data set of a second language of the target user; generating, by the speech synthesizer, multi-lingual speech data corresponding to the text to be broadcast by the target user according to a request sent by the client to broadcast the text in multiple languages. comprises:

18. A news broadcasting method characterized by comprising: sending, to the server, a request to broadcast text in multiple languages; playing multi-lingual speech data corresponding to the text to be broadcast by the target user sent back by the server; the multi-lingual speech data is generated in the following manner: the server generates, by a cross-lingual speech conversion algorithm, a second speech data set of at least one first language having a target user voice color according to a first speech data set of the at least one first language; ​ According to the second voice data set, and a third voice data set of a second language of the target user, a multi-lingual voice synthesizer of the target user is generated; And, through the voice synthesizer, multi-lingual voice data corresponding to the text to be broadcast is generated and broadcasted by the target user.

19. A method of constructing a speech synthesizer, characterized by, Comprising: According to at least one first voice data set of at least one first language of at least one second user, a second voice data set of the first language with the first user's voice tone is generated through a cross-language voice conversion algorithm; According to the second voice data set, and a third voice data set of a second language of the target user, a multi-lingual voice synthesizer of the target user is generated; 20. A cross-dialect speech generation system, comprising: Comprising: A terminal device is configured to determine a text to be processed, send a voice generation request of the text read by a first user in a first dialect to a server, and play first voice data of the text read by the first user in the first dialect returned by the server; A server is configured to generate a third voice data set of the first dialect with the first user's voice tone according to a second voice data set of the first dialect of a second user through a cross-dialect voice conversion algorithm; According to the third voice data set, and a fourth voice data set of a second dialect of the first user, a multi-dialect voice synthesizer of the first user is generated; and for the request, first voice data is generated through the voice synthesizer.

21. A method for cross-dialect speech generation, the method comprising: Comprising: According to a second voice data set of a first dialect of a second user, a third voice data set of the first dialect with the first user's voice tone is generated through a cross-dialect voice conversion algorithm; According to the third voice data set, and a fourth voice data set of a second dialect of the first user, a multi-dialect voice synthesizer of the first user is generated; For a voice generation request of a text read by a first user in a first dialect sent by a client, first voice data is generated through the voice synthesizer.

22. A method of constructing a speech synthesizer, characterized by, Comprising: According to a second voice data set of a first dialect of a second user, a third voice data set of the first dialect with the first user's voice tone is generated through a cross-dialect voice conversion algorithm; According to the third voice data set, and a fourth voice data set of a second dialect of the first user, a multi-dialect voice synthesizer of the first user is generated.

23. A speech synthesis apparatus characterized by comprising: Comprising: A training data generation unit is configured to generate a second voice data set of a first language of a first user with the first user's voice tone according to a first voice data set of a first language of a second user through a cross-language voice conversion algorithm; A voice synthesizer training unit is configured to generate a multi-lingual voice synthesizer of the first user according to the second voice data set, and a third voice data set of a second language of the first user; A voice synthesis unit is configured to generate voice synthesis data of the first user corresponding to a first multi-lingual mixed text through the voice synthesizer.

24. An electronic device, comprising: Comprising: A processor; And a memory for storing a program implementing a speech synthesis method, the device, upon being powered on and running the program of the method by the processor, performs the following steps: generating, by a cross-lingual speech conversion algorithm, a second speech data set of a first language of a first user with a first user's voice tone according to a first speech data set of the first language of the second user; generating, according to the second speech data set and a third speech data set of a second language of the first user, a multi-lingual speech synthesizer of the first user; and generating, by the speech synthesizer, speech synthesis data of the first user corresponding to a first multi-lingual mixed text.

25. A speech synthesizer construction apparatus characterized by comprising: comprising: a training data generation unit for generating, by a cross-lingual speech conversion algorithm, a second speech data set of a first language of at least one second user with a first user's voice tone according to a first speech data set of the first language; a speech synthesizer training unit for generating, according to the second speech data set and a third speech data set of a second language of the first user, a multi-lingual speech synthesizer of the first user.

26. An electronic device, comprising: comprising: a processor; and a memory for storing a program implementing a speech synthesizer construction method, the device, upon being powered on and running the program of the method by the processor, performs the following steps: generating, by a cross-lingual speech conversion algorithm, a second speech data set of a first language of at least one second user with a first user's voice tone according to a first speech data set of the first language; generating, according to the second speech data set and a third speech data set of a second language of the first user, a multi-lingual speech synthesizer of the first user.

27. A speech synthesizer construction apparatus characterized by comprising: comprising: a training data generation unit for generating, by a cross-dialect speech conversion algorithm, a third speech data set of a first dialect of a second user with a first user's voice tone according to a second speech data set of a second dialect of the second user; a speech synthesizer training unit for generating, according to the third speech data set and a fourth speech data set of the second dialect of the first user, a multi-dialect speech synthesizer of the first user.

28. An electronic device, comprising: comprising: a processor; and a memory for storing a program implementing a speech synthesizer construction method, the device, upon being powered on and running the program of the method by the processor, performs the following steps: generating, by a cross-dialect speech conversion algorithm, a third speech data set of a first dialect of a second user with a first user's voice tone according to a second speech data set of a second dialect of the second user; and generating, according to the third speech data set and a fourth speech data set of the second dialect of the first user, a multi-dialect speech synthesizer of the first user.

Citation Information

Patent Citations

  • Voice conversion method and device, file generation method and device, broadcasting method and device, voice processing method and device and medium

    CN110970014A