Voice conversion related methods, systems and devices
By using the speech posterior probability graph (PPG) feature extractor and speech synthesis model, the problem of low cross-language speech conversion quality is solved, and efficient cross-language speech synthesis is achieved to meet users' personalized needs.
Patent Information
- Application Number
- CN202010602602.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-06-28
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2040-06-28
AI Technical Summary
The low quality of cross-language speech conversion in existing technologies, especially the unstable speech synthesis and encoding quality between different languages, limits the application of cross-language speech conversion.
The posterior probability graph (PPG) feature extractor and speech synthesis model are used to extract the voiceprint information and speech content information of the speech data through the PPG feature extractor, and combined with the speech synthesis model of the target user to generate speech data with the timbre of the target speaker.
It improves the quality of cross-language speech conversion, realizes efficient speech synthesis between different languages, and meets the personalized needs of users.
Smart Images

Figure CN113851140B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of speech conversion technology, and specifically to a speech conversion method and device, a voice message sending system, method and device, a film and television dubbing system, method and device, a news broadcasting system, method and device, a speech interaction system, method and device, a speech library construction method and device, a speech translation system, method and device, a cross-language reading system, method and device, a cross-dialect reading system, method and device, a speech voice changing system, a question-and-answer system, method and device, and an electronic device. Background Art
[0002] The fundamental task of speech conversion is to modify the source speaker's voice characteristics to make it sound like the target speaker's timbre while preserving the spoken content, thus meeting the user's personalized needs in voice interaction applications. Among them, cross-language speech conversion has a richer range of applications in voice interaction.
[0003] A typical cross-language speech conversion system employs a technical solution that first converts the source speaker's speech signal into corresponding text information. This text information is then combined with the target speaker's speech characteristics to generate a speech signal with the target speaker's timbre. However, because this solution requires extracting text information from the speech before speech synthesis, this text information is language-dependent and cannot be used interchangeably across languages. The speech synthesis module, which relies on text information, is also language-dependent. This solution faces difficulties in training models for target speakers of different languages within a unified framework, thus limiting its cross-language speech conversion capabilities.
[0004] Another typical cross-language speech conversion system employs a technical solution that uses speech representation learning and style transfer techniques. First, the source speaker's speech is encoded to extract textual content and prosodic features. An encoder is then used to extract the target speaker's voiceprint features. These features are converted into acoustic features with the target speaker's timbre through an attention mechanism and decoder. Finally, a WaveRNN vocoder is used to synthesize the acoustic features into speech. However, while this speech feature encoding scheme can simultaneously encode speech in multiple languages, its encoding quality is relatively unstable, making it difficult to ensure consistent encoding quality across different languages. This limits its application in cross-language speech conversion.
[0005] However, in the process of implementing the present invention, the inventors found that the technical solution has at least the following problems: the problem of low quality of cross-language speech conversion. In summary, how to improve the quality of cross-language speech conversion has become an urgent problem that those skilled in the art need to solve. Summary of the Invention
[0006] This application provides a cross-language speech conversion method to address the low quality of cross-language speech conversion in existing technologies. This application also provides a voice message sending system, method, and device; a film and television dubbing system, method, and device; a news broadcasting system, method, and device; a cross-language speech conversion device; and an electronic device.
[0007] The present application provides a voice message sending system, comprising:
[0008] The first terminal device is configured to collect first voice data of a first user and send a request to the server to change the first voice data into the voice of a second user;
[0009] The server is configured to construct a speech synthesis model for each user; and to construct a speech posterior probability graph (PPG) feature extractor; determine, by the PPG feature extractor, PPG feature data of the first speech data based on first acoustic feature data of the first speech data, including voiceprint information and speech content information of the first user; generate, by the speech synthesis model of the second user, second speech data of the second user corresponding to the first speech data, based on the PPG feature data and second acoustic feature data of the first speech data, including prosody information; and transmit the second speech data to the second terminal device;
[0010] The second terminal device is used to play the second voice data.
[0011] This application also provides a film and television dubbing system, including:
[0012] The terminal device is configured to collect first voice data of a first user for a film and television dialogue text, and send a request to a server to convert the first voice data into a dubbing of a second user;
[0013] The server is used to construct a speech posterior probability graph PPG feature extractor and a speech synthesis model for each user; through the PPG feature extractor, based on the first acoustic feature data of the first speech data including the first user's voiceprint information and speech content information, the PPG feature data of the first speech data is determined; through the speech synthesis model of the second user, based on the PPG feature data and the second acoustic feature data of the first speech data including prosody information, the second speech data of the second user for the film and television dialogue text is generated.
[0014] This application also provides a news broadcast system, including:
[0015] The terminal device is configured to collect first voice data of a first user of a news text to be broadcast in multiple languages, and send a request to a server to broadcast the news text to be broadcast in multiple languages by a second user voice;
[0016] The server is used to construct a speech posterior probability graph PPG feature extractor and a speech synthesis model for each user; for the first speech data in each language, the PPG feature extractor is used to determine the PPG feature data of the first speech data based on the first acoustic feature data of the first speech data including the first user's voiceprint information and speech content information; and the speech synthesis model of the second user is used to generate second speech data in multiple languages broadcast by the second user based on the PPG feature data and the second acoustic feature data of the first speech data including prosody information.
[0017] This application also provides a voice conversion method, including:
[0018] Build a speech posterior probability graph (PPG) feature extractor and a speech synthesis model for each user;
[0019] Determining, by the PPG feature extractor, PPG feature data of the first voice data according to first acoustic feature data of the first voice data of the first user, wherein the first acoustic feature data includes voiceprint information and voice content information of the first user;
[0020] The second voice data of the second user corresponding to the first voice data is generated through the voice synthesis model of the second user according to the PPG feature data and the second acoustic feature data of the first voice data; the second acoustic feature data includes prosody information.
[0021] Optionally, generating, by the speech synthesis model of the second user, second speech data of the second user corresponding to the first speech data according to the PPG feature data and second acoustic feature data of the first speech data includes:
[0022] Determining, by a feature synthesis network included in the speech synthesis model, third acoustic feature data having a second user timbre based on the PPG feature data and the second acoustic feature data;
[0023] The second speech data is generated according to the third acoustic feature data by a vocoder included in the speech synthesis model.
[0024] Optionally, also include:
[0025] Determine the sampling rate of the PPG feature data as the PPG feature sampling rate of the feature synthesis network;
[0026] determining, by a feature synthesis network, third acoustic feature data having a second user timbre and corresponding to the sampled PPG feature data based on the PPG feature data and the second acoustic feature data;
[0027] The second speech data having a speech playback speed corresponding to the sampling rate is generated according to the third acoustic feature data through the vocoder included in the speech synthesis model.
[0028] Optionally, determining the sampling rate of the PPG feature data includes:
[0029] determining a playback speed of the second voice data;
[0030] The sampling rate is determined according to the playback speed.
[0031] Optionally, the structure of the feature synthesis network includes: a FastSpeech model.
[0032] Optionally, the FastSpeech model does not include: a predictor and a length regulator.
[0033] Optionally, also include:
[0034] For each user, the feature synthesis network is learned from the corresponding relationship set among the PPS feature data, the second acoustic feature data and the third acoustic feature data of the user's speech data.
[0035] Optionally, also include:
[0036] For each user, the vocoder is learned from a set of correspondences between the third acoustic feature data of the user's speech data and the speech synthesis data.
[0037] Optionally, the speech synthesis model of each user is constructed in the following manner:
[0038] For each user, a speech synthesis model of the user is learned based on the correspondence between the PPS feature data, the second acoustic feature data and the speech synthesis data of the user's speech data.
[0039] Optionally, the speech posterior probability graph PPG feature extractor is constructed in the following manner:
[0040] The extractor is obtained by learning the correspondence between the first acoustic feature data of the speech data of multiple users and the pronunciation annotation information.
[0041] This application also provides a method for sending a voice message, comprising:
[0042] collecting first voice data of a first user;
[0043] A request is sent to the server to change the first voice data into the voice of the second user, so that the server constructs a voice posterior probability graph PPG feature extractor and a voice synthesis model for each user; through the PPG feature extractor, the PPG feature data of the first voice data is determined according to the first acoustic feature data of the first voice data; through the voice synthesis model of the second user, the second voice data of the second user corresponding to the first voice data is generated according to the PPG feature data and the second acoustic feature data of the first voice data; and the second voice data is sent to the second terminal device.
[0044] This application also provides a film and television dubbing method, including:
[0045] Collecting first voice data of a first user for a film and television dialogue text in a first language;
[0046] A request is sent to the server to convert the first voice data into the dubbing of the second user, so that the server constructs a voice posterior probability graph PPG feature extractor and a voice synthesis model for each user; through the PPG feature extractor, the PPG feature data of the first voice data is determined according to the first acoustic feature data of the first voice data; through the voice synthesis model of the second user, the second voice data of the second user for the film and television dialogue text is generated according to the PPG feature data and the second acoustic feature data of the first voice data.
[0047] This application also provides a news broadcasting method, including:
[0048] collecting first voice data of a first user of a news text to be broadcast in multiple languages;
[0049] A request is sent to the server for the second user to voice-broadcast the news text in multiple languages, so that the server constructs a speech posterior probability graph PPG feature extractor and a speech synthesis model for each user; for the first speech data in each language, the PPG feature extractor is used to determine the PPG feature data of the first speech data based on the first acoustic feature data of the first speech data; and the speech synthesis model of the second user is used to generate the second speech data in multiple languages broadcast by the second user based on the PPG feature data and the second acoustic feature data of the first speech data.
[0050] This application also provides a voice interaction method, including:
[0051] The first terminal device is configured to determine interaction information between the first user and the second user, and send the interaction information to the server;
[0052] The server is configured to construct speech synthesis models for multiple speaking styles of the first user; determine target speaking style information of the first user corresponding to the interaction information; generate speech data of the interaction information having the target speaking style using the speech synthesis model for the target speaking style of the first user, and transmit the speech data to a second terminal device of the second user;
[0053] The second terminal device is used to play the voice data.
[0054] This application also provides a voice interaction method, including:
[0055] Constructing speech synthesis models for various speaking styles of the first user;
[0056] Determining target speaking style information of the first user based on interaction information between the first user and the second user;
[0057] The speech data of the interactive information with the target speaking mode is generated by using the speech synthesis model of the target speaking mode of the first user, and the speech data is sent to the terminal device of the second user.
[0058] Optionally, determining the target speaking style information of the first user includes:
[0059] The target speaking manner information is determined according to the relationship information between the second user and the first user.
[0060] Optionally, determining the target speaking style information of the first user includes:
[0061] Determining the domain to which the interactive information belongs;
[0062] The target speaking mode information is determined according to the domain information and domains of the first user and the second user respectively.
[0063] Optionally, the interaction information includes: voice interaction information;
[0064] The method further comprises:
[0065] Construct a speech posterior probability graph PPG feature extractor;
[0066] Determining PPG feature data of the voice interaction information by the PPG feature extractor;
[0067] The speech data is generated according to the PPG feature data using a speech synthesis model of the target speaking style.
[0068] This application also provides a voice interaction method, including:
[0069] Determining interaction information of the first user with the second user;
[0070] The interaction information is sent to the server so that the server can determine the target speaking mode information of the first user corresponding to the interaction information; the speech data with the target speaking mode of the interaction information is generated through the speech synthesis model of the target speaking mode of the first user, and the speech data is sent to the terminal device of the second user.
[0071] This application also provides a voice interaction method, including:
[0072] receiving voice data of a first user having a target speaking manner sent by a server;
[0073] Play the voice data.
[0074] This application also provides a voice interaction system, including:
[0075] The first terminal device is configured to send a request for assistance reply for the interaction information of the second user to the first user to the server;
[0076] The server is used to build a voice library for each user; determine the target field corresponding to the interaction information, and generate voice data with the first user's timbre based on the third user's knowledge of the target field based on the voice library of the third user corresponding to the target field, and send the voice data to the terminal device of the second user.
[0077] This application also provides a voice interaction method, including:
[0078] Build a voice library for each user;
[0079] In response to a request for assistance in replying to a request for interaction information of a second user with a first user, sent by the first terminal device, determining a target domain corresponding to the interaction information;
[0080] Based on the voice library of the third user corresponding to the target domain, voice data with the first user's timbre and originating from the third user's knowledge of the target domain is generated, and the voice data is sent to the terminal device of the second user.
[0081] This application also provides a voice interaction method, including:
[0082] receiving interaction information from the second user to the first user;
[0083] Sending a request for help replying to the interaction information to the server.
[0084] This application also provides a voice interaction method, including:
[0085] Sending the interaction information of the second user to the first user to the server;
[0086] receiving voice reply data sent by the server in response to the interaction information and having the first user's timbre and originating from the third user's knowledge in the target domain;
[0087] Play the voice reply data.
[0088] This application also provides a method for constructing a voice library, including:
[0089] Collect multiple voice data of users;
[0090] Determining speaking style information of each voice data;
[0091] A voice database is generated according to the correspondence between the user, the voice data and the speaking manner.
[0092] This application also provides a method for constructing a voice library, including:
[0093] Collect multiple voice data of users;
[0094] Determine domain information of each voice data;
[0095] A voice database is generated according to the correspondence between the user, the voice data and the domain.
[0096] This application also provides a speech translation system, including:
[0097] The terminal device is configured to collect first voice data in a source language of a first user, send a request for translation of the first voice data by a second user to a server, and play second voice data in a target language having a timbre of the second user, which is sent back by the server and corresponds to the first voice data;
[0098] The server is configured to construct a speech conversion model for the second user; generate, in response to the request, third speech data in the target language of the third user corresponding to the first speech data; and generate, through the speech conversion model for the second user, second speech data in the target language having the timbre of the second user corresponding to the third speech data as the second speech data.
[0099] This application also provides a speech translation method, comprising:
[0100] collecting first speech data in a source language of a first user;
[0101] Sending a request for translation of the first voice data by the second user to the server;
[0102] The second voice data in the target language having the second user's timbre and corresponding to the first voice data sent back by the server is played.
[0103] This application also provides a speech translation method, comprising:
[0104] Building a voice conversion model for the second user;
[0105] In response to a request from a second user for translation of first voice data in a source language of a first user sent by a client, generating third voice data in a target language of a third user corresponding to the first voice data;
[0106] Second speech data in the target language having the second user's timbre corresponding to the third speech data is generated through the speech conversion model of the second user, and is sent to the client as the second speech data.
[0107] This application also provides a cross-language speech generation system, comprising:
[0108] The terminal device is configured to determine a text in a first language, send a request to a server for generating a voice message for a first user to read the text aloud; and play first voice data sent back by the server, the first user reading the text aloud; the first user's native language is a second language;
[0109] The server is used to construct a speech posterior probability graph PPG feature extractor and a speech synthesis model for each user; and, in response to the request, determine the second speech data of the second user reading the text, and determine the PPG feature data of the second speech data through the PPG feature extractor based on the first acoustic feature data of the second speech data including the second user's voiceprint information and speech content information; and generate the first speech data through the speech synthesis model of the first user based on the PPG feature data and the second acoustic feature data of the second speech data including prosody information.
[0110] This application also provides a cross-language speech generation method, comprising:
[0111] Identify the text in the first language;
[0112] A request for generating a voice for the first user to read the text aloud is sent to the server.
[0113] This application also provides a cross-language speech generation method, comprising:
[0114] In response to a request for generating a voice message for a first user to read aloud a text in a first language, the client determines second voice data of a second user reading aloud the text;
[0115] The second voice data is converted into first voice data of the first user reading the text through a voice conversion model of the first user.
[0116] Optionally, the first user's native language is the second language, and the second user's native language is the first language.
[0117] Optionally, the request includes: a second user identifier or first language dialect information.
[0118] This application also provides a cross-language speech generation method, comprising:
[0119] Identify the text in the first language;
[0120] Determining second voice data of a second user reading the text aloud;
[0121] The second voice data is converted into first voice data of the first user reading the text through a voice conversion model of the first user.
[0122] This application also provides a cross-dialect speech generation system, including:
[0123] The terminal device is configured to determine a target text and a target dialect, send a request to a server for a first user to read the text in the target dialect, and play first voice data sent back by the server, which is the first user reading the text in the target dialect.
[0124] The server is used to construct a speech posterior probability graph (PPG) feature extractor and a speech synthesis model for each user; and, in response to the request, determine second speech data of the second user reading the text in the target dialect, and determine the PPG feature data of the second speech data through the PPG feature extractor based on the first acoustic feature data of the second speech data including the second user's voiceprint information and speech content information; and generate the first speech data through the speech synthesis model of the first user based on the PPG feature data and the second acoustic feature data of the second speech data including prosody information.
[0125] This application also provides a cross-dialect speech generation method, including:
[0126] Determine the target text and target dialect;
[0127] First voice data of the text read aloud by the first user in the target dialect is sent to the server.
[0128] This application also provides a cross-dialect speech generation method, including:
[0129] In response to a request by a first user to read a target text in a target dialect, determining second speech data of a second user reading a target text in the target dialect;
[0130] The second voice data is converted into first voice data of the first user reading the text in the target dialect through the voice conversion model of the first user.
[0131] This application also provides a cross-dialect speech generation method, including:
[0132] Determine the target text and target dialect;
[0133] determining second voice data of a second user reading the target text in the target dialect;
[0134] The second voice data is converted into first voice data of the first user reading the text in the target dialect through the voice conversion model of the first user.
[0135] This application also provides a voice changing system, comprising:
[0136] The terminal device is configured to determine first voice data of a first user and send a request to a server to change the first voice data into a second user's voice;
[0137] The server is used to build a speech synthesis model for each user; and to build a speech posterior probability graph PPG feature extractor; through the PPG feature extractor, based on the first acoustic feature data of the first speech data including the voiceprint information of the first user and the speech content information, the PPG feature data of the first speech data is determined; through the speech synthesis model of the second user, based on the PPG feature data and the second acoustic feature data of the first speech data including prosody information, the second speech data of the second user corresponding to the first speech data is generated.
[0138] This application also provides a voice changing method, comprising:
[0139] determining first voice data of a first user;
[0140] Send a request to the server to change the first voice data into the second user's voice.
[0141] This application also provides a voice changing method, comprising:
[0142] Build a speech posterior probability graph (PPG) feature extractor and a speech synthesis model for each user;
[0143] In response to a request to morph first voice data of a first user into a voice of a second user, determining, by the PPG feature extractor, PPG feature data of the first voice data based on first acoustic feature data of the first voice data including first user voiceprint information and voice content information;
[0144] The second voice data of the second user corresponding to the first voice data is generated by the voice synthesis model of the second user based on the PPG feature data and the second acoustic feature data of the first voice data including prosody information.
[0145] This application also provides a voice changing method, comprising:
[0146] determining first voice data of a first user;
[0147] Determining, by a speech posterior probability graph (PPG) feature extractor, PPG feature data of the first speech data based on first acoustic feature data of the first speech data including first user voiceprint information and speech content information;
[0148] The second voice data of the second user corresponding to the first voice data is generated based on the PPG feature data and the second acoustic feature data of the first voice data including prosody information through the speech synthesis model of the second user.
[0149] This application also provides a cross-language voice interaction system, including:
[0150] a first terminal device, configured to collect first voice data in a source language of a first user, and send the first voice data to a server;
[0151] The server is configured to determine a target language text corresponding to the first voice data; generate second voice data in the target language of the second user corresponding to the target language text based on the target language voice library of the second user; generate third voice data in the target language having the timbre of the first user corresponding to the second voice data using a voice conversion model of the first user; and send the third voice data to the second terminal device;
[0152] The second terminal device is used to play the third voice data.
[0153] This application also provides a cross-language voice interaction method, including:
[0154] collecting first speech data in a source language of a first user;
[0155] The first voice data is sent to a server so that the server determines a target language text corresponding to the first voice data; second voice data in the target language of the second user corresponding to the target language text is generated based on the target language voice library of the second user; third voice data in the target language with the timbre of the first user is generated corresponding to the second voice data through a voice conversion model of the first user; the third voice data is sent to a second terminal device; and the second terminal device plays the third voice data.
[0156] Optionally, also include:
[0157] Receiving sixth voice data in a source language of a third user sent by a server; the server generating the sixth voice data by performing the following steps: determining a source language text corresponding to the fourth voice data in a target language of the third user sent by a second terminal device; generating fifth voice data in the source language of the fourth user corresponding to the source language text based on a source language speech library of the fourth user; and generating sixth voice data in the source language having the timbre of the third user corresponding to the fifth voice data using a speech conversion model of the third user;
[0158] Play the sixth voice data.
[0159] This application also provides a cross-language voice interaction method, including:
[0160] Determining, for first voice data in a source language of a first user sent by a first terminal device, a target language text corresponding to the first voice data;
[0161] generating second speech data in the target language of the second user corresponding to the target language text based on the target language speech library of the second user;
[0162] The third voice data in the target language having the timbre of the first user is generated corresponding to the second voice data through the voice conversion model of the first user, and the third voice data is sent to the second terminal device.
[0163] Optionally, also include:
[0164] determining, for fourth voice data in a target language of a third user sent by the second terminal device, a source language text corresponding to the fourth voice data;
[0165] generating, based on a source language speech library of a fourth user, fifth speech data in the source language of the fourth user corresponding to the source language text;
[0166] The sixth voice data in the source language having the timbre of the third user corresponding to the fifth voice data is generated by using the voice conversion model of the third user, and the sixth voice data is sent to the first terminal device.
[0167] This application also provides a cross-language voice interaction method, including:
[0168] receiving third voice data in the target language of the first user sent by the server;
[0169] Play the third voice data.
[0170] Optionally, also include:
[0171] collecting fourth speech data in the target language of a third user;
[0172] The fourth voice data is sent to the server so that the server determines the source language text corresponding to the fourth voice data; fifth voice data in the source language of the fourth user corresponding to the source language text is generated based on the source language voice library of the fourth user; sixth voice data in the source language with the timbre of the third user is generated corresponding to the fifth voice data through the voice conversion model of the third user, and the sixth voice data is sent to the first terminal device; and the first terminal device plays the sixth voice data.
[0173] This application also provides a question-answering system, including:
[0174] The terminal device is configured to collect first voice data of a first user, determine a target speaking mode, convert the first voice data into second voice data of the target speaking mode, and send the second voice data to a server;
[0175] The server is used to determine reply information based on the second voice data and send the reply information back to the terminal device.
[0176] This application also provides a question-and-answer method, including:
[0177] collecting first voice data of a first user and determining a target speaking mode;
[0178] Converting the first voice data into second voice data in a target speaking manner;
[0179] The second voice data is sent to the server, so that the server determines the reply information according to the second voice data and sends the reply information back to the terminal device.
[0180] This application also provides a question-and-answer method, including:
[0181] receiving second voice data of a target speaking mode of a first user;
[0182] determining a reply message according to the second voice data;
[0183] Send a reply message back to the terminal device.
[0184] Optionally, determining the reply information according to the second voice data includes:
[0185] determining speaking style information and text information of the second voice data;
[0186] The reply information is determined based on the speaking mode information and the text information.
[0187] This application also provides a voice interaction system, including:
[0188] The first terminal device is configured to collect first voice data of a first user, determine a target speaking mode, and send a request to send the first voice data to a server; and play second voice data of a second user sent back by a second terminal device via the server;
[0189] The server is configured to send a request for the first voice data, convert the first voice data into third voice data in the target speaking mode of the first user, and send the third voice data to the second terminal device; and send the second voice data sent by the second terminal device to the first terminal device;
[0190] The second terminal device is used to play the third voice data; and collect the second voice data and send a second voice data sending request to the server.
[0191] This application also provides a voice interaction method, including:
[0192] collecting first voice data of a first user, determining a target speaking mode, sending a first voice data sending request to a server, so that the server converts the first voice data into third voice data in the target speaking mode of the first user, and sending the third voice data to a second terminal device; and the second terminal device plays the third voice data;
[0193] Play the second voice data of the second user sent back by the second terminal device through the server.
[0194] This application also provides a voice interaction method, including:
[0195] In response to the first voice data sending request sent by the first terminal device, convert the first voice data of the first user into third voice data in the target speaking mode of the first user, and send the third voice data to the second terminal device;
[0196] In response to the second voice data sending request sent by the second terminal device, the second voice data sent by the second terminal device is sent to the first terminal device.
[0197] This application also provides a voice interaction method, including:
[0198] Playing the third voice data of the first user's target speaking mode sent by the server;
[0199] collecting second voice data;
[0200] Send a second voice data sending request to the server.
[0201] This application also provides a method for constructing a speech-to-speech conversion model, including:
[0202] Construct a speech synthesis model for each user and a speech posterior probability graph PPG feature extractor to form a speech conversion model; wherein the PPG feature extractor is used to determine the PPG feature data of the first speech data based on the first acoustic feature data of the first speech data of the first user; the first acoustic feature data includes the first user's voiceprint information and speech content information; the speech synthesis model is used to generate the second speech data of the second user corresponding to the first speech data based on the PPG feature data and the second acoustic feature data of the first speech data; the second acoustic feature data includes prosody information.
[0203] Optionally, constructing a speech synthesis model for each user includes:
[0204] For each user, a speech synthesis model of the user is learned based on the correspondence between the PPS feature data, the second acoustic feature data and the speech synthesis data of the user's speech data.
[0205] Optionally, the speech synthesis model includes: a feature synthesis network and a vocoder; wherein the feature synthesis network is used to determine third acoustic feature data having a second user's timbre based on the PPG feature data and the second acoustic feature data; and the vocoder is used to generate the second speech data based on the third acoustic feature data;
[0206] The method further comprises:
[0207] For each user, learning the feature synthesis network from a set of correspondences between the PPS feature data, the second acoustic feature data, and the third acoustic feature data of the user's speech data;
[0208] For each user, the vocoder is learned from a set of correspondences between the third acoustic feature data of the user's speech data and the speech synthesis data.
[0209] Optionally, the PPG feature extractor is constructed in the following manner:
[0210] The extractor is obtained by learning the correspondence between the first acoustic feature data of the speech data of multiple users and the pronunciation annotation information.
[0211] The present application also provides a voice conversion device, comprising:
[0212] A model building unit, configured to build a speech synthesis model for each user; and a speech posterior probability graph (PPG) feature extractor;
[0213] a PPG feature extraction unit, configured to determine, by the PPG feature extractor, PPG feature data of the first voice data based on first acoustic feature data of the first voice data of the first user, wherein the first acoustic feature data includes voiceprint information and voice content information of the first user;
[0214] A speech synthesis unit is used to generate second speech data of the second user corresponding to the first speech data based on the PPG feature data and second acoustic feature data of the first speech data through the speech synthesis model of the second user; the second acoustic feature data includes prosody information.
[0215] The present application also provides an electronic device, comprising:
[0216] processor; and
[0217] A memory for storing a program for implementing a voice conversion method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: constructing a voice synthesis model for each user; and constructing a voice posterior probability graph PPG feature extractor; determining, through the PPG feature extractor, the PPG feature data of the first voice data of the first user based on the first acoustic feature data of the first voice data; the first acoustic feature data includes the first user's voiceprint information and voice content information; generating, through the voice synthesis model of the second user, the second voice data of the second user corresponding to the first voice data based on the PPG feature data and the second acoustic feature data of the first voice data; the second acoustic feature data includes prosody information.
[0218] The present application also provides a voice message sending device, comprising:
[0219] A voice data collection unit, configured to collect first voice data of a first user;
[0220] A request sending unit is used to send a request to the server to change the first voice data into the voice of the second user, so that the server builds a voice posterior probability graph PPG feature extractor and a voice synthesis model for each user; through the PPG feature extractor, the PPG feature data of the first voice data is determined according to the first acoustic feature data of the first voice data; through the voice synthesis model of the second user, the second voice data of the second user corresponding to the first voice data is generated according to the PPG feature data and the second acoustic feature data of the first voice data; and the second voice data is sent to the second terminal device.
[0221] The present application also provides an electronic device, comprising:
[0222] processor; and
[0223] A memory is used to store a program for implementing a method for sending a voice message. After the device is powered on and the program of the method is run through the processor, the following steps are performed: collecting first voice data of a first user; sending a request to a server to change the first voice data into the voice of a second user, so that the server constructs a voice posterior probability graph PPG feature extractor and a voice synthesis model for each user; determining, through the PPG feature extractor, the PPG feature data of the first voice data based on the first acoustic feature data of the first voice data; generating, through the voice synthesis model of the second user, the second voice data of the second user corresponding to the first voice data based on the PPG feature data and the second acoustic feature data of the first voice data; and sending the second voice data to a second terminal device.
[0224] The present application also provides a voice message sending device, comprising:
[0225] The voice data receiving unit is configured to receive second voice data sent by the server; the second voice data is generated by constructing a posterior probability graph (PPG) feature extractor and a voice synthesis model for each user; in response to a request sent by a first terminal device to change the first voice data of a first user into the voice of a second user, determining, by the PPG feature extractor, PPG feature data of the first voice data based on first acoustic feature data of the first voice data including voiceprint information and voice content information of the first user; and generating, by the voice synthesis model of the second user, second voice data of the second user corresponding to the first voice data based on the PPG feature data and second acoustic feature data of the first voice data including prosody information;
[0226] The voice playing unit is configured to play the second voice data.
[0227] The present application also provides an electronic device, comprising:
[0228] processor; and
[0229] A memory for storing a program for implementing a method for sending a voice message. After the device is powered on and the program of the method is run through the processor, the following steps are performed: receiving second voice data sent by a server; the second voice data is generated by constructing a voice posterior probability graph (PPG) feature extractor and a voice synthesis model for each user; in response to a request sent by a first terminal device to change the first voice data of a first user into the voice of a second user, determining, through the PPG feature extractor, PPG feature data of the first voice data based on first acoustic feature data of the first voice data including voiceprint information and voice content information of the first user; generating, through the voice synthesis model of the second user, second voice data of the second user corresponding to the first voice data based on the PPG feature data and second acoustic feature data of the first voice data including prosody information; and playing the second voice data.
[0230] The present application also provides a film and television dubbing device, comprising:
[0231] A voice data collection unit, configured to collect first voice data of a first user for a film or television dialogue text in a first language;
[0232] A request sending unit is used to send a request to the server to convert the first voice data into the dubbing of the second user, so that the server builds a voice synthesis model for each user; and builds a voice posterior probability graph PPG feature extractor; through the PPG feature extractor, based on the first acoustic feature data of the first voice data, the PPG feature data of the first voice data is determined; through the voice synthesis model of the second user, based on the PPG feature data and the second acoustic feature data of the first voice data, the second voice data of the second user for the film and television dialogue text is generated.
[0233] The present application also provides an electronic device, comprising:
[0234] processor; and
[0235] A memory is used to store a program for implementing a film and television dubbing method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: collecting first voice data of a first user for a film and television dialogue text in a first language; sending a request to a server to convert the first voice data into a dubbing of a second user, so that the server builds a speech synthesis model for each user; and building a speech posterior probability graph PPG feature extractor; determining PPG feature data of the first voice data based on first acoustic feature data of the first voice data through the PPG feature extractor; and generating second voice data of the second user for the film and television dialogue text based on the PPG feature data and second acoustic feature data of the first voice data through the speech synthesis model of the second user.
[0236] The present application also provides a news broadcasting device, comprising:
[0237] A voice data collection unit, configured to collect first voice data of a first user of a news text to be broadcast in multiple languages;
[0238] A request sending unit is used to send a request to the server for the second user to voice broadcast the news text to be broadcast in multiple languages, so that the server builds a speech posterior probability graph PPG feature extractor and a speech synthesis model for each user; for the first speech data in each language, the PPG feature extractor is used to determine the PPG feature data of the first speech data based on the first acoustic feature data of the first speech data; and the speech synthesis model of the second user is used to generate the second speech data in multiple languages broadcast by the second user based on the PPG feature data and the second acoustic feature data of the first speech data.
[0239] The present application also provides an electronic device, comprising:
[0240] processor; and
[0241] The memory is used to store a program for implementing a news broadcasting method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: first voice data of a first user of multiple language versions of a news text to be broadcast is collected; a request is sent to a server for a second user to voice broadcast the multiple language versions of the news text to be broadcast, so that the server constructs a speech posterior probability graph PPG feature extractor and a speech synthesis model for each user; for the first voice data in each language, the PPG feature extractor determines the PPG feature data of the first voice data based on the first acoustic feature data of the first voice data; and the speech synthesis model of the second user generates second voice data in multiple languages broadcast by the second user based on the PPG feature data and the second acoustic feature data of the first voice data.
[0242] The present application also provides a voice interaction device, comprising:
[0243] a model building unit, configured to build speech synthesis models for various speaking styles of the first user;
[0244] a speaking mode determining unit, configured to determine target speaking mode information of the first user based on interaction information between the first user and the second user;
[0245] The speech synthesis unit is used to generate speech data of the interactive information with the target speaking mode through the speech synthesis model of the target speaking mode of the first user, and send the speech data to the terminal device of the second user.
[0246] The present application also provides an electronic device, comprising:
[0247] processor; and
[0248] The memory is used to store a program for implementing the voice interaction method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: constructing a voice synthesis model of multiple speaking styles of the first user; determining the target speaking style information of the first user based on the interaction information of the first user to the second user; generating voice data with the target speaking style of the interaction information through the voice synthesis model of the first user's target speaking style, and sending the voice data to the terminal device of the second user.
[0249] The present application also provides a voice interaction device, comprising:
[0250] an information determining unit, configured to determine interaction information between a first user and a second user;
[0251] An information sending unit is used to send the interaction information to the server so that the server can determine the target speaking mode information of the first user corresponding to the interaction information; generate voice data with the target speaking mode of the interaction information through the speech synthesis model of the target speaking mode of the first user, and send the voice data to the terminal device of the second user.
[0252] The present application also provides an electronic device, comprising:
[0253] processor; and
[0254] A memory is used to store a program for implementing a voice interaction method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: determining interaction information of a first user with a second user; sending the interaction information to a server so that the server can determine target speaking mode information of the first user corresponding to the interaction information; generating voice data with the target speaking mode of the interaction information through a speech synthesis model of the first user's target speaking mode, and sending the voice data to the terminal device of the second user.
[0255] The present application also provides a voice interaction device, comprising:
[0256] A data receiving unit, configured to receive voice data of a first user having a target speaking mode sent by a server;
[0257] The voice playing unit is used to play the voice data.
[0258] The present application also provides an electronic device, comprising:
[0259] processor; and
[0260] The memory is used to store a program for implementing the voice interaction method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: receiving voice data of the first user with a target speaking method sent by the server; and playing the voice data.
[0261] The present application also provides a voice interaction device, comprising:
[0262] A voice library construction unit, used to construct a voice library for each user;
[0263] a domain determining unit, configured to determine, in response to a request for assistance in replying to a request sent by the first terminal device for interaction information of the second user with the first user, a target domain corresponding to the interaction information;
[0264] The voice data sending unit is used to generate voice data with the first user's timbre based on the third user's knowledge of the target field through the voice library of the third user corresponding to the target field, and send the voice data to the terminal device of the second user.
[0265] The present application also provides an electronic device, comprising:
[0266] processor; and
[0267] The memory is used to store a program for implementing a voice interaction method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: building a voice library for each user; in response to a request for help replying to a request sent by a first terminal device for a second user's interaction information with the first user, determining a target field corresponding to the interaction information; generating voice data with the first user's timbre and derived from the third user's knowledge of the target field through a voice library of a third user corresponding to the target field, and sending the voice data to the second user's terminal device.
[0268] The present application also provides a voice interaction device, comprising:
[0269] An information receiving unit, configured to receive interaction information from a second user to a first user;
[0270] The request sending unit is configured to send a request for help replying to the interactive information to the server.
[0271] The present application also provides an electronic device, comprising:
[0272] processor; and
[0273] The memory is used to store a program for implementing the voice interaction method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: receiving interaction information of the second user to the first user; and sending a request for help reply to the interaction information to the server.
[0274] The present application also provides a voice interaction device, comprising:
[0275] A data sending unit, configured to send interaction information of the second user to the first user to the server;
[0276] a data receiving unit, configured to receive voice reply data sent by the server in response to the interaction information and having the first user's timbre and originating from the third user's knowledge in the target field;
[0277] The voice playing unit is used to play the voice reply data.
[0278] The present application also provides an electronic device, comprising:
[0279] processor; and
[0280] A memory is used to store a program for implementing a voice interaction method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: sending interaction information of the second user to the first user to the server; receiving voice reply data sent by the server for the interaction information with the first user's timbre and derived from the third user's knowledge in the target field; and playing the voice reply data.
[0281] The present application also provides a speech library construction device, comprising:
[0282] A voice data collection unit, used to collect multiple voice data of users;
[0283] a speaking mode determining unit, configured to determine speaking mode information of each voice data;
[0284] The voice library generating unit is configured to generate a voice library according to the corresponding relationship between the user, the voice data and the speaking mode.
[0285] The present application also provides an electronic device, comprising:
[0286] processor; and
[0287] The memory is used to store a program for implementing the voice library construction method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: collecting multiple voice data of the user; determining the speaking style information of each voice data; and generating a voice library based on the correspondence between the user, the voice data and the speaking style.
[0288] The present application also provides a speech library construction device, comprising:
[0289] A voice data collection unit, used to collect multiple voice data of users;
[0290] A domain determination unit, configured to determine domain information of each voice data;
[0291] The voice library generating unit is configured to generate a voice library according to the corresponding relationship between the user, the voice data and the domain.
[0292] The present application also provides an electronic device, comprising:
[0293] processor; and
[0294] The memory is used to store a program for implementing the voice library construction method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: collecting multiple voice data of the user; determining the domain information of each voice data; and generating a voice library based on the correspondence between the user, the voice data and the domain.
[0295] This application also provides a speech translation device, comprising:
[0296] A model building unit, configured to build a voice conversion model for a second user;
[0297] a first speech generating unit configured to generate, in response to a request from a second user for translation of first speech data in a source language of a first user sent by a client, third speech data in a target language of a third user corresponding to the first speech data;
[0298] The second speech generating unit is configured to generate second speech data in a target language corresponding to the third speech data and having the timbre of the second user through a speech conversion model of the second user, and send the second speech data to the client as the second speech data.
[0299] The present application also provides an electronic device, comprising:
[0300] processor; and
[0301] The memory is used to store a program for implementing a speech translation method. After the device is powered on and the program of the method is run by the processor, the device performs the following steps: constructing a speech conversion model for a second user; in response to a request for translation by the second user for first speech data in a source language of a first user sent by a client, generating third speech data in a target language of a third user corresponding to the first speech data; and generating second speech data in the target language with the timbre of the second user corresponding to the third speech data using the speech conversion model of the second user, and sending the second speech data to the client as the second speech data.
[0302] This application also provides a speech translation device, comprising:
[0303] A voice data collection unit, configured to collect first voice data in a source language of a first user;
[0304] a request sending unit, configured to send a request for translation by a second user for the first voice data to a server;
[0305] The voice playing unit is used to play the second voice data in the target language with the second user's timbre corresponding to the first voice data sent back by the server.
[0306] The present application also provides an electronic device, comprising:
[0307] processor; and
[0308] The memory is used to store a program for implementing a speech translation method. After the device is powered on and the program of the method is run by the processor, the device performs the following steps: collecting first speech data in a source language of a first user; sending a request for translation of the first speech data by a second user to a server; and playing second speech data in a target language with the second user's timbre, which is sent back by the server and corresponds to the first speech data.
[0309] The present application also provides a cross-language speech generation device, comprising:
[0310] a speech generating unit configured to determine, in response to a speech generation request sent by a client for a first user to read a text in a first language, second speech data of a second user reading the text;
[0311] The voice conversion unit is configured to convert the second voice data into first voice data of the first user reading the text through a voice conversion model of the first user.
[0312] The present application also provides an electronic device, comprising:
[0313] processor; and
[0314] The memory is used to store a program for implementing a cross-language speech generation method. After the device is powered on and the program of the method is run by the processor, the following steps are performed: in response to a speech generation request for a first user to read a text in a first language, determining second speech data of a second user reading the text; and converting the second speech data into first speech data of the first user reading the text through a speech conversion model of the first user.
[0315] The present application also provides a cross-language speech generation device, comprising:
[0316] a text determination unit for determining a text in a first language;
[0317] The request sending unit is used to send a speech generation request for the first user to read the text aloud to the server.
[0318] The present application also provides an electronic device, comprising:
[0319] processor; and
[0320] The memory is used to store a program for implementing a cross-language speech generation method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: determining a text in a first language; and sending a speech generation request to a server for a first user to read the text aloud.
[0321] The present application also provides a cross-language speech generation device, comprising:
[0322] a text determination unit for determining a text in a first language;
[0323] a voice determination unit, configured to determine second voice data of a second user reading the text;
[0324] The voice conversion unit is configured to convert the second voice data into first voice data of the first user reading the text through a voice conversion model of the first user.
[0325] The present application also provides an electronic device, comprising:
[0326] processor; and
[0327] The memory is used to store a program for implementing the cross-language speech generation method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: determining a text in a first language; determining second speech data of a second user reading the text; and converting the second speech data into first speech data of the first user reading the text through a speech conversion model of the first user.
[0328] The present application also provides a cross-dialect speech generation device, comprising:
[0329] a speech generating unit for determining second speech data for a second user to read aloud the target text in the target dialect in response to a request by a first user to read aloud the target text in the target dialect;
[0330] The speech conversion unit is configured to convert the second speech data into first speech data of the first user reading the text in the target dialect through a speech conversion model of the first user.
[0331] The present application also provides an electronic device, comprising:
[0332] processor; and
[0333] The memory is used to store a program for implementing a cross-dialect speech generation method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: in response to a request by a first user to read a target text in a target dialect, second speech data of a second user reading the target text in the target dialect is determined; and the second speech data is converted into first speech data of the first user reading the text in the target dialect through a speech conversion model of the first user.
[0334] The present application also provides a cross-dialect speech generation device, comprising:
[0335] a text determination unit for determining a target text and a target dialect;
[0336] The request sending unit is used to send first voice data of the text read aloud by the first user in the target dialect to the server.
[0337] The present application also provides an electronic device, comprising:
[0338] processor; and
[0339] The memory is used to store a program for implementing a cross-dialect speech generation method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: determining a target text and a target dialect; and sending first speech data of the text read aloud by a first user in the target dialect to a server.
[0340] The present application also provides a cross-dialect speech generation device, comprising:
[0341] a text determination unit for determining a target text and a target dialect;
[0342] a speech determination unit, configured to determine second speech data of a second user reading aloud the target text in a target dialect;
[0343] The speech conversion unit is configured to convert the second speech data into first speech data of the first user reading the text in the target dialect through a speech conversion model of the first user.
[0344] The present application also provides an electronic device, comprising:
[0345] processor; and
[0346] The memory is used to store a program for implementing a cross-dialect speech generation method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: determining a target text and a target dialect; determining second speech data of a second user reading the target text in the target dialect; and converting the second speech data into first speech data of the first user reading the text in the target dialect through a speech conversion model of the first user.
[0347] The present application also provides a voice changing device, comprising:
[0348] A model building unit, used to build a speech posterior probability graph (PPG) feature extractor and a speech synthesis model for each user;
[0349] a feature extraction unit configured to determine, by the PPG feature extractor, PPG feature data of the first voice data based on first acoustic feature data of the first voice data including first user voiceprint information and voice content information, in response to a request to morph first voice data of a first user into a second user's voice;
[0350] A speech generation unit is configured to generate second speech data of a second user corresponding to the first speech data based on the PPG feature data and second acoustic feature data including prosody information of the first speech data by using the speech synthesis model of the second user.
[0351] The present application also provides an electronic device, comprising:
[0352] processor; and
[0353] A memory is used to store a program for implementing a cross-dialect speech generation method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: a speech posterior probability graph (PPG) feature extractor and a speech synthesis model for each user are constructed; in response to a request to change the first speech data of a first user into the voice of a second user, the PPG feature extractor is used to determine the PPG feature data of the first speech data based on first acoustic feature data of the first speech data including the first user's voiceprint information and speech content information; and the speech synthesis model of the second user is used to generate second speech data of the second user corresponding to the first speech data based on the PPG feature data and second acoustic feature data of the first speech data including prosody information.
[0354] The present application also provides a voice changing device, comprising:
[0355] a voice determination unit, configured to determine first voice data of a first user;
[0356] The request sending unit is used to send a request to the server to change the first voice data into the second user's voice.
[0357] The present application also provides an electronic device, comprising:
[0358] processor; and
[0359] The memory is used to store a program for implementing the cross-dialect speech generation method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: determining first voice data of a first user; and sending a request to the server to change the first voice data into the voice of a second user.
[0360] The present application also provides a voice changing device, comprising:
[0361] a voice determination unit, configured to determine first voice data of a first user;
[0362] a feature extraction unit, configured to determine, by a posterior probability graph (PPG) feature extractor, PPG feature data of the first speech data based on first acoustic feature data of the first speech data including first user voiceprint information and speech content information;
[0363] The speech generation unit is configured to generate second speech data of the second user corresponding to the first speech data based on the PPG feature data and second acoustic feature data including prosody information of the first speech data by using a speech synthesis model of the second user.
[0364] The present application also provides an electronic device, comprising:
[0365] processor; and
[0366] The memory is used to store a program for implementing a cross-dialect speech generation method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: determining first speech data of a first user; using a speech posterior probability graph (PPG) feature extractor, based on first acoustic feature data of the first speech data including voiceprint information and speech content information of the first user, determining PPG feature data of the first speech data; using a speech synthesis model of the second user, based on the PPG feature data and second acoustic feature data of the first speech data including prosody information, generating second speech data of the second user corresponding to the first speech data.
[0367] The present application also provides a cross-language voice interaction device, comprising:
[0368] A target language text determination unit, configured to determine a target language text corresponding to first voice data in a source language of a first user sent by a first terminal device;
[0369] a second speech data generating unit, configured to generate second speech data in the target language of the second user corresponding to the target language text based on the target language speech library of the second user;
[0370] The third voice data generating unit is configured to generate third voice data in a target language corresponding to the second voice data and having the first user's timbre through the first user's voice conversion model, and send the third voice data to the second terminal device.
[0371] The present application also provides an electronic device, comprising:
[0372] processor; and
[0373] The memory is used to store a program for implementing a cross-language voice interaction method. After the device is powered on and the program of the method is run by the processor, the device performs the following steps: determining a target language text corresponding to first voice data in a source language of a first user sent by a first terminal device; generating second voice data in the target language of the second user corresponding to the target language text based on a target language voice library of a second user; generating third voice data in the target language with the timbre of the first user corresponding to the second voice data through a voice conversion model of the first user, and sending the third voice data to the second terminal device.
[0374] The present application also provides a cross-language voice interaction device, comprising:
[0375] A voice collection unit, configured to collect first voice data in a source language of a first user;
[0376] A request sending unit is configured to send the first voice data to a server so that the server determines a target language text corresponding to the first voice data; generate second voice data in the target language of the second user corresponding to the target language text based on a target language voice library of the second user; generate third voice data in the target language with the timbre of the first user corresponding to the second voice data through a voice conversion model of the first user; send the third voice data to a second terminal device; and the second terminal device plays the third voice data.
[0377] The present application also provides an electronic device, comprising:
[0378] processor; and
[0379] The memory is used to store a program for implementing a cross-language voice interaction method. After the device is powered on and the program of the method is run by the processor, the device performs the following steps: collecting first voice data in a source language of a first user; sending the first voice data to a server so that the server determines a target language text corresponding to the first voice data; generating second voice data in the target language of the second user corresponding to the target language text based on a target language voice library of the second user; generating third voice data in the target language with the timbre of the first user corresponding to the second voice data through a voice conversion model of the first user; sending the third voice data to a second terminal device; and the second terminal device playing the third voice data.
[0380] The present application also provides a cross-language voice interaction device, comprising:
[0381] a voice receiving unit, configured to receive third voice data in a target language of a first user sent by a server;
[0382] The voice playing unit is used to play the third voice data.
[0383] The present application also provides an electronic device, comprising:
[0384] processor; and
[0385] The memory is used to store a program for implementing the cross-language voice interaction method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: receiving third voice data in the target language of the first user sent by the server; and playing the third voice data.
[0386] This application also provides a question-answering device, comprising:
[0387] A voice receiving unit, configured to receive second voice data of a target speaking mode of a first user;
[0388] an information determining unit, configured to determine reply information based on the second voice data;
[0389] The information sending unit is used to send reply information back to the terminal device.
[0390] The present application also provides an electronic device, comprising:
[0391] processor; and
[0392] The memory is used to store a program for implementing the question-answering method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: receiving second voice data of the target speaking mode of the first user; determining reply information based on the second voice data; and sending the reply information back to the terminal device.
[0393] This application also provides a question-answering device, comprising:
[0394] a voice collection unit, configured to collect first voice data of a first user and determine a target speaking mode;
[0395] A voice data unit, configured to convert the first voice data into second voice data in a target speaking mode;
[0396] The voice sending unit is used to send the second voice data to the server, so that the server determines the reply information according to the second voice data and sends the reply information back to the terminal device.
[0397] The present application also provides an electronic device, comprising:
[0398] processor; and
[0399] The memory is used to store a program for implementing the question-answering method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: first voice data of the first user is collected and a target speaking mode is determined; the first voice data is converted into second voice data of the target speaking mode; the second voice data is sent to the server, so that the server determines reply information based on the second voice data and sends the reply information back to the terminal device.
[0400] The present application also provides a voice interaction device, comprising:
[0401] a first request processing unit, configured to send a request for the first voice data sent by the first terminal device, convert the first voice data of the first user into third voice data in a target speaking mode of the first user, and send the third voice data to the second terminal device;
[0402] The second request processing unit is configured to send the second voice data sent by the second terminal device to the first terminal device in response to a second voice data sending request sent by the second terminal device.
[0403] The present application also provides an electronic device, comprising:
[0404] processor; and
[0405] A memory is used to store a program for implementing a voice interaction method. After the device is powered on and runs the program of the method through the processor, the following steps are performed: in response to a first voice data sending request sent by a first terminal device, the first voice data of the first user is converted into third voice data of the first user's target speaking mode, and the third voice data is sent to a second terminal device; in response to a second voice data sending request sent by the second terminal device, the second voice data sent by the second terminal device is sent to the first terminal device.
[0406] The present application also provides a voice interaction device, comprising:
[0407] A voice collection unit, configured to collect first voice data of a first user;
[0408] a speaking style determination unit, configured to determine a target speaking style;
[0409] a request sending unit, configured to send a first voice data sending request to a server, so that the server converts the first voice data into third voice data in a target speaking mode of the first user, and sends the third voice data to a second terminal device; and the second terminal device plays the third voice data;
[0410] The voice playing unit is used to play the second voice data of the second user sent back by the second terminal device through the server.
[0411] The present application also provides an electronic device, comprising:
[0412] processor; and
[0413] A memory is used to store a program for implementing a voice interaction method. After the device is powered on and runs the program of the method through the processor, the following steps are performed: collecting first voice data of a first user and determining a target speaking mode, sending a first voice data sending request to a server, so that the server converts the first voice data into third voice data of the first user's target speaking mode, and sending the third voice data to a second terminal device; the second terminal device plays the third voice data; and playing the second voice data of the second user sent back by the second terminal device through the server.
[0414] The present application also provides a voice interaction device, comprising:
[0415] a voice playing unit, configured to play the third voice data of the first user's target speaking mode sent by the server;
[0416] A voice collecting unit, configured to collect second voice data;
[0417] The request sending unit is used to send a second voice data sending request to the server.
[0418] The present application also provides an electronic device, comprising:
[0419] processor; and
[0420] The memory is used to store a program for implementing the voice interaction method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: playing third voice data of the first user's target speaking method sent by the server; collecting second voice data; and sending a request to send the second voice data to the server.
[0421] The present application also provides a computer-readable storage medium, wherein the computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the above-mentioned various methods.
[0422] The present application also provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to perform the above methods.
[0423] Compared with the prior art, this application has the following advantages:
[0424] The speech conversion method provided by the embodiment of the present application constructs a speech synthesis model for each user; and constructs a speech posterior probability graph PPG feature extractor; through the PPG feature extractor, based on the first acoustic feature data of the first speech data of the first user, determines the PPG feature data of the first speech data; through the speech synthesis model of the second user, based on the PPG feature data and the second acoustic feature data of the first speech data, generates the second speech data of the second user corresponding to the first speech data; This processing method uses the PPG feature as the input of the speech synthesis module. The PPG feature retains acoustic information (such as prosody and pronunciation information). The PPG feature of each frame can be universal between different speakers and different languages, thus achieving cross-language, cross-dialect, and same-language speech conversion. In addition, this processing method also allows only the speech of one language to be used for model training and system construction, and can support speech conversion of input speech in different languages, and can obtain the same speech conversion quality in different languages. Furthermore, through model adaptive training, the requirement for the number of target speaker's speech can be reduced, and higher speech synthesis quality can be achieved using less corpus.
[0425] The voice message sending system provided in an embodiment of the present application collects first voice data of a first user through a first terminal device and sends a request to a server to change the first voice data into the voice of a second user, where the first user and the second user have different native languages. The server constructs a speech synthesis model for each user and also constructs a posterior probability graph (PPG) feature extractor for speech. The PPG feature extractor determines PPG feature data of the first voice data based on first acoustic feature data of the first voice data, including voiceprint information and voice content information of the first user. The speech synthesis model of the second user generates second voice data of the second user corresponding to the first voice data based on the PPG feature data and second acoustic feature data of the first voice data, including prosody information. The second voice data is sent to a second terminal device. The first terminal device plays the second voice data. This processing method uses PPG features as input to the speech synthesis module. The PPG features retain acoustic information (such as prosody and pronunciation information). The PPG features of each frame are universal across different speakers and different languages, thereby enabling cross-language, cross-dialect, and same-language speech conversion. Furthermore, this processing approach allows model training and system development using only speech from a single language to support speech conversion from input speech in different languages, achieving comparable speech conversion quality across all languages. Furthermore, adaptive model training reduces the required number of target speaker voices, enabling higher speech synthesis quality to be achieved with less corpus.
[0426] The film and television dubbing system provided in an embodiment of the present application collects first speech data of a first user for a film and television dialogue text through a terminal device, and sends a request to a server to convert the first speech data into a dubbing of a second user; the server builds a speech synthesis model for each user; and builds a speech posterior probability graph (PPG) feature extractor; the PPG feature extractor determines PPG feature data of the first speech data based on first acoustic feature data of the first speech data, including voiceprint information and speech content information of the first user; and the speech synthesis model of the second user generates second speech data of the second user for the film and television dialogue text based on the PPG feature data and second acoustic feature data of the first speech data, including prosody information. This processing method uses PPG features as input to the speech synthesis module. The PPG features retain acoustic information (such as prosody and pronunciation information). The PPG features of each frame are universal across different speakers and different languages, thereby enabling cross-language, cross-dialect, and same-language film and television dubbing. Furthermore, this processing approach allows model training and system development using only speech from a single language, while supporting voice input in different languages for film and television dubbing, achieving comparable dubbing quality across all languages. Furthermore, adaptive model training reduces the required number of target speaker voices, enabling higher speech synthesis quality with less corpus.
[0427] The news broadcasting system provided in an embodiment of the present application collects first speech data of a first user in multiple languages of a news text to be broadcast through a terminal device, and sends a request to a server for a second user to broadcast the news text to be broadcast in multiple languages. The server constructs a speech synthesis model for each user, and also constructs a speech posterior probability graph (PPG) feature extractor. For the first speech data in each language, the PPG feature extractor determines PPG feature data of the first speech data based on first acoustic feature data of the first speech data, including the first user's voiceprint information and speech content information. The speech synthesis model of the second user generates second speech data in multiple languages broadcast by the second user based on the PPG feature data and second acoustic feature data of the first speech data, including prosody information. This processing method uses PPG features as input to the speech synthesis module. The PPG features retain acoustic information (such as prosody and pronunciation information). The PPG features of each frame are universal across different speakers and different languages, thereby enabling cross-language, cross-dialect, and same-language news broadcasting. Furthermore, this processing approach allows the system to support news broadcasts in multiple languages using only speech from a single language for model training and system development, while maintaining consistent news broadcast quality across all languages. Furthermore, adaptive model training reduces the required number of target speaker voices, enabling higher speech synthesis quality to be achieved with less corpus.
[0428] The voice interaction system provided in the embodiment of the present application determines the interaction information of the first user to the second user through the first terminal device, and sends the interaction information to the server; the server constructs a voice synthesis model of multiple speaking styles of the first user; determines the target speaking style information of the first user corresponding to the interaction information; generates voice data with the target speaking style of the interaction information through the voice synthesis model of the target speaking style of the first user, and sends the voice data to the second terminal device of the second user; the second terminal device plays the voice data; this processing method enables the generation of voice data with appropriate speaking styles according to the characteristics of different environments; therefore, the accuracy of voice tonality can be effectively improved, thereby improving the user experience.
[0429] The voice interaction system provided in the embodiment of the present application sends a request for help replying to the interaction information to the server side through the first terminal device based on the interaction information of the second user to the first user; the server side builds a voice library of each user; determines the target field corresponding to the interaction information, and generates voice data with the timbre of the first user and the knowledge of the third user in the target field based on the voice library of the third user corresponding to the target field, and sends the voice data to the terminal device of the second user; this processing method enables, when the interaction information involves a field that the first user is not familiar with, the first user can refer to the voice data of other users in the field to determine the voice reply information of the first user to the second user; therefore, the accuracy of the voice reply can be effectively improved, thereby improving the user experience.
[0430] The voice library construction method provided in the embodiment of the present application collects multiple voice data of users; determines the speaking style information of each voice data; and generates a voice library based on the correspondence between the user, the voice data and the speaking style. This processing method enables the construction of a scene library of voice characteristics (speaking style) of different people. The voice library can be called in combination with different scenarios to solve problems in different scenarios; therefore, it can provide a data foundation for related voice processing.
[0431] The voice library construction method provided in the embodiment of the present application collects multiple voice data of users; determines the domain information of each voice data; and generates a voice library based on the correspondence between the user, the voice data and the domain. This processing method enables the construction of domain knowledge scenario libraries for different people, which can be called in combination with different scenarios to solve problems in different scenarios; therefore, it can provide a data foundation for related voice processing.
[0432] The speech translation system provided in the embodiment of the present application uses a terminal device to collect first speech data in the source language of a first user, and sends a request for translation of the first speech data by a second user to a server; and plays second speech data in the target language with the timbre of the second user corresponding to the first speech data sent back by the server; the server is used to build a speech conversion model for the second user; in response to the request, third speech data in the target language of a third user corresponding to the first speech data is generated; through the speech conversion model of the second user, speech data in the target language with the timbre of the second user corresponding to the third speech data is generated as the second speech data; this processing method makes it possible to generate speech data of the second user translating the source language speech into the target language even if the second user does not have the ability to translate the source language speech into the target language speech, that is, to realize personalized translation and achieve a translation effect of "speaking like the real thing".
[0433] The cross-language speech generation system provided in an embodiment of the present application determines text in a first language through a terminal device, sends a speech generation request for a first user to read the text aloud to a server, and plays first speech data of the first user reading the text sent back by the server; the first user's native language is a second language; the server constructs a speech posterior probability graph (PPG) feature extractor and a speech synthesis model for each user; and, in response to the request, determines second speech data of a second user reading the text aloud, and determines PPG feature data of the second speech data based on first acoustic feature data of the second speech data, including voiceprint information and speech content information of the second user, using the PPG feature extractor; and generates the first speech data based on the PPG feature data and second acoustic feature data of the second speech data, including prosody information, using the speech synthesis model of the first user. This processing method enables the generation of speech data of the first user reading text in a certain language even if the first user does not have the ability to read text in that language, thereby realizing cross-language text reading. The cross-dialect speech generation system provided by an embodiment of the present application determines a target text and a target dialect through a terminal device, and sends a request to a server for a first user to read the text in the target dialect; and plays first speech data of the first user reading the text in the target dialect sent back by the server; the server constructs a posterior probability graph (PPG) feature extractor for speech and a speech synthesis model for each user; and, in response to the request, determines second speech data of a second user reading the text in the target dialect, and uses the PPG feature extractor to determine PPG feature data of the second speech data based on first acoustic feature data including the second user's voiceprint information and speech content information; and uses the speech synthesis model (first speech synthesis model) of the first user to generate the first speech data based on the PPG feature data and second acoustic feature data including prosody information of the second speech data. This processing method enables the generation of speech data of the first user reading the text in the target dialect even if the first user does not have good dialect listening and speaking skills, thereby realizing cross-dialect text reading.
[0434] The speech voice changing system provided in the embodiment of the present application determines the first speech data of the first user through the terminal device and sends a request to the server to change the first speech data into the voice of the second user; the server builds a speech synthesis model for each user; and builds a speech posterior probability graph PPG feature extractor; through the PPG feature extractor, based on the first acoustic feature data of the first speech data including the first user's voiceprint information and speech content information, determines the PPG feature data of the first speech data; through the speech synthesis model of the second user, based on the PPG feature data and the second acoustic feature data of the first speech data including prosody information, generates the second speech data of the second user corresponding to the first speech data; this processing method uses the PPG feature as the input of the speech synthesis module, and the PPG feature retains acoustic information (such as prosody and pronunciation information). The PPG feature of each frame can be used between different speakers, thereby achieving the voice change of the first user's voice into the second user's voice; therefore, the accuracy of the speech voice change can be effectively improved. At the same time, since the system does not need to recognize the text information of the first speech data, the efficiency of the speech voice change can be effectively improved.
[0435] The cross-language voice interaction system provided in the embodiment of the present application collects first voice data of the first user in the source language through a first terminal device, and sends the first voice data to a server; the server determines the target language text corresponding to the first voice data; generates second voice data of the second user in the target language corresponding to the target language text based on the target language voice library of the second user; generates third voice data in the target language with the timbre of the first user corresponding to the second voice data through the voice conversion model of the first user; sends the third voice data to a second terminal device; the second terminal device is used to play the third voice data; this processing method enables users of different languages to directly use their respective native languages for voice interaction with each other; therefore, the efficiency of cross-language voice interaction can be effectively improved.
[0436] The voice interaction system provided in the embodiment of the present application is used to collect first voice data of a first user through a terminal device, determine a target speaking method, convert the first voice data into second voice data of the target speaking method, and send the second voice data to a server; the server is used to determine reply information based on the second voice data, and send the reply information back to the terminal device; this processing method enables the terminal device to send to the server voice data of the speaking method specified by the user, rather than the actual collected voice data of the user's real speaking method. The user can conceal his or her true emotions by specifying the speaking method, and the server can determine the reply information based on the voice data after the speaking method is changed. In this way, even for the same voice content, if the speaking method is different, the server reply information will be different; therefore, the accuracy of the question and answer reply information can be effectively improved, thereby improving the user experience.
[0437] The voice interaction system provided in the embodiment of the present application uses a first terminal device to collect first voice data of a first user, determine a target speaking method, and send a first voice data sending request to a server; and play the second voice data of the second user sent back by the second terminal device through the server; the server is used to send a request for the first voice data, convert the first voice data into third voice data of the first user's target speaking method, and send the third voice data to the second terminal device; and send the second voice data sent by the second terminal device to the first terminal device; the second terminal device is used to play the third voice data; and collect the second voice data and send a second voice data sending request to the server; this processing method allows the second terminal device to play the voice data of the first user's specified speaking method when the first user communicates with the second user by voice, rather than the actually collected voice data of the first user's real speaking method. The first user can conceal his true emotions from the second user by specifying the speaking method; therefore, the quality of voice interaction can be effectively improved, thereby improving the user experience.
[0438] The speech conversion model construction method provided in the embodiments of the present application forms a speech conversion model by constructing a speech synthesis model for each user and a posterior probability graph (PPG) feature extractor for speech. The PPG feature extractor is configured to determine PPG feature data for the first speech data of the first user based on first acoustic feature data of the first speech data of the first user. The first acoustic feature data includes the first user's voiceprint information and speech content information. The speech synthesis model is configured to generate second speech data of the second user corresponding to the first speech data based on the PPG feature data and second acoustic feature data of the first speech data. The second acoustic feature data includes prosody information. This processing method enables the use of PPG features as model training data to construct a model that supports cross-language, cross-dialect, and same-language speech conversion. Furthermore, this processing method allows the use of speech from only one language for model training to support speech conversion for input speech in different languages, and to achieve equivalent speech conversion quality across different languages. Furthermore, through model adaptive training, the requirement for the number of target speaker's speech can be reduced, and higher speech synthesis quality can be achieved using less corpus. BRIEF DESCRIPTION OF THE DRAWINGS
[0439] Figure 1 A schematic structural diagram of an embodiment of a voice message sending system provided by the present application;
[0440] Figure 2 A schematic diagram of a scenario of an embodiment of a voice message sending system provided by the present application;
[0441] Figure 3 A schematic diagram of device interaction in an embodiment of a voice message sending system provided by the present application;
[0442] Figure 4 A flowchart of an embodiment of a voice conversion method provided by the present application;
[0443] Figure 5 A schematic diagram of a model of an embodiment of a speech conversion method provided by the present application;
[0444] Figure 6 A schematic diagram of a feature synthesis module based on the Tacotron2 model structure in an embodiment of a speech conversion method provided by the present application;
[0445] Figure 7 A schematic diagram of a feature synthesis module based on a Transformer model structure in an embodiment of a speech conversion method provided by the present application;
[0446] Figure 8 A schematic diagram of a feature synthesis module based on the FastSpeech model structure in an embodiment of a cross-language speech conversion method provided in the application. DETAILED DESCRIPTION
[0447] The following description sets forth many specific details to facilitate a thorough understanding of the present application. However, the present application can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of the present application. Therefore, the present application is not limited to the specific implementations disclosed below.
[0448] This application provides a voice message sending system, method, and device, a film and television dubbing system, method, and device, a news broadcasting system, method, and device, a cross-language voice conversion method and device, and an electronic device. Each of these solutions is described in detail in the following embodiments.
[0449] First embodiment
[0450] Please refer to Figure 1 , which is a schematic diagram of an embodiment of a voice message sending system of the present application. The voice message sending system provided in this embodiment includes: a server 1, a first terminal device 2, and a second terminal device 3.
[0451] The server can be a server deployed on a cloud server, or a server dedicated to implementing a voice message sending system, which can be deployed in a data center.
[0452] The first terminal device and the second terminal device include but are not limited to mobile communication devices, namely, commonly known as mobile phones or smart phones, and also include personal computers, PADs, iPads and other terminal devices.
[0453] Please refer to Figure 2 , which is a scenario diagram of the voice message sending system of the present application. The server, the first terminal device and the second terminal device can be connected through a network, for example, the first terminal device and the second terminal device can be connected to the network through WIFI, etc. The voice information speaker (the first user) speaks a first voice with the timbre of the first user and specifies that the voice data be converted into the voice of the second user. The first voice data is uploaded to the server, and the server converts the first voice data into second voice data with the voice of the second user, and sends the second voice data to the device of the third user for playback.
[0454] For example, an input method tool is installed on the first terminal device, and a voice changing function is provided in the input method tool. When the first user sends a voice message, he can arbitrarily select the voice of another person. The person who receives the voice message on the other end actually hears the voice of the selected person when the voice message is played.
[0455] Please refer to Figure 3, which is a device schematic diagram of the voice message sending system of the present application. In this embodiment, the first terminal device is used to collect first voice data of the first user and send a request to the server to change the first voice data into the voice of the second user, where the first user and the second user have different native languages; the server is used to build a speech synthesis model for each user; and a speech posterior probability graph PPG feature extractor; the PPG feature extractor determines the PPG feature data of the first voice data based on the first acoustic feature data of the first voice data including the voiceprint information and voice content information of the first user; the speech synthesis model of the second user generates the second voice data of the second user corresponding to the first voice data based on the PPG feature data and the second acoustic feature data of the first voice data including prosody information; the second voice data is sent to the second terminal device; and the first terminal device is used to play the second voice data.
[0456] The first user and the second user may have the same native language or different native languages. If the two users have the same native language, the system can realize voice conversion of different users' timbre in the same language; if the two users have different native languages, the system can realize cross-language voice conversion.
[0457] The specific processing process of the server is detailed in the description of the second embodiment and will not be repeated here.
[0458] In a specific implementation, the first user can specify the second user through the first terminal device, and the request includes the second user identifier. The server can pre-build and store the second user's speech synthesis model, retrieve the second user's speech synthesis model parameters based on the second user identifier, and use the second user's speech synthesis model to convert the first user's first speech data into second speech data with the second user's timbre. Of course, the server can also store a second user that is fixedly associated with the first user. In this way, the first user does not need to specify the second user through the first terminal device. The server only converts the first user's first speech data into second speech data with the second user's timbre.
[0459] As can be seen from the above embodiments, the voice message sending system provided by the embodiments of the present application collects first voice data of a first user through a first terminal device, and sends a request to a server to change the first voice data into the voice of a second user, where the first user and the second user have different native languages; the server constructs a speech synthesis model for each user; and constructs a speech posterior probability graph (PPG) feature extractor; the PPG feature extractor determines PPG feature data of the first voice data based on first acoustic feature data of the first voice data, including voiceprint information and voice content information of the first user; the speech synthesis model of the second user generates second voice data of the second user corresponding to the first voice data based on the PPG feature data and second acoustic feature data of the first voice data, including prosody information; the second voice data is sent to the second terminal device; and the first terminal device plays the second voice data. This processing method uses PPG features as input to the speech synthesis module. The PPG features retain acoustic information (such as prosody and pronunciation information). The PPG features of each frame are universal across different speakers and different languages, thereby enabling cross-language, cross-dialect, and same-language speech conversion. Furthermore, this processing approach allows model training and system development using only speech from a single language to support speech conversion from input speech in different languages, achieving comparable speech conversion quality across all languages. Furthermore, adaptive model training reduces the required number of target speaker voices, enabling higher speech synthesis quality to be achieved with less corpus.
[0460] Second embodiment
[0461] Please refer to Figure 4 , which is a flow chart of an embodiment of the cross-language speech conversion method of the present application. The execution subject of the cross-language speech conversion method provided in this embodiment can be a server, etc. The method includes:
[0462] Step S101: Construct a speech synthesis model for each user; and construct a speech posterior probability graph PPG feature extractor.
[0463] The Phonetic Posterior Gram (PPG) is a time-to-class matrix that represents the posterior probability of each phonetic class for each specific time frame of an utterance. Phonetic classes can refer to words, phonemes, or phoneme states (senones).
[0464] A speech segment can include three aspects of information: speech content (what is said), speaker's timbre (voiceprint), and speaking style. Speaking style can include things like tone and rhythm. PPG features can include speech pronunciation information (such as the pronunciation of text content) and some information about speaking style.
[0465] The PPG feature extractor extracts the speaker-related acoustic features of each frame into a speaker-independent speech posterior graph. This posterior graph retains acoustic information, such as prosody and pronunciation. The posterior graph of each frame can be used across different speakers, and even across speakers of different languages (such as different native languages) and dialects.
[0466] In practice, the deep neural network (DNN)-based acoustic model in a speech recognition system can be used as a PPG extractor. This example uses a 5-layer GRU network structure and a 488-dimensional PPG vector. The output of the layer immediately preceding the DNN acoustic model's output layer is extracted as the PPG feature, which is then logarithmically scaled.
[0467] In one example, the extractor is learned from the correspondence between first acoustic feature data of speech data of multiple users and pronunciation annotation information. The first acoustic feature data includes, but is not limited to, Mel-spectral coefficient (MFCC) acoustic features. MFCC is an acoustic feature that contains voiceprint information and speech content information.
[0468] The PPG feature extractor is common to multiple users, and the speech synthesis model is dedicated to each user, that is, each user has his or her own speech synthesis model.
[0469] In one example, the construction of the speech synthesis model for each user can be achieved in the following manner: for each user, the speech synthesis model of the user is learned based on the correspondence between the PPS feature data, the second acoustic feature data and the speech synthesis data (waveform) of the user's speech data.
[0470] The second acoustic feature data may include prosody information, and in specific implementations, may be Log-F0 features, for example. MFCCs may include voiceprint information and speech content information, while Log-F0 may include prosody information. In this embodiment, the prosody information included in the PPG feature is insufficient, so Log-F0 is used to supplement the prosody information lost by the PPG feature.
[0471] like Figure 5As shown, in one example, the speech synthesis model includes: a feature synthesis network module (Synthesizer network) and a vocoder. The feature synthesis network is used to determine third acoustic feature data (such as LPCNet acoustic features) having the second user's timbre based on the PPG feature data and the second acoustic feature data (such as Log-F0); the vocoder is used to generate the second speech data (waveform) based on the third acoustic feature data.
[0472] The third acoustic feature data may include: information about the first user's speech content, the second user's timbre (voiceprint), and the first user's speaking style (such as tone and accent).
[0473] In specific implementation, the feature synthesis network module can synthesize the 20-dimensional LPCNet acoustic features of the target speaker by combining the PPG features extracted by the PPG extractor and the Log-F0 extracted from the first speech data. The design of this module is based on the existing speech synthesis models such as Tacotron2, Transformer, and FastSpeech. Figures 6 to 8 The improved Tacotron2, Transformer, and FastSpeech are shown. This example replaces the embedding and output layers of the speech synthesis module. The original input layer data of the speech synthesis module is character or phoneme embeddings, and the output layer data is an 80-dimensional spectrogram vector.
[0474] It should be noted that the PPG input features used to train the three feature synthesis networks described above can be cross-lingual feature data. This means that there is no language / speech distinction at the PPG level, so the feature synthesis network models trained using PPG features as input can be applied to other languages. This processing approach allows only speech in one language to be used for model training and system construction, while supporting speech conversion for input speech in different languages and achieving equivalent speech conversion quality across different languages.
[0475] The vocoder can be an LPCNet vocoder, a WaveRNN vocoder, or other similar vocoder. Because the WaveRNN vocoder has limitations in synthesis quality and real-time performance, lagging behind LPCNet-based vocoders, this embodiment uses the LPCNet vocoder. The LPCNet vocoder synthesizes the acoustic features output by the feature synthesis module into an audio file. For example, the LPCNet vocoder uses 20-dimensional acoustic features as input data and directly outputs the speech waveform.
[0476] In this embodiment, three modules are trained independently: the PPG extractor, the feature synthesis module, and the LPCNet vocoder. The PPG extractor can be trained using a speech training library from a speech recognition system, either a public or private library. In practice, Mel-spectrogram frequency coefficient (MFCC) acoustic features and aligned pronunciation annotations are first extracted from the library. Cross-entropy loss is then used for training. The PPG extractor is speaker-independent.
[0477] In one example, for each user, the feature synthesis network is learned from a set of correspondences between the PPS feature data, the second acoustic feature data, and the third acoustic feature data of the user's speech data.
[0478] In practice, the feature synthesis module is trained using a speech library of a selected target speaker. Log-F0 features and LPCNet acoustic features are extracted from the speech library, and posterior spectral features are extracted using a PPG extractor. The posterior spectral and Log-F0 features serve as input features for the network, while the LPCNet acoustic features serve as output features. MSELoss can be used for network training.
[0479] In one example, the structure of the feature synthesis network includes a FastSpeech model; the FastSpeech model does not include: a predictor and a length regulator. Figure 8 As shown, the improved FastSpeech model removes the Duration Predictor and Length Regulator modules. It also improves the FastSpeech model training process, eliminating its reliance on the Teacher model's Duration output. Using a feature synthesis network based on this improved FastSpeech model allows for adjustable speech rate during the speech conversion phase, thereby enhancing speech synthesis flexibility.
[0480] exist Figure 5 In the figure, the Encoder encodes the PPG feature input to generate a higher-level information representation, which is beneficial for the subsequent Decoder to decode the input information. The decoder uses the encoder's output information and Log-F0 to generate the target LPCNet feature.
[0481] In one example, for each user, the vocoder is learned from a set of correspondences between the third acoustic feature data of the user's speech data and the speech synthesis data.
[0482] In practice, the LPCNet vocoder can be used. This vocoder is trained using a speech corpus of the same target speaker, with LPCNet acoustic features as input and the speech waveform as output. It can be trained using MSELoss. It's important to note that the feature synthesis module and the LPCNet vocoder are speaker-dependent.
[0483] Step S101 is the model construction phase (model training). After all modules (PPS feature extractor, feature synthesis network, and vocoder) are trained, they are connected to enable speech conversion. In practice, this is accomplished by simply connecting the output of one module to the input of the next, without requiring any special processing.
[0484] Step S103: Determine, by the PPG feature extractor, the PPG feature data of the first voice data according to the first acoustic feature data of the first voice data of the first user.
[0485] In one example, the first user is British and reads a paragraph of English text in authentic English pronunciation. The first acoustic feature data, such as MFC features, can be extracted from the speech using existing technologies.
[0486] In another example, the way a first user (e.g., an English speaker) reads a text in a first language is first learned from a speech library of the user to form a speech synthesis model of the first user. Then, for a text in the first language to be spoken by a second user (e.g., a Chinese speaker who does not speak English or speaks English poorly), a voice file of the text with the timbre of the first user is generated using the speech synthesis model of the first user. Then, using the method provided in this embodiment, the voice file is converted into a voice file with the same reading method but with the timbre of the second user, thereby enabling a person who does not speak English well or does not speak English at all to speak authentic English.
[0487] Step S105: Generate second voice data of the second user corresponding to the first voice data according to the PPG feature data and the second acoustic feature data of the first voice data through the voice synthesis model of the second user.
[0488] Step S103 and step S105 are the speech conversion stages.
[0489] In this embodiment, step S103 may include the following sub-steps: 1) determining, through the feature synthesis network included in the speech synthesis model, third acoustic feature data having the timbre of the second user based on the PPG feature data and the second acoustic feature data; 2) generating, through the vocoder included in the speech synthesis model, the second speech data based on the third acoustic feature data.
[0490] For a given speech to be converted, the feature synthesis network and vocoder of the target speaker are selected. Mel-spectrogram coefficient features (MFCC) and normalized Log-F0 features can be extracted from the given speech. PPG features are extracted from the Mel-spectrogram coefficient features using a PPG extractor. The PPG features and Log-F0 features are then concatenated and input into the feature synthesis module to generate LPCNet acoustic features. Finally, the LPCNet vocoder synthesizes the LPCNet acoustic features into the speech file of the target speaker.
[0491] In one example, the structure of the feature synthesis network includes the above-mentioned improved FastSpeech model; the FastSpeech model does not include: a predictor and a length adjuster; accordingly, the method may also include the following steps: determining the sampling rate of the PPG feature data as the PPG feature sampling rate of the feature synthesis network based on the FastSpeech model; accordingly, through the feature synthesis network based on the FastSpeech model, according to the PPG feature data and the second acoustic feature data, determining third acoustic feature data corresponding to the sampled PPG feature data with a second user timbre; through the vocoder included in the speech synthesis model, generating the second speech data with a speech playback speed corresponding to the sampling rate according to the third acoustic feature data.
[0492] In the FastSpeech model, one frame of PPG feature data corresponds to one frame of tertiary acoustic feature data—that is, one frame of PPG features corresponds to one frame of speech output. Users can control speech speed by adjusting the sampling rate of the PPG features. When accelerating, the PPG features are downsampled. This means that one frame of PPG features is removed at regular intervals, and the synthesized speech in the same frame is correspondingly removed. The end result is a reduction in the number of synthesized speech frames. At the same playback speed, the user will perceive the speech as playing faster. When decelerating, the sampling rate of the PPG features is increased, and the user will perceive the speech as playing slower. This processing approach allows the speed of synthesized speech to be adjusted, effectively improving the flexibility of speech synthesis.
[0493] In specific implementations, determining the sampling rate of the PPG feature data can include the following sub-steps: determining the playback speed of the second voice data; and determining the sampling rate based on the playback speed. This approach allows the user to control the playback speed of the converted voice during voice conversion. Especially for longer voices, faster playback can save time, while slower playback allows for clearer listening. This speed-changing function does not alter the voice's pitch, resulting in the perception of regular speech. In specific implementations, the user can adjust the voice playback speed. The method can determine the PPG feature sampling rate based on the user-specified speed, enabling real-time voice speed-changing.
[0494] In one example, the first user and the second user may be two users with different native languages, and the method can achieve cross-language speech conversion. For example, the second user wants to read an English text, but the user is Chinese and cannot speak English or speaks English poorly. In order to achieve the effect of the second user speaking English fluently, a British person can be designated as the first user. The first user can first read the English text in authentic English as the first speech data. Then, the method provided in the embodiment of the application is used to generate second speech data corresponding to the first speech data with the second user's timbre, as if the second user has a good English proficiency.
[0495] In another example, the first user and the second user may have different dialects, and the method can achieve cross-dialect speech conversion. For example, the second user wants to speak a paragraph in Cantonese, but the user does not speak Cantonese or speaks Cantonese poorly. In order to achieve the effect of the second user speaking Cantonese fluently, a Cantonese person can be designated as the first user. The first user can first speak the paragraph in fluent Cantonese as the first voice data, and then the method provided in the embodiment of the present application is used to generate second voice data with the second user's timbre corresponding to the first voice data, as if the second user has a good level of Cantonese.
[0496] In another example, the method provided in the embodiments of the present application can also be used to perform voice conversion processing on a first user and a second user who speak the same language to achieve a change in timbre. For example, the second user wants to speak a paragraph with clear pronunciation like a host, but the user's language level is average. In order to achieve the broadcasting effect of the second user like a host, a host can be designated as the first user. The host can first speak the paragraph as the first voice data. Then, using the method provided in the embodiments of the present application, second voice data with the second user's timbre corresponding to the first voice data is generated, as if the second user has a better language level.
[0497] In practice, adaptive model training can reduce the required number of target speaker voices. Experiments have shown that using only 500 target speaker voices, the synthesized speech quality is not significantly reduced compared to training with over 10,000 voices. The principle of adaptive training is to first train the model using other data sources, and then iterate on a smaller amount of target data. Because the model performs essentially the same task, only minor adjustments to the model parameters are needed to adapt it to the target data, thus eliminating the need for extensive target data iteration.
[0498] As can be seen from the above embodiments, the speech conversion method provided by the embodiments of the present application constructs a speech synthesis model for each user; and constructs a speech posterior probability graph (PPG) feature extractor; the PPG feature extractor determines PPG feature data of the first speech data of the first user based on first acoustic feature data of the first speech data of the first user; and the speech synthesis model of the second user generates second speech data of the second user corresponding to the first speech data based on the PPG feature data and second acoustic feature data of the first speech data. Furthermore, speech interaction processing is performed based on entity information. This processing method uses PPG features as input to the speech synthesis module. These PPG features retain acoustic information (such as prosody and pronunciation information). The PPG features of each frame are universal across different speakers and different languages, thus enabling cross-language, cross-dialect, and same-language speech conversion. Furthermore, this processing method supports speech conversion for input speech in different languages using only speech in one language for model training and system development, and achieves equivalent speech conversion quality across different languages. Furthermore, through model adaptive training, the requirement for the amount of target speaker speech can be reduced, and higher speech synthesis quality can be achieved using less corpus.
[0499] Third embodiment
[0500] In the above embodiment, a method for voice conversion is provided. Accordingly, this application also provides a device for voice conversion. This device corresponds to the embodiment of the above method. The parts of this embodiment that are identical to the first embodiment are not repeated here. Please refer to the corresponding parts in the first embodiment.
[0501] The present application provides a speech conversion device comprising:
[0502] A model building unit, configured to build a speech synthesis model for each user; and a speech posterior probability graph (PPG) feature extractor;
[0503] a PPG feature extraction unit, configured to determine, by the PPG feature extractor, PPG feature data of the first voice data based on first acoustic feature data of the first voice data of the first user, wherein the first acoustic feature data includes voiceprint information and voice content information of the first user;
[0504] A speech synthesis unit is used to generate second speech data of the second user corresponding to the first speech data based on the PPG feature data and second acoustic feature data of the first speech data through the speech synthesis model of the second user; the second acoustic feature data includes prosody information.
[0505] Fourth embodiment
[0506] This application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple. For relevant details, please refer to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0507] An electronic device according to this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a speech conversion method. After the device is powered on and the program of the method is run by the processor, the following steps are performed: constructing a speech synthesis model for each user; and constructing a speech posterior probability graph (PPG) feature extractor; determining, by the PPG feature extractor, PPG feature data of the first speech data of the first user based on first acoustic feature data of the first speech data; the first acoustic feature data includes the first user's voiceprint information and speech content information; generating, by the speech synthesis model of the second user, second speech data of the second user corresponding to the first speech data based on the PPG feature data and second acoustic feature data of the first speech data; the second acoustic feature data includes prosody information.
[0508] The electronic device may be a smart mobile device, a personal computer, a smart speaker, a chat robot, etc.
[0509] Fifth embodiment
[0510] In the above embodiment, a voice message sending system is provided. Correspondingly, this application also provides a voice message sending method, which can be performed by a first terminal device. This method corresponds to the embodiment of the above system. The parts of this embodiment that are identical to the first embodiment are not repeated here; please refer to the corresponding parts in Example 1.
[0511] A method for sending a voice message provided in this application may include the following steps:
[0512] Step 1: Collect first voice data of a first user;
[0513] Step 2: Send a request to the server to change the first voice data into the voice of the second user, so that the server builds a voice synthesis model for each user; and builds a voice posterior probability graph PPG feature extractor; through the PPG feature extractor, determine the PPG feature data of the first voice data according to the first acoustic feature data of the first voice data; through the voice synthesis model of the second user, generate the second voice data of the second user corresponding to the first voice data according to the PPG feature data and the second acoustic feature data of the first voice data; and send the second voice data to the second terminal device.
[0514] In one example, the first terminal device is loaded with an input method device, which is used to execute the method.
[0515] In another example, the first terminal device is loaded with an instant messaging device, which is used to execute the method.
[0516] Sixth embodiment
[0517] In the above embodiment, a method for sending a voice message is provided. Accordingly, this application also provides a device for sending a voice message. This device corresponds to the embodiment of the above method. The parts of this embodiment that are identical to the first embodiment are not repeated here. Please refer to the corresponding parts in the first embodiment.
[0518] The present application provides a voice message sending device comprising:
[0519] A voice data collection unit, configured to collect first voice data of a first user;
[0520] A request sending unit is used to send a request to the server to change the first voice data into the voice of the second user, so that the server builds a voice posterior probability graph PPG feature extractor and a voice synthesis model for each user; through the PPG feature extractor, the PPG feature data of the first voice data is determined according to the first acoustic feature data of the first voice data; through the voice synthesis model of the second user, the second voice data of the second user corresponding to the first voice data is generated according to the PPG feature data and the second acoustic feature data of the first voice data; and the second voice data is sent to the second terminal device.
[0521] Seventh embodiment
[0522] This application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple. For relevant details, please refer to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0523] An electronic device of this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a voice message sending method. After the device is powered on and the program of the method is run by the processor, the following steps are performed: collecting first voice data of a first user; sending a request to a server to change the first voice data into the voice of a second user, so that the server constructs a voice posterior probability graph PPG feature extractor and a speech synthesis model for each user; determining PPG feature data of the first voice data according to first acoustic feature data of the first voice data through the PPG feature extractor; generating second voice data of the second user corresponding to the first voice data according to the PPG feature data and second acoustic feature data of the first voice data through the speech synthesis model of the second user; and sending the second voice data to a second terminal device.
[0524] Eighth embodiment
[0525] In the above embodiment, a voice message sending system is provided. Accordingly, this application also provides a voice message sending method, which can be performed by a second terminal device, etc. This method corresponds to the embodiment of the above system. The parts of this embodiment that are identical to the first embodiment are not repeated here; please refer to the corresponding parts in the first embodiment.
[0526] A method for sending a voice message provided in this application may include the following steps:
[0527] Step 1: Receive second voice data sent by the server; the second voice data is generated by: constructing a posterior probability graph (PPG) feature extractor and a speech synthesis model for each user; in response to a request sent by a first terminal device to change the first voice data of a first user into the voice of a second user, determining PPG feature data of the first voice data based on first acoustic feature data of the first voice data including voiceprint information and voice content information of the first user through the PPG feature extractor; and generating second voice data of the second user corresponding to the first voice data based on the PPG feature data and second acoustic feature data of the first voice data including prosody information through the speech synthesis model of the second user;
[0528] Step 2: Play the second voice data.
[0529] Ninth embodiment
[0530] In the above embodiment, a method for sending a voice message is provided. Accordingly, this application also provides a device for sending a voice message. This device corresponds to the embodiment of the above method. The parts of this embodiment that are identical to the first embodiment are not repeated here. Please refer to the corresponding parts in the first embodiment.
[0531] The present application provides a voice message sending device comprising:
[0532] The voice data receiving unit is configured to receive second voice data sent by the server; the second voice data is generated by constructing a posterior probability graph (PPG) feature extractor and a voice synthesis model for each user; in response to a request sent by a first terminal device to change the first voice data of a first user into the voice of a second user, determining, by the PPG feature extractor, PPG feature data of the first voice data based on first acoustic feature data of the first voice data including voiceprint information and voice content information of the first user; and generating, by the voice synthesis model of the second user, second voice data of the second user corresponding to the first voice data based on the PPG feature data and second acoustic feature data of the first voice data including prosody information;
[0533] The voice playing unit is configured to play the second voice data.
[0534] Tenth embodiment
[0535] This application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple. For relevant details, please refer to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0536] An electronic device according to the present embodiment includes: a processor and a memory; the memory is used to store a program for implementing a method for sending a voice message. After the device is powered on and the program of the method is run by the processor, the device performs the following steps: receiving second voice data sent by a server; the second voice data is generated by constructing a voice posterior probability graph (PPG) feature extractor and a voice synthesis model for each user; in response to a request sent by a first terminal device to change the first voice data of a first user into the voice of a second user, determining PPG feature data of the first voice data based on first acoustic feature data of the first voice data including voiceprint information and voice content information of the first user through the PPG feature extractor; generating second voice data of the second user corresponding to the first voice data based on the PPG feature data and second acoustic feature data of the first voice data including prosody information through the voice synthesis model of the second user; and playing the second voice data.
[0537] Eleventh embodiment
[0538] In the above embodiment, a speech conversion method is provided. Accordingly, this application also provides a film and television dubbing system. This system corresponds to the embodiment of the above method. The server-side processing in this embodiment, which is identical to that in the second embodiment, will not be repeated here. Please refer to the corresponding section in the second embodiment.
[0539] The present application provides a film and television dubbing system including: a terminal device and a server.
[0540] Among them, the terminal device is used to collect the first voice data of the first user for the film and television dialogue text, and send a request to the server to convert the first voice data into the dubbing of the second user; the server is used to build a speech synthesis model for each user; and build a speech posterior probability graph PPG feature extractor; through the PPG feature extractor, the PPG feature data of the first voice data is determined according to the first acoustic feature data of the first voice data including the voiceprint information of the first user and the voice content information; through the speech synthesis model of the second user, the second voice data of the second user for the film and television dialogue text is generated according to the PPG feature data and the second acoustic feature data of the first voice data including rhythm information.
[0541] The first user and the second user may have the same native language or different native languages. If the two users have the same native language, the system can realize dubbing of movies and TV shows in the same language; if the two users have different native languages, the system can realize dubbing of movies and TV shows in different languages.
[0542] In one example, a terminal device collects first voice data from a first user for a film or TV script in a first language, while a second user's native language is a second language. For example, in a movie dubbing scenario, if actor A (the second user) is a native Chinese speaker but the movie requires actor A to speak English, cross-language speech conversion can replace a sentence spoken by English actor B (the first user) (the first voice data) with actor A's voice, thereby generating actor A's voice speaking English (the second voice data).
[0543] In another example, a terminal device collects first speech data of a first user for a film or TV script in a first language, and a second user's native language is also the first language, such as Chinese. For example, in a movie dubbing scenario, if actor A (the second user) is a native Chinese speaker and speaks standard Mandarin, but the movie requires actor A to speak Cantonese, cross-language speech conversion can replace a sentence spoken by Cantonese actor B (the first user) (the first speech data) with actor A's voice, thereby generating actor A's voice speaking Cantonese (the second speech data).
[0544] As can be seen from the above embodiments, the film and television dubbing system provided by the embodiments of the present application collects first voice data of a first user for a film and television dialogue text through a terminal device, and sends a request to a server to convert the first voice data into a dubbing of a second user; the server builds a speech synthesis model for each user; and builds a speech posterior probability graph PPG feature extractor; the PPG feature extractor determines the PPG feature data of the first voice data based on the first acoustic feature data of the first voice data including the voiceprint information and voice content information of the first user; the speech synthesis model of the second user generates the second voice data of the second user for the film and television dialogue text based on the PPG feature data and the second acoustic feature data of the first voice data including prosody information; this processing method uses the PPG feature as the input of the speech synthesis module. The PPG feature retains acoustic information (such as prosody and pronunciation information). The PPG feature of each frame can be universal between different speakers and different languages, thereby realizing cross-language, cross-dialect, and same-language film and television dubbing. Furthermore, this processing approach allows model training and system development using only speech from a single language, while supporting voice input in different languages for film and television dubbing, achieving comparable dubbing quality across all languages. Furthermore, adaptive model training reduces the required number of target speaker voices, enabling higher speech synthesis quality with less corpus.
[0545] Twelfth embodiment
[0546] In the above-mentioned embodiment, a film and television dubbing system is provided. Correspondingly, this application also provides a film and television dubbing method, which can be performed by a terminal device, etc. This method corresponds to the embodiment of the above-mentioned system. The parts of this embodiment that are identical to the eleventh embodiment will not be repeated here. Please refer to the corresponding parts in the eleventh embodiment.
[0547] The present application provides a method for dubbing a film or television program, which may include the following steps:
[0548] Step 1: collecting first voice data of a first user for a film and television dialogue text in a first language;
[0549] Step 2: Send a request to the server to convert the first voice data into the dubbing of the second user, so that the server builds a voice synthesis model for each user; and builds a voice posterior probability graph PPG feature extractor; through the PPG feature extractor, determine the PPG feature data of the first voice data according to the first acoustic feature data of the first voice data; through the voice synthesis model of the second user, generate the second voice data of the second user for the film and television dialogue text according to the PPG feature data and the second acoustic feature data of the first voice data, and the second user's native language can be a second language.
[0550] Thirteenth embodiment
[0551] In the above embodiment, a method for dubbing a film or television program is provided. Accordingly, this application also provides a device for dubbing a film or television program. This device corresponds to the embodiment of the above method. The parts of this embodiment that are identical to the first embodiment are not repeated here. Please refer to the corresponding parts in the first embodiment.
[0552] The present application provides a device for dubbing a film or television program, comprising:
[0553] A voice data collection unit, configured to collect first voice data of a first user for a film or television dialogue text in a first language;
[0554] A request sending unit is used to send a request to the server to convert the first voice data into the dubbing of the second user, so that the server builds a voice synthesis model for each user; and builds a voice posterior probability graph PPG feature extractor; through the PPG feature extractor, based on the first acoustic feature data of the first voice data, the PPG feature data of the first voice data is determined; through the voice synthesis model of the second user, based on the PPG feature data and the second acoustic feature data of the first voice data, the second voice data of the second user for the film and television dialogue text is generated.
[0555] Fourteenth embodiment
[0556] This application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple. For relevant details, please refer to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0557] An electronic device according to the present embodiment includes: a processor and a memory; the memory is used to store a program for implementing a film and television dubbing method. After the device is powered on and the program of the method is run by the processor, the device performs the following steps: collecting first voice data of a first user for a film and television dialogue text in a first language; sending a request to a server to convert the first voice data into a dubbing of a second user, so that the server builds a speech synthesis model for each user; and building a speech posterior probability graph (PPG) feature extractor; determining PPG feature data of the first voice data based on first acoustic feature data of the first voice data by the PPG feature extractor; and generating second voice data of the second user for the film and television dialogue text based on the PPG feature data and second acoustic feature data of the first voice data by the speech synthesis model of the second user.
[0558] Fifteenth embodiment
[0559] In the above embodiment, a speech conversion method is provided. Accordingly, this application also provides a news broadcast system. This system corresponds to the embodiment of the above method. The server-side processing in this embodiment, which is identical to that in the second embodiment, will not be repeated here. Please refer to the corresponding portion in the second embodiment.
[0560] The present application provides a news broadcasting system comprising: a terminal device and a server.
[0561] Among them, the terminal device is used to collect the first voice data of the first user of multiple language versions of the news text to be broadcast, and send a request to the server for the second user to broadcast the news text to be broadcast in multiple languages. The second user's native language may be a language other than the multiple languages; the server is used to construct a speech synthesis model for each user; and, to construct a speech posterior probability graph PPG feature extractor; for the first voice data of each language, the PPG feature extractor is used to determine the PPG feature data of the first voice data based on the first acoustic feature data of the first voice data including the voiceprint information and voice content information of the first user; and the speech synthesis model of the second user is used to generate the second voice data in multiple languages broadcast by the second user based on the PPG feature data and the second acoustic feature data of the first voice data including rhythmic information.
[0562] In the news broadcasting scenario, we hope to use one person's voice to synthesize voices in multiple languages (such as using the voice of a famous announcer to broadcast news in multiple languages). However, it is usually unlikely that one person will be proficient in multiple languages at the same time. In this case, cross-language speech conversion can be used to use the voice of one person to generate voices in different languages. Using these voices, we can build a multilingual speech synthesis system (TTS).
[0563] In one example, a Chinese announcer is required to report a piece of news in multiple languages, including Chinese, English, Japanese, and French, but his or her proficiency is limited to Chinese and English. Through cross-language speech conversion, the Japanese version of the news reported by a Japanese announcer can be replaced with the voice of the Chinese announcer, and the French version of the news reported by a French announcer can be replaced with the voice of the Chinese announcer, thereby generating the French and Japanese voices of the Chinese announcer.
[0564] In another example, a Chinese standard speaking announcer is required to announce multiple language versions of a piece of news, including Cantonese, Northeastern dialect, and Minnan dialect, but the announcer's broadcasting level is limited to Chinese standard speaking. Through cross-language speech conversion, the Cantonese version of the news announced by the Cantonese announcer can be replaced with the voice of the Chinese standard speaking announcer, and the Minnan dialect version of the news announced by the Minnan dialect announcer can be replaced with the voice of the Chinese standard speaking announcer, thereby generating the voice of the Chinese announcer speaking the Cantonese, Northeastern dialect, and Minnan dialect versions.
[0565] During specific implementation, different language versions of the first voice data may be broadcast by different first users, such as a British person broadcasting the English version of the news, a French person broadcasting the French version of the news, a Cantonese person broadcasting the Cantonese version of the news, and so on.
[0566] As can be seen from the above embodiments, the news broadcasting system provided by the embodiments of the present application collects first voice data of a first user in multiple languages of a news text to be broadcast through a terminal device, and sends a request to a server for a second user to broadcast the news text to be broadcast in multiple languages. The server constructs a speech synthesis model for each user; and constructs a speech posterior probability graph (PPG) feature extractor. For the first voice data in each language, the PPG feature extractor determines PPG feature data of the first voice data based on first acoustic feature data of the first voice data, including voiceprint information and voice content information of the first user. The speech synthesis model of the second user generates second voice data in multiple languages broadcast by the second user based on the PPG feature data and second acoustic feature data of the first voice data, including prosody information. This processing method uses PPG features as input to the speech synthesis module. The PPG features retain acoustic information (such as prosody and pronunciation information). The PPG features of each frame are universal across different speakers and different languages, thereby enabling cross-language, cross-dialect, and same-language news broadcasting. Furthermore, this processing approach allows the system to support news broadcasts in multiple languages using only speech from a single language for model training and system development, while maintaining consistent news broadcast quality across all languages. Furthermore, adaptive model training reduces the required number of target speaker voices, enabling higher speech synthesis quality to be achieved with less corpus.
[0567] Sixteenth embodiment
[0568] In the above embodiment, a news broadcast system is provided. Accordingly, this application also provides a news broadcast method, which can be performed by a terminal device, etc. This method corresponds to the embodiment of the above system. The parts of this embodiment that are identical to the fifteenth embodiment will not be repeated here. Please refer to the corresponding parts of the fifteenth embodiment.
[0569] A news broadcasting method provided by this application may include the following steps:
[0570] Step 1: collecting first voice data of a first user of a news text to be broadcast in multiple languages;
[0571] Step 2: Send a request to the server for the second user to voice-broadcast the news text in multiple languages, so that the server constructs a speech posterior probability graph PPG feature extractor and a speech synthesis model for each user; for the first speech data in each language, the PPG feature extractor is used to determine the PPG feature data of the first speech data according to the first acoustic feature data of the first speech data; and the speech synthesis model of the second user is used to generate the second speech data in multiple languages broadcast by the second user according to the PPG feature data and the second acoustic feature data of the first speech data.
[0572] Seventeenth embodiment
[0573] In the above embodiment, a news broadcast method is provided. Accordingly, this application also provides a news broadcast device. This device corresponds to the embodiment of the above method. The parts of this embodiment that are identical to the first embodiment are not repeated here. Please refer to the corresponding parts in the first embodiment.
[0574] The present application provides a news broadcasting device comprising:
[0575] A voice data collection unit, configured to collect first voice data of a first user of a news text to be broadcast in multiple languages;
[0576] A request sending unit is used to send a request to the server for the second user to voice broadcast the news text to be broadcast in multiple languages, so that the server builds a speech posterior probability graph PPG feature extractor and a speech synthesis model for each user; for the first speech data in each language, the PPG feature extractor is used to determine the PPG feature data of the first speech data based on the first acoustic feature data of the first speech data; and the speech synthesis model of the second user is used to generate the second speech data in multiple languages broadcast by the second user based on the PPG feature data and the second acoustic feature data of the first speech data.
[0577] Eighteenth embodiment
[0578] This application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple. For relevant details, please refer to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0579] An electronic device of this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a news broadcasting method, and after the device is powered on and the program of the method is run by the processor, the following steps are performed: first voice data of a first user of multiple language versions of a news text to be broadcast is collected; a request for a second user to voice broadcast the multiple language versions of the news text to be broadcast to a server, so that the server constructs a speech posterior probability graph PPG feature extractor and a speech synthesis model for each user; for the first voice data in each language, the PPG feature extractor determines the PPG feature data of the first voice data based on the first acoustic feature data of the first voice data; and the speech synthesis model of the second user generates second voice data in multiple languages broadcast by the second user based on the PPG feature data and the second acoustic feature data of the first voice data.
[0580] Nineteenth embodiment
[0581] In the above embodiment, a speech conversion method is provided. Accordingly, this application also provides a speech interaction system. This system corresponds to the embodiment of the above method. The server-side processing in this embodiment, which is identical to that in the second embodiment, will not be repeated here. Please refer to the corresponding section in the second embodiment.
[0582] The present application provides a voice interaction system comprising: a first terminal device, a server, and a second terminal device.
[0583] Among them, the first terminal device is used to determine the interaction information of the first user to the second user, and send the interaction information to the server; the server is used to build a speech synthesis model of multiple speaking methods of the first user; determine the target speaking method information of the first user corresponding to the interaction information; through the speech synthesis model of the target speaking method of the first user, generate voice data with the target speaking method of the interaction information, and send the voice data to the second terminal device of the second user; the second terminal device is used to play the voice data.
[0584] In one example, both the first terminal device and the second terminal device are loaded with an instant messaging tool (such as DingTalk), and the first user communicates with the second user through the instant messaging tool, such as sending voice messages; the first user can input text-based interaction information or voice-based interaction information; the server can synthesize the interaction information into voice data in the target speaking mode of the first user, and the second terminal device of the second user receives and plays the voice data through the instant messaging tool.
[0585] The speaking style may be gentle, sonorous, etc. Table 1 below shows the correspondence between user relationships and speaking styles stored in the server in this embodiment, and Table 2 shows the relationships between users stored in the server in this embodiment.
[0586] User Relationship The first user's speaking style The second user is a customer of the first user Sincerity, professionalism and respect The first user is a customer of the second user presence The first user is the leader of the second user Authoritative and serious The first user is a subordinate of the second user Respect and humility The first user and the second user are colleagues friendly The first user is the parent of the second user Care The first user is the child of the second user respect
[0587] Table 1. Correspondence between user relationships and speaking styles
[0588] First user ID Second user ID User Relationship User A User B The second user is a customer of the first user User A User C The first user is a customer of the second user … …
[0589] Table 2. Relationships between users
[0590] The server can determine the target speaking mode of the first user based on the correspondence between the user relationship and the speaking mode shown in Table 1 and the relationship between users shown in Table 2.
[0591] The speech synthesis model can be a universal model for different speaking styles. The model input data includes interaction information and speaking style information, and the output data includes speech data of the target speaking style. If the interaction information is textual information, the model input data is the corresponding word embedding data; if the interaction information is speech information, the model input data is the corresponding acoustic feature data.
[0592] The speech synthesis model may also include different models corresponding to different speaking styles, and the model input data may include interaction information. The server may construct speech synthesis models for the first user's multiple speaking styles by learning corresponding speech synthesis models for each speaking style from the first user's speech data corresponding to that speaking style. The corresponding speech synthesis models may include information such as the voice tonality of the speaking style.
[0593] In a specific implementation, the speaking style of each piece of speech data can be annotated first. The speech data and the annotated speaking style information can be used as training data. Then, a network structure for the speech synthesis model can be constructed. The network is trained based on the training data, and the training results include network parameters. The network structure can be a neural network structure, including a DNN, CNN, and other structures.
[0594] In one example, the first user specifies target speaking style information through the first terminal device and sends the information to the server.
[0595] In one example, the server determines the target speaking style information for the first user by specifically adopting the following method: determining the target speaking style information based on the relationship information between the second user and the first user. For example, if the second user is a customer of the first user, the target speaking style may include sincerity, professionalism, and respect. If the first user is a customer of the second user, the target speaking style may include presence, etc., so that the consumer feels more present when speaking to the merchant. If the first user is the second user's leader, the target speaking style may include authority and seriousness, so that the leader speaks more seriously to his subordinates. If the first user is the second user's subordinate, the target speaking style may include respect and humility, so that the subordinate speaks more respectfully to the leader. If the first user and the second user are colleagues, the target speaking style may include friendliness, etc., so that colleagues speak more friendly. Parents speak more lovingly to their children, while children speak more respectfully to their parents. This processing method generates voice data with an appropriate speaking style based on the relationship with the conversation partner, effectively improving the accuracy of voice / language intonation and thus enhancing the user experience.
[0596] In one example, the server determines the target speaking style for the first user by: determining the domain to which the interaction information belongs; and determining the target speaking style based on the domain information and domains of the first and second users. These domains include, but are not limited to, e-commerce operations, law, and biopharmaceuticals. For example, if the second user is a child education expert and the first user is a parent, and the interaction is about child education, the first user's speaking style should be more respectful, with target speaking styles including respectfulness and politeness.
[0597] In one example, the interaction information includes voice interaction information; the server is further configured to construct a speech posterior probability graph (PPG) feature extractor; the PPG feature extractor determines PPG feature data for the voice interaction information; and the speech data is generated based on the PPG feature data using a speech synthesis model for the target speaking style. This processing method ensures that the PPG feature does not include information about the original speaking style of the speech, but rather includes information about the speech content. The speech synthesis model is then used to resynthesize the speech data for the first user's target speaking style, allowing the first user to interact directly with the second user through speech. This effectively improves the user experience.
[0598] As can be seen from the above embodiments, the voice interaction system provided in the embodiments of the present application determines the interaction information of the first user to the second user through the first terminal device, and sends the interaction information to the server; the server constructs a voice synthesis model of multiple speaking methods of the first user; determines the target speaking method information of the first user corresponding to the interaction information; generates voice data with the target speaking method of the interaction information through the voice synthesis model of the target speaking method of the first user, and sends the voice data to the second terminal device of the second user; the second terminal device plays the voice data; this processing method enables the generation of voice data with appropriate speaking methods according to the characteristics of different environments; therefore, the accuracy of voice tonality can be effectively improved, thereby improving the user experience.
[0599] Twentieth embodiment
[0600] In the above embodiment, a voice interaction system is provided. Accordingly, this application also provides a voice interaction method, which may be executed by a server, etc. This method corresponds to the embodiment of the above system. The parts of this embodiment that are identical to the nineteenth embodiment will not be repeated here. Please refer to the corresponding parts of the nineteenth embodiment.
[0601] The present application provides a voice interaction method, which may include the following steps:
[0602] Step 1: Construct a speech synthesis model for the first user's multiple speaking styles;
[0603] Step 2: Determine target speaking style information of the first user based on interaction information between the first user and the second user;
[0604] Step 3: Generate voice data of the interactive information in the target speaking mode through the speech synthesis model of the target speaking mode of the first user, and send the voice data to the terminal device of the second user.
[0605] In one example, the target speaking manner information of the first user may be determined in the following manner: the target speaking manner information is determined according to relationship information between the second user and the first user.
[0606] In one example, the target speaking mode information of the first user may be determined by: determining the domain to which the interaction information belongs; and determining the target speaking mode information according to the domain information and domain to which the first user and the second user belong respectively.
[0607] In one example, the interaction information includes: speech interaction information; the method further includes the following steps: constructing a speech posterior probability graph PPG feature extractor; determining the PPG feature data of the speech interaction information through the PPG feature extractor; and generating the speech data based on the PPG feature data through the speech synthesis model of the target speaking mode.
[0608] Twenty-first embodiment
[0609] In the above embodiment, a voice interaction method is provided. Correspondingly, this application also provides a voice interaction device. This device corresponds to the embodiment of the above method. The parts of this embodiment that are the same as the first embodiment are not repeated here. Please refer to the corresponding parts in the first embodiment.
[0610] The present application provides a voice interaction device comprising:
[0611] a model building unit, configured to build speech synthesis models for various speaking styles of the first user;
[0612] a speaking mode determining unit, configured to determine target speaking mode information of the first user based on interaction information between the first user and the second user;
[0613] The speech synthesis unit is used to generate speech data of the interactive information with the target speaking mode through the speech synthesis model of the target speaking mode of the first user, and send the speech data to the terminal device of the second user.
[0614] Twenty-second embodiment
[0615] This application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple. For relevant details, please refer to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0616] An electronic device of this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a voice interaction method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: constructing a voice synthesis model of multiple speaking styles of a first user; determining the target speaking style information of the first user based on the interaction information of the first user to the second user; generating voice data with the target speaking style of the interaction information through the voice synthesis model of the first user's target speaking style, and sending the voice data to the terminal device of the second user.
[0617] Twenty-third embodiment
[0618] In the above embodiment, a voice interaction system is provided. Accordingly, this application also provides a voice interaction method, which can be performed by a first terminal device, etc. This method corresponds to the embodiment of the above system. The parts of this embodiment that are identical to the nineteenth embodiment are not repeated here. Please refer to the corresponding parts of the nineteenth embodiment.
[0619] The present application provides a voice interaction method, which may include the following steps:
[0620] Step 1: Determine interaction information of the first user with the second user;
[0621] Step 2: Send the interaction information to the server so that the server can determine the target speaking method information of the first user corresponding to the interaction information; generate voice data with the target speaking method of the interaction information through the speech synthesis model of the first user's target speaking method, and send the voice data to the terminal device of the second user.
[0622] Twenty-fourth embodiment
[0623] In the above embodiment, a voice interaction method is provided. Correspondingly, this application also provides a voice interaction device. This device corresponds to the embodiment of the above method. The parts of this embodiment that are the same as the first embodiment are not repeated here. Please refer to the corresponding parts in the first embodiment.
[0624] The present application provides a voice interaction device comprising:
[0625] an information determining unit, configured to determine interaction information between a first user and a second user;
[0626] An information sending unit is used to send the interaction information to the server so that the server can determine the target speaking mode information of the first user corresponding to the interaction information; generate voice data with the target speaking mode of the interaction information through the speech synthesis model of the target speaking mode of the first user, and send the voice data to the terminal device of the second user.
[0627] Twenty-fifth embodiment
[0628] This application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple. For relevant details, please refer to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0629] An electronic device of this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a voice interaction method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: determining interaction information of a first user with a second user; sending the interaction information to a server so that the server can determine target speaking mode information of the first user corresponding to the interaction information; generating voice data with the target speaking mode of the interaction information through a voice synthesis model of the first user's target speaking mode, and sending the voice data to the terminal device of the second user.
[0630] Twenty-sixth embodiment
[0631] In the above embodiment, a voice interaction system is provided. Accordingly, this application also provides a voice interaction method, which can be performed by a second terminal device, etc. This method corresponds to the embodiment of the above system. The parts of this embodiment that are identical to the nineteenth embodiment are not repeated here. Please refer to the corresponding parts of the nineteenth embodiment.
[0632] The present application provides a voice interaction method, which may include the following steps:
[0633] Step 1: receiving voice data of a first user having a target speaking style sent by a server;
[0634] Step 2: Play the voice data.
[0635] Twenty-seventh embodiment
[0636] In the above embodiment, a voice interaction method is provided. Correspondingly, this application also provides a voice interaction device. This device corresponds to the embodiment of the above method. The parts of this embodiment that are the same as the first embodiment are not repeated here. Please refer to the corresponding parts in the first embodiment.
[0637] The present application provides a voice interaction device comprising:
[0638] A data receiving unit, configured to receive voice data of a first user having a target speaking mode sent by a server;
[0639] The voice playing unit is used to play the voice data.
[0640] Twenty-eighth embodiment
[0641] This application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple. For relevant details, please refer to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0642] An electronic device of this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a voice interaction method. After the device is powered on and the program of the method is run through the processor, the device performs the following steps: receiving voice data of a first user with a target speaking method sent by a server; and playing the voice data.
[0643] Twenty-ninth embodiment
[0644] In the above embodiment, a speech conversion method is provided. Accordingly, this application also provides a speech interaction system. This system corresponds to the embodiment of the above method. The server-side processing in this embodiment, which is identical to that in the second embodiment, will not be repeated here. Please refer to the corresponding section in the second embodiment.
[0645] The present application provides a voice interaction system comprising: a first terminal device, a server, and a second terminal device.
[0646] Among them, the first terminal device is used to send a request for help reply to the interaction information of the second user to the first user to the server; the server is used to build a voice library of each user; determine the target field corresponding to the interaction information, and generate voice data with the first user's timbre based on the third user's knowledge of the target field based on the voice library of the third user corresponding to the target field, and send the voice data to the terminal device of the second user; the second terminal device is used to play the voice data.
[0647] The server establishes a voice library for each user. If the interactive information of the second user to the first user involves a field that the first user is not familiar with, the server can send a request for help reply through the first terminal device to the server. The server determines the voice reply information of the first user to the second user based on the voice data of other users in the field. The fields include but are not limited to: computer technology, law, e-commerce, patents, etc. This processing method allows a user to get help from other people's voice libraries for fields that he is not familiar with. For example, user A and user B communicate about issues in the field of children's education. Although user B is more familiar with this field, he is not familiar with a certain sub-field. At this time, user B can draw on the relevant knowledge in the voice library of user C, an expert in this sub-field, to reply to some questions in this sub-field raised by user A.
[0648] The second user's interactive information with the first user can be text-based or voice-based. In this embodiment, both the first and second terminal devices are equipped with an instant messaging tool (such as DingTalk). The first user communicates with the second user via the instant messaging tool, such as sending a voice message. The second user's second terminal device can receive and play the voice data generated by the server via the instant messaging tool. In specific implementations, the first and second users can also communicate via other voice interaction tools.
[0649] In this embodiment, the server stores the corresponding records between users and voice data to form a voice library for each user. Based on the specific content of the voice data, the server can determine the domain associated with each piece of voice data. These domains can include primary domains (e.g., patent domains) and secondary domains (e.g., infringement determination sub-domains, patent operation sub-domains), etc.
[0650] In a specific implementation, in response to the request sent by the first terminal device, the server may first determine the text content of the interaction information using voice recognition technology; then, based on the text content, determine the associated target domain. Alternatively, the first user may specify the target domain in the request, or the second user may specify the target domain when sending the interaction information to the first user.
[0651] After determining the target domain, the server can determine the information that the first user replies to the second user based on the voice library of the third user (one user or multiple users) involved in the field, which is derived from the third user's knowledge of the target domain, and generate voice reply data with the first user's timbre through the first user's speech synthesis model.
[0652] In one example, the first user can specify a third user, who can then draw upon the domain knowledge of that third user's voice database to communicate with the second user. This approach allows the user to draw upon the domain knowledge of a highly specialized expert, rather than the less specialized domain knowledge of a general practitioner. This effectively improves the accuracy of responses and enhances the user experience.
[0653] In another example, the server provides assistance to the first user based on the voice data of multiple third users. This allows the server to draw on the relevant domain knowledge of more users' voice libraries to communicate with the second user. This approach allows the server to draw on the domain knowledge of more relevant personnel, rather than just the domain knowledge of a single expert, effectively improving the accuracy of response information and thus enhancing the user experience.
[0654] In another example, the server can identify a third user based on the second user's attribute information (such as gender, age group, etc.) and / or preference information, and provide assistance to the first user based on the voice library of the automatically determined third user, thereby generating a reply message suitable for the second user. This processing method generates a reply message that is more suitable for the second user, which may not be applicable to other second users; therefore, it can effectively improve the accuracy of the reply message, thereby enhancing the user experience.
[0655] In one example, the server can pre-convert speech data into knowledge in the relevant field based on the user's speech library. This can be done by generating a knowledge base for each user or a general knowledge base based on the speech libraries of all users. The server then determines the response information corresponding to the interaction information based on this knowledge. The speech synthesis model for the first user can be trained based on the first user's speech data and will not be further described here.
[0656] After the server generates the voice reply data of the first user, it sends the data to the second terminal device of the second user for playback so that the second user can view it.
[0657] It can be seen from the above embodiments that the voice interaction system provided in the embodiments of the present application sends a request for help replying to the interaction information to the server through the first terminal device based on the interaction information of the second user to the first user; the server builds a voice library of each user; determines the target field corresponding to the interaction information, and generates voice data with the timbre of the first user and the knowledge of the third user in the target field based on the voice library of the third user corresponding to the target field, and sends the voice data to the terminal device of the second user; this processing method enables, when the interaction information of the second user to the first user involves a field that the first user is not familiar with, the first user can refer to the voice data of other users in this field to determine the voice reply information of the first user to the second user; therefore, the accuracy of the voice reply can be effectively improved, thereby improving the user experience.
[0658] Thirtieth embodiment
[0659] In the above embodiment, a voice interaction system is provided. Correspondingly, this application also provides a voice interaction method, which can be executed by a server, etc. This method corresponds to the embodiment of the above system. The parts of this embodiment that are identical to the twenty-ninth embodiment are not repeated here. Please refer to the corresponding parts of the twenty-ninth embodiment.
[0660] The present application provides a voice interaction method, which may include the following steps:
[0661] Step 1: Build a voice library for each user;
[0662] Step 2: determining a target domain corresponding to the interactive information sent by the first terminal device for the second user's request for assistance in replying to the interactive information of the first user;
[0663] Step 3: Generate voice data with the first user's timbre based on the third user's knowledge of the target domain through the third user's voice library corresponding to the target domain, and send the voice data to the second user's terminal device.
[0664] Thirty-first embodiment
[0665] In the above embodiment, a voice interaction method is provided. Correspondingly, this application also provides a voice interaction device. This device corresponds to the embodiment of the above method. The parts of this embodiment that are the same as the first embodiment are not repeated here. Please refer to the corresponding parts in the first embodiment.
[0666] The present application provides a voice interaction device comprising:
[0667] A voice library construction unit, used to construct a voice library for each user;
[0668] a domain determining unit, configured to determine, in response to a request for assistance in replying to a request sent by the first terminal device for interaction information of the second user with the first user, a target domain corresponding to the interaction information;
[0669] The voice data sending unit is used to generate voice data with the first user's timbre based on the third user's knowledge of the target field through the voice library of the third user corresponding to the target field, and send the voice data to the terminal device of the second user.
[0670] Thirty-second embodiment
[0671] This application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple. For relevant details, please refer to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0672] An electronic device of this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a voice interaction method. After the device is powered on and the program of the method is run by the processor, the following steps are performed: building a voice library for each user; in response to a request for help replying to a request sent by a first terminal device for a second user's interaction information with the first user, determining a target field corresponding to the interaction information; generating voice data with the first user's timbre based on the third user's knowledge of the target field through a voice library of a third user corresponding to the target field, and sending the voice data to the second user's terminal device.
[0673] Thirty-third embodiment
[0674] In the above embodiment, a voice interaction system is provided. Accordingly, this application also provides a voice interaction method, which can be performed by a first terminal device, etc. This method corresponds to the embodiment of the above system. The parts of this embodiment that are identical to those in the twenty-ninth embodiment are not repeated here. Please refer to the corresponding parts in the twenty-ninth embodiment.
[0675] The present application provides a voice interaction method, which may include the following steps:
[0676] Step 1: Receive interaction information from the second user to the first user;
[0677] Step 2: Send a help reply request for the interaction information to the server.
[0678] Thirty-fourth embodiment
[0679] In the above embodiment, a voice interaction method is provided. Correspondingly, this application also provides a voice interaction device. This device corresponds to the embodiment of the above method. The parts of this embodiment that are the same as the first embodiment are not repeated here. Please refer to the corresponding parts in the first embodiment.
[0680] The present application provides a voice interaction device comprising:
[0681] An information receiving unit, configured to receive interaction information from a second user to a first user;
[0682] The request sending unit is configured to send a request for help replying to the interactive information to the server.
[0683] Thirty-fifth embodiment
[0684] This application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple. For relevant details, please refer to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0685] An electronic device of this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a voice interaction method. After the device is powered on and the program of the method is run through the processor, the device performs the following steps: receiving interaction information from a second user to a first user; and sending a request for help reply to the interaction information to a server.
[0686] Thirty-sixth embodiment
[0687] In the above embodiment, a voice interaction system is provided. Accordingly, this application also provides a voice interaction method, which can be performed by a second terminal device, etc. This method corresponds to the embodiment of the above system. The parts of this embodiment that are identical to the nineteenth embodiment are not repeated here. Please refer to the corresponding parts of the nineteenth embodiment.
[0688] The present application provides a voice interaction method, which may include the following steps:
[0689] Step 1: Send the interaction information of the second user to the first user to the server;
[0690] Step 2: receiving voice reply data sent by the server in response to the interaction information and having the first user's timbre and based on the third user's knowledge in the target field;
[0691] Step 3: Play the voice reply data.
[0692] Thirty-seventh embodiment
[0693] In the above embodiment, a voice interaction method is provided. Correspondingly, this application also provides a voice interaction device. This device corresponds to the embodiment of the above method. The parts of this embodiment that are the same as the first embodiment are not repeated here. Please refer to the corresponding parts in the first embodiment.
[0694] The present application provides a voice interaction device comprising:
[0695] A data sending unit, configured to send interaction information of the second user to the first user to the server;
[0696] a data receiving unit, configured to receive voice reply data sent by the server in response to the interaction information and having the first user's timbre and originating from the third user's knowledge in the target field;
[0697] The voice playing unit is used to play the voice reply data.
[0698] Thirty-eighth embodiment
[0699] This application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple. For relevant details, please refer to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0700] An electronic device of this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a voice interaction method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: sending the interaction information of the second user to the first user to the server; receiving voice reply data sent by the server for the interaction information with the first user's timbre and derived from the third user's knowledge in the target field; and playing the voice reply data.
[0701] Thirty-ninth embodiment
[0702] In the above-mentioned embodiment, a voice interaction system is provided. Correspondingly, this application also provides a voice interaction method, which can be executed by a server, etc. This method corresponds to the above-mentioned system embodiment. The parts of this embodiment that are identical to the system embodiment will not be repeated here; please refer to the corresponding parts in the system embodiment.
[0703] The present application provides a method for constructing a speech library, which may include the following steps:
[0704] Step 1: Collect multiple voice data of the user;
[0705] Step 2: Determine the speaking style information of each voice data;
[0706] Step 3: Generate a voice library based on the correspondence between the user, the voice data and the speaking style.
[0707] In one example, the method may further include the following step: constructing speech synthesis models for various speaking styles of the user based on the speech library.
[0708] The speech synthesis model can be a universal model for different speaking styles. The model input data may include interaction information and speaking style information, and the output data includes speech data of the target speaking style. If the interaction information is textual information, the model input data is the corresponding word embedding data; if the interaction information is speech information, the model input data may include the corresponding acoustic feature data.
[0709] The speech synthesis model may also correspond to different speaking styles, and the model input data may include interaction information. To construct speech synthesis models for a user's various speaking styles, the following approach may be employed: for each speaking style, a corresponding speech synthesis model is learned from the user's speech data corresponding to that speaking style. The model may include information such as the voice tonality of that speaking style.
[0710] As can be seen from the above embodiments, the voice library construction method provided in the embodiments of the present application collects multiple voice data of users; determines the speaking style information of each voice data; and generates a voice library based on the correspondence between the user, the voice data and the speaking style. This processing method enables the construction of a scene library of voice characteristics (speaking style) of different people, and the voice library can be called in combination with different scenarios to solve problems in different scenarios. Therefore, it can provide a data basis for related voice processing.
[0711] The 40th embodiment
[0712] In the above embodiment, a method for constructing a speech library is provided. Accordingly, this application also provides a device for constructing a speech library. This device corresponds to the embodiment of the above method. The parts of this embodiment that are identical to the thirty-ninth embodiment are not repeated here. Please refer to the corresponding parts of the thirty-ninth embodiment.
[0713] The present application provides a speech library construction device comprising:
[0714] A voice data collection unit, used to collect multiple voice data of users;
[0715] a speaking mode determining unit, configured to determine speaking mode information of each voice data;
[0716] The voice library generating unit is configured to generate a voice library according to the corresponding relationship between the user, the voice data and the speaking mode.
[0717] Forty-first embodiment
[0718] This application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple. For relevant details, please refer to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0719] An electronic device of this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a voice interaction method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: collecting multiple voice data of the user; determining speaking style information of each voice data; and generating a voice library based on the correspondence between the user, the voice data and the speaking style.
[0720] Example 42
[0721] In the above-mentioned embodiment, a voice interaction system is provided. Correspondingly, this application also provides a method for constructing a voice library, which can be executed by a server, for example. This method corresponds to the above-mentioned system embodiment. The parts of this embodiment that are identical to the system embodiment will not be repeated here; please refer to the corresponding parts in the system embodiment.
[0722] The present application provides a method for constructing a speech library, which may include the following steps:
[0723] Step 1: Collect multiple voice data of the user.
[0724] Step 2: Determine the domain information of each voice data.
[0725] The said fields may include primary fields (such as patent fields and biopharmaceutical fields) and secondary fields (such as infringement determination sub-fields and patent operation sub-fields).
[0726] Step 3: Generate a voice library based on the correspondence between the user, the voice data and the domain.
[0727] As can be seen from the above embodiments, the voice library construction method provided in the embodiments of the present application collects multiple voice data of users; determines the domain information of each voice data; and generates a voice library based on the correspondence between the user, the voice data and the domain. This processing method enables the construction of domain knowledge scenario libraries for different people, and the voice library can be called in combination with different scenarios to solve problems in different scenarios. Therefore, it can provide a data basis for related voice processing.
[0728] Forty-third embodiment
[0729] In the above embodiment, a method for constructing a speech library is provided. Accordingly, this application also provides a device for constructing a speech library. This device corresponds to the embodiment of the above method. The parts of this embodiment that are identical to the thirty-ninth embodiment are not repeated here. Please refer to the corresponding parts of the thirty-ninth embodiment.
[0730] The present application provides a speech library construction device comprising:
[0731] A voice data collection unit, used to collect multiple voice data of users;
[0732] A domain determination unit, configured to determine domain information of each voice data;
[0733] The voice library generating unit is configured to generate a voice library according to the corresponding relationship between the user, the voice data and the domain.
[0734] Forty-fourth embodiment
[0735] This application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple. For relevant details, please refer to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0736] An electronic device of this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a method for building a voice library. After the device is powered on and the program of the method is run through the processor, the device performs the following steps: collecting multiple voice data of a user; determining domain information of each voice data; and generating a voice library based on the correspondence between the user, the voice data and the domain.
[0737] Forty-fifth embodiment
[0738] In the above embodiment, a speech conversion method is provided. Accordingly, this application also provides a speech translation system. This system corresponds to the embodiment of the above method. The server-side processing in this embodiment, which is identical to that in the second embodiment, will not be repeated here. Please refer to the corresponding portion in the second embodiment.
[0739] The present application provides a speech translation system including: a terminal device and a server.
[0740] Among them, the terminal device is used to collect first voice data in the source language of the first user, and send a request for translation of the first voice data by the second user to the server; and play the second voice data in the target language with the timbre of the second user corresponding to the first voice data sent back by the server; the server is used to build a voice conversion model for the second user; in response to the request, generate third voice data in the target language of the third user corresponding to the first voice data; through the voice conversion model of the second user, generate voice data in the target language with the timbre of the second user corresponding to the third voice data as the second voice data.
[0741] In one example, the server is specifically configured to determine the source language text of the first speech data using a speech recognition algorithm; determine the target language text corresponding to the source language text using a speech translation algorithm; and generate third speech data in the target language of the third user corresponding to the first speech data using a speech synthesis model for the third user. The speech synthesis model for the third user can be learned from a speech database of the third user. The input data of the model can be the target language text, and the output data can be speech data of the target language text read aloud by the third user.
[0742] For example, at the conference site, the second user collects the voice data (first voice data) of the English speech (source language) of the conference speaker (first user) through the terminal device, and sends the English voice data to the server; the server can first determine the Chinese translation (target language text) of the English voice through the voice recognition algorithm and the voice translation algorithm; then, based on the Chinese voice library of the third user or the Chinese voice synthesis model of the third user, it can generate Chinese voice data (third voice data) of the third user reading the Chinese translation; through the voice conversion model of the second user (the model can execute the voice conversion method of Example 2), Chinese voice data with the timbre of the second user is generated, and the Chinese translated voice data of the second user can be played at the conference site or other places, thereby achieving the effect of simultaneous interpretation using the voice of the second user.
[0743] In one example, the second user's speech conversion model may include the PPS feature extractor and the second user's speech synthesis model in Example 2, and the speech conversion model performs the speech conversion method in Example 2. For related descriptions, please refer to Example 2 and will not be repeated here.
[0744] In specific implementation, the second user can specify a third user, such as a favorite English announcer, and use the announcer's broadcasting method to read the translation; or the server can determine a third user and provide translation broadcasting services to multiple first users through the third user's broadcasting method.
[0745] As can be seen from the above embodiments, the speech translation system provided in the embodiments of the present application uses a terminal device to collect first speech data in the source language of a first user, and sends a request for translation of the first speech data by a second user to the server; and plays the second speech data in the target language with the timbre of the second user corresponding to the first speech data sent back by the server; the server is used to build a speech conversion model for the second user; in response to the request, third speech data in the target language of a third user corresponding to the first speech data is generated; through the speech conversion model of the second user, speech data in the target language with the timbre of the second user corresponding to the third speech data is generated as the second speech data; this processing method makes it possible to generate speech data for the second user to translate the source language speech into the target language even if the second user does not have the ability to translate the source language speech into the target language speech, that is, to realize personalized translation and achieve a translation effect of "speaking like the real thing".
[0746] Forty-sixth embodiment
[0747] In the above embodiment, a speech translation system is provided. Accordingly, this application also provides a speech translation method, which may be performed by a server, etc. This method corresponds to the embodiment of the above system. The parts of this embodiment that are identical to the forty-fifth embodiment will not be repeated here; please refer to the corresponding parts of the forty-fifth embodiment.
[0748] The present application provides a speech translation method, which may include the following steps:
[0749] Step 1: Build a voice conversion model for the second user;
[0750] Step 2: generating third voice data in a target language of a third user corresponding to the first voice data in the source language of the first user in response to a request from the second user for translation sent by the client;
[0751] Step 3: Generate second speech data in the target language with the second user's timbre corresponding to the third speech data through the second user's speech conversion model, and send the second speech data to the client as the second speech data.
[0752] In one example, the method also includes: constructing a speech conversion model for the second user; the speech conversion model for the second user may be constructed in the following manner: constructing a speech synthesis model for the second user and a speech posterior probability graph PPG feature extractor; accordingly, step 3 may include the following sub-steps: 1) determining the PPG feature data of the third speech data based on the first acoustic feature data of the third speech data through the PPG feature extractor included in the speech conversion model; 2) generating the second speech data based on the PPG feature data and the second acoustic feature data of the third speech data through the speech synthesis model of the second user included in the speech conversion model; the second acoustic feature data includes prosody information.
[0753] Forty-seventh embodiment
[0754] In the above embodiment, a speech translation method is provided. Accordingly, this application also provides a speech translation device. This device corresponds to the embodiment of the above method. The parts of this embodiment that are identical to the first embodiment are not repeated here. Please refer to the corresponding parts in the first embodiment.
[0755] The present application provides a speech translation device comprising:
[0756] A model building unit, configured to build a voice conversion model for a second user;
[0757] a first speech generating unit configured to generate, in response to a request from a second user for translation of first speech data in a source language of a first user sent by a client, third speech data in a target language of a third user corresponding to the first speech data;
[0758] The second speech generating unit is configured to generate second speech data in a target language corresponding to the third speech data and having the timbre of the second user through a speech conversion model of the second user, and send the second speech data to the client as the second speech data.
[0759] Forty-eighth embodiment
[0760] This application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple. For relevant details, please refer to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0761] An electronic device according to this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a speech translation method. After the device is powered on and the program of the method is run by the processor, the device performs the following steps: constructing a speech conversion model for a second user; generating third speech data in a target language of a third user corresponding to a request for translation by the second user for first speech data in a source language of a first user sent by a client; and generating second speech data in the target language having the timbre of the second user corresponding to the third speech data using the speech conversion model of the second user, and sending the second speech data to the client as the second speech data.
[0762] Forty-ninth embodiment
[0763] In the above embodiment, a speech translation system is provided. Accordingly, this application also provides a speech translation method, which can be performed by a terminal device, etc. This method corresponds to the embodiment of the above system. The parts of this embodiment that are identical to the 45th embodiment will not be repeated here. Please refer to the corresponding parts of the 45th embodiment.
[0764] The present application provides a speech translation method, which may include the following steps:
[0765] Step 1: Collect first speech data of a first user in a source language;
[0766] Step 2: Sending a translation request for the first voice data by the second user to the server;
[0767] Step 3: Play the second voice data in the target language with the second user's timbre corresponding to the first voice data sent back by the server.
[0768] The 50th embodiment
[0769] In the above embodiment, a speech translation method is provided. Accordingly, this application also provides a speech translation device. This device corresponds to the embodiment of the above method. The parts of this embodiment that are identical to the first embodiment are not repeated here. Please refer to the corresponding parts in the first embodiment.
[0770] The present application provides a speech translation device comprising:
[0771] A voice data collection unit, configured to collect first voice data in a source language of a first user;
[0772] a request sending unit, configured to send a request for translation by a second user for the first voice data to a server;
[0773] The voice playing unit is used to play the second voice data in the target language with the second user's timbre corresponding to the first voice data sent back by the server.
[0774] Fifty-first embodiment
[0775] This application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple. For relevant details, please refer to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0776] An electronic device according to this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a speech translation method. After the device is powered on and the program of the method is run by the processor, the device performs the following steps: collecting first speech data in a source language of a first user; sending a request for translation of the first speech data by a second user to a server; and playing second speech data in a target language with the timbre of the second user, which is sent back by the server and corresponds to the first speech data.
[0777] Fifty-second embodiment
[0778] In the above embodiment, a speech conversion method is provided. Accordingly, this application also provides a cross-language speech generation system. This system corresponds to the embodiment of the above method. The server-side processing in this embodiment, which is identical to that in the second embodiment, will not be repeated here. Please refer to the corresponding part in the second embodiment.
[0779] The present application provides a cross-language speech generation system including: a terminal device and a server.
[0780] Among them, the terminal device is used to determine the text in the first language, and send a speech generation request to the server for the first user to read the text aloud; and play the first speech data of the first user reading the text sent back by the server; the first user's native language is the second language; the server is used to construct a speech posterior probability graph PPG feature extractor, and a speech synthesis model (first speech synthesis model) for each user; and, in response to the request, determine the second speech data of the second user reading the text aloud, and determine the PPG feature data of the second speech data through the PPG feature extractor based on the first acoustic feature data of the second speech data including the second user's voiceprint information and speech content information; and generate the first speech data through the speech synthesis model of the first user based on the PPG feature data and the second acoustic feature data of the second speech data including prosody information.
[0781] The first user's native language is the second language, and the second user's native language is the first language. For example, the first user is Chinese and cannot speak English or speaks English poorly. To achieve the effect of the first user speaking English fluently, a British person can be designated as the second user. The second user can first read the English text in authentic English as the second voice data. Then, the first voice data with the first user's timbre is generated corresponding to the second voice data, as if the first user has a good command of English.
[0782] In one example, the server uses the text as input data for a second speech synthesis model of the second user. The model can generate speech data of the text in the second user's reading style, with the speech data having not only the second user's timbre but also characteristics such as the second user's reading intonation. The second speech synthesis model can be learned from a speech library of the second user.
[0783] It should be noted that the speech synthesis model of the second user is different from the first speech synthesis model of each user constructed by the server. The differences include:
[0784] 1) Different input data: The input data of the first speech synthesis model includes PPS feature data, while the input data of the second user's speech synthesis model includes text;
[0785] 2) The models have different functions: the first speech synthesis model generates speech data with the corresponding user's timbre based on PPS feature data, while the second speech synthesis model generates corresponding speech data with the corresponding user's timbre and reading style based on the input text;
[0786] 3) Different output data: The speech data output by the first speech synthesis model has nothing to do with the speaking style of the user corresponding to the model, but is related to the speaking style of other users; the speech data output by the second speech synthesis model has nothing to do with the speech data of other users.
[0787] In one example, the first user can specify the second user through the terminal device, such as specifying an English announcer as the second user; the first user can also specify the first language dialect information through the terminal device, such as if the first language is English, the specified first language dialect is American English or British English, or, for example, if the first language is Chinese, the specified first language dialect is Standard Chinese or Shanghainese; the server can also arbitrarily determine a user of the first language as the second user.
[0788] As can be seen from the above embodiments, the cross-language speech generation system provided in the embodiments of the present application determines a text in a first language through a terminal device, sends a speech generation request for a first user to read the text aloud to a server, and plays first speech data of the first user reading the text sent back by the server; the first user's native language is a second language; the server constructs a speech posterior probability graph (PPG) feature extractor and a speech synthesis model for each user; and, in response to the request, determines second speech data of a second user reading the text aloud, and determines PPG feature data of the second speech data based on first acoustic feature data of the second speech data including voiceprint information and speech content information of the second user through the PPG feature extractor; and generates the first speech data based on the PPG feature data and second acoustic feature data of the second speech data including prosody information through the speech synthesis model of the first user. This processing method enables the generation of speech data of the first user reading text in a certain language even if the first user does not have the ability to read text in that language, thereby realizing cross-language text reading.
[0789] Fifty-third embodiment
[0790] In the above embodiment, a cross-language speech generation system is provided. Correspondingly, this application also provides a cross-language speech generation method, which can be executed by a server, etc. This method corresponds to the embodiment of the above system. The parts of this embodiment that are identical to the fifty-second embodiment are not repeated here. Please refer to the corresponding parts of the fifty-second embodiment.
[0791] The present application provides a cross-language speech generation method, which may include the following steps:
[0792] Step 1: In response to a speech generation request for a first user to read a text in a first language, determining second speech data of a second user reading the text;
[0793] Step 2: Using the first user's voice conversion model, convert the second voice data into first voice data of the first user reading the text.
[0794] In one example, the method may further include the following steps: constructing a speech posterior probability graph PPG feature extractor and a speech synthesis model for each user; accordingly, step 2 may include the following sub-steps: determining, through the PPG feature extractor, the PPG feature data of the second speech data based on the first acoustic feature data of the second speech data including the voiceprint information of the second user and the speech content information; and generating, through the speech synthesis model of the first user, the first speech data of the first user reading the text based on the PPG feature data and the second acoustic feature data of the second speech data including prosody information.
[0795] During specific implementation, step 2 may also adopt other speech conversion models, and the method does not limit the specific implementation of the speech conversion model.
[0796] Fifty-fourth embodiment
[0797] In the above embodiment, a cross-language speech generation method is provided. Correspondingly, this application also provides a cross-language speech generation device. This device corresponds to the embodiment of the above method. The parts of this embodiment that are the same as the first embodiment are not repeated here. Please refer to the corresponding parts in the first embodiment.
[0798] The present application provides a cross-language speech generation device comprising:
[0799] a speech generating unit configured to determine, in response to a speech generation request sent by a client for a first user to read a text in a first language, second speech data of a second user reading the text;
[0800] The voice conversion unit is configured to convert the second voice data into first voice data of the first user reading the text through a voice conversion model of the first user.
[0801] In one example, the apparatus further comprises:
[0802] A model building unit, used to build a speech posterior probability graph (PPG) feature extractor and a speech synthesis model for each user;
[0803] The speech conversion unit includes a PPG feature extraction subunit and a second speech generation subunit;
[0804] A PPG feature extraction subunit is configured to determine, through the PPG feature extractor, PPG feature data of the second voice data based on the first acoustic feature data of the second voice data including the second user voiceprint information and voice content information;
[0805] The second speech generation subunit is used to generate first speech data of the first user reading the text according to the PPG feature data and the second acoustic feature data of the second speech data including prosody information through the speech synthesis model of the first user.
[0806] Fifty-fifth embodiment
[0807] This application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple. For relevant details, please refer to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0808] An electronic device of this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a cross-language speech generation method. After the device is powered on and the program of the method is run by the processor, the following steps are performed: in response to a speech generation request for a first user to read a text in a first language, determining second speech data of a second user reading the text; and converting the second speech data into first speech data of the first user reading the text through a speech conversion model of the first user.
[0809] Fifty-sixth embodiment
[0810] In the above-mentioned embodiment, a cross-language speech generation system is provided. Correspondingly, this application also provides a cross-language speech generation method, which can be performed by a terminal device, etc. This method corresponds to the embodiment of the above-mentioned system. The parts of this embodiment that are identical to the fifty-second embodiment are not repeated here. Please refer to the corresponding parts of the fifty-second embodiment.
[0811] The present application provides a cross-language speech generation method, which may include the following steps:
[0812] Step 1: Identify the text in the first language;
[0813] Step 2: Sending a request for generating a voice for the first user to read the text to the server.
[0814] Fifty-seventh embodiment
[0815] In the above embodiment, a cross-language speech generation method is provided. Correspondingly, this application also provides a cross-language speech generation device. This device corresponds to the embodiment of the above method. The parts of this embodiment that are the same as the first embodiment are not repeated here. Please refer to the corresponding parts in the first embodiment.
[0816] The present application provides a cross-language speech generation device comprising:
[0817] a text determination unit for determining a text in a first language;
[0818] The request sending unit is used to send a speech generation request for the first user to read the text aloud to the server.
[0819] Fifty-eighth embodiment
[0820] This application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple. For relevant details, please refer to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0821] An electronic device of this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a cross-language speech generation method. After the device is powered on and the program of the method is run by the processor, the device performs the following steps: determining text in a first language; and sending a speech generation request to a server for a first user to read the text aloud.
[0822] Fifty-ninth embodiment
[0823] In the above-mentioned embodiment, a cross-language speech generation system is provided. Correspondingly, this application also provides a cross-language speech generation method, which can be performed by a terminal device, etc. This method corresponds to the embodiment of the above-mentioned system. The parts of this embodiment that are identical to the fifty-second embodiment are not repeated here. Please refer to the corresponding parts of the fifty-second embodiment.
[0824] The present application provides a cross-language speech generation method, which may include the following steps:
[0825] Step 1: Identify the text in the first language;
[0826] Step 2: Determine second voice data of a second user reading the text;
[0827] Step 3: Using the first user's voice conversion model, convert the second voice data into first voice data of the first user reading the text.
[0828] In one example, step 3 may include the following sub-steps: 3.1) determining the PPG feature data of the second voice data based on the first acoustic feature data of the second voice data including the voiceprint information of the second user and the voice content information through a speech posterior probability graph PPG feature extractor; 3.2) generating the first voice data of the first user reading the text based on the PPG feature data and the second acoustic feature data of the second voice data including prosody information through a speech synthesis model of the first user.
[0829] Sixtieth embodiment
[0830] In the above embodiment, a cross-language speech generation method is provided. Correspondingly, this application also provides a cross-language speech generation device. This device corresponds to the embodiment of the above method. The parts of this embodiment that are the same as the first embodiment are not repeated here. Please refer to the corresponding parts in the first embodiment.
[0831] The present application provides a cross-language speech generation device comprising:
[0832] a text determination unit for determining a text in a first language;
[0833] a voice determination unit, configured to determine second voice data of a second user reading the text;
[0834] The voice conversion unit is configured to convert the second voice data into first voice data of the first user reading the text through a voice conversion model of the first user.
[0835] Sixty-first embodiment
[0836] This application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple. For relevant details, please refer to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0837] An electronic device of this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a cross-language speech generation method. After the device is powered on and the program of the method is run through the processor, the device performs the following steps: determining text in a first language; determining second speech data of a second user reading the text aloud; and converting the second speech data into first speech data of the first user reading the text aloud through a speech conversion model of the first user.
[0838] Sixty-second embodiment
[0839] In the above embodiment, a speech conversion method is provided. Accordingly, this application also provides a cross-dialect speech generation system. This system corresponds to the embodiment of the above method. The server-side processing in this embodiment, which is identical to that in the second embodiment, will not be repeated here. Please refer to the corresponding section in the second embodiment.
[0840] The present application provides a cross-dialect speech generation system including: a terminal device and a server.
[0841] Among them, the terminal device is used to determine the target text and target dialect, and send a request to the server for the first user to read the text in the target dialect; and play the first voice data of the first user reading the text in the target dialect sent back by the server; the server is used to construct a speech posterior probability graph PPG feature extractor and a speech synthesis model for each user; and, in response to the request, determine the second voice data of the second user reading the text in the target dialect, and determine the PPG feature data of the second voice data through the PPG feature extractor based on the first acoustic feature data of the second voice data including the second user's voiceprint information and voice content information; and generate the first voice data through the speech synthesis model of the first user (first speech synthesis model) based on the PPG feature data and the second acoustic feature data of the second voice data including prosody information.
[0842] The first and second users may be native speakers of the same language, but may be from different regions and speak different dialects. For example, if the first user wants to read a text in Cantonese, but the user does not speak Cantonese or speaks Cantonese poorly, in order to achieve the effect of the first user speaking Cantonese fluently, a Cantonese speaker can be designated as the second user. The second user can first speak the text in fluent Cantonese as the second voice data, and then the first voice data with the first user's timbre is generated corresponding to the second voice data, as if the first user has a good Cantonese proficiency.
[0843] In one example, the server uses the text as input data for a second speech synthesis model of the second user. The model can generate speech data of the text in the second user's reading style, with the speech data having not only the second user's timbre but also characteristics such as the second user's reading intonation. The second speech synthesis model can be learned from a speech library of the second user.
[0844] It should be noted that the speech synthesis model of the second user is different from the first speech synthesis model of each user constructed by the server. The differences include:
[0845] 1) Different input data: The input data of the first speech synthesis model includes PPS feature data, while the input data of the second user's speech synthesis model includes text;
[0846] 2) The models have different functions: the first speech synthesis model generates speech data with the corresponding user's timbre based on PPS feature data, while the second speech synthesis model generates corresponding speech data with the corresponding user's timbre and reading style based on the input text;
[0847] 3) Different output data: The speech data output by the first speech synthesis model has nothing to do with the speaking style of the user corresponding to the model, but is related to the speaking style of other users; the speech data output by the second speech synthesis model has nothing to do with the speech data of other users.
[0848] During specific implementation, the first user may specify the second user through the terminal device, or the server may arbitrarily determine a user of the target dialect as the second user.
[0849] In one example, a first user specifies a second user through the terminal device used by the first user, and the server sends a target text to the terminal device of the designated second user, and collects second voice data of the target text read aloud by the second user in the target dialect through the terminal device of the second user. The server determines the PPG feature data of the second voice data through the PPG feature extractor based on the first acoustic feature data of the second voice data including the voiceprint information and voice content information of the second user; and generates the first voice data through the speech synthesis model of the first user based on the PPG feature data and the second acoustic feature data of the second voice data including prosody information.
[0850] As can be seen from the above embodiments, the cross-dialect speech generation system provided by the embodiments of the present application determines the target text and target dialect through a terminal device, and sends a request to the server for a first user to read the text in the target dialect; and plays the first speech data of the first user reading the text in the target dialect sent back by the server; the server constructs a speech posterior probability graph (PPG) feature extractor and a speech synthesis model for each user; and, in response to the request, determines second speech data of a second user reading the text in the target dialect, and determines PPG feature data of the second speech data based on first acoustic feature data including the second user's voiceprint information and speech content information of the second speech data through the PPG feature extractor; and generates the first speech data based on the PPG feature data and second acoustic feature data including prosody information of the second speech data through the speech synthesis model (first speech synthesis model) of the first user. This processing method enables the generation of speech data of the first user reading the text in the target dialect even if the first user does not have good dialect listening and speaking skills, thereby realizing cross-dialect text reading.
[0851] Sixty-third embodiment
[0852] In the above-mentioned embodiment, a cross-dialect speech generation system is provided. Correspondingly, this application also provides a cross-dialect speech generation method, which can be executed by a server, etc. This method corresponds to the embodiment of the above-mentioned system. The parts of this embodiment that are identical to the sixty-second embodiment are not repeated here. Please refer to the corresponding parts of the sixty-second embodiment.
[0853] The present application provides a method for generating cross-dialect speech, which may include the following steps:
[0854] Step 1: In response to a request by a first user to read a target text in a target dialect, determining second speech data of a second user reading a target text in the target dialect;
[0855] Step 2: Using the first user's speech conversion model, convert the second speech data into first speech data of the first user reading the text in the target dialect.
[0856] In one example, the method may further include the following steps: constructing a speech posterior probability graph PPG feature extractor and a speech synthesis model for each user; accordingly, step 2 may include the following sub-steps: determining the PPG feature data of the second speech data through the PPG feature extractor based on the first acoustic feature data of the second speech data including the voiceprint information of the second user and the speech content information; generating the first speech data through the speech synthesis model of the first user based on the PPG feature data and the second acoustic feature data of the second speech data including prosody information.
[0857] During specific implementation, step 2 may also adopt other speech conversion models, and the method does not limit the specific implementation of the speech conversion model.
[0858] Sixty-fourth embodiment
[0859] In the above embodiment, a method for generating cross-dialect speech is provided. Accordingly, this application also provides a device for generating cross-dialect speech. This device corresponds to the embodiment of the above method. The parts of this embodiment that are identical to the first embodiment are not repeated here. Please refer to the corresponding parts in the first embodiment.
[0860] The present application provides a cross-dialect speech generation device comprising:
[0861] a speech generating unit for determining second speech data for a second user to read aloud the target text in the target dialect in response to a request by a first user to read aloud the target text in the target dialect;
[0862] The speech conversion unit is configured to convert the second speech data into first speech data of the first user reading the text in the target dialect through a speech conversion model of the first user.
[0863] In one example, the apparatus further comprises:
[0864] A model building unit, used to build a speech posterior probability graph (PPG) feature extractor and a speech synthesis model for each user;
[0865] The speech conversion unit includes a PPG feature extraction subunit and a second speech generation subunit;
[0866] A PPG feature extraction subunit is configured to determine, through the PPG feature extractor, PPG feature data of the second voice data based on the first acoustic feature data of the second voice data including the second user voiceprint information and voice content information;
[0867] The second speech generation subunit is used to generate the first speech data according to the PPG feature data and the second acoustic feature data of the second speech data including prosody information through the speech synthesis model of the first user.
[0868] Sixty-fifth embodiment
[0869] This application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple. For relevant details, please refer to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0870] An electronic device of this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a cross-dialect speech generation method. After the device is powered on and the program of the method is run by the processor, the following steps are performed: in response to a request by a first user to read a target text in a target dialect, second speech data of a second user reading a target text in the target dialect is determined; and the second speech data is converted into first speech data of the first user reading the text in the target dialect through a speech conversion model of the first user.
[0871] Sixty-sixth embodiment
[0872] In the above embodiment, a cross-dialect speech generation system is provided. Correspondingly, this application also provides a cross-language speech generation method, which can be executed by a terminal device, etc. This method corresponds to the embodiment of the above system. The parts of this embodiment that are identical to the sixty-second embodiment are not repeated here. Please refer to the corresponding parts of the sixty-second embodiment.
[0873] The present application provides a method for generating cross-dialect speech, which may include the following steps:
[0874] Step 1: Determine the target text and target dialect;
[0875] Step 2: Sending first voice data of the text read aloud by the first user in the target dialect to the server.
[0876] Sixty-seventh embodiment
[0877] In the above embodiment, a method for generating cross-dialect speech is provided. Accordingly, this application also provides a device for generating cross-dialect speech. This device corresponds to the embodiment of the above method. The parts of this embodiment that are identical to the first embodiment are not repeated here. Please refer to the corresponding parts in the first embodiment.
[0878] The present application provides a cross-dialect speech generation device comprising:
[0879] a text determination unit for determining a target text and a target dialect;
[0880] The request sending unit is used to send first voice data of the text read aloud by the first user in the target dialect to the server.
[0881] Sixty-eighth embodiment
[0882] This application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple. For relevant details, please refer to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0883] An electronic device of this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a cross-dialect speech generation method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: determining a target text and a target dialect; and sending first speech data of the text read aloud by a first user in the target dialect to a server.
[0884] Sixty-ninth embodiment
[0885] In the above-mentioned embodiment, a cross-dialect speech generation system is provided. Correspondingly, this application also provides a cross-dialect speech generation method, which can be performed by a terminal device, etc. This method corresponds to the embodiment of the above-mentioned system. The parts of this embodiment that are identical to the sixty-second embodiment are not repeated here. Please refer to the corresponding parts of the sixty-second embodiment.
[0886] The present application provides a method for generating cross-dialect speech, which may include the following steps:
[0887] Step 1: Determine the target text and target dialect;
[0888] Step 2: determining second voice data of a second user reading the target text in the target dialect;
[0889] Step 3: Using the first user's speech conversion model, convert the second speech data into first speech data of the first user reading the text in the target dialect.
[0890] In one example, step 3 may include the following sub-steps: 3.1) determining the PPG feature data of the second voice data based on the first acoustic feature data of the second voice data including the voiceprint information of the second user and the voice content information through a speech posterior probability graph PPG feature extractor; 3.2) generating the first voice data based on the PPG feature data and the second acoustic feature data of the second voice data including prosody information through a speech synthesis model of the first user.
[0891] Seventieth embodiment
[0892] In the above embodiment, a method for generating cross-dialect speech is provided. Accordingly, this application also provides a device for generating cross-dialect speech. This device corresponds to the embodiment of the above method. The parts of this embodiment that are identical to the first embodiment are not repeated here. Please refer to the corresponding parts in the first embodiment.
[0893] The present application provides a cross-dialect speech generation device comprising:
[0894] a text determination unit for determining a target text and a target dialect;
[0895] a speech determination unit, configured to determine second speech data of a second user reading aloud the target text in a target dialect;
[0896] The speech conversion unit is configured to convert the second speech data into first speech data of the first user reading the text in the target dialect through a speech conversion model of the first user.
[0897] Seventy-first embodiment
[0898] This application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple. For relevant details, please refer to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0899] An electronic device of this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a cross-dialect speech generation method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: determining a target text and a target dialect; determining second speech data of a second user reading the target text in the target dialect; and converting the second speech data into first speech data of the first user reading the text in the target dialect through a speech conversion model of the first user.
[0900] Seventy-second embodiment
[0901] In the above embodiment, a voice conversion method is provided. Accordingly, this application also provides a voice conversion system. This system corresponds to the embodiment of the above method. The server-side processing in this embodiment is not repeated for the parts that are identical to those in the second embodiment. Please refer to the corresponding parts in the second embodiment.
[0902] The present application provides a voice changing system comprising: a terminal device and a server.
[0903] Among them, the terminal device is used to determine the first voice data of the first user, and send a request to the server to change the first voice data into the voice of the second user; the server builds a voice synthesis model for each user; and builds a voice posterior probability graph PPG feature extractor; through the PPG feature extractor, the PPG feature data of the first voice data is determined according to the first acoustic feature data of the first voice data including the voiceprint information of the first user and the voice content information; through the voice synthesis model of the second user, the second voice data of the second user corresponding to the first voice data is generated according to the PPG feature data and the second acoustic feature data of the first voice data including rhythm information.
[0904] The first and second users may be native speakers of the same language or dialect. For example, if the second user wants to imitate a certain speech segment of the first user, the second user can use the speech segment as the first speech data. The method can then generate second speech data corresponding to the speech segment with the second user's timbre.
[0905] As can be seen from the above embodiment, the voice changing system provided by the embodiment of the present application determines the first voice data of the first user through the terminal device, and sends a request to the server to change the first voice data into the voice of the second user; the server builds a voice synthesis model for each user; and builds a voice posterior probability graph PPG feature extractor; through the PPG feature extractor, based on the first acoustic feature data of the first voice data including the first user's voiceprint information and voice content information, determines the PPG feature data of the first voice data; through the voice synthesis model of the second user, based on the PPG feature data and the second acoustic feature data of the first voice data including prosody information, generates the second voice data of the second user corresponding to the first voice data; this processing method enables the use of PPG features as the input of the voice synthesis module. The PPG features retain acoustic information (such as prosody and pronunciation information), and the PPG features of each frame can be used between different speakers, thereby achieving the voice change of the first user's voice into the second user's voice; therefore, the accuracy of the voice change can be effectively improved. At the same time, since the system does not need to recognize the text information of the first voice data, the efficiency of the voice change can be effectively improved.
[0906] Seventy-third embodiment
[0907] In the above-mentioned embodiment, a voice changing system is provided. Accordingly, this application also provides a voice changing method, which may be performed by a server, for example. This method corresponds to the embodiment of the above-mentioned system. The parts of this embodiment that are identical to those in the seventy-second embodiment are not repeated here. Please refer to the corresponding parts in the seventy-second embodiment.
[0908] The present application provides a method for voice modification, which may include the following steps:
[0909] Step 1: Build a speech posterior probability graph (PPG) feature extractor and a speech synthesis model for each user.
[0910] Step 2: In response to a request to change the first voice data of the first user into the voice of the second user, the PPG feature extractor determines, based on first acoustic feature data of the first voice data including the first user's voiceprint information and voice content information, PPG feature data of the first voice data;
[0911] Step 3: Generate second speech data of the second user corresponding to the first speech data based on the PPG feature data and second acoustic feature data of the first speech data including prosody information through the speech synthesis model of the second user.
[0912] Seventy-fourth embodiment
[0913] In the above embodiment, a method for voice modification is provided. Accordingly, the present application also provides a device for voice modification. This device corresponds to the embodiment of the above method. The parts of this embodiment that are identical to the first embodiment are not repeated here. Please refer to the corresponding parts in the first embodiment.
[0914] The present application provides a voice changing device comprising:
[0915] A model building unit, used to build a speech posterior probability graph (PPG) feature extractor and a speech synthesis model for each user;
[0916] a feature extraction unit configured to determine, by the PPG feature extractor, PPG feature data of the first voice data based on first acoustic feature data of the first voice data including first user voiceprint information and voice content information, in response to a request to morph first voice data of a first user into a second user's voice;
[0917] A speech generation unit is configured to generate second speech data of a second user corresponding to the first speech data based on the PPG feature data and second acoustic feature data including prosody information of the first speech data by using the speech synthesis model of the second user.
[0918] Seventy-fifth embodiment
[0919] This application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple. For relevant details, please refer to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0920] An electronic device according to the present embodiment includes: a processor and a memory; the memory is used to store a program for implementing a voice changing method. After the device is powered on and the program of the method is run by the processor, the following steps are performed: a speech posterior probability graph (PPG) feature extractor and a speech synthesis model for each user are constructed; in response to a request to change first speech data of a first user into a second user's voice, PPG feature data of the first speech data is determined by the PPG feature extractor based on first acoustic feature data of the first speech data including voiceprint information and speech content information of the first user; and second speech data of the second user corresponding to the first speech data is generated by the speech synthesis model of the second user based on the PPG feature data and second acoustic feature data of the first speech data including prosody information.
[0921] Seventy-sixth embodiment
[0922] In the above-mentioned embodiment, a voice changing system is provided. Accordingly, this application also provides a voice changing method, which can be performed by a terminal device, etc. This method corresponds to the embodiment of the above-mentioned system. The parts of this embodiment that are identical to the seventy-second embodiment are not repeated here. Please refer to the corresponding parts of the seventy-second embodiment.
[0923] The present application provides a method for voice modification, which may include the following steps:
[0924] Step 1: Determine first voice data of a first user;
[0925] Step 2: Send a request to the server to change the first voice data into the second user's voice.
[0926] Seventy-seventh embodiment
[0927] In the above embodiment, a method for voice modification is provided. Accordingly, the present application also provides a device for voice modification. This device corresponds to the embodiment of the above method. The parts of this embodiment that are identical to the first embodiment are not repeated here. Please refer to the corresponding parts in the first embodiment.
[0928] The present application provides a voice changing device comprising:
[0929] a voice determination unit, configured to determine first voice data of a first user;
[0930] The request sending unit is used to send a request to the server to change the first voice data into the second user's voice.
[0931] Seventy-eighth embodiment
[0932] This application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple. For relevant details, please refer to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0933] An electronic device of this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a voice changing method. After the device is powered on and the program of the method is run by the processor, the device performs the following steps: determining first voice data of a first user; and sending a request to a server to change the first voice data into the voice of a second user.
[0934] Seventy-ninth embodiment
[0935] In the above-mentioned embodiment, a voice changing system is provided. Accordingly, this application also provides a voice changing method, which can be performed by a terminal device, etc. This method corresponds to the embodiment of the above-mentioned system. The parts of this embodiment that are identical to the seventy-second embodiment are not repeated here. Please refer to the corresponding parts of the seventy-second embodiment.
[0936] The present application provides a method for voice modification, which may include the following steps:
[0937] Step 1: Determine first voice data of a first user;
[0938] Step 2: Determine PPG feature data of the first speech data based on the first acoustic feature data of the first speech data including the first user's voiceprint information and speech content information through a speech posterior probability graph (PPG) feature extractor;
[0939] Step 3: Generate second voice data of the second user corresponding to the first voice data based on the PPG feature data and second acoustic feature data of the first voice data including prosody information through the second user's voice synthesis model.
[0940] The eightieth embodiment
[0941] In the above embodiment, a method for voice modification is provided. Accordingly, the present application also provides a device for voice modification. This device corresponds to the embodiment of the above method. The parts of this embodiment that are identical to the first embodiment are not repeated here. Please refer to the corresponding parts in the first embodiment.
[0942] The present application provides a voice changing device comprising:
[0943] a voice determination unit, configured to determine first voice data of a first user;
[0944] a feature extraction unit, configured to determine, by a posterior probability graph (PPG) feature extractor, PPG feature data of the first speech data based on first acoustic feature data of the first speech data including first user voiceprint information and speech content information;
[0945] The speech generation unit is configured to generate second speech data of the second user corresponding to the first speech data based on the PPG feature data and second acoustic feature data including prosody information of the first speech data by using a speech synthesis model of the second user.
[0946] Eighty-first embodiment
[0947] This application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple. For relevant details, please refer to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0948] An electronic device according to this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a voice changing method. After the device is powered on and the program of the method is run by the processor, the following steps are performed: determining first voice data of a first user; using a posterior probability graph (PPG) feature extractor, based on first acoustic feature data of the first voice data including voiceprint information of the first user and voice content information, determining PPG feature data of the first voice data; and using a voice synthesis model of the second user, based on the PPG feature data and second acoustic feature data of the first voice data including prosody information, generating second voice data of the second user corresponding to the first voice data.
[0949] Eighty-second embodiment
[0950] In the above embodiment, a speech conversion method is provided. Accordingly, this application also provides a cross-language speech interaction system. This system corresponds to the embodiment of the above method. The server-side processing in this embodiment, which is identical to that in the second embodiment, will not be repeated here. Please refer to the corresponding section in the second embodiment.
[0951] The present application provides a cross-language voice interaction system including: a first terminal device, a second terminal device, and a server.
[0952] The first terminal device and the second terminal device include but are not limited to mobile communication devices, namely: commonly known as mobile phones or smart phones, and also include personal computers, PADs, iPads and other terminal devices.
[0953] Among them, the first terminal device is used to collect first voice data in the source language of the first user and send the first voice data to the server; the server is used to determine the target language text corresponding to the first voice data; based on the target language voice library of the second user, generate second voice data in the target language of the second user corresponding to the target language text; through the voice conversion model of the first user, generate third voice data in the target language with the timbre of the first user corresponding to the second voice data; send the third voice data to the second terminal device; the second terminal device is used to play the third voice data.
[0954] For example, user A (first user) whose native language is Chinese uses a first terminal device to have a voice call with user B (third user) whose native language is English and uses a second terminal device. The two users can chat directly with each other in their respective native languages, and the server performs voice translation and other processing.
[0955] In one example, the server is specifically configured to determine a source language text of the first voice data through a voice recognition algorithm; and determine a target language text corresponding to the source language text through a voice translation algorithm.
[0956] In one example, the server is specifically used to generate a target language speech synthesis model for the second user based on the target language speech library of the second user; and generate second speech data in the target language of the second user corresponding to the target language text through the target language speech synthesis model of the second user.
[0957] The target language speech synthesis model for the second user can generate second speech data having the second user's timbre and user reading style based on the input target language text. The target language speech synthesis model for the second user can be learned from a target language speech database for the second user. The input data for the model can be the target language text, and the output data can be speech data of the target language text in the second user's reading style.
[0958] In one example, the server is further configured to construct a speech posterior probability graph (PPG) feature extractor and a speech synthesis model for the first user. Specifically, the server is configured to determine PPG feature data for the second speech data using the PPG feature extractor based on first acoustic feature data of the second speech data, including the second user's voiceprint information and speech content information. Furthermore, the server is configured to generate, using the first user's speech synthesis model, third speech data corresponding to the second speech data and having the first user's timbre, representing a target language text, based on the PPG feature data and second acoustic feature data of the second speech data, including prosody information. Specifically, the server may employ other speech conversion models, and the system does not limit the specific implementation of the speech conversion model.
[0959] It should be noted that the target language speech synthesis model of the second user is different from the speech synthesis model of the first user. The differences include:
[0960] 1) Different input data: The input data of the first user's speech synthesis model includes PPS feature data, while the input data of the second user's target language speech synthesis model includes text;
[0961] 2) The models have different functions: the speech synthesis model for the first user generates speech data with the first user's timbre based on PPS feature data, while the target language speech synthesis model for the second user generates speech data with the second user's timbre and reading style based on input text;
[0962] 3) Different output data: The speech data output by the speech synthesis model of the first user may be independent of the first user's speaking style, but related to the second user's speaking style; the speech data output by the target language speech synthesis model of the second user is independent of the speech data of other users.
[0963] In specific implementation, the first user can specify a second user, such as specifying a favorite English announcer, and use the announcer's broadcasting method to read the translation; or the server can determine a second user and provide the target language text reading service to multiple first users through the second user's broadcasting method.
[0964] In one example, the second terminal device is also used to collect fourth voice data in the target language of the third user; send the fourth voice data to the server; accordingly, the server is also used to determine the source language text corresponding to the fourth voice data for the fourth voice data in the target language of the third user sent by the second terminal device; generate fifth voice data in the source language of the fourth user corresponding to the source language text based on the source language voice library of the fourth user; generate sixth voice data in the source language with the timbre of the third user corresponding to the fifth voice data through the voice conversion model of the third user, and send the sixth voice data to the first terminal device; accordingly; the first terminal device is also used to receive the sixth voice data in the source language of the third user sent by the server; the server generates the sixth voice data by the following steps: determine the source language text corresponding to the fourth voice data for the fourth voice data in the target language of the third user sent by the second terminal device; generate fifth voice data in the source language of the fourth user corresponding to the source language text based on the source language voice library of the fourth user; generate sixth voice data in the source language with the timbre of the third user corresponding to the fifth voice data through the voice conversion model of the third user; and play the sixth voice data.
[0965] Since the processing process of the second terminal device sending voice data to the first terminal device is the same as the processing process of the first terminal device sending voice data to the second terminal device, the specific implementation method will not be repeated here. Please refer to the above-mentioned processing process of the first terminal device sending voice data to the second terminal device.
[0966] As can be seen from the above embodiments, the cross-language voice interaction system provided by the embodiments of the present application collects first voice data of the source language of the first user through a first terminal device, and sends the first voice data to the server; the server determines the target language text corresponding to the first voice data; generates second voice data of the target language of the second user corresponding to the target language text based on the target language voice library of the second user; generates third voice data of the target language with the timbre of the first user corresponding to the second voice data through the voice conversion model of the first user; sends the third voice data to the second terminal device; the second terminal device is used to play the third voice data; this processing method allows users of different languages to directly use their respective native languages for voice interaction with each other; therefore, the efficiency of cross-language voice interaction can be effectively improved.
[0967] Eighty-third embodiment
[0968] In the above embodiment, a cross-language voice interaction system is provided. Correspondingly, this application also provides a cross-language voice interaction method, which can be executed by a server, etc. This method corresponds to the embodiment of the above system. The parts of this embodiment that are identical to the 82nd embodiment are not repeated here. Please refer to the corresponding parts of the 82nd embodiment.
[0969] The present application provides a cross-language voice interaction method, which may include the following steps:
[0970] Step 1: determining a target language text corresponding to first voice data in a source language of a first user sent by a first terminal device;
[0971] Step 2: generating second speech data in the target language of the second user corresponding to the target language text based on the target language speech database of the second user;
[0972] Step 3: Generate third voice data in the target language with the first user's timbre corresponding to the second voice data through the first user's voice conversion model, and send the third voice data to the second terminal device.
[0973] In one example, the method also includes: determining the source language text corresponding to the fourth voice data in the target language of the third user sent by the second terminal device; generating fifth voice data in the source language of the fourth user corresponding to the source language text based on the source language voice library of the fourth user; generating sixth voice data in the source language with the timbre of the third user corresponding to the fifth voice data through the voice conversion model of the third user, and sending the sixth voice data to the first terminal device.
[0974] Eighty-fourth embodiment
[0975] In the above embodiment, a cross-language voice interaction method is provided. Correspondingly, this application also provides a cross-language voice interaction device. This device corresponds to the embodiment of the above method. The parts of this embodiment that are the same as the first embodiment are not repeated here. Please refer to the corresponding parts in the first embodiment.
[0976] The present application provides a cross-language voice interaction device comprising:
[0977] A target language text determination unit, configured to determine a target language text corresponding to first voice data in a source language of a first user sent by a first terminal device;
[0978] a second speech data generating unit, configured to generate second speech data in the target language of the second user corresponding to the target language text based on the target language speech library of the second user;
[0979] The third voice data generating unit is configured to generate third voice data in a target language corresponding to the second voice data and having the first user's timbre through the first user's voice conversion model, and send the third voice data to the second terminal device.
[0980] Eighty-fifth embodiment
[0981] This application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple. For relevant details, please refer to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0982] An electronic device of this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a cross-language voice interaction method. After the device is powered on and the program of the method is run by the processor, the following steps are performed: determining a target language text corresponding to first voice data in a source language of a first user sent by a first terminal device; generating second voice data in the target language of the second user corresponding to the target language text based on a target language voice library of the second user; generating third voice data in the target language with the timbre of the first user corresponding to the second voice data through a voice conversion model of the first user, and sending the third voice data to the second terminal device.
[0983] Eighty-sixth embodiment
[0984] In the above embodiment, a cross-language voice interaction system is provided. Correspondingly, this application also provides a cross-language voice interaction method, which can be performed by a first terminal device, etc. This method corresponds to the embodiment of the above system. The parts of this embodiment that are identical to the 45th embodiment are not repeated here. Please refer to the corresponding parts in Example 82.
[0985] The present application provides a cross-language voice interaction method, which may include the following steps:
[0986] Step 1: Collect first speech data of a first user in a source language;
[0987] Step 2: Send the first voice data to the server so that the server determines the target language text corresponding to the first voice data; generate second voice data in the target language of the second user corresponding to the target language text based on the target language voice library of the second user; generate third voice data in the target language with the timbre of the first user corresponding to the second voice data through the voice conversion model of the first user; send the third voice data to the second terminal device; the second terminal device plays the third voice data.
[0988] In one example, the method also includes: receiving sixth voice data in the source language of the third user sent by the server; the server generates the sixth voice data by the following steps: determining the source language text corresponding to the fourth voice data in the target language of the third user sent by the second terminal device; generating fifth voice data in the source language of the fourth user corresponding to the source language text based on the source language voice library of the fourth user; generating sixth voice data in the source language with the timbre of the third user corresponding to the fifth voice data through the voice conversion model of the third user; and playing the sixth voice data.
[0989] Eighty-seventh embodiment
[0990] In the above embodiment, a cross-language voice interaction method is provided. Correspondingly, this application also provides a cross-language voice interaction device. This device corresponds to the embodiment of the above method. The parts of this embodiment that are the same as the first embodiment are not repeated here. Please refer to the corresponding parts in the first embodiment.
[0991] The present application provides a cross-language voice interaction device comprising:
[0992] A voice collection unit, configured to collect first voice data in a source language of a first user;
[0993] A request sending unit is configured to send the first voice data to a server so that the server determines a target language text corresponding to the first voice data; generate second voice data in the target language of the second user corresponding to the target language text based on a target language voice library of the second user; generate third voice data in the target language with the timbre of the first user corresponding to the second voice data through a voice conversion model of the first user; send the third voice data to a second terminal device; and the second terminal device plays the third voice data.
[0994] Eighty-eighth embodiment
[0995] This application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple. For relevant details, please refer to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0996] An electronic device of this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a cross-language voice interaction method. After the device is powered on and the program of the method is run by the processor, the following steps are performed: collecting first voice data in a source language of a first user; sending the first voice data to a server so that the server determines a target language text corresponding to the first voice data; generating second voice data in the target language of the second user corresponding to the target language text based on a target language voice library of the second user; generating third voice data in the target language with the timbre of the first user corresponding to the second voice data through a voice conversion model of the first user; sending the third voice data to a second terminal device; and the second terminal device playing the third voice data.
[0997] Eighty-ninth embodiment
[0998] In the above embodiment, a cross-language voice interaction system is provided. Correspondingly, this application also provides a cross-language voice interaction method, which can be performed by a second terminal device, etc. This method corresponds to the embodiment of the above system. The parts of this embodiment that are identical to the 82nd embodiment are not repeated here. Please refer to the corresponding parts of the 82nd embodiment.
[0999] The present application provides a cross-language voice interaction method, which may include the following steps:
[1000] Step 1: Receive third voice data in the target language of the first user sent by the server;
[1001] Step 2: Play the third voice data.
[1002] In one example, the method also includes: collecting fourth voice data in the target language of the third user; sending the fourth voice data to the server so that the server determines the source language text corresponding to the fourth voice data; generating fifth voice data in the source language of the fourth user corresponding to the source language text based on the source language voice library of the fourth user; generating sixth voice data in the source language with the timbre of the third user corresponding to the fifth voice data through the voice conversion model of the third user, and sending the sixth voice data to the first terminal device; the first terminal device plays the sixth voice data.
[1003] Ninetieth embodiment
[1004] In the above embodiment, a cross-language voice interaction method is provided. Correspondingly, this application also provides a cross-language voice interaction device. This device corresponds to the embodiment of the above method. The parts of this embodiment that are the same as the first embodiment are not repeated here. Please refer to the corresponding parts in the first embodiment.
[1005] The present application provides a cross-language voice interaction device comprising:
[1006] a voice receiving unit, configured to receive third voice data in a target language of a first user sent by a server;
[1007] The voice playing unit is used to play the third voice data.
[1008] Ninety-first embodiment
[1009] This application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple. For relevant details, please refer to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[1010] An electronic device of this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a cross-language voice interaction method. After the device is powered on and the program of the method is run through the processor, the device performs the following steps: receiving third voice data in a target language of a first user sent by a server; and playing the third voice data.
[1011] Ninety-second embodiment
[1012] In the above embodiment, a speech conversion method is provided. Accordingly, this application also provides a question-and-answer system. This system corresponds to the embodiment of the above method. The server-side processing in this embodiment, which is identical to that in the second embodiment, will not be repeated here. Please refer to the corresponding portion in the second embodiment.
[1013] The present application provides a question-answering system including a terminal device and a server.
[1014] The terminal device includes but is not limited to mobile communication devices, namely: commonly known as mobile phones or smart phones, and also includes personal computers, PADs, iPads and other terminal devices.
[1015] Among them, the terminal device is used to collect the first voice data of the first user, determine the target speaking method, convert the first voice data into the second voice data of the target speaking method, and send the second voice data to the server; the server is used to determine the reply information based on the second voice data, and send the reply information back to the terminal device. The terminal device can display the reply information for the user to view.
[1016] In one example, the server is specifically used to determine the speaking style information and text information of the second voice data; and determine the reply information based on the speaking style information and text information. Accordingly, the server can determine the reply information based on the speaking style information and text information of the second voice data. By adopting this processing method, the terminal device sends the voice data of the speaking style specified by the user to the server, rather than the actual speaking style of the user collected. The user can conceal his or her true emotions by specifying the speaking style. When determining the reply information, the server must consider not only the user's speaking content, but also the user's speaking style. In this way, even if the speaking content is the same, if the speaking style is different, the server's reply information will be different; therefore, the accuracy of the question and answer reply information can be effectively improved, thereby improving the user experience.
[1017] For example, a user uses a terminal device to consult a customer service robot deployed on the server about a certain issue. Since the user is in a low mood at the time, his or her natural speaking tone is relatively low. Therefore, the user can specify the speaking style of the consultation voice, such as happy, fluent, clear, powerful, etc. Accordingly, the server can determine the reply information A; if the user directly sends the same voice content in his or her natural speaking tone to the server, the server may determine the reply information B.
[1018] In one example, the server stores a set of correspondences between inquiry information, speaking patterns, and reply information; and determines reply information corresponding to the speaking pattern information and the text information based on the set of correspondences.
[1019] In one example, the terminal device determines the text sequence of the first voice data through a voice recognition algorithm, and generates second voice data with the target speaking style corresponding to the first voice data through a voice synthesis model of the first user's target speaking style.
[1020] The nineteenth embodiment will not be described in detail here.
[1021] The reply message may be a text message or a voice message.
[1022] It can be seen from the above embodiments that the voice interaction system provided in the embodiments of the present application uses a terminal device to collect first voice data of a first user, determine a target speaking method, convert the first voice data into second voice data of the target speaking method, and send the second voice data to the server; the server is used to determine reply information based on the second voice data, and send the reply information back to the terminal device; this processing method enables the terminal device to send to the server voice data of the speaking method specified by the user, rather than the actual collected voice data of the user's real speaking method. The user can conceal his or her true emotions by specifying the speaking method, and the server can determine the reply information based on the voice data after the speaking method is changed. In this way, even for the same voice content, if the speaking method is different, the server reply information will be different; therefore, the accuracy of the question and answer reply information can be effectively improved, thereby improving the user experience.
[1023] Ninety-third embodiment
[1024] In the above embodiment, a question-and-answer system is provided. Correspondingly, this application also provides a question-and-answer method, which can be executed by a server, for example. This method corresponds to the embodiment of the above system. The parts of this embodiment that are identical to the ninety-second embodiment are not repeated here; please refer to the corresponding parts of the ninety-second embodiment.
[1025] A question-and-answer method provided in this application may include the following steps:
[1026] Step 1: receiving second voice data of a first user in a target speaking mode;
[1027] Step 2: Determine the reply information based on the second voice data;
[1028] Step 3: Send a reply message to the terminal device.
[1029] In one example, step 2 may include the following sub-steps: 2.1) determining speaking style information and text information of the second voice data; 2.2) determining reply information based on the speaking style information and text information.
[1030] Ninety-fourth embodiment
[1031] In the above embodiment, a question-and-answer method is provided. Accordingly, this application also provides a question-and-answer device. This device corresponds to the embodiment of the above method. The parts of this embodiment that are identical to the first embodiment are not repeated here. Please refer to the corresponding parts in the first embodiment.
[1032] The present application provides a question-answering device comprising:
[1033] A voice receiving unit, configured to receive second voice data of a target speaking mode of a first user;
[1034] an information determining unit, configured to determine reply information based on the second voice data;
[1035] The information sending unit is used to send reply information back to the terminal device.
[1036] Ninety-fifth embodiment
[1037] This application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple. For relevant details, please refer to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[1038] An electronic device of this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a question-and-answer method. After the device is powered on and the program of the method is run through the processor, the device performs the following steps: receiving second voice data of a target speaking mode of a first user; determining reply information based on the second voice data; and sending the reply information back to the terminal device.
[1039] Ninety-sixth embodiment
[1040] In the above embodiment, a question-and-answer system is provided. Correspondingly, this application also provides a question-and-answer method, which can be performed by a terminal device, etc. This method corresponds to the embodiment of the above system. The parts of this embodiment that are identical to the ninety-second embodiment are not repeated here. Please refer to the corresponding parts of the ninety-second embodiment.
[1041] A question-and-answer method provided in this application may include the following steps:
[1042] Step 1: Collect first voice data of a first user and determine a target speaking style;
[1043] Step 2: converting the first voice data into second voice data in a target speaking mode;
[1044] Step 3: Send the second voice data to the server, so that the server determines the reply information according to the second voice data and sends the reply information back to the terminal device.
[1045] Ninety-seventh embodiment
[1046] In the above embodiment, a question-and-answer method is provided. Accordingly, this application also provides a question-and-answer device. This device corresponds to the embodiment of the above method. The parts of this embodiment that are identical to the first embodiment are not repeated here. Please refer to the corresponding parts in the first embodiment.
[1047] The present application provides a question-answering device comprising:
[1048] a voice collection unit, configured to collect first voice data of a first user and determine a target speaking mode;
[1049] A voice data unit, configured to convert the first voice data into second voice data in a target speaking mode;
[1050] The voice sending unit is used to send the second voice data to the server, so that the server determines the reply information according to the second voice data and sends the reply information back to the terminal device.
[1051] Ninety-eighth embodiment
[1052] This application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple. For relevant details, please refer to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[1053] An electronic device of this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a question-and-answer method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: first voice data of a first user is collected and a target speaking mode is determined; the first voice data is converted into second voice data of the target speaking mode; the second voice data is sent to a server, so that the server determines a reply message based on the second voice data and sends the reply message back to the terminal device.
[1054] Ninety-ninth embodiment
[1055] In the above embodiment, a speech conversion method is provided. Accordingly, this application also provides a speech interaction system. This system corresponds to the embodiment of the above method. The server-side processing in this embodiment, which is identical to that in the second embodiment, will not be repeated here. Please refer to the corresponding section in the second embodiment.
[1056] The present application provides a voice interaction system comprising: a first terminal device, a second terminal device and a server.
[1057] The terminal device includes but is not limited to mobile communication devices, namely: commonly known as mobile phones or smart phones, and also includes personal computers, PADs, iPads and other terminal devices.
[1058] Among them, the first terminal device is used to collect the first voice data of the first user, determine the target speaking method, and send a first voice data sending request to the server; and play the second voice data of the second user sent back by the second terminal device through the server; the server is used to send a request for the first voice data, convert the first voice data into third voice data of the target speaking method of the first user, and send the third voice data to the second terminal device; and send the second voice data sent by the second terminal device to the first terminal device; the second terminal device is used to play the third voice data; and collect the second voice data and send a second voice data sending request to the server.
[1059] For example, a customer service representative (first user) uses his / her first terminal device to communicate with a customer (second user) who uses a second terminal device by sending voice messages. Since the customer service representative is in a low mood at the time, his / her voice tone is naturally low. Out of professionalism, the user does not want the second user to perceive his / her emotions, so he / she can specify a voice speaking style, such as happy, fluent, clear, powerful, etc. Correspondingly, the server can convert the first user's voice data into voice data of the first user's target speaking style, and send the converted voice data to the second terminal device. The second terminal device plays the first user's voice data after the converted speaking style, which can effectively improve the quality of customer service communication.
[1060] In one example, the server determines the text sequence of the first voice data through a speech recognition algorithm, and generates third voice data with the target speaking style corresponding to the first voice data through a speech synthesis model of the first user's target speaking style.
[1061] The nineteenth embodiment will not be described in detail here.
[1062] In one example, the second user may also specify his or her speaking mode through the second terminal device, and the server may send the voice data of the second user's specified speaking mode corresponding to the second user's original voice data to the first terminal device for playback.
[1063] As can be seen from the above embodiments, the voice interaction system provided by the embodiments of the present application uses a first terminal device to collect first voice data of a first user, determine a target speaking method, and send a first voice data sending request to a server; and play the second voice data of a second user sent back by a second terminal device through the server; the server is used to send a request for the first voice data, convert the first voice data into third voice data of the first user's target speaking method, and send the third voice data to the second terminal device; and send the second voice data sent by the second terminal device to the first terminal device; the second terminal device is used to play the third voice data; and collect the second voice data and send a second voice data sending request to the server; this processing method enables the second terminal device to play the voice data of the first user's specified speaking method when the first user communicates with the second user by voice, rather than the actually collected voice data of the first user's real speaking method. The first user can conceal his true emotions from the second user by specifying the speaking method; therefore, the quality of voice interaction can be effectively improved, thereby improving the user experience.
[1064] The 100th embodiment
[1065] In the above-mentioned embodiment, a voice interaction system is provided. Correspondingly, this application also provides a voice interaction method, which can be executed by a server, etc. This method corresponds to the embodiment of the above-mentioned system. The parts of this embodiment that are identical to the ninety-ninth embodiment are not repeated here. Please refer to the corresponding parts of the ninety-ninth embodiment.
[1066] The present application provides a voice interaction method, which may include the following steps:
[1067] Step 1: In response to a first voice data sending request sent by a first terminal device, convert the first voice data of the first user into third voice data in the first user's target speaking mode, and send the third voice data to a second terminal device;
[1068] Step 2: In response to the second voice data sending request sent by the second terminal device, the second voice data sent by the second terminal device is sent to the first terminal device.
[1069] One Hundred and First Embodiment
[1070] In the above embodiment, a voice interaction method is provided. Correspondingly, this application also provides a voice interaction device. This device corresponds to the embodiment of the above method. The parts of this embodiment that are the same as the first embodiment are not repeated here. Please refer to the corresponding parts in the first embodiment.
[1071] The present application provides a voice interaction device comprising:
[1072] a first request processing unit, configured to send a request for the first voice data sent by the first terminal device, c...
Claims
1. A voice message sending system, characterized in that: include: The first terminal device is configured to collect first voice data of a first user and send a request to the server to change the first voice data into the voice of a second user; The server is configured to construct a speech synthesis model for each user; and to construct a speech posterior probability graph (PPG) feature extractor; determine, by the PPG feature extractor, PPG feature data of the first speech data based on first acoustic feature data of the first speech data, including voiceprint information and speech content information of the first user; generate, by the speech synthesis model of the second user, second speech data of the second user corresponding to the first speech data, based on the PPG feature data and second acoustic feature data of the first speech data, including prosody information; and transmit the second speech data to the second terminal device; The second terminal device is used to play the second voice data.
2. A film and television dubbing system, characterized in that: include: The terminal device is configured to collect first voice data of a first user for a film and television dialogue text, and send a request to a server to convert the first voice data into a dubbing of a second user; The server is used to construct a speech posterior probability graph PPG feature extractor and a speech synthesis model for each user; through the PPG feature extractor, based on the first acoustic feature data of the first speech data including the first user's voiceprint information and speech content information, the PPG feature data of the first speech data is determined; through the speech synthesis model of the second user, based on the PPG feature data and the second acoustic feature data of the first speech data including prosody information, the second speech data of the second user for the film and television dialogue text is generated.
3. A news broadcasting system, characterized in that: include: The terminal device is configured to collect first voice data of a first user of a news text to be broadcast in multiple languages, and send a request to a server to broadcast the news text to be broadcast in multiple languages by a second user voice; The server is used to construct a speech posterior probability graph PPG feature extractor and a speech synthesis model for each user; for the first speech data in each language, the PPG feature extractor is used to determine the PPG feature data of the first speech data based on the first acoustic feature data of the first speech data including the first user's voiceprint information and speech content information; and the speech synthesis model of the second user is used to generate second speech data in multiple languages broadcast by the second user based on the PPG feature data and the second acoustic feature data of the first speech data including prosody information.
4. A voice conversion method, characterized in that: include: Build a speech posterior probability graph (PPG) feature extractor and a speech synthesis model for each user; Determining, by the PPG feature extractor, PPG feature data of the first voice data according to first acoustic feature data of the first voice data of the first user, wherein the first acoustic feature data includes voiceprint information and voice content information of the first user; The second voice data of the second user corresponding to the first voice data is generated through the voice synthesis model of the second user according to the PPG feature data and the second acoustic feature data of the first voice data; the second acoustic feature data includes prosody information.
5. The method according to claim 4, characterized in that The generating, by the speech synthesis model of the second user, second speech data of the second user corresponding to the first speech data according to the PPG feature data and second acoustic feature data of the first speech data, includes: Determining, by a feature synthesis network included in the speech synthesis model, third acoustic feature data having a second user timbre based on the PPG feature data and the second acoustic feature data; The second speech data is generated according to the third acoustic feature data by a vocoder included in the speech synthesis model.
6. The method according to claim 5, characterized in that Also includes: Determine the sampling rate of the PPG feature data as the PPG feature sampling rate of the feature synthesis network; determining, by a feature synthesis network, third acoustic feature data having a second user timbre and corresponding to the sampled PPG feature data based on the PPG feature data and the second acoustic feature data; The second speech data having a speech playback speed corresponding to the sampling rate is generated according to the third acoustic feature data through the vocoder included in the speech synthesis model.
7. The method according to claim 6, characterized in that Determining the sampling rate of the PPG feature data includes: determining a playback speed of the second voice data; The sampling rate is determined according to the playback speed.
8. The method according to claim 6, characterized in that The structure of the feature synthesis network includes: a FastSpeech model.
9. The method according to claim 8, characterized in that The FastSpeech model does not include: predictor and length regulator.
10. The method according to claim 5, characterized in that Also includes: For each user, the feature synthesis network is learned from the corresponding relationship set between the PPG feature data, the second acoustic feature data and the third acoustic feature data of the user's voice data.
11. The method according to claim 5, characterized in that Also includes: For each user, the vocoder is learned from a set of correspondences between the third acoustic feature data of the user's speech data and the speech synthesis data.
12. The method according to claim 4, characterized in that The speech synthesis model of each user is constructed in the following way: For each user, a speech synthesis model of the user is learned based on the correspondence between the PPG feature data, the second acoustic feature data and the speech synthesis data of the user's speech data.
13. The method according to claim 4, characterized in that The speech posterior probability graph PPG feature extractor is constructed in the following way: The extractor is obtained by learning the correspondence between the first acoustic feature data of the speech data of multiple users and the pronunciation annotation information.
14. A method for sending a voice message, characterized in that: include: collecting first voice data of a first user; A request is sent to the server to change the first voice data into the voice of the second user, so that the server constructs a voice posterior probability graph PPG feature extractor and a voice synthesis model for each user; through the PPG feature extractor, the PPG feature data of the first voice data is determined according to the first acoustic feature data of the first voice data including the voiceprint information of the first user and the voice content information; through the voice synthesis model of the second user, the second voice data of the second user corresponding to the first voice data is generated according to the PPG feature data and the second acoustic feature data of the first voice data including prosody information; and the second voice data is sent to the second terminal device.
15. A film and television dubbing method, characterized in that: include: Collecting first voice data of a first user for a film and television dialogue text in a first language; A request is sent to the server to convert the first voice data into the dubbing of the second user, so that the server constructs a voice posterior probability graph PPG feature extractor and a voice synthesis model for each user; through the PPG feature extractor, the PPG feature data of the first voice data is determined based on the first acoustic feature data of the first voice data including the voiceprint information of the first user and the voice content information; through the voice synthesis model of the second user, the second voice data of the second user for the film and television dialogue text is generated based on the PPG feature data and the second acoustic feature data of the first voice data including prosody information.
16. A news broadcasting method, characterized in that: include: collecting first voice data of a first user of a news text to be broadcast in multiple languages; A request is sent to the server for the second user to voice-broadcast the news text in multiple languages, so that the server constructs a speech posterior probability graph PPG feature extractor and a speech synthesis model for each user; for the first speech data in each language, the PPG feature extractor is used to determine the PPG feature data of the first speech data based on the first acoustic feature data of the first speech data including the voiceprint information of the first user and speech content information; and the speech synthesis model of the second user is used to generate the second speech data in multiple languages broadcast by the second user based on the PPG feature data and the second acoustic feature data of the first speech data including prosody information.
17. A voice interaction system, characterized in that: include: The first terminal device is configured to determine voice interaction information of the first user to the second user and send the voice interaction information to the server; The server is configured to construct a speech synthesis model for multiple speaking styles of the first user; construct a speech posterior probability graph (PPG) feature extractor; determine PPG feature data of the speech interaction information using the PPG feature extractor; determine target speaking style information of the first user corresponding to the speech interaction information; generate speech data of the speech interaction information having the target speaking style based on the PPG feature data using the speech synthesis model for the target speaking style of the first user, and transmit the speech data to a second terminal device of the second user; The second terminal device is used to play the voice data.
18. A voice interaction method, characterized in that: include: Constructing a speech synthesis model for multiple speaking styles of the first user; and constructing a speech posterior probability graph (PPG) feature extractor; Determining PPG feature data of the voice interaction information by the PPG feature extractor; Determining target speaking mode information of the first user based on voice interaction information of the first user to the second user; Through the speech synthesis model of the target speaking mode of the first user, based on the PPG feature data, speech data with the target speaking mode of the speech interaction information is generated, and the speech data is sent to the terminal device of the second user.
19. The method according to claim 18, characterized in that The determining the target speaking style information of the first user includes: The target speaking manner information is determined according to the relationship information between the second user and the first user.
20. The method according to claim 18, wherein The determining the target speaking style information of the first user includes: Determining the domain to which the interactive information belongs; The target speaking mode information is determined according to the domain information and domains of the first user and the second user respectively.
21. A cross-dialect speech generation system, characterized in that: include: The terminal device is configured to determine a target text and a target dialect, send a request to a server for a first user to read the text in the target dialect, and play first voice data sent back by the server, which is the first user reading the text in the target dialect. The server is used to construct a speech posterior probability graph (PPG) feature extractor and a speech synthesis model for each user; and, in response to the request, determine second speech data of the second user reading the text in the target dialect, and determine the PPG feature data of the second speech data through the PPG feature extractor based on the first acoustic feature data of the second speech data including the second user's voiceprint information and speech content information; and generate the first speech data through the speech synthesis model of the first user based on the PPG feature data and the second acoustic feature data of the second speech data including prosody information.
22. A voice changing system, characterized in that: include: The terminal device is configured to determine first voice data of a first user and send a request to a server to change the first voice data into a second user's voice; The server is used to build a speech synthesis model for each user; and to build a speech posterior probability graph PPG feature extractor; through the PPG feature extractor, based on the first acoustic feature data of the first speech data including the voiceprint information of the first user and the speech content information, the PPG feature data of the first speech data is determined; through the speech synthesis model of the second user, based on the PPG feature data and the second acoustic feature data of the first speech data including prosody information, the second speech data of the second user corresponding to the first speech data is generated.
23. A method for changing voice, characterized in that: include: Build a speech posterior probability graph (PPG) feature extractor and a speech synthesis model for each user; In response to a request to morph first voice data of a first user into a voice of a second user, determining, by the PPG feature extractor, PPG feature data of the first voice data based on first acoustic feature data of the first voice data including first user voiceprint information and voice content information; The second voice data of the second user corresponding to the first voice data is generated by the voice synthesis model of the second user based on the PPG feature data and the second acoustic feature data of the first voice data including prosody information.
24. A method for changing voice, characterized in that: include: Determining first voice data of a first user; Determining, by a speech posterior probability graph (PPG) feature extractor, PPG feature data of the first speech data based on first acoustic feature data of the first speech data including first user voiceprint information and speech content information; The second voice data of the second user corresponding to the first voice data is generated based on the PPG feature data and the second acoustic feature data of the first voice data including prosody information through the speech synthesis model of the second user.
25. A method for constructing a speech conversion model, characterized in that: include: Construct a speech synthesis model for each user and a speech posterior probability graph PPG feature extractor to form a speech conversion model; wherein the PPG feature extractor is used to determine the PPG feature data of the first speech data based on the first acoustic feature data of the first speech data of the first user; the first acoustic feature data includes the first user's voiceprint information and speech content information; the speech synthesis model is used to generate the second speech data of the second user corresponding to the first speech data based on the PPG feature data and the second acoustic feature data of the first speech data; the second acoustic feature data includes prosody information.
26. The method according to claim 25, characterized in that The step of constructing a speech synthesis model for each user includes: For each user, a speech synthesis model of the user is learned based on the correspondence between the PPG feature data, the second acoustic feature data and the speech synthesis data of the user's speech data.
27. The method according to claim 25, characterized in that The speech synthesis model includes: a feature synthesis network and a vocoder; wherein the feature synthesis network is used to determine third acoustic feature data having a second user's timbre based on the PPG feature data and the second acoustic feature data; and the vocoder is used to generate the second speech data based on the third acoustic feature data; The method further comprises: For each user, learning the feature synthesis network from a set of correspondences between the PPG feature data, the second acoustic feature data, and the third acoustic feature data of the user's speech data; For each user, the vocoder is learned from a set of correspondences between the third acoustic feature data of the user's speech data and the speech synthesis data.
28. The method according to claim 25, characterized in that The PPG feature extractor is constructed as follows: The extractor is obtained by learning the correspondence between the first acoustic feature data of the speech data of multiple users and the pronunciation annotation information.
29. A voice conversion device, characterized in that: include: A model building unit, configured to build a speech synthesis model for each user; and a speech posterior probability graph (PPG) feature extractor; a PPG feature extraction unit, configured to determine, by the PPG feature extractor, PPG feature data of the first voice data based on first acoustic feature data of the first voice data of the first user, wherein the first acoustic feature data includes voiceprint information and voice content information of the first user; A speech synthesis unit is used to generate second speech data of the second user corresponding to the first speech data based on the PPG feature data and second acoustic feature data of the first speech data through the speech synthesis model of the second user; the second acoustic feature data includes prosody information.
Citation Information
Patent Citations
Cross-language emotional speech synthesis method and system
CN107103900A
Voice conversion method, device and equipment and readable storage medium
CN110223705A